arXiv ScienceSearch

arXiv subjects

Payel Das

Publications and source records attributed to Payel Das.

At least 19 recordsLinked to original sources

Can Vision-Language Models Reason about AI Edits in Images?

Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Language Models (VLMs) offer a promising alternative due to their strong visual understanding and reasoning capabilities; however, existing approaches typically rely on supervised finetuning with curated explanations rather than exploiting their inherent reasoning capabilities. In this work, we investigate whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision. Motivated by the success in Group Relative Policy Optimization (GRPO), an RL technique that incentivizes the model to reason by generating thinking traces prior to giving the final answer, we propose a GRPO-based training framework that utilizes simple accuracy and format rewards. Given an input image, the model produces a structured reasoning trace and predicts whether the image has been tampered with. A lightweight segmentation model is then guided by the reasoning output to generate pixel-level localization masks. Experiments across multiple image manipulation datasets demonstrate that our approach achieves competitive detection and localization performance compared to state-of-the-art image forgery detectors, despite requiring substantially weaker supervision. We introduce effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization. These results suggest that reinforcement learning provides an effective and scalable mechanism for teaching VLMs to reason about AI-generated content.

cs.CV

BioMamba: Domain-Adaptive Biomedical Language Models

Background. Biomedical language models should improve performance on biomedical text while retaining general-language-modeling fluency. For Mamba-based models, this trade-off has not been systematically studied across biomedical literature and clinical text. Methods. We developed BioMamba, a family of biomedical Mamba2 models at five scales obtained by continued pretraining of released public Mamba2 checkpoints on a balanced 80%/10%/10% mixture of PubMed abstracts, the Colossal Clean Crawled Corpus (C4), and Wikipedia. The contribution is the adaptation recipe and the accompanying open-weight checkpoints. Results. Across five scales, BioMamba consistently lowered PubMed perplexity, improved Wikipedia-style held-out perplexity by 1.46-4.72 PPL, and left C4 perplexity essentially unchanged. On six out-of-domain multiple-choice benchmarks, BioMamba stayed within +/-3 percentage points of Mamba2 with no systematic regression. After supervised fine-tuning, BioMamba+SFT matched or exceeded Mamba2+SFT on MIMIC-IV note completion and discharge summary generation at every evaluated scale, and improved PubMedQA at every scale. The strongest model (BioMamba-2.7B) reached a PubMed perplexity of 5.28 and accuracies of 90.24% and 73.00% on BioASQ and PubMedQA, respectively. Conclusions. A balanced domain-adaptive continued pretraining recipe strengthens Mamba2 language models on biomedical literature and clinical text while preserving general-language-modeling fluency.

cs.CL

Reconstructing chemical enrichment pathways in disc galaxies: A phylogenetic approach

Phylogenetic methods, traditionally used in biology to trace the evolutionary relationships among species, are emerging as a powerful framework to reconstruct evolutionary processes in galaxies from chemical information. We apply galactic phylogenetics to study the chemical evolution of stellar populations in distinct regions of a simulated disc galaxy, assessing its capability to unveil assembly histories. We used a high-resolution simulation that follows the chemical enrichment of an isolated disc galaxy, by different stellar progenitors. We track gas particles as they turn into stars and inherit their parent gas chemical composition. Target particles are selected to store the chemical history of each chemical element considered in the simulation. Two regions were analysed: an inner ring, influenced by early bar-driven inflows, and an outer ring, shaped by spiral arms. We built phylogenetic trees for stellar populations in each region and quantified their structure using the Corrected Colless index, a standard metric of tree balance used in biology. The inner ring tree reveals a compact clade of old stars enriched by rapid SNII feedback, followed by a hierarchical sequence with increasing SNIa and AGB contributions. In contrast, the outer ring exhibits more symmetric, caterpillar-like trees with smoother abundance gradients, consistent with more prolonged star formation and efficient local mixing. Chemical enrichment rates corroborate these trends, showing fast early enrichment in the inner ring and gradual, spatially extended enrichment in the outer disc. The structural indices differ significantly between the two regions and converge robustly even for modest stellar samples (NSSP = 100). Galactic phylogenetics provides a novel and complementary tool to decode the fossil record of galaxies.

astro-ph.GA

Disentangling chemical evolution histories with phylogenetic trees

Chemical abundances encode the fossil record of galaxy evolution in a complex and diverse way that requires innovative approaches to reconstruct galactic histories. We investigate the power of using phylogenetic methods to disentangle different evolutionary pathways in analytical chemical evolution models. We ran 1024 one-zone chemical evolution models using flexCE. The resulting chemical abundances are combined with those of two fiducial models, mw-fid and dw-fid, and then used both to determine which combinations produce two-branched phylogenetic trees, as well as how purely these trees split the two input models. We used random forests and Shapley analysis to predict which model combinations return well-separated trees and explain which input parameters are most important for this. We also studied the abundance patterns, as well as star formation rates, mass accumulation, and branch lengths. We found that η, the mass-loading outflow parameter in flexCE, had the largest impact in separating models into separate branches, due to its importance in driving the chemical enrichment rates and total abundances. Star formation rates and mass accumulation had some impact on η, but no direct relation between these quantities and the abundances was found. We also found that branches connected through the most metal rich tips in our trees, which is opposite to how phylogenetic trees connect in biological systems. Phylogenetic trees help to reconstruct histories when there is information that is inherited between generations, which is the case of the chemical elements in galaxy evolution. Branch topologies can provide information about the rates of evolutionary change of the various populations, and the connection between branches also contains information about their shared history. This work brings us a step further understanding galaxy evolution through cross-disciplinary research.

astro-ph.GA

Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models

We propose patching for large language models (LLMs) like software versions, a lightweight and modular approach for addressing safety vulnerabilities. While vendors release improved LLM versions, major releases are costly, infrequent, and difficult to tailor to customer needs, leaving released models with known safety gaps. Unlike full-model fine-tuning or major version updates, our method enables rapid remediation by prepending a compact, learnable prefix to an existing model. This "patch" introduces only 0.003% additional parameters, yet reliably steers model behavior toward that of a safer reference model. Across three critical domains (toxicity mitigation, bias reduction, and harmfulness refusal) policy patches achieve safety improvements comparable to next-generation safety-aligned models while preserving fluency. Our results demonstrate that LLMs can be "patched" much like software, offering vendors and practitioners a practical mechanism for distributing scalable, efficient, and composable safety updates between major model releases.

cs.AI

A method for constructing the joint mass function of binary stars

The initial mass function (IMF) describes the distribution of stellar masses in a population of newly born stars and is amongst the most fundamental concepts in astrophysics. It is not only the direct result of the star formation process but it also explains the evolution of galaxies' luminosities, metal yields, star-formation efficiencies, and supernova production rates. Because most stars exist in binary systems, however, a full statistical account of stellar mass requires not the IMF but rather the joint distribution of a binary population's primary- and secondary-star masses. This joint distribution must respect the IMF of the stars from which the population has been assembled as well as the distribution of mass ratios that results from the assembly mechanism. Despite its importance, this joint distribution is known only in the case of random pairing. Here we present a method for constructing it in the general case. We also illustrate the use of our method by recovering the known result for random pairing and by finding the previously unknown result for uniform pairing.

astro-ph.SR

Continuous-Time Modelling of Black Hole Binary Evolution with Neural ODEs

Pulsar timing arrays (PTAs) can detect the low-frequency stochastic gravitational-wave background (GWB) generated by an ensemble of supermassive black hole binaries (BHBs). Accurate determination of BHB merger timescales is essential for interpreting GWBs and constraining key astrophysical quantities such as black hole (BH) occupation fractions and galaxy coalescence rates. High-accuracy $N$-body codes such as \texttt{Griffin} can resolve sub-pc BHB dynamics but are too costly to explore a wide range of initial conditions, motivating the need for surrogate models that emulate their long-term evolution at much lower computational cost. We investigate neural ordinary differential equations (NODEs) as surrogates for the secular orbital evolution of BHBs. Our primary contribution is a parameterised NODE (PNODE) trained on an ensemble of $N$-body simulations of galaxy mergers spanning a two-dimensional parameter space defined by the initial orbital eccentricity and particle resolution $(e_i, N)$, with the learned vector field explicitly conditioned on these parameters. A single PNODE thereby learns a simulation-parameter-conditioned dynamical model for the coupled evolution of the BH pair's orbital state across the ensemble, yielding smooth trajectories from which stable hardening and eccentricity growth rates can be extracted. The PNODE accurately reproduces the secular evolution of the specific orbital energy and angular momentum, and the corresponding Keplerian orbital elements, for held-out trajectories, with modest generalisation to a partially unseen high-resolution case. Combining PNODE predictions with semi-analytical prescriptions for stellar hardening and gravitational-wave emission yields BHB merger timescales consistent with those obtained from direct $N$-body inputs within current theoretical uncertainties.

astro-ph.GA

GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance

The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology. We introduce in this paper an efficient training-free method for navigating and sampling from the molecular space with a generative Chemical Language Model (CLM), while using the molecular similarity to the target as a guide. Our method leverages the contextual representations learned from the CLM itself to estimate the molecular similarity, which is then used to adjust the autoregressive sampling strategy of the CLM. At each step of the decoding process, the method tracks the distance of the current generations from the target and updates the logits to encourage the preservation of similarity in generations. We implement the method using a recently proposed $\sim$47M parameter SMILES-based CLM, GP-MoLFormer, and therefore refer to the method as GP-MoLFormer-Sim, which enables a test-time update of the deep generative policy to reflect the contextual similarity to a set of guide molecules. The method is further integrated into a genetic algorithm (GA) and tested on a set of standard molecular optimization benchmarks involving property optimization, molecular rediscovery, and structure-based drug design. Results show that, GP-MoLFormer-Sim, combined with GA (GP-MoLFormer-Sim+GA) outperforms existing training-free baseline methods, when the oracle remains black-box. The findings in this work are a step forward in understanding and guiding the generative mechanisms of CLMs.

cs.LG

How much can we learn from resolved stellar kinematics of galactic haloes using action-based dynamical models?

Dynamical models are used to study dark matter (DM) in galaxies, how galaxies assemble through mergers, and to test galaxy formation models. Despite its widespread use, there has been no systematic study quantifying how much information can be obtained from just two on-sky positions and line-of-sight velocities, which are typically available for nearby external galaxies. In this work, we introduce axisymmetric, action-based dynamical models that use the positions and velocities of stellar halo stars to jointly constrain the total mass distribution of galaxies and the underlying DM component, as well as the stellar halo phase-space distribution. We rigorously test the method using both idealised equilibrium galaxy mocks and cosmological hydrodynamical simulations from the Auriga suite, systematically assessing how its performance degrades as the available phase-space information is progressively reduced. We further examine the impact of galaxy inclination, modelling assumptions, and methodological systematics on the recovered mass profiles. A crucial development in this work is the improved marginalisation of the model likelihood over missing phase-space dimensions. Our models successfully recover the total and DM mass distributions, as well as the kinematic properties of the stellar tracers, within the derived confidence intervals. However, we find that with limited (3D or 4D) phase-space information, the flattening of the DM halo cannot be constrained with any degree of certainty. Nevertheless, the recovered mass profile is insensitive to the flattening. This finding is independently validated by Schwarzschild modelling tests.

astro-ph.GA

Stellar velocity distributions in binary-rich ultrafaint dwarf galaxies

Ultrafaint dwarf (UFD) galaxies are dominated by dark matter, the distribution of which may be inferred from the kinematics of that galaxy's stellar population. Star-by-star observations are available for the satellite UFD galaxies of the Milky Way, making them uniquely good laboratories in which to test cosmological predictions at the smallest scales. However, the kinematics of these galaxies are complicated by the presence of binary stars, which alter the stellar velocity distribution. In particular these binary stars increase the galaxy's stellar velocity dispersion, which is related to the total galactic mass by the virial theorem. Without correctly eliminating or accounting for binary stars we may therefore overestimate the masses of UFD galaxies or even confuse globular clusters for UFD galaxies. Here we write down the probability density function for the observed line-of-sight (LOS) velocity of a stellar population containing both visual and spectroscopic binary stars, which we then use to determine the effect of those binary stars on the observed LOS velocity dispersion. For the coldest UFD galaxies the fractional increase in LOS velocity dispersion is of order one and for the coldest globular clusters is of order 100. However, if the stellar initial mass function is bottom light, as it may be for UFD galaxies and globular clusters, then both of these values increase by half a dex.

astro-ph.GA

EDGE: The emergence of dwarf galaxy scaling relations from cosmological radiation-hydrodynamics simulations

We present a new suite of EDGE (`Engineering Dwarfs at Galaxy formation's Edge') cosmological zoom simulations. The suite includes 15 radiation-hydrodynamical dwarf galaxies covering the ultra-faint to the dwarf irregular regime ($10^4 \leq M_{\star}(z=0) \leq 10^8 \, M_{\odot}$) to enable comparisons with observed scaling relations. Each object in the suite is evolved at high resolution ($\approx 3 \, \text{pc}$) and includes stellar radiation, winds and supernova feedback channels. We compare with previous \textsc{edge} simulations without radiation, finding that radiative feedback results in significantly weaker galactic outflows. This generalizes our previous findings to a wide mass range, and reveals that the effect is most significant at low $M_{\star}$. Despite this difference, stellar masses stay within a factor of two of each other, and key scaling relations of dwarf galaxies (size-mass, neutral gas-stellar mass, gas-phase mass-metallicity) emerge correctly in both simulation suites. Only the stellar mass -- stellar metallicity relation is strongly sensitive to the change in feedback. This highlights how obtaining statistical samples of dwarf galaxy stellar abundances with next-generation spectrographs will be key to probing and constraining the baryon cycle of dwarf galaxies.

astro-ph.GA

A North-South Metallicity Asymmetry in the Outer Galactic disk -- Evidence for the Pericentric Passage of the Sagittarius Dwarf Galaxy

We present maps of the mean metallicity distributions on the Galactocentric $R$--$Z$ plane at different azimuthal angles using red clump stars selected from the LAMOST and APOGEE surveys. In the inner disk ($R < $ 11\,kpc), the metallicity distribution is symmetric between the upper and lower disk. However, we find a North-South metallicity asymmetry in the outer disk ($R > 11$\,kpc), especially towards the anti-Galactic center ($-5^\circ < Φ< 15^\circ$) direction. By further dissecting the map in age space, we detect this asymmetry across all mono-age stellar populations. However, the asymmetry is less pronounced in older populations ($τ> 8$ Gyr) compared to younger ones ($τ< 6$\,Gyr). This reduced significance likely stems from three factors: larger age uncertainties, fewer stars in the outer disk, and the kinematically hotter nature of older populations. The observed metallicity asymmetry may be the consequence of the purturbation of the recent pericentric passage through the Galactic disk and tidal force of the well-known Sagittarius dwarf galaxy.

astro-ph.GA

Position: Theory of Mind Benchmarks are Broken for Large Language Models

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks for LLMs are overwhelmingly inspired by the methods used to test theory of mind in humans and fall victim to a fallacy of attributing human-like qualities to AI agents. We expect that humans will engage in a consistent reasoning process across various questions about a situation, but this is known to not be the case for current LLMs. Most theory of mind benchmarks only measure what we call literal theory of mind: the ability to predict the behavior of others. However, this type of metric is only informative when agents exhibit self-consistent reasoning. Thus, we introduce the concept of functional theory of mind: the ability to adapt to agents in-context following a rational response to their behavior. We find that many open source LLMs are capable of displaying strong literal theory of mind capabilities, but seem to struggle with functional theory of mind -- even with exceedingly simple partner policies. Simply put, strong literal theory of mind performance does not necessarily imply strong functional theory of mind performance or vice versa. Achieving functional theory of mind, particularly over long interaction horizons with a partner, is a significant challenge deserving a prominent role in any meaningful LLM theory of mind evaluation.

cs.AI

Combining Domain and Alignment Vectors to Achieve Better Knowledge-Safety Trade-offs in LLMs

There is a growing interest in training domain-expert LLMs that excel in specific technical fields compared to their general-purpose instruction-tuned counterparts. However, these expert models often experience a loss in their safety abilities in the process, making them capable of generating harmful content. As a solution, we introduce an efficient and effective merging-based alignment method called \textsc{MergeAlign} that interpolates the domain and alignment vectors, creating safer domain-specific models while preserving their utility. We apply \textsc{MergeAlign} on Llama3 variants that are experts in medicine and finance, obtaining substantial alignment improvements with minimal to no degradation on domain-specific benchmarks. We study the impact of model merging through model similarity metrics and contributions of individual models being merged. We hope our findings open new research avenues and inspire more efficient development of safe expert LLMs.

cs.AI

Aligning Protein Conformation Ensemble Generation with Physical Feedback

Protein dynamics play a crucial role in protein biological functions and properties, and their traditional study typically relies on time-consuming molecular dynamics (MD) simulations conducted in silico. Recent advances in generative modeling, particularly denoising diffusion models, have enabled efficient accurate protein structure prediction and conformation sampling by learning distributions over crystallographic structures. However, effectively integrating physical supervision into these data-driven approaches remains challenging, as standard energy-based objectives often lead to intractable optimization. In this paper, we introduce Energy-based Alignment (EBA), a method that aligns generative models with feedback from physical models, efficiently calibrating them to appropriately balance conformational states based on their energy differences. Experimental results on the MD ensemble benchmark demonstrate that EBA achieves state-of-the-art performance in generating high-quality protein ensembles. By improving the physical plausibility of generated structures, our approach enhances model predictions and holds promise for applications in structural biology and drug discovery.

q-bio.BM

PEEL the Layers and Find Yourself: Revisiting Inference-time Data Leakage for Residual Neural Networks

This paper explores inference-time data leakage risks of deep neural networks (NNs), where a curious and honest model service provider is interested in retrieving users' private data inputs solely based on the model inference results. Particularly, we revisit residual NNs due to their popularity in computer vision and our hypothesis that residual blocks are a primary cause of data leakage owing to the use of skip connections. By formulating inference-time data leakage as a constrained optimization problem, we propose a novel backward feature inversion method, \textbf{PEEL}, which can effectively recover block-wise input features from the intermediate output of residual NNs. The surprising results in high-quality input data recovery can be explained by the intuition that the output from these residual blocks can be considered as a noisy version of the input and thus the output retains sufficient information for input recovery. We demonstrate the effectiveness of our layer-by-layer feature inversion method on facial image datasets and pre-trained classifiers. Our results show that PEEL outperforms the state-of-the-art recovery methods by an order of magnitude when evaluated by mean squared error (MSE). The code is available at \href{https://github.com/Huzaifa-Arif/PEEL}{https://github.com/Huzaifa-Arif/PEEL}

cs.LG

GP-MoLFormer: A Foundation Model For Molecular Generation

Transformer-based models trained on large and general purpose datasets consisting of molecular strings have recently emerged as a powerful tool for successfully modeling various structure-property relations. Inspired by this success, we extend the paradigm of training chemical language transformers on large-scale chemical datasets to generative tasks in this work. Specifically, we propose GP-MoLFormer, an autoregressive molecular string generator that is trained on more than 1.1B (billion) chemical SMILES. GP-MoLFormer uses a 46.8M parameter transformer decoder model with linear attention and rotary positional encodings as the base architecture. GP-MoLFormer's utility is evaluated and compared with that of existing baselines on three different tasks: de novo generation, scaffold-constrained molecular decoration, and unconstrained property-guided optimization. While the first two are handled with no additional training, we propose a parameter-efficient fine-tuning method for the last task, which uses property-ordered molecular pairs as input. We call this new approach pair-tuning. Our results show GP-MoLFormer performs better or comparable with baselines across all three tasks, demonstrating its general utility for a variety of molecular generation tasks. We further report strong memorization of training data in GP-MoLFormer generations, which has so far remained unexplored for chemical language models. Our analyses reveal that training data memorization and novelty in generations are impacted by the quality and scale of the training data; duplication bias in training data can enhance memorization at the cost of lowering novelty. We further establish a scaling law relating inference compute and novelty in generations.

q-bio.BM

Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models

Fine-tuning Large Language Models (LLMs) on some task-specific datasets has been a primary use of LLMs. However, it has been empirically observed that this approach to enhancing capability inevitably compromises safety, a phenomenon also known as the safety-capability trade-off in LLM fine-tuning. This paper presents a theoretical framework for understanding the interplay between safety and capability in two primary safety-aware LLM fine-tuning strategies, providing new insights into the effects of data similarity, context overlap, and alignment loss landscape. Our theoretical results characterize the fundamental limits of the safety-capability trade-off in LLM fine-tuning, which are also validated by numerical experiments.

stat.ML