arXiv ScienceSearch

arXiv subjects

Ankush Kumar

Publications and source records attributed to Ankush Kumar.

9 recordsLinked to original sources

T2I-BiasBench: A Multi-Metric Framework for Auditing Demographic and Cultural Bias in Text-to-Image Models

Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in training data. We introduce T2I-BiasBench, a unified evaluation framework of thirteen complementary metrics that jointly captures demographic bias, element omission, and cultural collapse in diffusion models - the first framework to address all three dimensions simultaneously. We evaluate three open-source models - Stable Diffusion v1.5, BK-SDM Base, and Koala Lightning - against Gemini 2.5 Flash (RLHF-aligned) as a reference baseline. The benchmark comprises 1,574 generated images across five structured prompt categories. T2I-BiasBench integrates six established metrics with seven additional measures: four newly proposed (Composite Bias Score, Grounded Missing Rate, Implicit Element Missing Rate, Cultural Accuracy Ratio) and three adapted (Hallucination Score, Vendi Score, CLIP Proxy Score). Three key findings emerge: (1) Stable Diffusion v1.5 and BK-SDM exhibit bias amplification (>1.0) in beauty-related prompts; (2) contextual constraints such as surgical PPE substantially attenuate professional-role gender bias (Doctor CBS = 0.06 for SD v1.5); and (3) all models, including RLHF-aligned Gemini, collapse to a narrow set of cultural representations (CAS: 0.54-1.00), confirming that alignment techniques do not resolve cultural coverage gaps. T2I-BiasBench is publicly released to support standardized, fine-grained bias evaluation of generative models. The project page is available at: https://gyanendrachaubey.github.io/T2I-BiasBench/

cs.CV

State-of-the-art Small Language Coder Model: Mify-Coder

We present Mify-Coder, a 2.5B-parameter code model trained on 4.2T tokens using a compute-optimal strategy built on the Mify-2.5B foundation model. Mify-Coder achieves comparable accuracy and safety while significantly outperforming much larger baseline models on standard coding and function-calling benchmarks, demonstrating that compact models can match frontier-grade models in code generation and agent-driven workflows. Our training pipeline combines high-quality curated sources with synthetic data generated through agentically designed prompts, refined iteratively using enterprise-grade evaluation datasets. LLM-based quality filtering further enhances data density, enabling frugal yet effective training. Through disciplined exploration of CPT-SFT objectives, data mixtures, and sampling dynamics, we deliver frontier-grade code intelligence within a single continuous training trajectory. Empirical evidence shows that principled data and compute discipline allow smaller models to achieve competitive accuracy, efficiency, and safety compliance. Quantized variants of Mify-Coder enable deployment on standard desktop environments without requiring specialized hardware.

cs.SE

Attention via Synaptic Plasticity is All You Need: A Biologically Inspired Spiking Neuromorphic Transformer

Attention is the brain's ability to selectively focus on a few specific aspects while ignoring irrelevant ones. This biological principle inspired the attention mechanism in modern Transformers. Transformers now underpin large language models (LLMs) such as GPT, but at the cost of massive training and inference energy, leading to a large carbon footprint. While brain attention emerges from neural circuits, Transformer attention relies on dot-product similarity to weight elements in the input sequence. Neuromorphic computing, especially spiking neural networks (SNNs), offers a brain-inspired path to energy-efficient intelligence. Despite recent work on attention-based spiking Transformers, the core attention layer remains non-neuromorphic. Current spiking attention (i) relies on dot-product or element-wise similarity suited to floating-point operations, not event-driven spikes; (ii) keeps attention matrices that suffer from the von Neumann bottleneck, limiting in-memory computing; and (iii) still diverges from brain-like computation. To address these issues, we propose the Spiking STDP Transformer (S$^{2}$TDPT), a neuromorphic Transformer that implements self-attention through spike-timing-dependent plasticity (STDP), embedding query--key correlations in synaptic weights. STDP, a core mechanism of memory and learning in the brain and widely studied in neuromorphic devices, naturally enables in-memory computing and supports non-von Neumann hardware. On CIFAR-10 and CIFAR-100, our model achieves 94.35\% and 78.08\% accuracy with only four timesteps and 0.49 mJ on CIFAR-100, an 88.47\% energy reduction compared to a standard ANN Transformer. Grad-CAM shows that the model attends to semantically relevant regions, enhancing interpretability. Overall, S$^{2}$TDPT illustrates how biologically inspired attention can yield energy-efficient, hardware-friendly, and explainable neuromorphic models.

cs.NE

Memristive Nanowire Network for Energy Efficient Audio Classification: Pre-Processing-Free Reservoir Computing with Reduced Latency

Efficient audio feature extraction is critical for low-latency, resource-constrained speech recognition. Conventional preprocessing techniques, such as Mel Spectrogram, Perceptual Linear Prediction (PLP), and Learnable Spectrogram, achieve high classification accuracy but require large feature sets and significant computation. The low-latency and power efficiency benefits of neuromorphic computing offer a strong potential for audio classification. Here, we introduce memristive nanowire networks as a neuromorphic hardware preprocessing layer for spoken-digit classification, a capability not previously demonstrated. Nanowire networks extract compact, informative features directly from raw audio, achieving a favorable trade-off between accuracy, dimensionality reduction from the original audio size (data compression) , and training time efficiency. Compared with state-of-the-art software techniques, nanowire features reach 98.95% accuracy with 66 times data compression (XGBoost) and 97.9% accuracy with 255 times compression (Random Forest) in sub-second training latency. Across multiple classifiers nanowire features consistently achieve more than 90% accuracy with more than 62.5 times compression, outperforming features extracted by conventional state-of-the-art techniques such as MFCC in efficiency without loss of performance. Moreover, nanowire features achieve 96.5% accuracy classifying multispeaker audios, outperforming all state-of-the-art feature accuracies while achieving the highest data compression and lowest training time. Nanowire network preprocessing also enhances linear separability of audio data, improving simple classifier performance and generalizing across speakers. These results demonstrate that memristive nanowire networks provide a novel, low-latency, and data-efficient feature extraction approach, enabling high-performance neuromorphic audio classification.

cs.SD

Narrow Transformer: StarCoder-Based Java-LM For Desktop

This paper presents NT-Java-1.1B, an open-source specialized code language model built on StarCoderBase-1.1B, designed for coding tasks in Java programming. NT-Java-1.1B achieves state-of-the-art performance, surpassing its base model and majority of other models of similar size on MultiPL-E Java code benchmark. While there have been studies on extending large, generic pre-trained models to improve proficiency in specific programming languages like Python, similar investigations on small code models for other programming languages are lacking. Large code models require specialized hardware like GPUs for inference, highlighting the need for research into building small code models that can be deployed on developer desktops. This paper addresses this research gap by focusing on the development of a small Java code model, NT-Java-1.1B, and its quantized versions, which performs comparably to open models around 1.1B on MultiPL-E Java code benchmarks, making them ideal for desktop deployment. This paper establishes the foundation for specialized models across languages and sizes for a family of NT Models.

cs.SE

Dendritic organic electrochemical transistors grown by electropolymerization for 3D neuromorphic engineering

One of the major limitation of standard top-down technologies used in today's neuromorphic engineering is their inability to map the 3D nature of biological brains. Here, we show how bipolar electropolymerization can be used to engineer 3D networks of PEDOT:PSS dendritic fibers. By controlling the growth conditions of the electropolymerized material, we investigate how dendritic fibers can reproduce structural plasticity by creating structures of controllable shape. We demonstrate gradual topologies evolution in a multi-electrode configuration. We conduct a detail electrical characterization of the PEDOT:PSS dendrites through DC and impedance spectroscopy measurements and we show how organic electrochemical transistors (OECT) can be realized with these structures. These measurements reveal that quasi-static and transient response of OECTs can be adjust by controlling dendrites' morphologies. The unique properties of organic dendrites are used to demonstrate short-term, long-term and structural plasticity, which are essential features required for future neuromorphic hardware development.

cond-mat.dis-nn

Theoretical modeling of dendrite growth from conductive wire electropolymerization

Electropolymerization is a bottom-up materials engineering process of micro and nano-scale that utilizes electrical signals to deposit conducting dendrites' morphologies by a redox reaction in the liquid phase. It resembles synaptogenesis in the brain, in which electrical stimulation in the brain causes the formation of synapses from the cellular neural composites. The strategy has been recently explored for neuromorphic engineering by establishing link between the electrical signals and the dendrites' shapes. Since the geometry of these structures determines their electrochemical properties, understanding the mechanisms that regulate the polymer assembly under electrically programmed conditions is an important aspect. In this manuscript, we simulate this phenomenon using mesoscale simulations, taking into account the important features of spatial-temporal potential mapping based on the time-varying signal, the motion of charged particles in the liquid due to the electric field, and the attachment of particles on the electrode. The study helps in visualizing the motion of particles in different electrical conditions, which is not possible to probe experimentally. Consistent with the experiments, the higher AC frequency of electrical activities favors linear wire-like growth, while lower frequency leads to more dense and fractal dendrites growth, and voltage offset leads to asymmetrical growth. We find that dendrites' shape and growth process systematically depend on particle concentration and random scattering. We discover that the different dendrites' architectures are associated with different Laplace and diffusion fields, which govern the monomers trajectory and subsequent dendrites' growth. Such unconventional engineering routes could have a variety of applications from neuromorphic engineering to bottom-up computing strategies.

cond-mat.dis-nn

Analog Programing of Conducting-Polymer Dendritic Interconnections and Control of their Morphology

Although materials and processes are different from biological cells', brain mimicries led to tremendous achievements in massively parallel information processing via neuromorphic engineering. Inexistent in electronics, we describe how to emulate dendritic morphogenesis by electropolymerization in water, aiming in operando material modification for hardware learning. The systematic study of applied voltage-pulse parameters details on tuning independently morphological aspects of micrometric dendrites': as fractal number, branching degree, asymmetry, density or length. Time-lapse image processing of their growth shows the spatial features to be dynamically-dependent and expand distinctively before and after forming a conductive bridging of two electrochemically grown dendrites. Circuit-element analysis and electrochemical impedance spectroscopy confirms their morphological control to occur in temporal windows where the growth kinetics can be finely perturbed by the input signal frequency and duty cycle. By the emulation of one of the most preponderant mechanisms responsible for brain's long-term memory, its implementation in the vicinity of sensing arrays, neural probes or biochips shall greatly optimize computational costs and recognition performances required to classify high-dimensional patterns from complex aqueous environments.

cond-mat.dis-nn

Electrical percolation in metal wire network based strain sensors

Metal wire networks rely on percolation paths for electrical conduction, and by suitably introducing break-make junctions on a flexible platform, a network can be made to serve as a resistive strain sensor. Several experimental designs have been proposed using networks made of silver nanowires, carbon nanotubes and metal meshes with high sensitivities. However, there is limited theoretical understanding; the reported studies have taken the numerical approach and only consider rearrangement of nanowires with strain, while the critical break-make property of the sensor observed experimentally has largely been ignored. Herein, we propose a generic geometrical based model and study distortion, including the break-make aspect, and change in electrical percolation of the network on applying strain. The result shows that when a given strain is applied, wire segments below a critical angle with respect to the applied strain direction end up breaking, leading to increased resistance of the network. The percolation shows interesting attributes; the calculated resistance increases linearly in the beginning and at a higher rate for higher strains, consistent with the experimental findings. In a real scenario, the strain direction need not necessarily be in the direction of measurement, and therefore, strain value and its direction both are incorporated into the treatment. The study reveals interesting anisotropic conduction features; strain sensitivity is higher parallel to the strain, while strain range is wider for perpendicular measurement. The percolation is also investigated on direct microscopic images of metal networks to obtain resistance-strain characteristics and identification of current percolation pathways. The findings will be important for electrical percolation in general, particularly in predicting characteristics and improvising metal network-based strain sensors.

physics.app-ph