arXiv ScienceSearch

arXiv subjects

Sumeet Kumar Gupta

Publications and source records attributed to Sumeet Kumar Gupta.

At least 19 recordsLinked to original sources

BLINK: Batch Normalization-based Integrity Checkpoints for In-Situ Detection and Mitigation of Diverse Weight Corruptions in DNN Accelerators

In safety-critical deployments, AI hardware must remain reliable against a broad spectrum of threats such as aging, soft errors, hard faults, and adversarial attacks (e.g. progressive bit flip attack (PBFA)). All of these corrupt stored weights while the chip keeps producing confident but inaccurate predictions. Detecting and mitigating such weight perturbations is crucial for safety-critical platforms. To that end, we propose BLINK, an on-chip batch normalization (BN)-based on-the-fly detection and mitigation approach, which is based on continual sensing of the shift in the activation statistics, and targets a wide variety of weight corruptions (random and localized faults as well as adversarial bit flips). BLINK operates in two phases: (1) off-line pre-characterization of the relationship of the activation shifts with inference accuracy drop, and (2) on-chip runtime detection and mitigation of weight corruptions. Upon detection, the flagged layer is re-centered to bring it closer to its stored clean reference within the same forward pass. BLINK is fully autonomous, eliminating the need for host communication, operation halts, or access to fine-tuning data. If the residual shift after mitigation indicates that accuracy has fallen below a user-set floor, a held-out watcher aborts the inference. Evaluated on ResNet-20/50 and MobileNetV2 for CIFAR-10/100, BLINK detects harmful corruptions with >99% precision across all fault types. Further, it recovers accuracy from 10% to 85.88% under 0.5% random bit flips (Resnet-50/CIFAR-10), up to 84% for localized faults (MobileNetV2/CIFAR-10), and from random-guess accuracy to 80%-83% under PBFA (ResNet-20/CIFAR-10). Hardware overhead estimates indicate that BLINK incurs negligible costs, with less than a 2% increase in latency and only a 0.53% increase in computation overhead.

cs.AR

STEMPix: A Phase-Transition-Material-Based Pixel Sensor for Resolving Edge-Movement Direction

This paper proposes a spatio-temporal edge-movement direction pixel (STEMPix) for generating compact direction-aware edge movement information inside a CMOS-compatible image sensor array. The proposed design targets specialized sensing applications where local boundary movement is more important than full-frame intensity reconstruction. Instead of transferring full multi-bit frames for external processing, STEMPix generates a 3-bit local edge direction code (LEDC) by combining pixel-level temporal change information with neighboring-pixel spatial edge information. We design the architecture using a two-tier organization, where the photodiode layer is separated from the computation layer to preserve light-collection area while accommodating the additional in-array processing circuitry. The proposed circuit is evaluated through HSPICE transient simulations. The estimated implementation achieves a horizontal pitch of 1.73 μm, a vertical pitch of 2.36 μm, and a geometric fill factor of 95.47%. The average active switching energy is 0.465 fJ per LEDC operation across representative edge-movement cases. The proposed STEMPix operation also supports global-shutter capture and dynamic thresholding. These results indicate that STEMPix can provide a compact and scalable front-end representation for edge-movement-aware sensing systems.

cs.ET

ThermoPix: A High-Spatial-Resolution ElectronicPhotonic Temperature Sensor Array With Microsecond Row Readout

This paper presents ThermoPix, a CMOS-compatible electronic-photonic architecture for high-spatial-resolution temperature sensing. The proposed system converts temperature-induced wavelength shifts in a photonic interferometric sensor into timing information that can be processed by CMOS circuitry. We use a valley photonic crystal Mach-Zehnder interferometer (VPCMZI) as the sensing element, whose temperature-dependent spectral response is detected using an integrated waveguide photodetector and translated into a time-varying photocurrent. A CMOS readout circuit employing a phase-transition-material device performs threshold detection and generates a timing signal corresponding to the temperature-dependent crossing event. Circuit-level simulations demonstrate a temperature sensitivity of 3.15 ns/K, a row readout time of 2 us, and a sensing power-delay product (PDP) of 0.152 fJ. The required optical power per photonic cell is 150 nW, enabling energy-efficient array operation without requiring cooling or special environmental arrangements. We also present alternative photonic layer architectures for optical power distribution across the array. In one approach, we use different tap ratios along the row, while the other uses identical tap ratios with bidirectional excitation. The resulting average photonic cell pitches are 23.26 um and 38.52 um, respectively. The proposed ThermoPix architecture therefore provides a scalable platform for integrated temperature sensing arrays that combine photonic sensing elements with CMOS-compatible timing-based readout.

cs.ET

STRIDe: Cross-Coupled STT-MRAM Enabling Robust In-Memory-Computing for Deep Neural Network Accelerators

As deep neural network (DNN) models are growing exponentially in size, their deployment on resource-constrained edge platforms is becoming increasingly challenging. In-memory-computing (IMC) with non-volatile memories (NVMs) has emerged as a potential solution by virtue of its higher energy efficiency compared to standard DNN hardware platforms. Amongst various NVMs, STT-MRAM is highly promising owing to its high endurance and other benefits. However, their IMC implementation is challenging because of their inherently low distinguishability. This issue is exacerbated due to array non-idealities and process-variations, leading to poor IMC robustness and severe inference accuracy degradation. To address this problem, we propose STRIDe - STT-MRAM-based IMC leveraging cross-coupling action to boost the bitcell-level high-to-low current ratio to up to 8000. We propose two flavors of STRIDe designs, both offering robust IMC for inputs and weights in {-1, 1}(XNOR-IMC) and {0, 1}(AND-IMC) regime. Our evaluations for STRIDe arrays show up to 3.86x and 1.77x sense margin (SM) improvement for XNOR-IMC and AND-IMC, respectively, and up to 27.6% read disturb margin (RDM) improvement over standard MRAM-IMC designs. The enhanced robustness of STRIDe translates to near-software inference accuracies (considering crossbar non-idealities and process variations) for ResNet18 BNN and 4-bit DNN trained on CIFAR10 dataset. We observe accuracy improvements of up to 70% (for BNN) and up to 35%(for 4-bit DNN) over standard MRAM designs, albeit with some energy-area-latency penalty.

cs.ET

Lightweight True In-Pixel Encryption with FeFET Enabled Pixel Design for Secure Imaging

Ensuring end-to-end security in image sensors has become essential as visual data can be exposed through multiple stages of the imaging pipeline. Advanced protection requires encryption to occur before pixel values appear on any readout lines. This work introduces a secure pixel sensor (SecurePix), a compact CMOS-compatible pixel architecture that performs true in-pixel encryption using a symmetric key realized through programmable, non-volatile multidomain polarization states of a ferroelectric field-effect transistor. The pixel and array operations are designed and simulated in HSPICE, while a 45 nm CMOS process design kit is used for layout drawing. The resulting layout confirms a pixel pitch of 2.33 x 3.01 um^2. Each pixel's non-volatile programming level defines its analog transfer characteristic, enabling the photodiode voltage to be converted into an encrypted analog output within the pixel. Full-image evaluation shows that ResNet-18 recognition accuracy drops from 99.29 percent to 9.58 percent on MNIST and from 91.33 percent to 6.98 percent on CIFAR-10 after encryption, indicating strong resistance to neural-network-based inference. Lookup-table-based inverse mapping enables recovery for authorized receivers using the same symmetric key. Based on HSPICE simulation, the SecurePix achieves a per-pixel programming power-delay product of 17 uW us and a per-pixel sensing power-delay product of 1.25 uW us, demonstrating low-overhead hardware-level protection.

cs.CV

Weight Transformations in Bit-Sliced Crossbar Arrays for Fault Tolerant Computing-in-Memory: Design Techniques and Evaluation Framework

The deployment of deep neural networks (DNNs) on compute-in-memory (CiM) accelerators offers significant energy savings and speed-up by reducing data movement during inference. However, the reliability of CiM-based systems is challenged by stuck-at faults (SAFs) in memory cells, which corrupt stored weights and lead to accuracy degradation. While closest value mapping (CVM) has been shown to partially mitigate these effects for multibit DNNs deployed on bit-sliced crossbars, its fault tolerance is often insufficient under high SAF rates or for complex tasks. In this work, we propose two training-free weight transformation techniques, sign-flip and bit-flip, that enhance SAF tolerance in multi-bit DNNs deployed on bit-sliced crossbar arrays. Sign-flip operates at the weight-column level by selecting between a weight and its negation, whereas bit-flip provides finer granularity by selectively inverting individual bit slices. Both methods expand the search space for fault-aware mappings, operate synergistically with CVM, and require no retraining or additional memory. To enable scalability, we introduce a look-up-table (LUT)-based framework that accelerates the computation of optimal transformations and supports rapid evaluation across models and fault rates. Extensive experiments on ResNet-18, ResNet-50, and ViT models with CIFAR-100 and ImageNet demonstrate that the proposed techniques recover most of the accuracy lost under SAF injection. Hardware analysis shows that these methods incur negligible overhead, with sign-flip leading to negligible energy, latency, and area cost, and bit-flip providing higher fault resilience with modest overheads. These results establish sign-flip and bit-flip as practical and scalable SAF-mitigation strategies for CiM-based DNN accelerators.

cs.AR

Reverse Designing Ferroelectric Capacitors with Machine Learning-based Compact Modeling

Machine learning-based compact models provide a rapid and efficient approach for estimating device behavior across multiple input parameter variations. In this study, we introduce two reverse-design algorithms that utilize these compact models to identify device parameters corresponding to desired electrical characteristics. The algorithms effectively determine parameter sets, such as layer thicknesses, required to achieve specific device performance criteria. Significantly, the proposed methods are uniquely enabled by machine learning-based compact modeling; alternative computationally intensive approaches, such as phase-field modeling, would impose impractical time constraints for iterative design processes. Our comparative analysis demonstrates a substantial reduction in computation time when employing machine learning-based compact models compared to traditional phase-field methods, underscoring a clear and substantial efficiency advantage. Additionally, the accuracy and computational efficiency of both reverse-design algorithms are evaluated and compared, highlighting the practical advantages of machine learning-based compact modeling approaches.

cs.ET

Thickness Dependence of Coercive Field in Ferroelectric Doped-Hafnium Oxide

Ferroelectric hafnium oxide (${HfO_2}$) exhibits a thickness-dependent coercive field $(E_c)$ behavior that deviates from the trends observed in perovskites and the predictions of Janovec-Kay-Dunn (JKD) theory. Experiments reveal that, in thinner $HfO_2$ films ($<100\,nm$), $E_c$ increases with decreasing thickness but at a slower rate than predicted by the JKD theory. In thicker films, $E_c$ saturates and is independent of thickness. Prior studies attributed the thick film saturation to the thickness-independent grain size, which limits the domain growth. However, the reduced dependence in thinner films is poorly understood. In this work, we expound the reduced thickness dependence of $E_c$, attributing it to the anisotropic crystal structure of the polar orthorhombic (o) phase of $HfO_2$. This phase consists of continuous polar layers (CPL) along one in-plane direction and alternating polar and spacer layers (APSL) along the orthogonal direction. The spacer layers decouple adjacent polar layers along APSL, increasing the energy barrier for domain growth compared to CPL direction. As a result, the growth of nucleated domains is confined to a single polar plane in $HfO_2$, forming half-prolate elliptical cylindrical geometry rather than half-prolate spheroid geometry observed in perovskites. By modeling the nucleation and growth energetics of these confined domains, we derive a modified scaling law of $E_c \propto d^{-1/2}$ for $HfO_2$ that deviates from the classical JKD dependence of $E_c \propto d^{-2/3}$. The proposed scaling agrees well with the experimental trends in coercive field across various ferroelectric $HfO_2$ samples.

cond-mat.mtrl-sci

ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory based Ternary LLMs

Ternary large language models (LLMs), which utilize ternary precision weights and 8-bit activations, have demonstrated competitive performance while significantly reducing the high computational and memory requirements of full-precision LLMs. The energy efficiency and performance of Ternary LLMs can be further improved by deploying them on ternary computing-in-memory (TCiM) accelerators, thereby alleviating the von-Neumann bottleneck. However, TCiM accelerators are prone to memory stuck-at faults (SAFs) leading to degradation in the model accuracy. This is particularly severe for LLMs due to their low weight sparsity. To boost the SAF tolerance of TCiM accelerators, we propose ReTern that is based on (i) fault-aware sign transformations (FAST) and (ii) TCiM bit-cell reprogramming exploiting their natural redundancy. The key idea is to utilize FAST to minimize computations errors due to SAFs in +1/-1 weights, while the natural bit-cell redundancy is exploited to target SAFs in 0 weights (zero-fix). Our experiments on BitNet b1.58 700M and 3B ternary LLMs show that our technique furnishes significant fault tolerance, notably 35% reduction in perplexity on the Wikitext dataset in the presence of faults. These benefits come at the cost of < 3%, < 7%, and < 1% energy, latency and area overheads respectively.

cs.AR

Models for Spatially Resolved Conductivity of Rectangular Interconnects with Integrated Effect of Surface And Grain Boundary Scattering

Surface scattering and grain boundary scattering are two prominent mechanisms dictating the conductivity of interconnects and are traditionally modeled using the Fuchs-Sondheimer (FS) and Mayadas-Shatzkes (MS) theories, respectively. In addition to these approaches, modern interconnect structures need to capture the space-dependence of conductivity, for which a spatially resolved FS (SRFS) model was previously proposed to account for surface scattering based on Boltzmann transport equations (BTE). In this paper, we build upon the SRFS model to integrate grain-boundary scattering leading to a physics-based SRFS-MS model for the conductivity of rectangular interconnects. The effect of surface and grain scattering in our model is not merely added (as in several previous works) but is appropriately integrated following the original MS theory. Hence, the SRFS-MS model accounts for the interplay between surface scattering and grain boundary scattering in dictating the spatial dependence of conductivity. We also incorporate temperature (T) dependence into the SRFS-MS model. Further, we propose a circuit compatible conductivity model (SRFS-MS-C3), which captures the space-dependence and integration of surface and grain boundary scattering utilizing an analytical function and a few (three or four) invocations of the physical SRFS-MS model. We validate the SRFS-MS-C3 model across a wide range of physical parameters, demonstrating excellent agreement with the physical SRFS-MS model, with an error margin of less than 0.7%. The proposed SRFS-MS and SRFS-MS-C3 models explicitly relate the spatially resolved conductivity to physical parameters such as electron mean free path ($λ_0$), specularity of surface scattering (p), grain boundary reflectance coefficient (R), interconnect cross-section geometry and temperature (T).

physics.app-ph

Spatially Resolved Conductivity of Rectangular Interconnects considering Surface Scattering -- Part I: Physical Modeling

Accurate modeling of interconnect conductivity is important for performance evaluation of chips in advanced technologies. Surface scattering in interconnects is usually treated by using Fuchs-Sondheimer (FS) approach. While the FS model offer explicit inclusion of the physical parameters, it lacks spatial dependence of conductivity across the interconnect cross-section. To capture the space-dependency of conductivity, an empirical modeling approach based on "cosh" function has been proposed, but it lacks physical insights. In this work, we present a 2D spatially resolved FS (SRFS) model for rectangular interconnects derived from the Boltzmann transport equations. The proposed SRFS model for surface scattering offers both spatial dependence and explicit relation of conductivity to physical parameters such as mean free path and specularity of electrons and interconnect geometry. We highlight the importance of physics-based spatially resolved conductivity model by showing the differences in the spatial profiles between the proposed physical approach and the previous empirical approach. In Part II of this work, we build upon the SRFS approach to propose a compact model for spatially-resolved conductivity accounting for surface scattering in rectangular interconnects.

physics.app-ph

Spatially Resolved Conductivity of Rectangular Interconnects considering Surface Scattering -- Part II: Circuit-Compatible Modeling

Interconnect conductivity modeling is a critical aspect for modern chip design. Surface scattering -- an important scattering mechanism in scaled interconnects is usually captured using Fuchs-Sondheimer (FS) model which offers the average behavior of the interconnect. However, to support the modern interconnect structures (such as tapered geometries), modeling spatial dependency of conductivity becomes important. In Part I of this work, we presented a spatially resolved FS (SRFS) model for rectangular interconnects derived from the fundamental FS approach. While the proposed SRFS model offers both spatial-dependency of conductivity and its direct relationship with the physical parameters, its complex expression is not suitable for incorporation in circuit simulations. In this part, we build upon our SRFS model to propose a circuit-compatible conductivity model for rectangular interconnects accounting for 2D surface scattering. The proposed circuit-compatible model offers spatial resolution of conductivity as well as explicit dependence on the physical parameters such as electron mean free path ($λ_0$), specularity ($p$) and interconnect geometry. We validate our circuit-compatible model over a range of interconnect width/height (and $λ_0$) and p values and show a close match with the physical SRFS model proposed in Part I (with error < 0.7%). We also compare our circuit-compatible model with a previous spatially resolved analytical model (appropriately modified for a fair comparison) and show that our model captures the spatial resolution of conductivity and the dependence on physical parameters more accurately. Finally, we present a semi-analytical equation for the average conductivity based on our circuit-compatible model.

physics.app-ph

Experimental Investigation of Variations in Polycrystalline Hf0.5Zr0.5O2 (HZO)-based MFIM

Device-to-device variations in ferroelectric (FE) hafnium oxide (HfO2)-based devices pose a crucial challenge that limits the otherwise promising capabilities of this technology. Although previous simulation-based studies have identified polarization (P) domain nucleation and polycrystallinity as key contributors to variations in HfO2, experimental validation remains limited. Here, we experimentally investigate variations in remanent polarization (PR) of Hf0.5Zr0.5O2 (HZO)-based metal-ferroelectric-insulator-metal (MFIM) capacitors across different set voltages (VSET) and FE thicknesses (TFE). Our measurements reveal a non-monotonic behavior of the standard deviation of PR with VSET peaking around coercive voltage (VC), which is consistent with previous simulation-based predictions. In the low- and high-VSET regions, PR variations are primarily dictated by saturation polarization (PS) variations, mainly originating from charge trap effects at the interface between the FE-dielectric (DE) layer and the polycrystallinity of FE. On the other hand, in the mid-VSET region peak, the PR variations are attributed to the VC variation, which comes from a combined effect of multi-domain (MD) P switching and polycrystallinity. Notably, sharp P switching associated with domain nucleation amplifies the variations, resulting in a peak of PR variations in this VSET range. Further, we observe that as HZO thickness (TFE) is scaled, the non-monotonicity in variations with VSET is reduced, primarily due to reduced domain nucleation and smaller grain sizes. We experimentally establish a strong correlation of PR with PS in the low- and high-VSET regions and with VC in the mid-VSET region across various TFE. Finally, our experimental findings are corroborated with simulations using a 3D phase-field model.

physics.app-ph

Oxygen Vacancy-Induced Monoclinic Dead Layers in Ferroelectric $Hf_xZr_{1-x}O_2$ With Metal Electrodes

In this work, we analyze dead layer comprising non-polar monoclinic (m) phase in $Hf_xZr_{1-x}O_2$ (HZO)-based ferroelectric (FE) material using first principles analysis. We show that with widely used tungsten (W) metal electrode, the spatial distribution of the oxygen vacancy across the cross-section plays a key role in dictating the favorability of m- phase formation at the metal-HfO2 interface. The energetics are also impacted by the polarization direction as well as the depth of oxygen vacancy, i.e., position along the thickness. At the metal - $HfO_2$ interface, when polarization points towards the metal and vacancy forms at trigonally bonded O atomic site, both interfacial relaxation and m- phase formation can lead to dead layers. For vacancies at other oxygen atomic sites and polarization direction, dead layer is formed due to sole interfacial relaxation with polar phase. We also establish the relative favorability of the m-phase dead layer for different Zr concentrations (x=1 and x = 0.5) and metal electrodes. According to our analysis, 50% Zr doped $HfO_2$ exhibits less probability of m-phase dead layer formation compared to pure $HfO_2$. Moreover, with electrodes consisting of noble metal (Pt, Pd, Os, Ru, Rh), m-phase dead layer formation is less likely. Therefore, for these metals, dead layer forms mainly due to the interfacial relaxation with polar phase.

cond-mat.mtrl-sci

BinSparX: Sparsified Binary Neural Networks for Reduced Hardware Non-Idealities in Xbar Arrays

Compute-in-memory (CiM)-based binary neural network (CiM-BNN) accelerators marry the benefits of CiM and ultra-low precision quantization, making them highly suitable for edge computing. However, CiM-enabled crossbar (Xbar) arrays are plagued with hardware non-idealities like parasitic resistances and device non-linearities that impair inference accuracy, especially in scaled technologies. In this work, we first analyze the impact of Xbar non-idealities on the inference accuracy of various CiM-BNNs, establishing that the unique properties of CiM-BNNs make them more prone to hardware non-idealities compared to higher precision deep neural networks (DNNs). To address this issue, we propose BinSparX, a training-free technique that mitigates non-idealities in CiM-BNNs. BinSparX utilizes the distinct attributes of BNNs to reduce the average current generated during the CiM operations in Xbar arrays. This is achieved by statically and dynamically sparsifying the BNN weights and activations, respectively (which, in the context of BNNs, is defined as reducing the number of +1 weights and activations). This minimizes the IR drops across the parasitic resistances, drastically mitigating their impact on inference accuracy. To evaluate our technique, we conduct experiments on ResNet-18 and VGG-small CiM-BNNs designed at the 7nm technology node using 8T-SRAM and 1T-1ReRAM. Our results show that BinSparX is highly effective in alleviating the impact of non-idealities, recouping the inference accuracy to near-ideal (software) levels in some cases and providing accuracy boost of up to 77.25%. These benefits are accompanied by energy reduction, albeit at the cost of mild latency/area increase.

cs.AR

The Impact of TaS$_{2}$-Augmented Interconnects on Circuit Performance: A Temperature-Dependent Analysis

Monolayer TaS$_{2}$ is being explored as a future liner/barrier to circumvent the scalability issues of the state-of-the-art interconnects. However, its large vertical resistivity poses some concerns and mandates a comprehensive circuit analysis to understand the benefits and trade-offs of this technology. In this work, we present a detailed temperature-dependent modeling framework of TaS$_{2}$-augmented copper (Cu) interconnects and provide insights into their circuit implications. We build temperature-dependent 3D models for Cu-TaS$_{2}$ interconnect resistance capturing surface scattering and grain boundary scattering and integrate them in an RTL-GDSII design flow based on ASAP7 7nm process design kit. Using this framework, we perform synthesis and place-and-route (PnR) for advanced encryption standard (AES) circuit at different process and temperature corners and benchmark the circuit performance of Cu-TaS$_{2}$ interconnects against state-of-the-art interconnects. Our results show that Cu-TaS$_{2}$ interconnects yield an enhancement in the effective clock frequency of the AES circuit by 1%-10.6%. Considering the worst-case process-temperature corner, we further establish that the vertical resistivity of TaS$_{2}$ must be below 22 k$Ω$-nm to obtain performance benefits over conventional interconnects.

physics.app-ph

Memory Faults in Activation-sparse Quantized Deep Neural Networks: Analysis and Mitigation using Sharpness-aware Training

Improving the hardware efficiency of deep neural network (DNN) accelerators with techniques such as quantization and sparsity enhancement have shown an immense promise. However, their inference accuracy in non-ideal real-world settings (such as in the presence of hardware faults) is yet to be systematically analyzed. In this work, we investigate the impact of memory faults on activation-sparse quantized DNNs (AS QDNNs). We show that a high level of activation sparsity comes at the cost of larger vulnerability to faults, with AS QDNNs exhibiting up to 11.13% lower accuracy than the standard QDNNs. We establish that the degraded accuracy correlates with a sharper minima in the loss landscape for AS QDNNs, which makes them more sensitive to perturbations in the weight values due to faults. Based on this observation, we employ sharpness-aware quantization (SAQ) training to mitigate the impact of memory faults. The AS and standard QDNNs trained with SAQ have up to 19.50% and 15.82% higher inference accuracy, respectively compared to their conventionally trained equivalents. Moreover, we show that SAQ-trained AS QDNNs show higher accuracy in faulty settings than standard QDNNs trained conventionally. Thus, sharpness-aware training can be instrumental in achieving sparsity-related latency benefits without compromising on fault tolerance.

cs.LG

Reimagining Sense Amplifiers: Harnessing Phase Transition Materials for Current and Voltage Sensing

Energy-efficient sense amplifier (SA) circuits are essential for reliable detection of stored memory states in emerging memory systems. In this work, we present four novel sense amplifier (SA) topologies based on phase transition material (PTM) tailored for non-volatile memory applications. We utilize the abrupt switching and volatile hysteretic characteristics of PTMs which enables efficient and fast sensing operation in our proposed SA topologies. We provide comprehensive details of their functionality and assess how process variations impact their performance metrics. Our proposed sense amplifier topologies manifest notable performance enhancement. We achieve a ~67% reduction in sensing delay and a ~80% decrease in sensing power for current sensing. For voltage sensing, we achieve a ~75% reduction in sensing delay and a ~33% decrease in sensing power. Moreover, the proposed SA topologies exhibit improved variation robustness compared to conventional SAs. We also scrutinize the dependence of transistor mirroring window and PTM transition voltages on several device parameters to determine the optimum operating conditions and stance of tunability for each of the proposed SA topologies.

cs.ET