arXiv ScienceSearch

arXiv subjects

Ji Lin

Publications and source records attributed to Ji Lin.

At least 19 recordsLinked to original sources

Deep Clustering based Boundary-Decoder Net for Inter and Intra Layer Stress Prediction of Heterogeneous Integrated IC Chip

High stress occurs when 3D heterogeneous IC packages are subjected to thermal cycling at extreme temperatures. Stress mainly occurs at the interface between different materials. We investigate stress image using latent space representation which is based on using deep generative model (DGM). However, most DGM approaches are unsupervised, meaning they resort to image pairing (input and output) to train DGM. Instead, we rely on a recent boundary-decoder (BD) net, which uses boundary condition and image pairing for stress modeling. The boundary net maps material parameters to the latent space co-shared by its image counterpart. Because such a setup is dimensionally wise ill-posed, we further couple BD net with deep clustering. To access the performance of our proposed method, we simulate an IC chip dataset comprising of 1825 stress images. We compare our new approach using variants of BD net as well as a baseline approach. We show that our approach is able to outperform all the comparison in terms of train and test error reduction.

cs.LG

Gap solitons of the Wannier and Bloch types in spin-orbit-coupled Bose-Einstein condensates with a moir\'{e} lattice

Gap solitons (GSs) bifurcating from flat bands, which may be represented in terms of Wannier functions, have garnered significant interest due to their strong localization with extremely small norms. Moir\'{e} lattices (MLs), with multiple flat bands, offer an appropriate platform for creating such solitons. We explore the formation mechanism and stability of GSs in spin-1 Bose-Einstein condensates under the combined action of the Rashba spin-orbit coupling (SOC) and an ML potential. We identify five Wannier-type GS families bifurcating from the lowest five energy bands in the spectrum induced by the ML with sufficiently large period and depth. These fundamental GSs serve as basic elements for constructing more complex Wannier-type GS states. Reducing the lattice period and depth triggers a transition from the Wannier-type GSs to ones of the Bloch type, the latter exhibiting higher norm thresholds and pronounced spatial broadening near edges of the energy bands. In addition to tuning the lattice-potential parameters, adjusting the SOC strength can also modulate the flatness of energy bands and enhance the localization of gap solitons, enabling reversible transitions between the GSs of the Wannier and Bloch types. Distinctive properties of GSs in the quasiperiodic ML are uncovered too. Thus, we propose the theoretical foundation for the creation of and manipulations with strongly localized GSs.

cond-mat.quant-gas

gpt-oss-120b & gpt-oss-20b Model Card

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert transformer architecture and are trained using large-scale distillation and reinforcement learning. We optimize the models to have strong agentic capabilities (deep research browsing, python tool use, and support for developer-provided functions), all while using a rendered chat format that enables clear instruction following and role delineation. Both models achieve strong results on benchmarks ranging from mathematics, coding, and safety. We release the model weights, inference implementations, tool environments, and tokenizers under an Apache 2.0 license to enable broad use and further research.

cs.CL

From Green's formula to Derived Hall algebras

The aim of this note is to clarify the relationship between Green's formula and the associativity of multiplication for derived Hall algebra in the sense of To\"{e}n (Duke Math J 135(3):587-615, 2006), Xiao and Xu (Duke Math J 143(2):357-373, 2008) and Xu and Chen (Algebr Represent Theory 16(3):673-687, 2013). Let $\mathcal{A}$ be a finitary hereditary abelian category. It is known that the associativity of derived Hall algebra $\mathcal{D}\mathcal{H}_t(\mathcal{A})$ implies Green's formula. We show the converse statement holds. Namely, Green's formula implies the associativity of the derived Hall algebra $\mathcal{D}\mathcal{H}_t(\mathcal{A})$.

math.RT

Vector rogue waves in spin-1 Bose-Einstein condensates with spin-orbit coupling

We analytically and numerically study three-component rogue waves (RWs) in spin-1 Bose-Einstein condensates with Raman-induced spin-orbit coupling (SOC). Using the multiscale perturbative method, we obtain approximate analytical solutions for RWs with positive and negative effective masses, determined by the effective dispersion of the system. The solutions include RWs with smooth and striped shapes, as well as higher-order RWs. The analytical solutions demonstrate that the RWs in the three components of the system exhibit different velocities and their maximum peaks appear at the same spatiotemporal position, which is caused by SOC and interactions. The accuracy of the approximate analytical solutions is corroborated by comparison with direct numerical simulations of the underlying system. Additionally, we systematically explore existence domains for the RWs determined by the baseband modulational instability (BMI). Numerical simulations corroborate that, under the action of BMI, plane waves with random initial perturbations excite RWs, as predicted by the approximate analytical solutions.

cond-mat.quant-gas

A Seamless Phase II/III Design with Dose Optimization for Oncology Drug Development

The US FDA's Project Optimus initiative that emphasizes dose optimization prior to marketing approval represents a pivotal shift in oncology drug development. It has a ripple effect for rethinking what changes may be made to conventional pivotal trial designs to incorporate a dose optimization component. Aligned with this initiative, we propose a novel Seamless Phase II/III Design with Dose Optimization (SDDO framework). The proposed design starts with dose optimization in a randomized setting, leading to an interim analysis focused on optimal dose selection, trial continuation decisions, and sample size re-estimation (SSR). Based on the decision at interim analysis, patient enrollment continues for both the selected dose arm and control arm, and the significance of treatment effects will be determined at final analysis. The SDDO framework offers increased flexibility and cost-efficiency through sample size adjustment, while stringently controlling the Type I error. This proposed design also facilitates both Accelerated Approval (AA) and regular approval in a "one-trial" approach. Extensive simulation studies confirm that our design reliably identifies the optimal dosage and makes preferable decisions with a reduced sample size while retaining statistical power.

stat.ME

Construction of a new (3 + 1)-dimensional KdV equation and its closed-form solutions with solitary wave behaviour and conserved vectors

This paper discusses the construction of a new $(3+1)$-dimensional Korteweg-de Vries (KdV) equation. By employing the KdV's recursion operator, we extract two equations, and with elemental computation steps, the obtained result is $ 3u_{xyt}+3u_{xzt}-(u_{t}-6uu_{x}+u_{xxx})_{yz}-2\left(u_{x}\partial_{x}^{-1}u_{y}\right)_{xz}-2\left(u_{x}\partial_{x}^{-1}u_{z}\right)_{xy}=0 $. We then transform the new equation to a simpler one to avoid the appearance of the integral in the equation. Thereafter, we apply the Lie symmetry technique and gain a $7$-dimensional Lie algebra $L_7$ of point symmetries. The one-dimensional optimal system of Lie subalgebras is then computed and used in the reduction process to achieve seven exact solutions. These obtained solutions are graphically illustrated as 3D and 2D plots that show different propagations of solitary wave solutions such as breather, periodic, bell shape, and others. Finally, the conserved vectors are computed by invoking Ibragimov's method.

math-ph

Tiny Machine Learning: Progress and Futures

Tiny Machine Learning (TinyML) is a new frontier of machine learning. By squeezing deep learning models into billions of IoT devices and microcontrollers (MCUs), we expand the scope of AI applications and enable ubiquitous intelligence. However, TinyML is challenging due to hardware constraints: the tiny memory resource makes it difficult to hold deep learning models designed for cloud and mobile platforms. There is also limited compiler and inference engine support for bare-metal devices. Therefore, we need to co-design the algorithm and system stack to enable TinyML. In this review, we will first discuss the definition, challenges, and applications of TinyML. We then survey the recent progress in TinyML and deep learning on MCUs. Next, we will introduce MCUNet, showing how we can achieve ImageNet-scale AI applications on IoT devices with system-algorithm co-design. We will further extend the solution from inference to training and introduce tiny on-device training techniques. Finally, we present future directions in this area. Today's large model might be tomorrow's tiny model. The scope of TinyML should evolve and adapt over time.

cs.LG

VILA: On Pre-training for Visual Language Models

Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM with visual inputs, but lacks an in-depth study of the visual language pre-training process, where the model learns to perform joint modeling on both modalities. In this work, we examine the design options for VLM pre-training by augmenting LLM towards VLM through step-by-step controllable comparisons. We introduce three main findings: (1) freezing LLMs during pre-training can achieve decent zero-shot performance, but lack in-context learning capability, which requires unfreezing the LLM; (2) interleaved pre-training data is beneficial whereas image-text pairs alone are not optimal; (3) re-blending text-only instruction data to image-text data during instruction fine-tuning not only remedies the degradation of text-only tasks, but also boosts VLM task accuracy. With an enhanced pre-training recipe we build VILA, a Visual Language model family that consistently outperforms the state-of-the-art models, e.g., LLaVA-1.5, across main benchmarks without bells and whistles. Multi-modal pre-training also helps unveil appealing properties of VILA, including multi-image reasoning, enhanced in-context learning, and better world knowledge.

cs.CV

Rule-Guided Joint Embedding Learning over Knowledge Graphs

Recent studies on knowledge graph embedding focus on mapping entities and relations into low-dimensional vector spaces. While most existing models primarily exploit structural information, knowledge graphs also contain rich contextual and textual information that can enhance embedding effectiveness. In this work, we propose a novel model that integrates both contextual and textual signals into entity and relation embeddings through a graph convolutional network. To better utilize context, we introduce two metrics: confidence, computed via a rule-based method, and relatedness, derived from textual representations. These metrics enable more precise weighting of contextual information during embedding learning. Extensive experiments on two widely used benchmark datasets demonstrate the effectiveness of our approach, showing consistent improvements over strong baselines.

cs.CL

PockEngine: Sparse and Efficient Fine-tuning in a Pocket

On-device learning and efficient fine-tuning enable continuous and privacy-preserving customization (e.g., locally fine-tuning large language models on personalized data). However, existing training frameworks are designed for cloud servers with powerful accelerators (e.g., GPUs, TPUs) and lack the optimizations for learning on the edge, which faces challenges of resource limitations and edge hardware diversity. We introduce PockEngine: a tiny, sparse and efficient engine to enable fine-tuning on various edge devices. PockEngine supports sparse backpropagation: it prunes the backward graph and sparsely updates the model with measured memory saving and latency reduction while maintaining the model quality. Secondly, PockEngine is compilation first: the entire training graph (including forward, backward and optimization steps) is derived at compile-time, which reduces the runtime overhead and brings opportunities for graph transformations. PockEngine also integrates a rich set of training graph optimizations, thus can further accelerate the training cost, including operator reordering and backend switching. PockEngine supports diverse applications, frontends and hardware backends: it flexibly compiles and tunes models defined in PyTorch/TensorFlow/Jax and deploys binaries to mobile CPU/GPU/DSPs. We evaluated PockEngine on both vision models and large language models. PockEngine achieves up to 15 $\times$ speedup over off-the-shelf TensorFlow (Raspberry Pi), 5.6 $\times$ memory saving back-propagation (Jetson AGX Orin). Remarkably, PockEngine enables fine-tuning LLaMav2-7B on NVIDIA Jetson AGX Orin at 550 tokens/s, 7.9$\times$ faster than the PyTorch.

cs.LG

Novel Gravastar Solutions: Investigating Stability, Energy, and Entropy in the Presence of Cloud of Strings and Quintessence

Gravastars, theoretical alternatives to black holes, have captured the interest of scientists in astrophysics due to their unique properties. This paper aims to further investigate the exact solution of a novel gravastar model based on the Mazur-Mottola (2004) method within the framework of general relativity, specifically by incorporating the cloud of strings and quintessence. By analyzing the gravitational field and energy density of gravastars, valuable insights into the nature of compact objects in the universe can be gained. Understanding the stability of gravastars is also crucial for our comprehension of black holes and alternative compact objects. For this purpose, we present the Einstein field equations with the modified matter source and calculate the exact solutions for the inner and intermediate regions of gravastars. The exterior region is considered as a black hole surrounded by the cloud of strings and quintessence, and the spacetimes are matched using the Darmoise-Israel formalism. {An investigation is conducted on the stability of gravastars using linearized radial perturbation. Additionally, the proper length, energy content, and entropy of the shell are computed. The stability of gravastars is positively correlated with the enhancement of the cloud of strings parameter, while it is negatively correlated with the growth in the quintessence field parameter.} The paper concludes with a summary of the findings and their implications in the field of astrophysics and cosmology.

gr-qc

Magnetic lump motion in saturated ferromagnetic films

In this paper, we study in detail the nonlinear propagation of magnetic soliton in a ferromagnetic film. The sample is magnetized to saturation by an external field perpendicular to film plane. A new generalized (2+1)-dimensional short-wave asymptotic model is derived. The bilinear-like forms of this equation are constructed, and exact magnetic line soliton solutions are exhibited. It is observed that a series of stable lumps can be generated by an unstable magnetic soliton under Gaussian disturbance. Such magnetic lumps are highly stable and can maintain their shapes and velocities during evolution or collision. The interaction between lump and magnetic soliton, as well as interaction between two lumps, are numerically investigated. We further discuss the nonlinear motion of lumps in ferrites with Gilbert-damping and inhomogeneous exchange effects. The results show that the Gilbert-damping effects make the amplitude and velocity of the magnetic lump decay exponentially during propagation. And the shock waves are generated from a lump when quenching the strength of inhomogeneous exchange.

nlin.PS

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Large language models (LLMs) have transformed numerous AI applications. On-device LLM is becoming increasingly important: running LLMs locally on edge devices can reduce the cloud computing cost and protect users' privacy. However, the astronomical model size and the limited hardware resource pose significant deployment challenges. We propose Activation-aware Weight Quantization (AWQ), a hardware-friendly approach for LLM low-bit weight-only quantization. AWQ finds that not all weights in an LLM are equally important. Protecting only 1% salient weights can greatly reduce quantization error. To identify salient weight channels, we should refer to the activation distribution, not weights. To avoid the hardware-inefficient mix-precision quantization, we mathematically derive that scaling up the salient channels can reduce the quantization error. AWQ employs an equivalent transformation to scale the salient weight channels to protect them. The scale is determined by collecting the activation statistics offline. AWQ does not rely on any backpropagation or reconstruction, so it generalizes to different domains and modalities without overfitting the calibration set. AWQ outperforms existing work on various language modeling and domain-specific benchmarks (coding and math). Thanks to better generalization, it achieves excellent quantization performance for instruction-tuned LMs and, for the first time, multi-modal LMs. Alongside AWQ, we implement TinyChat, an efficient and flexible inference framework tailored for 4-bit on-device LLM/VLMs. With kernel fusion and platform-aware weight packing, TinyChat offers more than 3x speedup over the Huggingface FP16 implementation on both desktop and mobile GPUs. It also democratizes the deployment of the 70B Llama-2 model on mobile GPUs.

cs.CL

Stationary and moving bright solitons in Bose-Einstein condensates with spin-orbit coupling in a Zeeman field

With the discovery of various matter wave solitons in spin-orbit-coupled Bose-Einstein condensates (BECs), exploring their properties has become increasingly significant. We mainly study stationary and moving bright solitons in spin-orbit-coupled spin-1 BECs with or without a Zeeman field. The bright solitons correspond to the plane wave (PW) and standing wave (SW) phases. With the assistance of single-particle energy spectrum, we obtain the existence domains of PW and SW solitons by analytical and numerical methods. The results indicate that the interaction between atoms is also a key factor determining the existence of solitons. In addition, we systematically discuss the stability domains of PW and SW solitons, and investigate the impact of different parameters on the stability domains. We find that PW solitons are unstable when the linear Zeeman effect reaches a certain threshold, and the threshold is determined by other parameters. The linear Zeeman effect also leads to the alternating distribution of stable and unstable areas of SW solitons, and makes SW solitons stably exist in the area with stronger ferromagnetism. Finally, we analyze the collision dynamics of different types of stable solitons.

cond-mat.quant-gas

Semi-derived Ringel-Hall algebras and Hall algebras of odd-periodic relative derived categories

Let $t$ be a positive integer and $\mathcal{A}$ a hereditary abelian category satisfying some finiteness conditions. We define the semi-derived Ringel-Hall algebra of $\mathcal{A}$ from the category $\mathcal{C}_{\mathbb{Z}/t}(\mathcal{A})$ of $\mathbb{Z}/t$-graded complexes and obtain a natural basis of the semi-derived Ringel-Hall algebra. Moreover, we describe the semi-derived Ringel-Hall algebra by the generators and defining relations. In particular, if $t$ is an odd integer, we show that there is an embedding of derived Hall algebra of the odd-periodic relative derived category in the extended semi-derived Ringel-Hall algebra.

math.RT

A Multi-Arm Two-Stage (MATS) Design for Proof-of-Concept and Dose Optimization in Early-Phase Oncology Trials

The Project Optimus initiative by the FDA's Oncology Center of Excellence is widely viewed as a groundbreaking effort to change the $\textit{status quo}$ of conventional dose-finding strategies in oncology. Unlike in other therapeutic areas where multiple doses are evaluated thoroughly in dose ranging studies, early-phase oncology dose-finding studies are characterized by the practice of identifying a single dose, such as the maximum tolerated dose (MTD) or the recommended phase 2 dose (RP2D). Following the spirit of Project Optimus, we propose an Multi-Arm Two-Stage (MATS) design for proof-of-concept (PoC) and dose optimization that allows the evaluation of two selected doses from a dose-escalation trial. The design assess the higher dose first across multiple indications in the first stage, and adaptively enters the second stage for an indication if the higher dose exhibits promising anti-tumor activities. In the second stage, a randomized comparison between the higher and lower doses is conducted to achieve proof-of-concept (PoC) and dose optimization. A Bayesian hierarchical model governs the statistical inference and decision making by borrowing information across doses, indications, and stages. Our simulation studies show that the proposed MATS design yield desirable performance. An R Shiny application has been developed and made available at https://matsdesign.shinyapps.io/mats/.

stat.AP

Quantum Borcherds-Bozec algebras via semi-derived Ringel-Hall algebras II: braid group actions

Based on the realization of quantum Borcherds-Bozec algebra $\widetilde{\mathbf{U}}$ and quantum generalized Kac-Moody algebra ${}^B\widetilde{\mathbf{U}}$ via semi-derived Ringel-Hall algebra of a quiver with loops, we deduce the braid group actions of $\widetilde{\mathbf{U}}$ introduced by Fan and Tong recently and establish braid group actions for ${}^B\widetilde{\mathbf{U}}$ by applying the BGP reflection functors to semi-derived Ringel-Hall algebras.

math.RT