arXiv ScienceSearch

arXiv subjects

Jiaming Luo

Publications and source records attributed to Jiaming Luo.

At least 19 recordsLinked to original sources

TranslateGemma Technical Report

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the translation task, we employ a two-stage fine-tuning process. First, supervised fine-tuning is performed using a rich mixture of high-quality large-scale synthetic parallel data generated via state-of-the-art models and human-translated parallel data. This is followed by a reinforcement learning phase, where we optimize translation quality using an ensemble of reward models, including MetricX-QE and AutoMQM, targeting translation quality. We demonstrate the effectiveness of TranslateGemma with human evaluation on the WMT25 test set across 10 language pairs and with automatic evaluation on the WMT24++ benchmark across 55 language pairs. Automatic metrics show consistent and substantial gains over the baseline Gemma 3 models across all sizes. Notably, smaller TranslateGemma models often achieve performance comparable to larger baseline models, offering improved efficiency. We also show that TranslateGemma models retain strong multimodal capabilities, with enhanced performance on the Vistra image translation benchmark. The release of the open TranslateGemma models aims to provide the research community with powerful and adaptable tools for machine translation.

cs.CL

From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity

Psychiatric comorbidity is clinically significant yet challenging due to the complexity of multiple co-occurring disorders. To address this, we develop a novel approach integrating synthetic patient electronic medical record (EMR) construction and multi-agent diagnostic dialogue generation. We create 502 synthetic EMRs for common comorbid conditions using a pipeline that ensures clinical relevance and diversity. Our multi-agent framework transfers the clinical interview protocol into a hierarchical state machine and context tree, supporting over 130 diagnostic states while maintaining clinical standards. Through this rigorous process, we construct PsyCoTalk, the first large-scale dialogue dataset supporting comorbidity, containing 3,000 multi-turn diagnostic dialogues validated by psychiatrists. This dataset enhances diagnostic accuracy and treatment planning, offering a valuable resource for psychiatric comorbidity research. Compared to real-world clinical transcripts, PsyCoTalk exhibits high structural and linguistic fidelity in terms of dialogue length, token distribution, and diagnostic reasoning strategies. Licensed psychiatrists confirm the realism and diagnostic validity of the dialogues. This dataset enables the development and evaluation of models capable of multi-disorder psychiatric screening in a single conversational pass.

cs.AI

On the Log Hodge Theory of Toroidal Varieties and a Partial Proof of the Absolute Hodge Conjecture

In this paper, we establish an innovative framework in logarithmic Hodge theory for toroidal varieties, introducing weighted toroidal structures and developing a systematic obstruction theory for Hodge classes. Building upon recent advances in toroidal geometry, particularly the $E_1$-degeneration results of Wei (\cite{Wei24}), we construct a categorical Hodge correspondence and prove the Absolute Hodge conjecture for projective toroidal varieties equipped with rational-weighted structures satisfying specific compatibility conditions. Our approach provides new tools for understanding the fine structure of Hodge classes and their extension properties across boundaries, with potential applications in moduli space compactifications, non-abelian Hodge theory, and tropical geometry.

math.AG

Lattice-induced spin dynamics in Dirac magnet CoTiO3

Spin-lattice coupling is crucial for understanding the spin transport and dynamics for spintronics and magnonics applications. Recently, cobalt titanate (CoTiO3), an easy-plane antiferromagnet, has been found to host axial phonons with a large magnetic moment, which may originate from spin-lattice coupling. Here, we investigate the effect of light-driven lattice dynamics on the magnetic properties of CoTiO3 using time-resolved spectroscopy with a THz pump and a magneto-optic probe. We found resonantly driven Raman active phonons, phonon-polariton-induced excitation of the antiferromagnetic magnons, and a slow increase in the polarization rotation of the probe, all indicating symmetry breaking that is not intrinsic to the magnetic space group. The temperature dependence confirmed that the observed spin dynamics is related to the magnetic order, and we suggest surface effects as a possible mechanism. Our results of THz-induced spin-lattice dynamics signify that extrinsic symmetry breaking may contribute strongly and unexpectedly to light-driven phenomena in bulk complex oxides.

cond-mat.mtrl-sci

Derived Stratified-Microlocal Framework and Moduli Space Resolution for the Cheeger-Goresky-Macpherson Conjecture

In this paper, We define the stratified metric $\infty$-category $\mathbf{StratMet}_{\infty}$ and the middle perversity moduli stack $\mathscr{M}^{\mathrm{mid}}$. We construct a universal truncation complex $\Omega_{X,\mathrm{FS}}^{\bullet,\mathrm{univ}}$ for a projective variety $X\subseteq\mathbb{P}^N$. By introducing the stratified singular characteristic variety $\mathrm{SSH}_{\mathrm{strat}}$, we establish a microlocal correspondence between metric asymptotic behavior and topology, proving the natural isomorphism $$H_2^*(X_{\mathrm{reg}}, ds_{\mathrm{FS}}^2) \cong IH^*(X,\mathbb{C}).$$ This framework transcends transverse singularity constraints, achieves moduli space parametrized duality, and develops new paradigms for high-codimension singular topology, quantum singularity theory, and $p$-adic Hodge theory.

math.AG

Derived Stratifications and Arithmetic Intersection Theory for Varieties with Isolated Singularities

In this paper, We develop the stratified de Rham theory on singular spaces using modern tools including derived geometry and stratified structures. This work unifies and extends the de Rham theory, Hodge theory, and deformation theory of singular spaces into the frameworks of stratified geometry, $p$-adic geometry, and derived geometry. Additionally, we close a gap in Ohsawa's original proof, concerning the convergence of $L^2$ harmonic forms in the Cheeger-Goresky-MacPherson conjecture for varieties with isolated singularities. Indicating that harmonic forms converge strongly and the $L^2$-cohomology coincides with intersection cohomology.

math.AG

Stratified Interpretation for De Rham Cohomology and Non-Witt Spaces

In this paper, we mainly build up the theory of sheaf-correspondence filtered spaces and stratified de Rham complexes for studying singular spaces. We prove the finiteness of a stratified de Rham cohomology and obtain its isomorphism to intersection cohomology through establishing a proper duality theory. Additionally, we present the stratified Poincar\'e duality, the K\"unneth decomposition theorem and develop stratified structure theory to non-Witt spaces as an application of a theory of stratified mezzoperversities. Our results connect differential forms, sheaf theory and intersection homology and pave the way for new approaches to study singular geometries, as well as topological invariants on them. Extensions to conical singularities and fibration of complex curves provide examples of the power of this method. This development will be foundational to new tools in stratified calculus and a strengthening of Hodge theory, advancing research in the Cheeger-Goresky-MacPherson conjecture.

math.AG

Gemma 3 Technical Report

We introduce Gemma 3, a multimodal addition to the Gemma family of lightweight open models, ranging in scale from 1 to 27 billion parameters. This version introduces vision understanding abilities, a wider coverage of languages and longer context - at least 128K tokens. We also change the architecture of the model to reduce the KV-cache memory that tends to explode with long context. This is achieved by increasing the ratio of local to global attention layers, and keeping the span on local attention short. The Gemma 3 models are trained with distillation and achieve superior performance to Gemma 2 for both pre-trained and instruction finetuned versions. In particular, our novel post-training recipe significantly improves the math, chat, instruction-following and multilingual abilities, making Gemma3-4B-IT competitive with Gemma2-27B-IT and Gemma3-27B-IT comparable to Gemini-1.5-Pro across benchmarks. We release all our models to the community.

cs.CL

Leveraging Domain Knowledge at Inference Time for LLM Translation: Retrieval versus Generation

While large language models (LLMs) have been increasingly adopted for machine translation (MT), their performance for specialist domains such as medicine and law remains an open challenge. Prior work has shown that LLMs can be domain-adapted at test-time by retrieving targeted few-shot demonstrations or terminologies for inclusion in the prompt. Meanwhile, for general-purpose LLM MT, recent studies have found some success in generating similarly useful domain knowledge from an LLM itself, prior to translation. Our work studies domain-adapted MT with LLMs through a careful prompting setup, finding that demonstrations consistently outperform terminology, and retrieval consistently outperforms generation. We find that generating demonstrations with weaker models can close the gap with larger model's zero-shot performance. Given the effectiveness of demonstrations, we perform detailed analyses to understand their value. We find that domain-specificity is particularly important, and that the popular multi-domain benchmark is testing adaptation to a particular writing style more so than to a specific domain.

cs.CL

A Diverse and Effective Retrieval-Based Debt Collection System with Expert Knowledge

Designing effective debt collection systems is crucial for improving operational efficiency and reducing costs in the financial industry. However, the challenges of maintaining script diversity, contextual relevance, and coherence make this task particularly difficult. This paper presents a debt collection system based on real debtor-collector data from a major commercial bank. We construct a script library from real-world debt collection conversations, and propose a two-stage retrieval based response system for contextual relevance. Experimental results show that our system improves script diversity, enhances response relevance, and achieves practical deployment efficiency through knowledge distillation. This work offers a scalable and automated solution, providing valuable insights for advancing debt collection practices in real-world applications.

cs.IR

SMOL: Professionally translated parallel data for 115 under-represented languages

We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine translation for low-resource languages. SMOL has been translated into 124 (and growing) under-resourced languages (125 language pairs), including many for which there exist no previous public resources, for a total of 6.1M translated tokens. SMOL comprises two sub-datasets, each carefully chosen for maximum impact given its size: SMOLSENT, a set of sentences chosen for broad unique token coverage, and SMOLDOC, a document-level resource focusing on a broad topic coverage. They join the already released GATITOS for a trifecta of paragraph, sentence, and token-level content. We demonstrate that using SMOL to prompt or fine-tune Large Language Models yields robust chrF improvements. In addition to translation, we provide factuality ratings and rationales for all documents in SMOLDOC, yielding the first factuality datasets for most of these languages.

cs.CL

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

Data contamination -- the accidental consumption of evaluation examples within the pre-training data -- can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of contamination on language models at 1B and 8B scales on the machine translation task. Starting from a carefully decontaminated train-test split, we systematically introduce contamination at various stages, scales, and data formats to isolate its effect and measure its impact on performance metrics. Our experiments reveal that contamination with both source and target substantially inflates BLEU scores, and this inflation is 2.5 times larger (up to 30 BLEU points) for 8B compared to 1B models. In contrast, source-only and target-only contamination generally produce smaller, less consistent over-estimations. Finally, we study how the temporal distribution and frequency of contaminated samples influence performance over-estimation across languages with varying degrees of data resources.

cs.CL

How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and multilingual contexts. In this study, we subject the entirety of non-English Wikipedia to a data filtering procedure typically reserved for noisy web-text -- a process which removes a large percentage of the collection's data. In analysing the removed data, we reveal numerous systematic quality issues, such as script and language contamination, repeated template and placeholder articles, and a high concentration of bot-generated content. We consolidate these findings into a 4-level quality ranking of Wikipedia, which shows strong correspondence with alternative quality measures and heuristics. Lastly, we evaluate the downstream impact of quality filtering in three practical language modelling scenarios, showing that models trained on filtered data largely match or outperform those trained on raw Wikipedia, with the largest gains observed for lower-quality language editions. Ultimately, our experiments serve as a first step in establishing quality-aware best practices for Wikipedia utilization in NLP, laying groundwork that can inform future dataset creation and curation efforts.

cs.CL

Learning from others' mistakes: Finetuning machine translation models with span-level error annotations

Despite growing interest in incorporating feedback to improve language models, most efforts focus only on sequence-level annotations. In this work, we explore the potential of utilizing fine-grained span-level annotations from offline datasets to improve model quality. We develop a simple finetuning algorithm, called Training with Annotations (TWA), to directly train machine translation models on such annotated data. TWA utilizes targeted span-level error information while also flexibly learning what to penalize within a span. Moreover, TWA considers the overall trajectory of a sequence when deciding which non-error spans to utilize as positive signals. Experiments on English-German and Chinese-English machine translation show that TWA outperforms baselines such as Supervised FineTuning on sequences filtered for quality and Direct Preference Optimization on pairs constructed from the same data.

cs.CL

Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts

In this paper we present a step-by-step approach to long-form text translation, drawing on established processes in translation studies. Instead of viewing machine translation as a single, monolithic task, we propose a framework that engages language models in a multi-turn interaction, encompassing pre-translation research, drafting, refining, and proofreading, resulting in progressively improved translations. Extensive automatic evaluations using Gemini 1.5 Pro across ten language pairs show that translating step-by-step yields large translation quality improvements over conventional zero-shot prompting approaches and earlier human-like baseline strategies, resulting in state-of-the-art results on WMT2024.

cs.CL

Chip-Scale Aligned Chiral Carbon Nanotubes Exhibiting Giant Second Harmonic Generation

Chiral carbon nanotubes (CNTs) are direct-gap semiconductors with optical properties governed by one-dimensional excitons with enormous oscillator strengths. Each species of chiral CNTs has an enantiomeric pair of left- and right-handed CNTs with nearly identical properties, but enantiomer-dependent phenomena can emerge, especially in nonlinear optical processes. Theoretical studies have predicted strong second-order nonlinearities in chiral CNTs, but no experimental quantitative verification has been reported due to the lack of macroscopically ordered assemblies of single-enantiomer chiral CNTs. Here, we report the synthesis of centimeter-scale, densely packed, aligned single-enantiomer chiral CNT films that are microfabrication-compatible. We observe giant second harmonic generation (SHG) emission from the chiral CNT film, which originates from the intrinsic chirality and inversion symmetry breaking of the atomic structure of chiral CNTs. The observed nonlinear susceptibility of the as-fabricated film reaches $4.9\times 10^2$\,pm/V at a pump wavelength of 1030\,nm, corresponding to the lowest-energy excitonic resonance, indicating $\chi_{xyz} = 1.6\times 10^3$\,pm/V for a perfectly aligned CNT crystal. Our calculations based on many-body theory correctly estimate the spectrum and magnitude of such excitonically enhanced optical nonlinearity. These results are promising for the development of scalable chiral-CNT electronics and nonlinear photonics.

physics.app-ph

To Diverge or Not to Diverge: A Morphosyntactic Perspective on Machine Translation vs Human Translation

We conduct a large-scale fine-grained comparative analysis of machine translations (MT) against human translations (HT) through the lens of morphosyntactic divergence. Across three language pairs and two types of divergence defined as the structural difference between the source and the target, MT is consistently more conservative than HT, with less morphosyntactic diversity, more convergent patterns, and more one-to-one alignments. Through analysis on different decoding algorithms, we attribute this discrepancy to the use of beam search that biases MT towards more convergent patterns. This bias is most amplified when the convergent pattern appears around 50% of the time in training data. Lastly, we show that for a majority of morphosyntactic divergences, their presence in HT is correlated with decreased MT performance, presenting a greater challenge for MT systems.

cs.CL

Layer-dependent exciton polarizability and the brightening of dark excitons in few-layer black phosphorus

The evolution of excitons from 2D to 3D is of great importance in photo-physics, yet the layer-dependent exciton polarizability has not been investigated in 2D semiconductors. Here, we determine the exciton polarizabilities for 3- to 11-layer black phosphorus-a direct bandgap semiconductor regardless of the thickness-through frequency-resolved photocurrent measurements on dual-gate devices and unveil the carrier screening effect in relatively thicker samples. By taking advantage of the broadband photocurrent spectra, we are also able to reveal the exciton response for higher-index subbands under the gate electrical field. Surprisingly, dark excitons are brightened with intensity even stronger than the allowed transitions above certain electrical field. Our study not only sheds light on the exciton evolution with sample thickness, but also paves a way for optoelectronic applications of few-layer BP in modulators, tunable photodetectors, emitters and lasers.

cond-mat.mes-hall