arXiv ScienceSearch

arXiv subjects

Hyunho Yang

Publications and source records attributed to Hyunho Yang.

3 recordsLinked to original sources

A.X K2 Technical Report

We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percentage points on some benchmarks, reflecting large gains in token efficiency. To support long contexts efficiently, we introduce Sparse Gated Attention (SGA), which combines sparse attention with gated attention, and adopt Gated Norm (GN) to stabilize large-scale training. SGA is trained natively at 128K through a \emph{sparse} indexer warmup that optimizes the indexer against its own sparse top-$k$ selection rather than the dense attention distribution, making adaptation markedly cheaper: each query reads only 2,048 positions, yet long-context quality is unchanged and A.X K2 scores 94.6 on RULER out to 256K. The outlier suppression of GN in turn keeps 4-bit NVFP4 serving within one point of FP8 accuracy. A simple yet effective Think-Fusion recipe further lets users switch between thinking and non-thinking modes within a single unified model. Extensive evaluations show that A.X K2 performs competitively against strong open-weight baselines, matching or exceeding them on math and Korean-language benchmarks.

cs.AI

A.X K1 Technical Report

We introduce A.X K1, a 519B-parameter Mixture-of-Experts (MoE) language model trained from scratch. Our design leverages scaling laws to optimize training configurations and vocabulary size under fixed computational budgets. A.X K1 is pre-trained on a corpus of approximately 10T tokens, curated by a multi-stage data processing pipeline. Designed to bridge the gap between reasoning capability and inference efficiency, A.X K1 supports explicitly controllable reasoning to facilitate scalable deployment across diverse real-world scenarios. We propose a simple yet effective Think-Fusion training recipe, enabling user-controlled switching between thinking and non-thinking modes within a single unified model. Extensive evaluations demonstrate that A.X K1 achieves performance competitive with leading open-source models, while establishing a distinctive advantage in Korean-language benchmarks.

cs.CL

Rapid low-temperature synthesis of graphene-coated SiC substrates for remote and van der Waals epitaxy

Non-conventional epitaxial techniques, such as van der Waals epitaxy (vdWE) and remote epitaxy, have attracted substantial attention in the semiconductor research community for their capability to repeatedly produce high-quality free-standing films from a single mother wafer. Successful implementation of these epitaxial techniques depends on creating a robust, uniform two-dimensional (2D) material surface. The conventional method for fabricating graphene on silicon carbide (SiC) is high-temperature graphitization. However, the extremely high temperature required for silicon sublimation (typically above 1500 {\deg}C) causes step-bunching of the SiC surface, forming non-uniform multilayer graphene stripes and an unfavorable surface morphology for epitaxial growth. Here, we developed a wafer-scale graphitization technique that allows fast synthesis of single-crystalline graphene at ultra-low temperatures by metal-assisted graphitization (MAG). We found annealing conditions that enable SiC dissociation while avoiding silicide formation, producing uniform single-crystalline graphene while maintaining the surface morphology of the substrate. The graphene thickness can be controlled by varying the metal thickness or annealing temperature, enabling remote epitaxy or vdWE. We successfully produced freestanding single-crystalline III-N (AlN, GaN) films on graphene/SiC via the 2D material-based layer transfer technique. Our results show that low-temperature graphene synthesis via MAG offers a promising route to producing large-scale ultra-wide bandgap free-standing crystalline membranes.

cond-mat.mtrl-sci