arXiv ScienceSearch

arXiv subjects

Jongho Park

Publications and source records attributed to Jongho Park.

At least 19 recordsLinked to original sources

Scaling Discovery through Test-Time Communication

Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

cs.LG

Subspace correction as a general framework for convex optimization algorithms

This paper shows that subspace correction methods provide a common algorithmic and theoretical foundation for several classes of convex optimization algorithms, including operator splitting, alternating projection, and multiplier methods. The underlying principle is to decompose a problem into smaller subproblems and combine their solutions, a strategy that appears throughout iterative algorithms. The main tool is an iterate-level formalism, which we call dualization, that abstracts classical primal--dual correspondences and relates a subspace correction method for a dual problem to an algorithm for the corresponding primal problem via a primal--dual consistency relation. At the algorithmic level, dualizing successive subspace correction yields the Peaceman--Rachford and Douglas--Rachford splitting methods; the von Neumann and Dykstra alternating projection algorithms arise as special cases. Dualizing parallel subspace correction yields a parallel splitting method. Multiplier methods, including the alternating direction method of multipliers (ADMM), are connected to subspace correction through the same mechanism together with equivalent block formulations. In particular, a multi-block ADMM-type algorithm is obtained by dualizing a splitting method derived from successive subspace correction. At the theoretical level, the primal--dual consistency relation transfers convergence estimates from subspace correction to the derived algorithms under suitable assumptions. Thus, these operator splitting, alternating projection, and multiplier algorithms can be derived from subspace correction at both the algorithmic and theoretical levels.

math.OC

Bayesian polarization calibration and imaging in very long baseline interferometry

Extracting polarimetric information from very long baseline interferometry (VLBI) data is demanding but vital for understanding the synchrotron radiation process and the magnetic fields of celestial objects, such as active galactic nuclei (AGNs). However, conventional CLEAN-based calibration and imaging methods provide suboptimal resolution without uncertainty estimation of calibration solutions, while requiring manual steering from an experienced user. We present a Bayesian polarization calibration and imaging method using Bayesian imaging software resolve for VLBI data sets, that explores the posterior distribution of antenna-based gains, polarization leakages, and polarimetric images jointly from pre-calibrated data. We demonstrate our calibration and imaging method with observations of the quasar 3C273 with the VLBA at 15 GHz and the blazar OJ287 with the GMVA+ALMA at 86 GHz. Compared to the CLEAN method, our approach provides physically realistic images that satisfy positivity of flux and polarization constraints and can reconstruct complex source structures composed of various spatial scales. Our method systematically accounts for calibration uncertainties in the final images and provides uncertainties of Stokes images and calibration solutions. The automated Bayesian approach for calibration and imaging will be able to obtain high-fidelity polarimetric images using high-quality data from next-generation radio arrays. The pipeline developed for this work is publicly available.

astro-ph.IM

Optimal block preconditioners for a mass-conserving mixed stress formulation of Stokes flow

We present optimal block diagonal and triangular preconditioners for a mass-conserving mixed stress formulation of Stokes flow. The algebraic formulation leads to a double saddle point system with unknowns corresponding to discrete stress, velocity, vorticity, and pressure. MINRES equipped with a block diagonal preconditioner for an augmented Lagrangian formulation of this system is analyzed and shown to be optimal, in the sense that the convergence rate is independent of key parameters such as mesh size and kinematic viscosity. GMRES equipped with a block triangular preconditioner is also analyzed using a field-of-values approach. Finally, we present numerical results for both two- and three-dimensional model problems to validate the parameter robustness of the proposed preconditioners.

math.NA

The Capella Program: Toward A Space-only High-frequency Radio VLBI Network Formed by Small Satellites in Low Earth Orbits

Very long baseline radio interferometry (VLBI) with ground-based observatories is limited by the size of Earth, the geographic distribution of antennas, and the transparency of the atmosphere. In this whitepaper, we present a design for a space-to-space VLBI program composed of two missions: Mimosa, a pathfinder, and Capella, a science-grade VLBI observatory. Mimosa is a two-element space-to-space radio interferometer composed of two small (250 kg) satellites on co-planar polar circular low Earth orbits. Using single-band, single-circular polarization heterodyne HEMT receivers operating at frequencies around 100 GHz, the interferometer is able to achieve a near-perfect visibility plane coverage and an angular resolution of approximately 35 microarcsec. Capella comprises four small (500 kg) satellites in two orthogonal polar low-Earth orbit planes. With single-band heterodyne receivers operating at frequencies around 690 GHz, the interferometer is able to achieve angular resolutions of approximately 7 microarcsec. Within a total observing time of three days, a near-complete uv plane coverage can be reached. The technology for all key components required - radio telescope, receiver, sampler, recorder, frequency standard, positioning system, data downlink, and pointing control system - is already available, partially off-the-shelf. Capella will be able to address a range of science cases, including: the shadows of supermassive black holes; the acceleration and collimation zones of plasma jets emitted from the vicinity of supermassive black holes; the chemical composition of accretion flows into active galactic nuclei through observations of molecular absorption lines; mapping supermassive binary black holes; the magnetic activity of stars; and nova eruptions of symbiotic binary stars -- and, like any substantially new observing technique, has the potential for unexpected discoveries.

astro-ph.IM

A polynomial dimension-dependence analysis of Bramble--Pasciak--Xu preconditioners

We investigate the dimension dependence of Bramble--Pasciak--Xu (BPX) preconditioners for high-dimensional partial differential equations and establish that the condition numbers of BPX-preconditioned systems grow only polynomially with the spatial dimension. Our analysis requires a careful derivation of the dimension dependence of several fundamental tools in the theory of finite element methods, including elliptic regularity, the Bramble--Hilbert lemma, trace inequalities, and inverse inequalities. We further analyze an averaged Scott--Zhang-type quasi-interpolation operator, and show that its associated constants scale polynomially with the dimension. Building on these ingredients, we prove a multilevel norm equivalence theorem and derive a BPX preconditioner with explicit polynomial bounds on its dimensional dependence. The analysis is motivated in part by recent tensor and quantum finite element methods, where dimension-explicit conditioning estimates for BPX preconditioners play an important role.

math.NA

Parallel multilevel methods for solving the Darcy--Forchheimer model based on a nearly semicoercive formulation

High-velocity fluid flow through porous media is modeled by prescribing a nonlinear relationship between the flow rate and the pressure gradient, called the Darcy--Forchheimer equation. This paper is concerned with the analysis of parallel multilevel methods for solving the Darcy--Forchheimer model. We begin by reformulating the Darcy--Forchheimer model as a nearly semicoercive convex optimization problem via the augmented Lagrangian method. Building on this formulation, we develop a parallel multilevel method, also known as a multilevel additive Schwarz method, within the framework of subspace correction for nearly semicoercive convex problems, yielding a theoretically supported and computationally efficient solver for the Darcy--Forchheimer model. The convergence analysis establishes robustness with respect to the augmented Lagrangian parameter $ε$. To further enhance convergence, we incorporate a backtracking line search and a full approximation scheme. Numerical results support the theoretical findings and demonstrate the effectiveness of the proposed approach.

math.NA

Looped Diffusion Language Models

Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDMs remains underexplored. In this paper, we show that selectively looping the early-middle transformer layers significantly improves both training efficiency and model performance in MDMs. We call this approach LoopMDM(Looped Masked Diffusion Model), which brings two key benefits: looping layers at training-time yields a depth-scaling effect without adding parameters, while varying the number of loops at inference-time enables flexible compute scaling. Despite the simplicity, the results are striking: across multiple pre-training corpora, LoopMDM matches the performance of same-size MDMs with up to 3.3 fewer training FLOPs, while its final performance outperforms them on various reasoning benchmarks, including up to 8.5 points on GSM8K. It even surpasses deeper non-looped MDMs trained with comparable per-step compute, indicating that selective looping is more effective than naive depth scaling. Furthermore, LoopMDM can scale inference-time compute by increasing the number of loops. Adaptively adjusting the number of loops throughout the sampling process further yields additional gains in compute efficiency while maintaining performance. Lastly, with attention analysis, we provide evidence that looping is effective in MDMs by promoting interactions among masked positions. Our code and weights will be publicly released.

cs.LG

Beyond RLHF: A Unified Theoretical Framework of Alignment

Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However, existing theories do not provide strong justification for the RLHF objective itself and do not allow comparisons of the guarantees between various methods because different methods are often analyzed under different frameworks. Toward a unified framework for alignment, we ask under what assumptions can we derive existing or new training objectives and obtain theoretical guarantees. To this end, we reframe alignment as distribution learning from pairwise preferences, which makes a probabilistic assumption describing how preferences reveal information about the target LM. This leads us to propose three principled alignment objectives: preference maximum likelihood estimation, preference distillation, and reverse KL minimization. We prove that they all enjoy strong non-asymptotic $O(1/n)$ convergence to the target LM, naturally avoiding degeneracy. In particular, reverse KL highly resembles the RLHF objective, providing strong justification for RLHF. Furthermore, our theory explains, for the first time, the empirical finding that on-policy objectives (e.g., RLHF) typically outperform likelihood-style objectives (e.g., DPO). Finally, empirical results indicate that the proposed objectives are competitive with strong baselines across several tasks and models.

cs.LG

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritizes compliance, making models vulnerable to harmful prompts. To mitigate this safety degradation, recent approaches rely on external teacher distillation, yet this introduces a distributional discrepancy that degrades native reasoning. We formalize safety realignment as a KL projection onto the safe simplex and prove that the student's own safety-filtered distribution is the unique KL-optimal target, while any external teacher incurs an irreducible excess KL penalty. Guided by this analysis, we propose ThinkSafe, a self-generated alignment framework that restores safety without external teachers. Our key insight is that while compliance suppresses safety mechanisms, models often retain latent knowledge to identify harm. ThinkSafe unlocks this via lightweight refusal steering, which preserves the KL-optimal target while increasing the acceptance rate. Experiments on DeepSeek-R1-Distill and Qwen3 show ThinkSafe significantly improves safety while preserving reasoning proficiency, and achieves superior safety and comparable reasoning to GRPO with roughly an order of magnitude less compute. Code, models, and datasets are available at https://github.com/seanie12/ThinkSafe and https://huggingface.co/Seanie-lee/collections.

cs.AI

Randomized subspace correction methods for convex optimization

This paper introduces an abstract framework for randomized subspace correction methods for convex optimization, which unifies and generalizes a broad class of existing algorithms, including domain decomposition, multigrid, and block coordinate descent methods. We provide a convergence rate analysis ranging from minimal assumptions to more practical settings, such as sharpness and strong convexity. While most existing studies on block coordinate descent methods focus on nonoverlapping decompositions and smooth or strongly convex problems, our framework extends to more general settings involving arbitrary space decompositions, inexact local solvers, and problems with weaker smoothness or convexity assumptions. The proposed framework is broadly applicable to convex optimization problems arising in areas such as nonlinear partial differential equations, imaging, and data science.

math.OC

Unified analysis of saddle point problems via auxiliary space theory

We present sharp estimates for the extremal eigenvalues of the Schur complements arising in saddle point problems. These estimates are derived using the auxiliary space theory, in which a given iterative method is interpreted as an equivalent but more elementary iterative method on an auxiliary space, enabling us to obtain sharp convergence estimates. The proposed framework improves or refines several existing results, which can be recovered as corollaries of our results. To demonstrate the versatility of the framework, we present various applications from scientific computing: the augmented Lagrangian method, mixed finite element methods, and nonoverlapping domain decomposition methods. In all these applications, the condition numbers of the corresponding Schur complements can be estimated in a straightforward manner using the proposed framework.

math.NA

Verification of the Polarimetric Capability of the East Asia VLBI Network

The East Asia VLBI Network (EAVN) has recently enabled dual-polarization observations at $22$ and $43\,\mathrm{GHz}$. We present the first systematic verification of its polarimetric performance using EAVN observations of M87, 3C 279, 3C 273, and OJ 287, calibrated with the GPCAL pipeline and evaluated against near-contemporaneous VLBA images at comparable frequencies. Most stations show stable polarimetric leakages with amplitudes of $5$-$10\%$ over monthly timescales. While several VERA stations exhibit D-term phase variations between epochs, we attribute these to field-rotator (FR) offsets and demonstrate that phase stability is restored after applying the analytically derived FR corrections. The resulting linear-polarization morphologies and EVPAs broadly agree with the VLBA results within uncertainties; fractional polarization measured by the EAVN tends to be slightly higher near polarization peaks. Although exact one-to-one comparisons are limited by moderate frequency and epoch differences, the combined evidence indicates robust EAVN polarimetric calibration and imaging capabilities at $22$ and $43\,\mathrm{GHz}$. These results support the scientific capability of EAVN polarimetry and lay the groundwork for expanded, higher-fidelity polarimetric studies in East Asia.

astro-ph.IM

A Unified Origin of Faraday Rotation toward 3C 84: The Circumnuclear Ambient Medium within the Parsec-Scale Bondi Radius of the Host Galaxy NGC 1275

We present multi-frequency polarimetric observations of 3C 84 obtained with the Korean VLBI Network at 43-141 GHz, the Very Long Baseline Array at 43 GHz, and the High Sensitivity Array at 8 GHz from 2015 to 2024. We find that the Faraday rotation measure (RM) decreases systematically with distance from the black hole over 1-8 pc, following a single power-law trend of RM proportional to r^{-2.7+/-0.2}. Notably, RM measurements from earlier studies across the same distance range follow the same relation. This consistency across epochs, frequencies, and independent datasets indicates a common and stable external Faraday screen. These results naturally identify the circumnuclear ambient medium within the parsec-scale Bondi radius of the host galaxy NGC 1275 as the origin of the Faraday rotation, thereby resolving a long-standing question about its physical origin. From the RM profile, we derive radial distributions of the electron density and magnetic-field strength in the circumnuclear ambient medium that are consistent with independent constraints. The derived density lies below that of the free-free absorption disk and, when extrapolated inward, remains below the density of the broad-line region. The magnetic-field strength gradually increases from 0.1-1.5 microgauss at the Bondi radius to milligauss-to-gauss levels toward the black hole, providing the first spatially resolved constraint on the magnetic-field strength at parsec-scale distances in an elliptical galaxy. Together, these results present a spatially resolved and physically consistent picture of the circumnuclear environment in NGC 1275.

astro-ph.GA

Transverse Oscillations and Wave Propagation in the Magnetically Dominated M87 Jet

We present an in-depth analysis of transverse oscillations in the M87 jet, as identified in our previous study (Ro et al. 2023a), which reported oscillatory patterns with a characteristic period of $\sim$1 year in the edge-brightened jet structure extending up to 12\,mas from the core. This work is based on high-cadence KaVA 22\,GHz observations conducted from December 2013 to June 2016. By analyzing the transverse velocity profiles and the spatial evolution of the oscillations, we find that the oscillations propagate downstream along the jet, with a wavelength of $\sim9-10$\,mas. A single-mode sinusoidal wave model applied to the ridge lines successfully reproduces the observed transverse oscillations and yields superluminal wave speeds of $\sim2.7-2.9\,c$, consistent with the bulk jet velocity in this region. These findings suggest that the transverse oscillations may be interpreted either as transverse MHD waves -- possibly excited by jet precession, nutation, or quasi-periodic magnetic flux eruptions near the central engine -- or as manifestations of jet instabilities, such as current-driven instabilities (CDIs). Further investigation is required to distinguish between these scenarios and to clarify the dominant physical mechanism.

astro-ph.HE

Locating the missing large-scale emission in the jet of M87* with short EHT baselines

In Very-Long Baseline Interferometric arrays, nearly co-located stations probe the largest scales and typically cannot resolve the observed source. In the absence of large-scale structure, closure phases constructed with these stations are zero and, since they are independent of station-based errors, they can be used to probe data issues. Here, we show with an expansion about co-located stations, how these trivial closure phases become non-zero with brightness distribution on smaller scales than their short baseline would suggest. When applied to sources that are made up of a bright compact and large-scale diffuse component, the trivial closure phases directly measure the centroid relative to the compact source and higher-order image moments. We present a technique to measure these image moments with minimal model assumptions and validate it on synthetic Event Horizon Telescope (EHT) data. We then apply this technique to 2017 and 2018 EHT observations of M87* and find a weak preference for extended emission in the direction of the large-scale jet. We also apply it to 2021 EHT data and measure the source centroid about 1 mas northwest of the compact ring, consistent with the jet observed at lower frequencies.

astro-ph.HE

A high-order augmented Lagrangian method with arbitrarily fast convergence

We propose a high-order version of the augmented Lagrangian method for solving convex optimization problems with linear constraints, which achieves arbitrarily fast -- and even superlinear -- convergence rates. First, we analyze the convergence rates of the high-order proximal point method under certain uniform convexity assumptions on the energy functional. We then introduce the high-order augmented Lagrangian method and analyze its convergence by leveraging the convergence results of the high-order proximal point method. Finally, we present applications of the high-order augmented Lagrangian method to various problems arising in the sciences, including data fitting, flow in porous media, and scientific machine learning.

math.OC

Probing jet base emission of M87* with the 2021 Event Horizon Telescope observations

We investigate the presence and spatial characteristics of the jet base emission in M87* at 230 GHz, enabled by the enhanced uv coverage in the 2021 Event Horizon Telescope (EHT) observations. The addition of the 12-m Kitt Peak Telescope and NOEMA provides two key intermediate-length baselines to SMT and the IRAM 30-m, giving sensitivity to emission structures at scales of $\sim250~μ$as and $\sim2500~μ$as (0.02 pc and 0.2 pc). Without these baselines, earlier EHT observations lacked the capability to constrain emission on large scales, where a "missing flux" of order $\sim1$ Jy is expected. To probe these scales, we analyzed closure phases, robust against station-based gain errors, and modeled the jet base emission using a simple Gaussian offset from the compact ring emission at separations $>100~μ$as. Our analysis reveals a Gaussian feature centered at ($Δ$RA $\approx320~μ$as, $Δ$Dec $\approx60~μ$as), a projected separation of $\approx5500$ AU, with a flux density of only $\sim60$ mJy, implying that most of the missing flux in previous studies must arise from larger scales. Brighter emission at these scales is ruled out, and the data do not favor more complex models. This component aligns with the inferred direction of the large-scale jet and is consistent with emission from the jet base. While our findings indicate detectable jet base emission at 230 GHz, coverage from only two intermediate baselines limits reconstruction of its morphology. We therefore treat the recovered Gaussian as an upper limit on the jet base flux density. Future EHT observations with expanded intermediate-baseline coverage will be essential to constrain the structure and nature of this component.

astro-ph.HE