arXiv ScienceSearch

arXiv subjects

Sohyun Park

Publications and source records attributed to Sohyun Park.

At least 19 recordsLinked to original sources

PEEM: Prompt Engineering Evaluation Metrics for Interpretable Joint Evaluation of Prompts and Responses

Prompt design is a primary control interface for large language models (LLMs), yet standard evaluations largely reduce performance to answer correctness, obscuring why a prompt succeeds or fails and providing little actionable guidance. We propose PEEM (Prompt Engineering Evaluation Metrics), a unified framework for joint and interpretable evaluation of both prompts and responses. PEEM defines a structured rubric with 9 axes: 3 prompt criteria (clarity/structure, linguistic quality, fairness) and 6 response criteria (accuracy, coherence, relevance, objectivity, clarity, conciseness), and uses an LLM-based evaluator to output (i) scalar scores on a 1-5 Likert scale and (ii) criterion-specific natural-language rationales grounded in the rubric. Across 7 benchmarks and 5 task models, PEEM's accuracy axis strongly aligns with conventional accuracy while preserving model rankings (aggregate Spearman rho about 0.97, Pearson r about 0.94, p < 0.001). A multi-evaluator study with four models shows consistent relative judgments (pairwise rho = 0.68-0.85), supporting evaluator-agnostic deployment. Beyond alignment, PEEM captures complementary linguistic failure modes and remains informative under prompt perturbations: prompt-quality trends track downstream accuracy under iterative rewrites, semantic adversarial manipulations induce clear score degradation, and meaning-preserving paraphrases yield high stability (robustness rate about 76.7-80.6%). Finally, using only PEEM scores and rationales as feedback, a zero-shot prompt rewriting loop improves downstream accuracy by up to 11.7 points, outperforming supervised and RL-based prompt-optimization baselines. Overall, PEEM provides a reproducible, criterion-driven protocol that links prompt formulation to response behavior and enables systematic diagnosis and optimization of LLM interactions.

cs.CL

KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding

We introduce KFinEval-Pilot, a benchmark suite specifically designed to evaluate large language models (LLMs) in the Korean financial domain. Addressing the limitations of existing English-centric benchmarks, KFinEval-Pilot comprises over 1,000 curated questions across three critical areas: financial knowledge, legal reasoning, and financial toxicity. The benchmark is constructed through a semi-automated pipeline that combines GPT-4-generated prompts with expert validation to ensure domain relevance and factual accuracy. We evaluate a range of representative LLMs and observe notable performance differences across models, with trade-offs between task accuracy and output safety across different model families. These results highlight persistent challenges in applying LLMs to high-stakes financial applications, particularly in reasoning and safety. Grounded in real-world financial use cases and aligned with the Korean regulatory and linguistic context, KFinEval-Pilot serves as an early diagnostic tool for developing safer and more reliable financial AI systems.

cs.CL

Label-aware Hard Negative Sampling Strategies with Momentum Contrastive Learning for Implicit Hate Speech Detection

Detecting implicit hate speech that is not directly hateful remains a challenge. Recent research has attempted to detect implicit hate speech by applying contrastive learning to pre-trained language models such as BERT and RoBERTa, but the proposed models still do not have a significant advantage over cross-entropy loss-based learning. We found that contrastive learning based on randomly sampled batch data does not encourage the model to learn hard negative samples. In this work, we propose Label-aware Hard Negative sampling strategies (LAHN) that encourage the model to learn detailed features from hard negative samples, instead of naive negative samples in random batch, using momentum-integrated contrastive learning. LAHN outperforms the existing models for implicit hate speech detection both in- and cross-datasets. The code is available at https://github.com/Hanyang-HCC-Lab/LAHN

cs.CL

HearHere: Mitigating Echo Chambers in News Consumption through an AI-based Web System

Considerable efforts are currently underway to mitigate the negative impacts of echo chambers, such as increased susceptibility to fake news and resistance towards accepting scientific evidence. Prior research has presented the development of computer systems that support the consumption of news information from diverse political perspectives to mitigate the echo chamber effect. However, existing studies still lack the ability to effectively support the key processes of news information consumption and quantitatively identify a political stance towards the information. In this paper, we present HearHere, an AI-based web system designed to help users accommodate information and opinions from diverse perspectives. HearHere facilitates the key processes of news information consumption through two visualizations. Visualization 1 provides political news with quantitative political stance information, derived from our graph-based political classification model, and users can experience diverse perspectives (Hear). Visualization 2 allows users to express their opinions on specific political issues in a comment form and observe the position of their own opinions relative to pro-liberal and pro-conservative comments presented on a map interface (Here). Through a user study with 94 participants, we demonstrate the feasibility of HearHere in supporting the consumption of information from various perspectives. Our findings highlight the importance of providing political stance information and quantifying users' political status as a means to mitigate political polarization. In addition, we propose design implications for system development, including the consideration of demographics such as political interest and providing users with initiatives.

cs.HC

The Application of Affective Measures in Text-based Emotion Aware Recommender Systems

This paper presents an innovative approach to address the problems researchers face in Emotion Aware Recommender Systems (EARS): the difficulty and cumbersome collecting voluminously good quality emotion-tagged datasets and an effective way to protect users' emotional data privacy. Without enough good-quality emotion-tagged datasets, researchers cannot conduct repeatable affective computing research in EARS that generates personalized recommendations based on users' emotional preferences. Similarly, if we fail to fully protect users' emotional data privacy, users could resist engaging with EARS services. This paper introduced a method that detects affective features in subjective passages using the Generative Pre-trained Transformer Technology, forming the basis of the Affective Index and Affective Index Indicator (AII). Eliminate the need for users to build an affective feature detection mechanism. The paper advocates for a separation of responsibility approach where users protect their emotional profile data while EARS service providers refrain from retaining or storing it. Service providers can update users' Affective Indices in memory without saving their privacy data, providing Affective Aware recommendations without compromising user privacy. This paper offers a solution to the subjectivity and variability of emotions, data privacy concerns, and evaluation metrics and benchmarks, paving the way for future EARS research.

cs.IR

KHAN: Knowledge-Aware Hierarchical Attention Networks for Accurate Political Stance Prediction

The political stance prediction for news articles has been widely studied to mitigate the echo chamber effect -- people fall into their thoughts and reinforce their pre-existing beliefs. The previous works for the political stance problem focus on (1) identifying political factors that could reflect the political stance of a news article and (2) capturing those factors effectively. Despite their empirical successes, they are not sufficiently justified in terms of how effective their identified factors are in the political stance prediction. Motivated by this, in this work, we conduct a user study to investigate important factors in political stance prediction, and observe that the context and tone of a news article (implicit) and external knowledge for real-world entities appearing in the article (explicit) are important in determining its political stance. Based on this observation, we propose a novel knowledge-aware approach to political stance prediction (KHAN), employing (1) hierarchical attention networks (HAN) to learn the relationships among words and sentences in three different levels and (2) knowledge encoding (KE) to incorporate external knowledge for real-world entities into the process of political stance prediction. Also, to take into account the subtle and important difference between opposite political stances, we build two independent political knowledge graphs (KG) (i.e., KG-lib and KG-con) by ourselves and learn to fuse the different political knowledge. Through extensive evaluations on three real-world datasets, we demonstrate the superiority of DASH in terms of (1) accuracy, (2) efficiency, and (3) effectiveness.

cs.CL

Medium-modifications of $g\to c \bar{c}$ splitting

We study medium-modifications of the gluon splitting into a quark and anti-quark pair. Applying the Baier-Dokshitzer-Mueller-Peign\'e-Schiff and Zakharov (BDMPS-Z) formalism, we derive a path-integral formula for the in-medium $g\to q \bar{q}$ splitting function in the close-to-eikonal limit. Our analysis shows that there are two qualitatively different medium effects: transverse momentum broadening of $q \bar{q}$ pairs and enhanced production of such pairs. We note that both effects are numerically sizeable if the average momentum transfer from the medium to the parton is at the quark mass scale. In ultra-relativistic heavy-ion collisions, this condition is realized by charm quarks, therefore we focus our numerical analysis on the medium-modifications of $g\to c \bar{c}$ splitting.

hep-ph

Medium-enhanced $c\bar{c}$ radiation

We show that the same QCD formalism that accounts for the suppression of high-$p_T$ hadron and jet spectra in heavy-ion collisions predicts medium-enhanced production of $c\bar{c}$ pairs in jets.

hep-ph

Opportunities with ultra-soft photons: Bremsstrahlung from stopping

We compute the spectra of bremsstrahlung photons for different stopping scenarios which give rise to different initial charge-rapidity distributions. In the light of novel experimental opportunities that may arise with a new heavy-ion detector ALICE-3 at the CERN LHC, we discuss how to discriminate between these stopping scenarios and how to disentangle bremsstrahlung photons from other photon sources.

hep-ph

The medium-modified $g\to c\bar{c}$ splitting function in the BDMPS-Z formalism

The formalism of Baier-Dokshitzer-Mueller-Peign\'e-Schiff and Zakharov determines the modifications of parton splittings in the QCD plasma that arise from medium-induced gluon radiation. Here, we study medium-modifications of the gluon splitting into a quark--anti-quark pair in this BDMPS-Z formalism. We derive a compact path-integral formulation that resums effects from an arbitrary number of interactions with the medium to leading order in the $1/N_c^2$ expansion. Analyses in the $N=1$ opacity and the saddle point approximations reveal two phenomena: a medium-induced momentum broadening of the relative quark--anti-quark pair momentum that increases the invariant mass of quark--anti-quark pairs, and a medium-enhanced production of such pairs. We note that both effects are numerically sizeable if the average momentum transfer from the medium is comparable to the quark mass. In ultra-relativistic heavy-ion collisions, this condition is satisfied for charm quarks. We therefore focus our numerical analysis on the medium modification of $g\to c\bar{c}$, although our derivation applies equally well to $g\to b\bar{b}$ and to gluons splitting into light-flavoured quark--anti-quark pairs.

hep-ph

Bremsstrahlung photons from stopping in heavy-ion collisions

We examine the spectrum of bremsstrahlung photons that results from the stopping of the initial net charge distributions in ultrarelativistic nucleus-nucleus collisions at the CERN Large Hadron Collier (LHC). This effect has escaped detection so far since it becomes sizable only at very low transverse momentum and at sufficiently forward rapidity. We argue that it may be within reach of the next-generation LHC heavy-ion detector ALICE-3 that is currently under study, and we comment on the physics motivation for measuring it.

hep-ph

Black hole quasinormal modes and isospectrality in Deser-Woodard nonlocal gravity

We investigate the gravitational perturbations of the Schwarzschild black hole in the nonlocal gravity model recently proposed by Deser and Woodard (DW-II). The analysis is performed in the localized version in which the nonlocal corrections are represented by some auxiliary scalar fields. We find that the nonlocal corrections do not affect the axial gravitational perturbations, and hence the axial modes are completely identical to those in General Relativity (GR). However, the polar modes get different from their GR counterparts when the scalar fields are excited at the background level. In such a case the polar modes are sourced by an additional massless scalar mode and, as a result, the isospectrality between the axial and the polar modes breaks down. We also perform a similar analysis for the predecessor of this model (DW-I) and arrive at the same conclusion for it.

gr-qc

Primordial bouncing cosmology in the Deser-Woodard nonlocal gravity

The Deser-Woodard (DW) nonlocal gravity model has been proposed in order to describe the late-time acceleration of the universe without introducing dark energy. In this paper we focus, however, on the early stage of the universe and demonstrate how a primordial bounce in the vacuum spacetime can be realized in the framework of the DW nonlocal model. We reconstruct the nonlocal distortion function, which encodes all the modifications to the Einstein-Hilbert action, in order to generate bouncing solutions to solve the initial singularity problem. We show that the initial conditions can be chosen in such a way that the distortion function and its first order derivative approach zero after the bounce and the standard cosmological solution described by general relativity is recovered afterwards. We also study the evolution of anisotropies near the bounce. It turns out that the shear density defined by the anisotropy grows towards the bounce, but due to the presence of nonlocal effects, it grows in a milder manner compared with that in Einstein gravity.

gr-qc

Observational Constraints in Nonlocal Gravity: the Deser-Woodard Case

We study the cosmology of a specific class of nonlocal model of modified gravity, the so-called Deser-Woodard (DW) model, modifying the Einstein-Hilbert action by a term $\sim R f(\Box^{-1}R)$, where $f$ is a free function. Choosing $f$ so as to reproduce the $\Lambda{\rm CDM}$ cosmological background expansion history within the nonlocal model, we implement the model in a cosmological linear Einstein--Boltzmann solver and study the deviations to GR the model induces in the scalar and tensor perturbations. We observe that the DW nonlocal model describes a modified propagation for the gravitational waves, as well as a lower linear growth rate and a stronger lensing power as compared to $\Lambda{\rm CDM}$, up to several percents. Such prominent growth and lensing features lead to the inference of a significantly smaller value of $\sigma_8$ with respect to the one in $\Lambda{\rm CDM}$, given \textit{Planck} CMB+lensing data. The prediction for the linear growth rate $f \sigma_8$ within the DW model is therefore significantly smaller than the one in $\Lambda{\rm CDM}$ and the addition of growth rate data $f \sigma_8$ from Redshift-space distortion measurements to \textit{Planck} CMB+lensing, opens a (dominant) tension between Redshift-space distortion data and the reconstructed \textit{Planck} CMB lensing potential. However, model selection issues only result in "weak" evidences for $\Lambda{\rm CDM}$ against the DW model given the data. Such a fact shows that the datasets we consider are not constraining enough for distinguishing between the models. As we discuss, the addition of galaxy WL data or cosmological constraints from future galaxy clustering, weak lensing surveys, but also third generation gravitational wave interferometers, prove to be useful for discriminating modified gravity models such as the DW one from $\Lambda{\rm CDM}$, within the close future.

astro-ph.CO

Does nonlocal gravity yield divergent gravitational energy-momentum fluxes?

Energy-momentum conservation requires the associated gravitational fluxes on an asymptotically flat spacetime to scale as $1/r^2$, as $r \to \infty$, where $r$ is the distance between the observer and the source of the gravitational waves. We expand the equations-of-motion for the Deser-Woodard nonlocal gravity model up to quadratic order in metric perturbations, to compute its gravitational energy-momentum flux due to an isolated system. The contributions from the nonlocal sector contains $1/r$ terms proportional to the acceleration of the Newtonian energy of the system, indicating such nonlocal gravity models may not yield well-defined energy fluxes at infinity. In the case of the Deser-Woodard model, this divergent flux can be avoided by requiring the first and second derivatives of the nonlocal distortion function $f[X]$ at $X=0$ to be zero, i.e., $f'[0] = 0 = f''[0]$. It would be interesting to investigate whether other classes of nonlocal models not involving such an arbitrary function can avoid divergent fluxes.

gr-qc

Do the Spirits Rise?

A nonlocal gravity model based on $\frac1{\square} R$ achieves the phenomenological goals of generating cosmic acceleration without dark energy and of suppressing the growth of perturbations compared to the $\Lambda$CDM model. Although the localized version of this model possesses a scalar ghost, the nonlocal version does not suffer from any obvious problem with ghosts. Here we study the possibility that the scalar ghost mode might be excited through time evolution, even though it is initially absent. We present strong evidence that this does not happen.

gr-qc

Revival of the Deser-Woodard nonlocal gravity model: Comparison of the original nonlocal form and a localized formulation

We examine the origin of two opposite results for the growth of perturbations in the Deser-Woodard (DW) nonlocal gravity model. One group previously analyzed the model in its original nonlocal form and showed that the growth of structure in the DW model is enhanced compared to general relativity (GR) and thus concluded that the model was ruled out. Recently, however, another group has reanalyzed it by localizing the model and found that the growth in their localized version is suppressed even compared to the one in GR. The question was whether the discrepancy originates from an intrinsic difference between the nonlocal and localized formulations or is due to their different implementations of the sub-horizon limit. We show that the nonlocal and local formulations give the same solutions for the linear perturbations as long as the initial conditions are set the same. The different implementations of the sub-horizon limit lead to different transient behaviors of some perturbation variables; however, they do not affect the growth of matter perturbations at the sub-horizon scale much. In the meantime, we also report an error in the numerical calculation code of the former group and verify that after fixing the error the nonlocal version also gives the suppressed growth. Finally, we discuss two alternative definitions of the effective gravitational constant taken by the two groups and some open problems.

gr-qc

One loop corrected conformally coupled scalar mode equations during inflation

We employ a fully renormalized computation of the one loop contribution to the self-mass-squared of the conformally coupled (CC) scalar interacting with gravitons during inflation to study how inflationary produced gravitons affect the CC scalar evolution equation. The quantum corrected scalar mode functions turn out to get a secular growth effect, proportional to a logarithm of the scale factor at late times.

gr-qc