arXiv ScienceSearch

arXiv subjects

Keyu Zhu

Publications and source records attributed to Keyu Zhu.

14 recordsLinked to original sources

A pilot study on the CSST astrometric capability: Detecting astrometric binaries with Gaia synergy via simulated data

Context. The China Space-station Survey Telescope (CSST) will provide deep, wide-field epoch astrometry during its 10-year mission. Astrometric binary orbits constrain the masses of stellar and compact-object components. Orbital recovery depends on astrometric precision and temporal coverage. Combining CSST and Gaia data extends the baseline and improves binary detection. Aims. We evaluate CSST, Gaia, and joint astrometry for binary-candidate selection and 12-parameter (12p) orbit fitting at faint magnitudes ($g>17.8$). We also test how regular CSST cadences affect the yield of 12p fits satisfying our criteria. Methods. We constructed a mock catalog, simulated CSST and Gaia epoch astrometry, and fitted five-parameter (5p) single-star models to derive astrometric diagnostics, proper-motion anomaly features, and observational-sampling features. A four-stage histogram-based gradient-boosting classifier used these features to select candidates for 12p orbit fitting and assessment. Results. On the independent test set, the classifier reaches a precision of 0.802 and a recall of 0.181 among eligible true binaries. In the scenario-specific fitted samples, joint astrometry raises the fiducial fraction from 6.76% for Gaia alone to 10.46%; for fitted binaries with $P_{\rm true}>15{\rm yr}$, it rises from 2.37% to 6.78%. The current CSST schedule yields few fiducial fits, while idealized regular cadences increase the yield mainly at $g\lesssim21$. Conclusions. In the simulation, joint CSST and Gaia epoch astrometry yields higher fractions of fitted unresolved binaries satisfying the stated criteria than Gaia-only solution. A practical strategy is to select candidates from 5p diagnostics and astrometric anomalies, obtain more regular CSST follow-up observations, and then fit 12p orbital models and apply the selection criteria.

astro-ph.IM

Optical-morphology-based assessment of astrometric quality in Gaia-CRF3 quasars

Context. Several studies have shown that host-galaxy structure or extended optical morphology in AGNs can induce spurious parallaxes and proper motions in Gaia DR3. However, it remains unclear whether source morphology also introduces systematic errors into the celestial reference frame constructed from Gaia data. Aims. We aim to provide a Gaia-independent external morphological indicator for Gaia-CRF3 sources and to use it to quantify the astrometric systematics associated with source morphology. Methods. Using morphological parameters derived from DESI, SDSS, and SkyMapper, together with the PS1-PSC point-source score as a common reference scale, we used XGBoost to infer external morphological scores for Gaia-CRF3 sources. We then developed a multi-survey fusion scheme to combine the four survey-based point-source scores into a single composite score that measures the degree to which each source departs from the morphology of an ideal point source. Results. We obtained morphological scores for 1,607,490 Gaia-CRF3 sources, corresponding to a completeness of 99.59\% with respect to the full Gaia-CRF3 catalogue. The score ranges from 0 to 1 and remains reliable for sources with $G<20.85$ mag. Based on this indicator, we find that AGNs with strongly non-point-like morphology induce a parallax zero-point shift of about $-43.7\,\mu$as, which cannot be effectively removed by the current parallax zero-point correction model. We also find that reference-source subsamples selected in different score ranges exhibit significantly different all-sky proper motion fields. For the high-purity point-source subsample with \texttt{point\_score} > 0.95, the total frame spin amplitude is reduced by 15.8\% relative to that of the full Gaia-CRF3 sample.

astro-ph.GA

Refining the Gaia DR3 Parallax Zero-point: A Hybrid Approach Combining Global Parametric Correction with Local Refinement

The Gaia Data Release 3 (GDR3) parallaxes are affected by a complex bias that depends on stellar magnitude, color, and celestial position, with amplitudes reaching tens of microarcseconds ($\mu$as). Standard global parametric models (e.g., Lindegren et al. 2021, hereafter L21) effectively remove large-scale trends but struggle to resolve small-scale spatial systematics due to functional rigidity. We aim to construct a flexible, data-driven calibration map that eliminates these residual local systematics without imposing rigid functional forms. We propose a "Global Pre-correction + Local Refinement" hybrid strategy. First, we utilize the L21 model as a baseline to remove the dominant magnitude and color-dependent biases. Second, we model the residual zero-point using a Local Non-parametric method based on a Sliding Window technique. This approach fits local trends using k-nearest neighbors from quasars (for faint stars, G>18) and wide binaries combined with Large Magellanic Cloud (LMC) (for bright stars, G < 18). Our hybrid model demonstrates significant improvements over the standard L21 solution. Validation against different samples reveals a remarkably flat residual map with near-zero bias across the full sky. Our mathematical attempt at calibrating the parallax zero-point is expected to provide a useful reference for the zero-point correction in future Gaia DR4, and to help move towards a physical resolution of this issue.

astro-ph.IM

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the prohibitive cost and scalability limits of human-in-the-loop collection paradigms. To break this barrier, we introduce Robust Autonomous Data Acquisition for Robotics (RADAR), a fully autonomous, closed-loop data generation engine that completely removes human intervention from the collection cycle. RADAR elegantly divides the cognitive load into a four-module pipeline. Anchored by 2-5 3D human demonstrations as geometric priors, a Vision-Language Model first orchestrates scene-relevant task generation via precise semantic object grounding and skill retrieval. Next, a Graph Neural Network policy translates these subtasks into physical actions via in-context imitation learning. Following execution, the VLM performs automated success evaluation using a structured Visual Question Answering pipeline. Finally, to shatter the bottleneck of manual resets, a Finite State Machine orchestrates an autonomous environment reset and asymmetric data routing mechanism. Driven by simultaneous forward-reverse planning with a strict Last-In, First-Out causal sequence, the system seamlessly restores unstructured workspaces and robustly recovers from execution failures. This continuous brain-cerebellum synergy transforms data collection into a self-sustaining process. Extensive evaluations highlight RADAR's exceptional versatility. In simulation, our framework achieves up to 90% success rates on complex, long-horizon tasks, effortlessly solving challenges where traditional baselines plummet to near-zero performance. In real-world deployments, the system reliably executes diverse, contact-rich skills (e.g., deformable object manipulation) via few-shot adaptation without domain-specific fine-tuning, providing a highly scalable paradigm for robotic data acquisition.

cs.RO

High Q and high gradient performance of the first medium-temperature baking 1.3 GHz cryomodule

World's first 1.3 GHz cryomodule containing eight 9-cell superconducting radio-frequency (RF) cavities treated by medium-temperature furnace baking (mid-T bake) was developed, assembled and tested at IHEP for the Dalian Advanced Light Source (DALS) and CEPC R&D. The 9-cell cavities in the cryomodule achieved an unprecedented highest average Q0 of 3.8E10 at 16 MV/m and 3.6E10 at 21 MV/m in the horizontal test. The cryomodule can operate stably up to a total CW RF voltage greater than 191 MV, with an average cavity CW accelerating gradient of more than 23 MV/m. The results significantly exceed the specifications of CEPC, DALS and the other high repetition rate free electron laser facilities (LCLS-II, LCLS-II-HE, SHINE, S3FEL). There is evidence that the mid-T bake cavity may not require fast cool-down or long processing time in the cryomodule. This paper reviews the cryomodule performance and discusses some important issues in cryomodule assembly and testing.

physics.acc-ph

Impacts of Differential Privacy on Fostering more Racially and Ethnically Diverse Elementary Schools

In the face of increasingly severe privacy threats in the era of data and AI, the US Census Bureau has recently adopted differential privacy, the de facto standard of privacy protection for the 2020 Census release. Enforcing differential privacy involves adding carefully calibrated random noise to sensitive demographic information prior to its release. This change has the potential to impact policy decisions like political redistricting and other high-stakes practices, partly because tremendous federal funds and resources are allocated according to datasets (like Census data) released by the US government. One under-explored yet important application of such data is the redrawing of school attendance boundaries to foster less demographically segregated schools. In this study, we ask: how differential privacy might impact diversity-promoting boundaries in terms of resulting levels of segregation, student travel times, and school switching requirements? Simulating alternative boundaries using differentially-private student counts across 67 Georgia districts, we find that increasing data privacy requirements decreases the extent to which alternative boundaries might reduce segregation and foster more diverse and integrated schools, largely by reducing the number of students who would switch schools under boundary changes. Impacts on travel times are minimal. These findings point to a privacy-diversity tradeoff local educational policymakers may face in forthcoming years, particularly as computational methods are increasingly poised to facilitate attendance boundary redrawings in the pursuit of less segregated schools.

cs.CY

Privacy and Bias Analysis of Disclosure Avoidance Systems

Disclosure avoidance (DA) systems are used to safeguard the confidentiality of data while allowing it to be analyzed and disseminated for analytic purposes. These methods, e.g., cell suppression, swapping, and k-anonymity, are commonly applied and may have significant societal and economic implications. However, a formal analysis of their privacy and bias guarantees has been lacking. This paper presents a framework that addresses this gap: it proposes differentially private versions of these mechanisms and derives their privacy bounds. In addition, the paper compares their performance with traditional differential privacy mechanisms in terms of accuracy and fairness on US Census data release and classification tasks. The results show that, contrary to popular beliefs, traditional differential privacy techniques may be superior in terms of accuracy and fairness to differential private counterparts of widely used DA mechanisms.

cs.CR

Fairness Increases Adversarial Vulnerability

The remarkable performance of deep learning models and their applications in consequential domains (e.g., facial recognition) introduces important challenges at the intersection of equity and security. Fairness and robustness are two desired notions often required in learning models. Fairness ensures that models do not disproportionately harm (or benefit) some groups over others, while robustness measures the models' resilience against small input perturbations. This paper shows the existence of a dichotomy between fairness and robustness, and analyzes when achieving fairness decreases the model robustness to adversarial samples. The reported analysis sheds light on the factors causing such contrasting behavior, suggesting that distance to the decision boundary across groups as a key explainer for this behavior. Extensive experiments on non-linear models and different architectures validate the theoretical findings in multiple vision domains. Finally, the paper proposes a simple, yet effective, solution to construct models achieving good tradeoffs between fairness and robustness.

cs.LG

SF-PATE: Scalable, Fair, and Private Aggregation of Teacher Ensembles

A critical concern in data-driven processes is to build models whose outcomes do not discriminate against some demographic groups, including gender, ethnicity, or age. To ensure non-discrimination in learning tasks, knowledge of the group attributes is essential. However, in practice, these attributes may not be available due to legal and ethical requirements. To address this challenge, this paper studies a model that protects the privacy of the individuals' sensitive information while also allowing it to learn non-discriminatory predictors. A key characteristic of the proposed model is to enable the adoption of off-the-selves and non-private fair models to create a privacy-preserving and fair model. The paper analyzes the relation between accuracy, privacy, and fairness, and the experimental evaluation illustrates the benefits of the proposed models on several prediction tasks. In particular, this proposal is the first to allow both scalable and accurate training of private and fair models for very large neural networks.

cs.LG

Differential Privacy and Fairness in Decisions and Learning Tasks: A Survey

This paper surveys recent work in the intersection of differential privacy (DP) and fairness. It reviews the conditions under which privacy and fairness may have aligned or contrasting goals, analyzes how and why DP may exacerbate bias and unfairness in decision problems and learning tasks, and describes available mitigation measures for the fairness issues arising in DP systems. The survey provides a unified understanding of the main challenges and potential risks arising when deploying privacy-preserving machine-learning or decisions-making tasks under a fairness lens.

cs.LG

Post-processing of Differentially Private Data: A Fairness Perspective

Post-processing immunity is a fundamental property of differential privacy: it enables arbitrary data-independent transformations to differentially private outputs without affecting their privacy guarantees. Post-processing is routinely applied in data-release applications, including census data, which are then used to make allocations with substantial societal impacts. This paper shows that post-processing causes disparate impacts on individuals or groups and analyzes two critical settings: the release of differentially private datasets and the use of such private datasets for downstream decisions, such as the allocation of funds informed by US Census data. In the first setting, the paper proposes tight bounds on the unfairness of traditional post-processing mechanisms, giving a unique tool to decision-makers to quantify the disparate impacts introduced by their release. In the second setting, this paper proposes a novel post-processing mechanism that is (approximately) optimal under different fairness metrics, either reducing fairness issues substantially or reducing the cost of privacy. The theoretical analysis is complemented with numerical simulations on Census data.

cs.CR

Bias and Variance of Post-processing in Differential Privacy

Post-processing immunity is a fundamental property of differential privacy: it enables the application of arbitrary data-independent transformations to the results of differentially private outputs without affecting their privacy guarantees. When query outputs must satisfy domain constraints, post-processing can be used to project the privacy-preserving outputs onto the feasible region. Moreover, when the feasible region is convex, a widely adopted class of post-processing steps is also guaranteed to improve accuracy. Post-processing has been applied successfully in many applications including census data-release, energy systems, and mobility. However, its effects on the noise distribution is poorly understood: It is often argued that post-processing may introduce bias and increase variance. This paper takes a first step towards understanding the properties of post-processing. It considers the release of census data and examines, both theoretically and empirically, the behavior of a widely adopted class of post-processing functions.

cs.LG

Differential Privacy of Hierarchical Census Data: An Optimization Approach

This paper is motivated by applications of a Census Bureau interested in releasing aggregate socio-economic data about a large population without revealing sensitive information about any individual. The released information can be the number of individuals living alone, the number of cars they own, or their salary brackets. Recent events have identified some of the privacy challenges faced by these organizations. To address them, this paper presents a novel differential-privacy mechanism for releasing hierarchical counts of individuals. The counts are reported at multiple granularities (e.g., the national, state, and county levels) and must be consistent across all levels. The core of the mechanism is an optimization model that redistributes the noise introduced to achieve differential privacy in order to meet the consistency constraints between the hierarchical levels. The key technical contribution of the paper shows that this optimization problem can be solved in polynomial time by exploiting the structure of its cost functions. Experimental results on very large, real datasets show that the proposed mechanism provides improvements of up to two orders of magnitude in terms of computational efficiency and accuracy with respect to other state-of-the-art techniques.

cs.DB

Optimal Pricing For MHR and $\lambda$-Regular Distributions

We study the performance of anonymous posted-price selling mechanisms for a standard Bayesian auction setting, where $n$ bidders have i.i.d. valuations for a single item. We show that for the natural class of Monotone Hazard Rate (MHR) distributions, offering the same, take-it-or-leave-it price to all bidders can achieve an (asymptotically) optimal revenue. In particular, the approximation ratio is shown to be $1+O(\ln \ln n/\ln n)$, matched by a tight lower bound for the case of exponential distributions. This improves upon the previously best-known upper bound of $e/(e-1)\approx 1.58$ for the slightly more general class of regular distributions. In the worst case (over $n$), we still show a global upper bound of $1.35$. We give a simple, closed-form description of our prices which, interestingly enough, relies only on minimal knowledge of the prior distribution, namely just the expectation of its second-highest order statistic. Furthermore, we extend our techniques to handle the more general class of $\lambda$-regular distributions that interpolate between MHR ($\lambda=0$) and regular ($\lambda=1$). Our anonymous pricing rule now results in an asymptotic approximation ratio that ranges smoothly, with respect to $\lambda$, from $1$ (MHR distributions) to $e/(e-1)$ (regular distributions). Finally, we explicitly give a class of continuous distributions that provide matching lower bounds, for every $\lambda$.

cs.GT