arXiv ScienceSearch

arXiv subjects

Han Yu

Publications and source records attributed to Han Yu.

At least 217 records · Page 12Linked to original sources

GTG-Shapley: Efficient and Accurate Participant Contribution Evaluation in Federated Learning

Federated Learning (FL) bridges the gap between collaborative machine learning and preserving data privacy. To sustain the long-term operation of an FL ecosystem, it is important to attract high quality data owners with appropriate incentive schemes. As an important building block of such incentive schemes, it is essential to fairly evaluate participants' contribution to the performance of the final FL model without exposing their private data. Shapley Value (SV)-based techniques have been widely adopted to provide fair evaluation of FL participant contributions. However, existing approaches incur significant computation costs, making them difficult to apply in practice. In this paper, we propose the Guided Truncation Gradient Shapley (GTG-Shapley) approach to address this challenge. It reconstructs FL models from gradient updates for SV calculation instead of repeatedly training with different combinations of FL participants. In addition, we design a guided Monte Carlo sampling approach combined with within-round and between-round truncation to further reduce the number of model reconstructions and evaluations required, through extensive experiments under diverse realistic data distribution settings. The results demonstrate that GTG-Shapley can closely approximate actual Shapley values, while significantly increasing computational efficiency compared to the state of the art, especially under non-i.i.d. settings.

cs.AI

Topology-aware Differential Privacy for Decentralized Image Classification

In this paper, we design Top-DP, a novel solution to optimize the differential privacy protection of decentralized image classification systems. The key insight of our solution is to leverage the unique features of decentralized communication topologies to reduce the noise scale and improve the model usability. (1) We enhance the DP-SGD algorithm with this topology-aware noise reduction strategy, and integrate the time-aware noise decay technique. (2) We design two novel learning protocols (synchronous and asynchronous) to protect systems with different network connectivities and topologies. We formally analyze and prove the DP requirement of our proposed solutions. Experimental evaluations demonstrate that our solution achieves a better trade-off between usability and privacy than prior works. To the best of our knowledge, this is the first DP optimization work from the perspective of network topologies.

cs.CR

Personalised Federated Learning: A Combinational Approach

Federated learning (FL) is a distributed machine learning approach involving multiple clients collaboratively training a shared model. Such a system has the advantage of more training data from multiple clients, but data can be non-identically and independently distributed (non-i.i.d.). Privacy and integrity preserving features such as differential privacy (DP) and robust aggregation (RA) are commonly used in FL. In this work, we show that on common deep learning tasks, the performance of FL models differs amongst clients and situations, and FL models can sometimes perform worse than local models due to non-i.i.d. data. Secondly, we show that incorporating DP and RA degrades performance further. Then, we conduct an ablation study on the performance impact of different combinations of common personalization approaches for FL, such as finetuning, mixture-of-experts ensemble, multi-task learning, and knowledge distillation. It is observed that certain combinations of personalization approaches are more impactful in certain scenarios while others always improve performance, and combination approaches are better than individual ones. Most clients obtained better performance with combined personalized FL and recover from performance degradation caused by non-i.i.d. data, DP, and RA.

cs.LG

Latent-Optimized Adversarial Neural Transfer for Sarcasm Detection

The existence of multiple datasets for sarcasm detection prompts us to apply transfer learning to exploit their commonality. The adversarial neural transfer (ANT) framework utilizes multiple loss terms that encourage the source-domain and the target-domain feature distributions to be similar while optimizing for domain-specific performance. However, these objectives may be in conflict, which can lead to optimization difficulties and sometimes diminished transfer. We propose a generalized latent optimization strategy that allows different losses to accommodate each other and improves training dynamics. The proposed method outperforms transfer learning and meta-learning baselines. In particular, we achieve 10.02% absolute performance gain over the previous state of the art on the iSarcasm dataset.

cs.LG

On sets containing a unit distance in every direction

We investigate the box dimensions of compact sets in $\mathbb{R}^2$ that contain a unit distance in every direction (such sets may have zero Hausdorff dimension). Among other results, we show that the lower box dimension must be at least $\frac{4}{7}$ and can be as low as $\frac{2}{3}$. This quantifies in a certain sense how far the unit circle is from being a difference set.

math.CA

Dimensions of the popcorn graph

The 'popcorn function' isThe `popcorn function' is a well-known and important example in real analysis with many interesting features. We prove that the box dimension of the graph of the popcorn function is 4/3, as well as computing the Assouad dimension and Assouad spectrum. The main ingredients include Duffin-Schaeffer type estimates from Diophantine approximation and the Chung-Erdős inequality from probability theory.

math.MG

Forecasting Health and Wellbeing for Shift Workers Using Job-role Based Deep Neural Network

Shift workers who are essential contributors to our society, face high risks of poor health and wellbeing. To help with their problems, we collected and analyzed physiological and behavioral wearable sensor data from shift working nurses and doctors, as well as their behavioral questionnaire data and their self-reported daily health and wellbeing labels, including alertness, happiness, energy, health, and stress. We found the similarities and differences between the responses of nurses and doctors. According to the differences in self-reported health and wellbeing labels between nurses and doctors, and the correlations among their labels, we proposed a job-role based multitask and multilabel deep learning model, where we modeled physiological and behavioral data for nurses and doctors simultaneously to predict participants' next day's multidimensional self-reported health and wellbeing status. Our model showed significantly better performances than baseline models and previous state-of-the-art models in the evaluations of binary/3-class classification and regression prediction tasks. We also found features related to heart rate, sleep, and work shift contributed to shift workers' health and wellbeing.

cs.LG

An improvement on Furstenberg's intersection problem

In this paper, we study a problem posed by Furstenberg on intersections between $\times 2, \times 3$ invariant sets. We present here a direct geometrical counting argument to revisit a theorem of Wu and Shmerkin. This argument can be used to obtain further improvements. For example, we show that if $A_2,A_3\subset [0,1]$ are closed and $\times 2, \times 3$ invariant respectively, assuming that $\dim A_2+\dim A_3<1$ then $A_2\cap (uA_3+v)$ is sparse (defined in this paper) and has box dimension zero uniformly with respect to the real parameters $u,v$ such that $u$ and $u^{-1}$ are both bounded away from $0$.

math.NT

Additive properties of numbers with restricted digits

In this paper, we consider some additive properties of integers with restricted digit expansions. Let $b\geq 3$ be an integer and $B_b$ be the set of integers whose base $b$ expansions have only digits $\{0,1\}.$ Let $a,b,c$ be three integers greater than $2.$ We give some estimates on the size of $(B_{a}+B_{b})\cap B_{c}.$ In particular, under mild conditions, $(B_{a}+B_{b})\cap B_{c}$ is a very thin set in the following sense that for each $ε>0,$ as $N\to\infty,$ \[ \#((B_{a}+B_{b})\cap B_{c} \cap [1,N])=O(N^ε). \]

math.DS

Towards Cost-Optimal Policies for DAGs to Utilize IaaS Clouds with Online Learning

Premier cloud service providers (CSPs) offer two types of purchase options, namely on-demand and spot instances, with time-varying features in availability and price. Users like startups have to operate on a limited budget and similarly others hope to reduce their costs. While interacting with a CSP, central to their concerns is the process of cost-effectively utilizing different purchase options possibly in addition to self-owned instances. A job in data-intensive applications is typically represented by a directed acyclic graph which can further be transformed into a chain of tasks. The key to achieving cost efficiency is determining the allocation of a specific deadline to each task, as well as the allocation of different types of instances to the task. In this paper, we propose a framework that determines the optimal allocation of deadlines to tasks. The framework also features an optimal policy to determine the allocation of spot and on-demand instances in a predefined time window, and a near-optimal policy for allocating self-owned instances. The policies are designed to be parametric to support the usage of online learning to infer the optimal values against the dynamics of cloud markets. Finally, several intuitive heuristics are used as baselines to validate the cost improvement brought by the proposed solutions. We show that the cost improvement over the state-of-the-art is up to 24.87% when spot and on-demand instances are considered and up to 59.05% when self-owned instances are considered.

cs.PF

Topological Pilot Assignment in Large-Scale Distributed MIMO Networks

We consider the pilot assignment problem in large-scale distributed multi-input multi-output (MIMO) networks, where a large number of remote radio head (RRH) antennas are randomly distributed in a wide area, and jointly serve a relatively smaller number of users (UE) coherently. By artificially imposing topological structures on the UE-RRH connectivity, we model the network by a partially-connected interference network, so that the pilot assignment problem can be cast as a topological interference management problem with multiple groupcast messages. Building upon such connection, we formulate the topological pilot assignment (TPA) problem in two different ways with respect to whether or not the to-be-estimated channel connectivity pattern is known a priori. When it is known, we formulate the TPA problem as a low-rank matrix completion problem that can be solved by a simple alternating projection algorithm. Otherwise, we formulate it as a sequential maximum weight induced matching problem that can be solved by either a mixed integer linear program or a simple yet efficient greedy algorithm. With respect to two different formulations of the TPA problem, we evaluate the efficiency of the proposed algorithms under the cell-free massive MIMO setting.

cs.IT

Downlink Precoding for DP-UPA FDD Massive MIMO via Multi-Dimensional Active Channel Sparsification

In this paper, we consider user selection and downlink precoding for an over-loaded single-cell massive multiple-input multiple-output (MIMO) system in frequency division duplexing (FDD) mode, where the base station is equipped with a dual-polarized uniform planar array (DP-UPA) and serves a large number of single-antenna users. Due to the absence of uplink-downlink channel reciprocity and the high-dimensionality of channel matrices, it is extremely challenging to design downlink precoders using closed-loop channel probing and feedback with limited spectrum resource. To address these issues, a novel methodology -- active channel sparsification (ACS) -- has been proposed recently in the literature for uniform linear array (ULA) to design sparsifying precoders, which boosts spectral efficiency for multi-user downlink transmission with substantially reduced channel feedback overhead. Pushing forward this line of research, we aim to facilitate the potential deployment of ACS in practical FDD massive MIMO systems, by extending it from ULA to DP-UPA with explicit user selection and making the current ACS implementation simplified. To this end, by leveraging Toeplitz structure of channel covariance matrices, we extend the original ACS using scale-weight bipartite graph representation to the matrix-weight counterpart. Building upon this, we propose a multi-dimensional ACS (MD-ACS) method, which is a generalization of original ACS formulation and is more suitable for DP-UPA antenna configurations. The nonlinear integer program formulation of MD-ACS can be classified as a generalized multi-assignment problem (GMAP), for which we propose a simple yet efficient greedy algorithm to solve it. Simulation results demonstrate the performance improvement of the proposed MD-ACS with greedy algorithm over the state-of-the-art methods based on the QuaDRiGa channel models.

cs.IT

Fourier decay of self-similar measures and self-similar sets of uniqueness

In this paper, we investigate the Fourier transform of self-similar measures on R. We provide quantitative decay rates of Fourier transform of some self-similar measures. Our method is based on random walks on lattices and Diophantine approximation in number fields. We also completely identify all self-similar sets which are sets of uniqueness. This generalizes a classical result of Salem and Zygmund.

math.CA

Digit expansions of numbers in different bases

A folklore conjecture in number theory states that the only integers whose expansions in base $3,4$ and $5$ contain solely binary digits are $0, 1$ and $82000$. In this paper, we present the first progress on this conjecture. Furthermore, we investigate the density of the integers containing only binary digits in their base $3$ or $4$ expansion, whereon an exciting transition in behaviour is observed. Our methods shed light on the reasons for this, and relate to several well-known questions, such as Graham's problem and a related conjecture of Pomerance. Finally, we generalise this setting and prove that the set of numbers in $[0, 1]$ who do not contain some digit in their $b$-expansion for all $b \geq 3$ has zero Hausdorff dimension.

math.NT

Noise-resistant Deep Metric Learning with Ranking-based Instance Selection

The existence of noisy labels in real-world data negatively impacts the performance of deep learning models. Although much research effort has been devoted to improving robustness to noisy labels in classification tasks, the problem of noisy labels in deep metric learning (DML) remains open. In this paper, we propose a noise-resistant training technique for DML, which we name Probabilistic Ranking-based Instance Selection with Memory (PRISM). PRISM identifies noisy data in a minibatch using average similarity against image features extracted by several previous versions of the neural network. These features are stored in and retrieved from a memory bank. To alleviate the high computational cost brought by the memory bank, we introduce an acceleration method that replaces individual data points with the class centers. In extensive comparisons with 12 existing approaches under both synthetic and real-world label noise, PRISM demonstrates superior performance of up to 6.06% in Precision@1.

cs.CV

Advances and Open Problems in Federated Learning

Federated learning (FL) is a machine learning setting where many clients (e.g. mobile devices or whole organizations) collaboratively train a model under the orchestration of a central server (e.g. service provider), while keeping the training data decentralized. FL embodies the principles of focused data collection and minimization, and can mitigate many of the systemic privacy risks and costs resulting from traditional, centralized machine learning and data science approaches. Motivated by the explosive growth in FL research, this paper discusses recent advances and presents an extensive collection of open problems and challenges.

cs.LG

On the metric theory of inhomogeneous Diophantine approximation: An Erdős-Vaaler type result

In 1958, Szüsz proved an inhomogeneous version of Khintchine's theorem on Diophantine approximation. Szüsz's theorem states that for any non-increasing approximation function $ψ:\mathbb{N}\to (0,1/2)$ with $\sum_q ψ(q)=\infty$ and any number $γ,$ the following set \[ W(ψ,γ)=\{x\in [0,1]: |qx-p-γ|< ψ(q) \text{ for infinitely many } q,p\in\mathbb{N}\} \] has full Lebesgue measure. Since then, there are very few results in relaxing the monotonicity condition. In this paper, we show that if $γ$ is can not be approximate by rational numbers too well, then the monotonicity condition can be replaced by the upper bound condition $ψ(q)=O((q(\log\log q)^2)^{-1}).$ In particular, this covers the case when $γ$ is not Liouville, for example $π,e,\ln 2, \sqrt{2}.$ In general, if $γ$ is irrational, $ψ(q)=O(q^{-1}(\log\log q)^{-2})$ and in addition, \[ \left(\liminf_{Q\to\infty} \sum_{q=Q}^{Q^{(\log Q)^{1/8} }}ψ(q)\right)=\infty, \] then $W(ψ,γ)$ has full Lebesgue measure. Our proof is based on a quantitative study of the discrepancy for irrational rotations.

math.NT

Rational points near self-similar sets

In this paper, we consider a problem of counting rational points near self-similar sets. Let $n\geq 1$ be an integer. We shall show that for some self-similar measures on $\mathbb{R}^n$, the set of rational points $\mathbb{Q}^n$ is 'equidistributed' in a sense that will be introduced in this paper. This implies that an inhomogeneous Khinchine convergence type result can be proved for those measures. In particular, for $n=1$ and large enough integers $p,$ the above holds for the middle-$p$th Cantor measure, i.e. the natural Hausdorff measure on the set of numbers whose base $p$ expansions do not have digit $[(p-1)/2].$ Furthermore, we partially proved a conjecture of Bugeaud and Durand for the middle-$p$th Cantor set and this also answers a question posed by Levesley, Salp and Velani. Our method includes a fine analysis of the Fourier coefficients of self-similar measures together with an Erdős-Kahane type argument. We will also provide a numerical argument to show that $p>10^7$ is sufficient for the above conclusions. In fact, $p\geq 15$ is already enough for most of the above conclusions.

math.NT