arXiv ScienceSearch

arXiv subjects

Lu Yu

Publications and source records attributed to Lu Yu.

At least 109 records · Page 6Linked to original sources

Oracle Inequalities for High-dimensional Prediction

The abundance of high-dimensional data in the modern sciences has generated tremendous interest in penalized estimators such as the lasso, scaled lasso, square-root lasso, elastic net, and many others. In this paper, we establish a general oracle inequality for prediction in high-dimensional linear regression with such methods. Since the proof relies only on convexity and continuity arguments, the result holds irrespective of the design matrix and applies to a wide range of penalized estimators. Overall, the bound demonstrates that generic estimators can provide consistent prediction with any design matrix. From a practical point of view, the bound can help to identify the potential of specific estimators, and they can help to get a sense of the prediction accuracy in a given application.

math.ST

Stochastic Tools for Network Intrusion Detection

With the rapid development of Internet and the sharp increase of network crime, network security has become very important and received a lot of attention. We model security issues as stochastic systems. This allows us to find weaknesses in existing security systems and propose new solutions. Exploring the vulnerabilities of existing security tools can prevent cyber-attacks from taking advantages of the system weaknesses. We propose a hybrid network security scheme including intrusion detection systems (IDSs) and honeypots scattered throughout the network. This combines the advantages of two security technologies. A honeypot is an activity-based network security system, which could be the logical supplement of the passive detection policies used by IDSs. This integration forces us to balance security performance versus cost by scheduling device activities for the proposed system. By formulating the scheduling problem as a decentralized partially observable Markov decision process (DEC-POMDP), decisions are made in a distributed manner at each device without requiring centralized control. The partially observable Markov decision process (POMDP) is a useful choice for controlling stochastic systems. As a combination of two Markov models, POMDPs combine the strength of hidden Markov Model (HMM) (capturing dynamics that depend on unobserved states) and that of Markov decision process (MDP) (taking the decision aspect into account). Decision making under uncertainty is used in many parts of business and science.We use here for security tools.We adopt a high-quality approximation solution for finite-space POMDPs with the average cost criterion, and their extension to DEC-POMDPs. We show how this tool could be used to design a network security framework.

cs.CR

Using Markov Models and Statistics to Learn, Extract, Fuse, and Detect Patterns in Raw Data

Many systems are partially stochastic in nature. We have derived data driven approaches for extracting stochastic state machines (Markov models) directly from observed data. This chapter provides an overview of our approach with numerous practical applications. We have used this approach for inferring shipping patterns, exploiting computer system side-channel information, and detecting botnet activities. For contrast, we include a related data-driven statistical inferencing approach that detects and localizes radiation sources.

cs.CR

TARN: A SDN-based Traffic Analysis Resistant Network Architecture

Destination IP prefix-based routing protocols are core to Internet routing today. Internet autonomous systems (AS) possess fixed IP prefixes, while packets carry the intended destination AS's prefix in their headers, in clear text. As a result, network communications can be easily identified using IP addresses and become targets of a wide variety of attacks, such as DNS/IP filtering, distributed Denial-of-Service (DDoS) attacks, man-in-the-middle (MITM) attacks, etc. In this work, we explore an alternative network architecture that fundamentally removes such vulnerabilities by disassociating the relationship between IP prefixes and destination networks, and by allowing any end-to-end communication session to have dynamic, short-lived, and pseudo-random IP addresses drawn from a range of IP prefixes rather than one. The concept is seemingly impossible to realize in todays Internet. We demonstrate how this is doable today with three different strategies using software defined networking (SDN), and how this can be done at scale to transform the Internet addressing and routing paradigms with the novel concept of a distributed software defined Internet exchange (SDX). The solution works with both IPv4 and IPv6, whereas the latter provides higher degrees of IP addressing freedom. Prototypes based on OpenvSwitches (OVS) have been implemented for experimentation across the PEERING BGP testbed. The SDX solution not only provides a technically sustainable pathway towards large-scale traffic analysis resistant network (TARN) support, it also unveils a new business model for customer driven, customizable and trustable end-to-end network services.

cs.NI

Provenance Threat Modeling

Provenance systems are used to capture history metadata, applications include ownership attribution and determining the quality of a particular data set. Provenance systems are also used for debugging, process improvement, understanding data proof of ownership, certification of validity, etc. The provenance of data includes information about the processes and source data that leads to the current representation. In this paper we study the security risks provenance systems might be exposed to and recommend security solutions to better protect the provenance information.

cs.CR

Stealthy Malware Traffic - Not as Innocent as It Looks

Malware is constantly evolving. Although existing countermeasures have success in malware detection, corresponding counter-countermeasures are always emerging. In this study, a counter-countermeasure that avoids network-based detection approaches by camouflaging malicious traffic as an innocuous protocol is presented. The approach includes two steps: Traffic format transformation and side-channel massage (SCM). Format transforming encryption (FTE) translates protocol syntax to mimic another innocuous protocol while SCM obscures traffic side-channels. The proposed approach is illustrated by transforming Zeus botnet (Zbot) Command and Control (C&C) traffic into smart grid Phasor Measurement Unit (PMU) data. The experimental results show that the transformed traffic is identified by Wireshark as synchrophasor protocol, and the transformed protocol fools current side-channel attacks. Moreover, it is shown that a real smart grid Phasor Data Concentrator (PDC) accepts the false PMU data.

cs.CR

A Covert Data Transport Protocol

Both enterprise and national firewalls filter network connections. For data forensics and botnet removal applications, it is important to establish the information source. In this paper, we describe a data transport layer which allows a client to transfer encrypted data that provides no discernible information regarding the data source. We use a domain generation algorithm (DGA) to encode AES encrypted data into domain names that current tools are unable to reliably differentiate from valid domain names. The domain names are registered using (free) dynamic DNS services. The data transmission format is not vulnerable to Deep Packet Inspection (DPI).

cs.CR

Identifying the Academic Rising Stars

Predicting the fast-rising young researchers (Academic Rising Stars) in the future provides useful guidance to the research community, e.g., offering competitive candidates to university for young faculty hiring as they are expected to have success academic careers. In this work, given a set of young researchers who have published the first first-author paper recently, we solve the problem of how to effectively predict the top k% researchers who achieve the highest citation increment in Δt years. We explore a series of factors that can drive an author to be fast-rising and design a novel impact increment ranking learning (IIRL) algorithm that leverages those factors to predict the academic rising stars. Experimental results on the large ArnetMiner dataset with over 1.7 million authors demonstrate the effectiveness of IIRL. Specifically, it outperforms all given benchmark methods, with over 8% average improvement. Further analysis demonstrates that the prediction models for different research topics follow the similar pattern. We also find that temporal features are the best indicators for rising stars prediction, while venue features are less relevant.

cs.DL

Pairing symmetry of heavy fermion superconductivity in the two-dimensional Kondo-Heisenberg lattice model

In the two-dimensional Kondo-Heisenberg lattice model away from half-filled, the local antiferromagnetic exchange coupling can provide the pairing mechanism of quasiparticles via the Kondo screening effect, leading to the heavy fermion superconductivity. We find that the pairing symmetry \textit{strongly} depends on the Fermi surface (FS) structure in the normal metallic state. When $J_{H}/J_{K}$ is very small, the FS is a small hole-like circle around the corner of the Brillouin zone, and the s-wave pairing symmetry has a lower ground state energy. For the intermediate coupling values of $J_{H}/J_{K}$, the extended s-wave pairing symmetry gives the favored ground state. However, when $J_{H}/J_{K}$ is larger than a critical value, the FS transforms into four small hole pockets crossing the boundary of the magnetic Brillouin zone, and the d-wave pairing symmetry becomes more favorable. In that regime, the resulting superconducting state is characterized by either nodal d-wave or nodeless d-wave state, depending on the conduction electron filling factor as well. A continuous phase transition exists between these two states. This result may be related to the phase transition of the nodal d-wave state to a fully gapped state, which is recently observed in Yb doped CeCoIn$_{5}$.

cond-mat.supr-con

Phase evolution of the two-dimensional Kondo lattice model near half-filling

Within a mean-field approximation, the ground state and finite temperature phase diagrams of the two-dimensional Kondo lattice model have been carefully studied as functions of the Kondo coupling $J$ and the conduction electron concentration $n_{c}$. In addition to the conventional hybridization between local moments and itinerant electrons, a staggered hybridization is proposed to characterize the interplay between the antiferromagnetism and the Kondo screening effect. As a result, a heavy fermion antiferromagnetic phase is obtained and separated from the pure antiferromagnetic ordered phase by a first-order Lifshitz phase transition, while a continuous phase transition exists between the heavy fermion antiferromagnetic phase and the Kondo paramagnetic phase. We have developed a efficient theory to calculate these phase boundaries. As $n_{c}$ decreases from the half-filling, the region of the heavy fermion antiferromagnetic phase shrinks and finally disappears at a critical point $n_{c}^{*}=0.8228$, leaving a first-order critical line between the pure antiferromagnetic phase and the Kondo paramagnetic phase for $n_{c}<n_{c}^{* }$. At half-filling limit, a finite temperature phase diagram is also determined on the Kondo coupling and temperature ($J$-$T$) plane. Notably, as the temperature is increased, the region of the heavy fermion antiferromagnetic phase is reduced continuously, and finally converges to a single point, together with the pure antiferromagnetic phase and the Kondo paramagnetic phase. The phase diagrams with such triple point may account for the observed phase transitions in related heavy fermion materials.

cond-mat.str-el

Efficient Channel-Hopping Rendezvous Algorithm Based on Available Channel Set

In cognitive radio networks, rendezvous is a fundamental operation by which two cognitive users establish a communication link on a commonly-available channel for communications. Some existing rendezvous algorithms can guarantee that rendezvous can be completed within finite time and they generate channel-hopping (CH) sequences based on the whole channel set. However, some channels may not be available (e.g., they are being used by the licensed users) and these existing algorithms would randomly replace the unavailable channels in the CH sequence. This random replacement is not effective, especially when the number of unavailable channels is large. In this paper, we design a new rendezvous algorithm that attempts rendezvous on the available channels only for faster rendezvous. This new algorithm, called Interleaved Sequences based on Available Channel set (ISAC), constructs an odd sub-sequence and an even sub-sequence and interleaves these two sub-sequences to compose a CH sequence. We prove that ISAC provides guaranteed rendezvous (i.e., rendezvous can be achieved within finite time). We derive the upper bound on the maximum time-to-rendezvous (MTTR) to be O(m) (m is not greater than Q) under the symmetric model and O(mn) (n is not greater than Q) under the asymmetric model, where m and n are the number of available channels of two users and Q is the total number of channels (i.e., all potentially available channels). We conduct extensive computer simulation to demonstrate that ISAC gives significantly smaller MTTR than the existing algorithms.

cs.NI

ZOS: A Fast Rendezvous Algorithm Based on Set of Available Channels for Cognitive Radios

Most of existing rendezvous algorithms generate channel-hopping sequences based on the whole channel set. They are inefficient when the set of available channels is a small subset of the whole channel set. We propose a new algorithm called ZOS which uses three types of elementary sequences (namely, Zero-type, One-type, and S-type) to generate channel-hopping sequences based on the set of available channels. ZOS provides guaranteed rendezvous without any additional requirements. The maximum time-to-rendezvous of ZOS is upper-bounded by O(m1*m2*log2M) where M is the number of all channels and m1 and m2 are the numbers of available channels of two users.

cs.NI

Multi-Linear Interactive Matrix Factorization

Recommender systems, which can significantly help users find their interested items from the information era, has attracted an increasing attention from both the scientific and application society. One of the widest applied recommendation methods is the Matrix Factorization (MF). However, most of MF based approaches focus on the user-item rating matrix, but ignoring the ingredients which may have significant influence on users' preferences on items. In this paper, we propose a multi-linear interactive MF algorithm (MLIMF) to model the interactions between the users and each event associated with their final decisions. Our model considers not only the user-item rating information but also the pairwise interactions based on some empirically supported factors. In addition, we compared the proposed model with three typical other methods: user-based collaborative filtering (UCF), item-based collaborative filtering (ICF) and regularized MF (RMF). Experimental results on two real-world datasets, \emph{MovieLens} 1M and \emph{MovieLens} 100k, show that our method performs much better than other three methods in the accuracy of recommendation. This work may shed some light on the in-depth understanding of modeling user online behaviors and the consequent decisions.

cs.IR

ILCR: Item-based Latent Factors for Sparse Collaborative Retrieval

Interactions between search and recommendation have recently attracted significant attention, and several studies have shown that many potential applications involve with a joint problem of producing recommendations to users with respect to a given query, termed $Collaborative$ $Retrieval$ (CR). Successful algorithms designed for CR should be potentially flexible at dealing with the sparsity challenges since the setup of collaborative retrieval associates with a given $query$ $\times$ $user$ $\times$ $item$ tensor instead of traditional $user$ $\times$ $item$ matrix. Recently, several works are proposed to study CR task from users' perspective. In this paper, we aim to sufficiently explore the sophisticated relationship of each $query$ $\times$ $user$ $\times$ $item$ triple from items' perspective. By integrating item-based collaborative information for this joint task, we present an alternative factorized model that could better evaluate the ranks of those items with sparse information for the given query-user pair. In addition, we suggest to employ a recently proposed scalable ranking learning algorithm, namely BPR, to optimize the state-of-the-art approach, $Latent$ $Collaborative$ $Retrieval$ model, instead of the original learning algorithm. The experimental results on two real-world datasets, (i.e. \emph{Last.fm}, \emph{Yelp}), demonstrate the efficiency and effectiveness of our proposed approach.

cs.IR

Information Filtering via Collaborative User Clustering Modeling

The past few years have witnessed the great success of recommender systems, which can significantly help users find out personalized items for them from the information era. One of the most widely applied recommendation methods is the Matrix Factorization (MF). However, most of researches on this topic have focused on mining the direct relationships between users and items. In this paper, we optimize the standard MF by integrating the user clustering regularization term. Our model considers not only the user-item rating information, but also takes into account the user interest. We compared the proposed model with three typical other methods: User-Mean (UM), Item-Mean (IM) and standard MF. Experimental results on a real-world dataset, \emph{MovieLens}, show that our method performs much better than other three methods in the accuracy of recommendation.

cs.IR

Weak ferromagnetism with the Kondo screening effect in the Kondo lattice systems

We carefully consider the interplay between ferromagnetism and the Kondo screening effect in the conventional Kondo lattice systems at finite temperatures. Within an effective mean-field theory for small conduction electron densities, a complete phase diagram has been determined. In the ferromagnetic ordered phase, there is a characteristic temperature scale to indicate the presence of the Kondo screening effect. We further find two distinct ferromagnetic long-range ordered phases coexisting with the Kondo screening effect: spin fully polarized and partially polarized states. A continuous phase transition exists to separate the partially polarized ferromagnetic ordered phase from the paramagnetic heavy Fermi liquid phase. These results may be used to explain the weak ferromagnetism observed recently in the Kondo lattice materials.

cond-mat.str-el

Accurate determination of the Gaussian transition in spin-1 chains with single-ion anisotropy

The Gaussian transition in the spin-one Heisenberg chain with single-ion anisotropy is extremely difficult to treat, both analytically and numerically. We introduce an improved DMRG procedure with strict error control, which we use to access very large systems. By considering the bulk entropy, we determine the Gaussian transition point to 4-digit accuracy, $D_{c}/J = 0.96845(8)$, resolving a long-standing debate in quantum magnetism. With this value, we obtain high-precision data for the critical behavior of quantities including the ground-state energy, gap, and transverse string-order parameter, and for the critical exponent, $ν= 1.472(2)$. Applying our improved technique at $J_{z} = 0.5$ highlights essential differences in critical behavior along the Gaussian transition line.

cond-mat.str-el

Lifshitz transitions in a heavy-Fermion liquid driven by short-range antiferromagnetic correlations in the two-dimensional Kondo lattice model

The heavy-Fermion liquid with short-range antiferromagnetic correlations is carefully considered in the two-dimensional Kondo-Heisenberg lattice model. As the ratio of the local Heisenberg superexchange $J_{H}$ to the Kondo coupling $J_{K}$ increases, Lifshitz transitions are anticipated, where the topology of the Fermi surface (FS) of the heavy quasiparticles changes from a hole-like circle to four kidney-like pockets centered around $(π,π)$. In-between these two limiting cases, a first-order quantum phase transition is identified at $J_{H}/J_{K}=0.1055$ where a small circle begins to emerge within the large deformed circle. When $J_{H}/J_{K}=0.1425$, the two deformed circles intersect each other and then decompose into four kidney-like Fermi pockets via a second-order quantum phase transition. As $J_{H}/J_{K}$ increases further, the Fermi pockets are shifted along the direction ($π,π$) to ($π/2,π/2$), and the resulting FS is consistent with the FS obtained recently using the quantum Monte Carlo cluster approach to the Kondo lattice system in the presence of the antiferrmagnetic order.

cond-mat.str-el