arXiv Science⌕ Search

arXiv subjects

Chen

Publications and source records attributed to Chen.

At least 91 records · Page 5Linked to original sources

RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation

Source code authorship attribution is an important problem often encountered in applications such as software forensics, bug fixing, and software quality analysis. Recent studies show that current source code authorship attribution methods can be compromised by attackers exploiting adversarial examples and coding style manipulation. This calls for robust solutions to the problem of code authorship attribution. In this paper, we initiate the study on making Deep Learning (DL)-based code authorship attribution robust. We propose an innovative framework called Robust coding style Patterns Generation (RoPGen), which essentially learns authors' unique coding style patterns that are hard for attackers to manipulate or imitate. The key idea is to combine data augmentation and gradient augmentation at the adversarial training phase. This effectively increases the diversity of training examples, generates meaningful perturbations to gradients of deep neural networks, and learns diversified representations of coding styles. We evaluate the effectiveness of RoPGen using four datasets of programs written in C, C++, and Java. Experimental results show that RoPGen can significantly improve the robustness of DL-based code authorship attribution, by respectively reducing 22.8% and 41.0% of the success rate of targeted and untargeted attacks on average.

cs.CR↗

An Unconditionally Stable Conformal LOD-FDTD Method For Curved PEC Objects and Its Application to EMC Problems

The traditional finite-difference time-domain (FDTD) method is constrained by the Courant-Friedrich-Levy (CFL) condition and suffers from the notorious staircase error in electromagnetic simulations. This paper proposes a three-dimensional conformal locally-one-dimensional FDTD (CLOD-FDTD) method to address the two issues for modeling perfectly electrical conducting (PEC) objects. By considering the partially filled cells, the proposed CLOD-FDTD method can significantly improve the accuracy compared with the traditional LOD-FDTD method and the FDTD method. At the same time, the proposed method preserves unconditional stability, which is analyzed and numerically validated using the Von-Neuman method. Significant gains in Central Processing Unit (CPU) time are achieved by using large time steps without sacrificing accuracy. Two numerical examples include a PEC cylinder and a missile are used to verify its accuracy and efficiency with different meshes and time steps. It can be found from these examples, the CLOD-FDTD method show better accuracy and can improve the efficiency compared with those of the traditional FDTD method and the traditional LOD-FDTD method.

cs.CE↗

A Stable FDTD Subgridding Scheme with SBP-SAT for Transient Electromagnetic Analysis

We proposed a provably stable FDTD subgridding method for accurate and efficient transient electromagnetic analysis. In the proposed method, several field components are properly added to the boundaries of Yee's grid to make sure that the discrete operators meet the summation-by-parts (SBP) property. Then, by incorporating the simultaneous approximation terms (SATs) into the finite-difference time-domain (FDTD) method, the proposed FDTD subgridding method mimics the energy estimate of the continuous Maxwell's equations at the semi-discrete level to guarantee its stability. Further, to couple multiple mesh blocks with different mesh sizes, the interpolation matrices are also derived. The proposed FDTD subgridding method is accurate, efficient, easy to implement and be integrated into the existing FDTD codes with only simple modifications. At last, three numerical examples with fine structures are carried out to validate the effectiveness of the proposed method.

cs.CE↗

A Hybrid SIE-PDE Formulation Without Boundary Condition Requirement for Transverse Magnetic Electromagnetic Analysis

A hybrid surface integral equation partial differential equation (SIE-PDE) formulation without the boundary condition requirement is proposed to solve the transverse magnetic (TM) electromagnetic problems. In the proposed formulation, the computational domain is decomposed into two overlapping domains: the SIE and PDE domains. In the SIE domain, complex structures with piecewise homogeneous media, e.g., highly conductive media, are included. An equivalent model for those structures is constructed by replacing them with the background medium and introducing a surface equivalent electric current density on an enclosed boundary to represent their electromagnetic effects. The remaining computational domain and homogeneous background medium replaced domain consist of the PDE domain, in which inhomogeneous or non-isotropic media are included. Through combining the surface equivalent electric current density and the inhomogeneous Helmholtz equation, a hybrid SIE-PDE formulation is derived. It requires no boundary conditions, and is mathematically equivalent to the original physical model. Through careful construction of basis functions to expand electric fields and the equivalent current density, the discretized formulation is made compatible with the SIE and PDE domain interface. The accuracy and efficiency are validated through two numerical examples. Results show that the proposed SIE-PDE formulation can obtain accurate results, and significant performance improvements in terms of CPU time and memory consumption compared with the FEM are achieved.

math.NA↗

FogROS: An Adaptive Framework for Automating Fog Robotics Deployment

As many robot automation applications increasingly rely on multi-core processing or deep-learning models, cloud computing is becoming an attractive and economically viable resource for systems that do not contain high computing power onboard. Despite its immense computing capacity, it is often underused by the robotics and automation community due to lack of expertise in cloud computing and cloud-based infrastructure. Fog Robotics balances computing and data between cloud edge devices. We propose a software framework, FogROS, as an extension of the Robot Operating System (ROS), the de-facto standard for creating robot automation applications and components. It allows researchers to deploy components of their software to the cloud with minimal effort, and correspondingly gain access to additional computing cores, GPUs, FPGAs, and TPUs, as well as predeployed software made available by other researchers. FogROS allows a researcher to specify which components of their software will be deployed to the cloud and to what type of computing hardware. We evaluate FogROS on 3 examples: (1) simultaneous localization and mapping (ORB-SLAM2), (2) Dexterity Network (Dex-Net) GPU-based grasp planning, and (3) multi-core motion planning using a 96-core cloud-based server. In all three examples, a component is deployed to the cloud and accelerated with a small change in system launch configuration, while incurring additional latency of 1.2 s, 0.6 s, and 0.5 s due to network communication, the computation speed is improved by 2.6x, 6.0x and 34.2x, respectively. Code, videos, and supplementary material can be found at https://github.com/BerkeleyAutomation/FogROS.

cs.RO↗

Vector Single-Source Surface Integral Equation for TE Scattering From Cylindrical Multilayered Objects

A single-source surface integral equation (SS-SIE) for transverse electric (TE) scattering from cylindrical multilayered objects is proposed in this paper. By incorporating the differential surface admittance operator (DSAO) and recursively applying the surface equivalence theorem from innermost to outermost boundaries, an equivalent model with only electric current density on the outermost boundary can be obtained. In addition, an integration approach is proposed, where the small argument expansion of the Hankel function is used to evaluate the singular and nearly singular integrals. Compared with other SIEs, such as the Poggio-Miller-Chang-Harrington-Wu-Tsai (PMCHWT) formulation, the computational expenditure is reduced for multilayered structures because only a single source is needed on the outermost boundary. As shown in the numerical results, the proposed method generates only 19% of unknowns, uses 26% of memory, and requires 29% of the CPU time of the PMCHWT formulation.

cs.CE↗

Single-Source SIE for Two-Dimensional Arbitrarily Connected Penetrable and PEC Objects with Nonconformal Meshes

We proposed a simple and efficient modular single-source surface integral equation (SS-SIE) formulation for electromagnetic analysis of arbitrarily connected penetrable and perfectly electrical conductor (PEC) objects in two-dimensional space. In this formulation, a modular equivalent model for each penetrable object consisting of the composite structure is first independently constructed through replacing it by the background medium, no matter whether it is surrounded by the background medium, other media, or partially connected objects, and enforcing an equivalent electric current density on the boundary to remain fields in the exterior region unchanged. Then, by combining all the modular models and any possible PEC objects together, an equivalent model for the composite structure can be derived. The troublesome junction handling techniques are not needed and non-conformal meshes are intrinsically supported. The proposed SS-SIE formulation is simple to implement, efficient, and flexible, which shows significant performance improvement in terms of CPU time compared with the original SS-SIE formulation and the Poggio-Miller-Chang-Harrington-Wu-Tsai (PMCHWT) formulation. Several numerical examples including the coated dielectric cuboid, the large lossy objects, the planar layered dielectric structure, and the partially connected dielectric and PEC structure are carried out to validate its accuracy, efficiency and robustness.

cs.CE↗

Studying Duplicate Logging Statements and Their Relationships with Code Clones

In this paper, we focus on studying duplicate logging statements, which are logging statements that have the same static text message. We manually studied over 4K duplicate logging statements and their surrounding code in five large-scale open source systems. We uncovered five patterns of duplicate logging code smells. For each instance of the duplicate logging code smell, we further manually identify the potentially problematic and justifiable cases. Then, we contact developers to verify our manual study result. We integrated our manual study result and the feedback of developers into our automated static analysis tool, DLFinder, which automatically detects problematic duplicate logging code smells. We evaluated DLFinder on the five manually studied systems and three additional systems. In total, combining the results of DLFinder and our manual analysis, we reported 91 problematic duplicate logging code smell instances to developers and all of them have been fixed. We further study the relationship between duplicate logging statements, including the problematic instances of duplicate logging code smells, and code clones. We find that 83% of the duplicate logging code smell instances reside in cloned code, but 17% of them reside in micro-clones that are difficult to detect using automated clone detection tools. We also find that more than half of the duplicate logging statements reside in cloned code snippets, and a large portion of them reside in very short code blocks which may not be effectively detected by existing code clone detection tools. Our study shows that, in addition to general source code that implements the business logic, code clones may also result in bad logging practices that could increase maintenance difficulties.

cs.SE↗

Establishing Secrecy Region for Directional Modulation Scheme with Random Frequency Diverse Array

Random frequency diverse array (RFDA) based directional modulation (DM) was proposed as a promising technology in secure communications to achieve a precise transmission of confidential messages, and artificial noise (AN) was considered as an important helper in RFDA-DM. Compared with previous works that only focus on the spot of the desired receiver, in this work, we investigate a secrecy region around the desired receiver, that is, a specific range and angle resolution around the desired receiver. Firstly, the minimum number of antennas and the bandwidth needed to achieve a secrecy region are derived. Moreover, based on the lower bound of the secrecy capacity in RFDA-DM-AN scheme, we investigate the performance impact of AN on the secrecy capacity. From this work, we conclude that: 1) AN is not always beneficial to the secure transmission. Specifically, when the number of antennas is sufficiently large and the transmit power is smaller than a specified value, AN will reduce secrecy capacity due to the consumption of limited transmit power. 2) Increasing bandwidth will enlarge the set for randomly allocating frequencies and thus lead to a higher secrecy capacity. 3) The minimum number of antennas increases as the predefined secrecy transmission rate increases.

cs.IT↗

Future Physics Programme of BESIII

There has recently been a dramatic renewal of interest in the subjects of hadron spectroscopy and charm physics. This renaissance has been driven in part by the discovery of a plethora of charmonium-like $XYZ$ states at BESIII and $B$ factories, and the observation of an intriguing proton-antiproton threshold enhancement and the possibly related $X(1835)$ meson state at BESIII, as well as the threshold measurements of charm mesons and charm baryons. We present a detailed survey of the important topics in tau-charm physics and hadron physics that can be further explored at BESIII over the remaining lifetime of BEPCII operation. This survey will help in the optimization of the data-taking plan over the coming years, and provides physics motivation for the possible upgrade of BEPCII to higher luminosity.

hep-ex↗

Measuring similarity between two mixture trees using mixture distance metric and algorithms

Ancestral mixture model, proposed by Chen and Lindsay (2006), is an important model to build a hierarchical tree from high dimensional binary sequences. Mixture trees created from ancestral mixture models involve in the inferred evolutionary relationships among various biological species. Moreover, it contains the information of time when the species mutates. Tree comparison metric, an essential issue in bioinformatics, is to measure the similarity between trees. However, to our knowledge, the approach to the comparison between two mixture trees is still under development. In this paper, we propose a new metric, named mixture distance metric, to measure the similarity of two mixture trees. It uniquely considers the factor of evolutionary times between trees. In addition, we also further develop two algorithms to compute the mixture distance between two mixture trees. One requires O(n^2) and the other requires O(nh) computation time with O(n) preprocessing time, where n denotes the number of leaves in the two mixture trees, and h denotes the minimum height of these two trees.

cs.DS↗

An Empirical Study of Obsolete Answers on Stack Overflow

Stack Overflow accumulates an enormous amount of software engineering knowledge. However, as time passes, certain knowledge in answers may become obsolete. Such obsolete answers, if not identified or documented clearly, may mislead answer seekers and cause unexpected problems (e.g., using an out-dated security protocol). In this paper, we investigate how the knowledge in answers becomes obsolete and identify the characteristics of such obsolete answers. We find that: 1) More than half of the obsolete answers (58.4%) were probably already obsolete when they were first posted. 2) When an obsolete answer is observed, only a small proportion (20.5%) of such answers are ever updated. 3) Answers to questions in certain tags (e.g., node.js, ajax, android, and objective-c) are more likely to become obsolete. Our findings suggest that Stack Overflow should develop mechanisms to encourage the whole community to maintain answers (to avoid obsolete answers) and answer seekers are encouraged to carefully go through all information (e.g., comments) in answer threads.

cs.SE↗

Whole-Slide Image Focus Quality: Automatic Assessment and Impact on AI Cancer Detection

Digital pathology enables remote access or consults and powerful image analysis algorithms. However, the slide digitization process can create artifacts such as out-of-focus (OOF). OOF is often only detected upon careful review, potentially causing rescanning and workflow delays. Although scan-time operator screening for whole-slide OOF is feasible, manual screening for OOF affecting only parts of a slide is impractical. We developed a convolutional neural network (ConvFocus) to exhaustively localize and quantify the severity of OOF regions on digitized slides. ConvFocus was developed using our refined semi-synthetic OOF data generation process, and evaluated using real whole-slide images spanning 3 different tissue types and 3 different stain types that were digitized by two different scanners. ConvFocus's predictions were compared with pathologist-annotated focus quality grades across 514 distinct regions representing 37,700 35x35 $μ$m image patches, and 21 digitized "z-stack" whole-slide images that contain known OOF patterns. When compared to pathologist-graded focus quality, ConvFocus achieved Spearman rank coefficients of 0.81 and 0.94 on two scanners, and reproduced the expected OOF patterns from z-stack scanning. We also evaluated the impact of OOF on the accuracy of a state-of-the-art metastatic breast cancer detector and saw a consistent decrease in performance with increasing OOF. Comprehensive whole-slide OOF categorization could enable rescans prior to pathologist review, potentially reducing the impact of digitization focus issues on the clinical workflow. We show that the algorithm trained on our semi-synthetic OOF data generalizes well to real OOF regions across tissue types, stains, and scanners. Finally, quantitative OOF maps can flag regions that might otherwise be misclassified by image analysis algorithms, preventing OOF-induced errors.

cs.CV↗

Nucleon resonance production in the $γp \to pηϕ$ reaction

In this work, we perform a study of nucleon resonance production in the $γp \to pηϕ$ reaction within an effective Lagrangian approach. In our model, we consider the excitation of the $N^*(1535)$, $N^*(1650)$, $N^*(1710)$ and $N^*(1720)$ in the intermediate state and the background term. We find that this reaction is dominated by the excitation of the $N^*(1535)$ in the near threshold region. Especially, we study the possible role of the scalar meson exchange in this reaction. It is found that the $f_0(980)$ exchange may give a significant contribution and the parity asymmetry can be used to identify its role in this reaction.

hep-ph↗

Development and Validation of a Deep Learning Algorithm for Improving Gleason Scoring of Prostate Cancer

For prostate cancer patients, the Gleason score is one of the most important prognostic factors, potentially determining treatment independent of the stage. However, Gleason scoring is based on subjective microscopic examination of tumor morphology and suffers from poor reproducibility. Here we present a deep learning system (DLS) for Gleason scoring whole-slide images of prostatectomies. Our system was developed using 112 million pathologist-annotated image patches from 1,226 slides, and evaluated on an independent validation dataset of 331 slides, where the reference standard was established by genitourinary specialist pathologists. On the validation dataset, the mean accuracy among 29 general pathologists was 0.61. The DLS achieved a significantly higher diagnostic accuracy of 0.70 (p=0.002) and trended towards better patient risk stratification in correlations to clinical follow-up data. Our approach could improve the accuracy of Gleason scoring and subsequent therapy decisions, particularly where specialist expertise is unavailable. The DLS also goes beyond the current Gleason system to more finely characterize and quantitate tumor morphology, providing opportunities for refinement of the Gleason system itself.

cs.CV↗

Spatial-Temporal Inference of Urban Traffic Emissions Based on Taxi Trajectories and Multi-Source Urban Data

Vehicle trajectory data collected via GPS-enabled devices have played increasingly important roles in estimating network-wide traffic, given their broad spatial-temporal coverage and representativeness of traffic dynamics. This paper exploits taxi GPS data, license plate recognition (LPR) data and geographical information for reconstructing the spatial and temporal patterns of urban traffic emissions. Vehicle emission factor models are employed to estimate emissions based on taxi trajectories. The estimated emissions are then mapped to spatial grids of urban areas to account for spatial heterogeneity. To extrapolate emissions from the taxi fleet to the whole vehicle population, we use Gaussian process regression models supported by geographical features to estimate the spatially heterogeneous traffic volume and fleet composition. Unlike previous studies, this paper utilizes the taxi GPS data and LPR data to disaggregate vehicle and emission characteristics through space and time in a large-scale urban network. The results of a case study in Hangzhou, China, reveal high-resolution spatiotemporal patterns of traffic flows and emissions and identify emission hotspots at different locations. This study provides an accessible means of inferring the environmental impact of urban traffic with multi-source data that are now widely available in urban areas.

physics.soc-ph↗

A Survey of Intrusion Detection Systems Leveraging Host Data

This survey focuses on intrusion detection systems (IDS) that leverage host-based data sources for detecting attacks on enterprise network. The host-based IDS (HIDS) literature is organized by the input data source, presenting targeted sub-surveys of HIDS research leveraging system logs, audit data, Windows Registry, file systems, and program analysis. While system calls are generally included in audit data, several publicly available system call datasets have spawned a flurry of IDS research on this topic, which merits a separate section. Similarly, a section surveying algorithmic developments that are applicable to HIDS but tested on network data sets is included, as this is a large and growing area of applicable literature. To accommodate current researchers, a supplementary section giving descriptions of publicly available datasets is included, outlining their characteristics and shortcomings when used for IDS evaluation. Related surveys are organized and described. All sections are accompanied by tables concisely organizing the literature and datasets discussed. Finally, challenges, trends, and broader observations are throughout the survey and in the conclusion along with future directions of IDS research.

cs.CR↗

Short-Term Forecasting of Passenger Demand under On-Demand Ride Services: A Spatio-Temporal Deep Learning Approach

Short-term passenger demand forecasting is of great importance to the on-demand ride service platform, which can incentivize vacant cars moving from over-supply regions to over-demand regions. The spatial dependences, temporal dependences, and exogenous dependences need to be considered simultaneously, however, which makes short-term passenger demand forecasting challenging. We propose a novel deep learning (DL) approach, named the fusion convolutional long short-term memory network (FCL-Net), to address these three dependences within one end-to-end learning architecture. The model is stacked and fused by multiple convolutional long short-term memory (LSTM) layers, standard LSTM layers, and convolutional layers. The fusion of convolutional techniques and the LSTM network enables the proposed DL approach to better capture the spatio-temporal characteristics and correlations of explanatory variables. A tailored spatially aggregated random forest is employed to rank the importance of the explanatory variables. The ranking is then used for feature selection. The proposed DL approach is applied to the short-term forecasting of passenger demand under an on-demand ride service platform in Hangzhou, China. Experimental results, validated on real-world data provided by DiDi Chuxing, show that the FCL-Net achieves better predictive performance than traditional approaches including both classical time-series prediction models and neural network based algorithms (e.g., artificial neural network and LSTM). This paper is one of the first DL studies to forecast the short-term passenger demand of an on-demand ride service platform by examining the spatio-temporal correlations.

cs.LG↗