arXiv ScienceSearch

arXiv subjects

Borching Su

Publications and source records attributed to Borching Su.

18 recordsLinked to original sources

Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework

This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying the listener's intended speaker in multi-talker environments. GG-AVSE addresses this issue by exploiting gaze direction as a supervisory cue for target-speaker selection. Specifically, we propose the GG-VM module, which combines gaze signals with a YOLO5Face detector to extract the target speaker's facial features and integrates them with the pretrained AVSEMamba model through two strategies: zero-shot merging and partial visual fine-tuning. For evaluation, we introduce the AVSEC2-Gaze dataset. Experimental results show that GG-AVSE achieves substantial performance gains over gaze-free baselines: a 10.08% improvement in PESQ (2.370 to 2.609), a 5.18% improvement in STOI (0.8802 to 0.9258), and a 23.69% improvement in SI-SDR (9.16 to 11.33). These results confirm that gaze provides an effective cue for resolving target-speaker ambiguity and highlight the scalability of GG-AVSE for real-world applications.

eess.AS

Unimodular Waveform Design that Minimizes PSL of Ambiguity Function over A Continuous Doppler Frequency Shift Region of Interest

In active sensing systems, waveforms with ambiguity functions (AFs) of low peak sidelobe levels (PSLs) across a time delay and Doppler frequency shift plane (delay-Doppler plane) of interest are desirable for reducing false alarms. Additionally, unimodular waveforms are preferred due to hardware limitations. In this paper, a new method is proposed to design unimodular waveforms with PSL suppression over a continuous Doppler frequency shift region, based on the discrete-time ambiguity function (DTAF). Compared with existing methods that suppress PSL over grid points in the delay-Doppler plane by using the discrete ambiguity function (DAF), we regard the DTAF optimization problem as of more practical interest because the Doppler frequency shifts observed in echo signals reflected from targets are inherently continuous rather than discrete. The problem of interest is formulated as an optimization problem with infinite constraints along with unimodular constraints. To the best of the authors' knowledge, such a problem has not been studied yet. We propose to reformulate a non-convex semi-infinite programming (SIP) to a semidefinite programming (SDP) with a finite number of constraints and a rank-one constraint, which is then solved by the sequential rank-one constraint relaxation (SROCR) algorithm. Simulation results demonstrate that the proposed method outperforms existing methods in achieving a lower PSL of AF over a continuous Doppler frequency shift region of interest. Moreover, the designed waveform can effectively prevent false alarms when detecting a target with an arbitrary velocity.

eess.SP

Broadened-beam Uniform Rectangular Array Coefficient Design in LEO SatComs Under Quality of Service and Constant Modulus Constraints

Satellite communications (SatComs) are anticipated to deliver global Internet access. Low Earth orbit (LEO) satellites (SATs) offer the advantage of higher downlink capacity due to their reduced link budget compared to medium Earth orbit (MEO) and geostationary Earth orbit (GEO) SATs. In this paper, beam broadening methods for uniform rectangular arrays (URAs) in LEO SatComs were studied. The proposed method is the first of its kind to jointly consider path loss variation from SAT to the user terminal (UT) due to the Earth's curvature to guarantee the quality of service (QoS), constant modulus constraints (CMCs) favored for maximizing power amplifier (PA) efficiency, and out-of-beam radiation suppression to avoid interference. A broadened-beam URA coefficient design problem is formulated and decomposed into two uniform linear array (ULA) design subproblems utilizing Kronecker product beamforming. With this decomposition, the number of beamforming coefficients that need to be optimized is significantly reduced compared to the original URA design problem. The non-convex ULA subproblems are addressed using the semidefinite relaxation (SDR) technique and a convex iterative algorithm. Simulation results reveal the advantages of the proposed method for suppressing the out-of-beam radiation and achieving the design criteria. In addition, channel capacity evaluations are carried out. It demonstrates that the proposed "broadened-beam" beamformers can offer capacities that are at least four times greater than those of beamformers employing an array steering vector when the beam transition time is considered. The proposed method holds potential for LEO SAT broadcasting applications, such as digital video broadcasting (DVB).

eess.SP

IANS: Intelligibility-aware Null-steering Beamforming for Dual-Microphone Arrays

Beamforming techniques are popular in speech-related applications due to their effective spatial filtering capabilities. Nonetheless, conventional beamforming techniques generally depend heavily on either the target's direction-of-arrival (DOA), relative transfer function (RTF) or covariance matrix. This paper presents a new approach, the intelligibility-aware null-steering (IANS) beamforming framework, which uses the STOI-Net intelligibility prediction model to improve speech intelligibility without prior knowledge of the speech signal parameters mentioned earlier. The IANS framework combines a null-steering beamformer (NSBF) to generate a set of beamformed outputs, and STOI-Net, to determine the optimal result. Experimental results indicate that IANS can produce intelligibility-enhanced signals using a small dual-microphone array. The results are comparable to those obtained by null-steering beamformers with given knowledge of DOAs.

eess.AS

Waveform Design for Optimal PSL Under Spectral and Unimodular Constraints via Alternating Minimization

In an active sensing system, waveforms with good auto-correlations are preferred for accurate parameter estimation. Furthermore, spectral compatibility is required to avoid mutual interference between devices as the electromagnetic environment becomes increasingly crowded. Waveforms should also be unimodular due to hardware limits. In this paper, a new approach to generating a unimodular sequence with an approximately optimal peak side-lobe level (PSL) in auto-correlation and adjustable stopband attenuation is proposed. The proposed method is based on alternating minimization (AM) and numerical results suggest that it outperforms existing methods in terms of PSL. We also develop a theoretical lower bound for the PSL minimization problem under spectral constraints and unimodular constraints, which can be used for the evaluation of the results in various works about this waveform design problem. It is observed in the numerical results that the PSL of the proposed algorithm is close to the derived lower bound.

eess.SP

Speech Enhancement Based on Cyclegan with Noise-informed Training

Cycle-consistent generative adversarial networks (CycleGAN) were successfully applied to speech enhancement (SE) tasks with unpaired noisy-clean training data. The CycleGAN SE system adopted two generators and two discriminators trained with losses from noisy-to-clean and clean-to-noisy conversions. CycleGAN showed promising results for numerous SE tasks. Herein, we investigate a potential limitation of the clean-to-noisy conversion part and propose a novel noise-informed training (NIT) approach to improve the performance of the original CycleGAN SE system. The main idea of the NIT approach is to incorporate target domain information for clean-to-noisy conversion to facilitate a better training procedure. The experimental results confirmed that the proposed NIT approach improved the generalization capability of the original CycleGAN SE system with a notable margin.

eess.AS

Frequency Reversal Alamouti Code-Based FBMC with Resilience to Inter-Antenna Frequency Offsets

Transmit diversity schemes for filter bank multicarrier (FBMC) are known to be challenging. No existing schemes have considered the presence of inter-antenna frequency offset (IAFO), which will result in performance degradation. In this letter, a new transmit scheme based on the frequency reversal Alamouti code (FRAC)-based structure to address the issue of IAFO is proposed and is proven to inherently cancel the inter-antenna inter-carrier interference (ICI) while preserving spatial diversity. Moreover, the proposed FRAC structure is applicable in frequency-selective channels. Numerical results show that the proposed scheme undergoes negligible bit error rate (BER) degradation even with considerable IAFOs.

cs.IT

Downlink SCMA Codebook Design with Low Error Rate by Maximizing Minimum Euclidean Distance of Superimposed Codewords

Sparse code multiple access (SCMA), as a codebook-based non-orthogonal multiple access (NOMA) technique, has received research attention in recent years. The codebook design problem for SCMA has also been studied to some extent since codebook choices are highly related to the system's error rate performance. In this paper, we approach the SCMA codebook design problem by formulating an optimization problem to maximize the minimum Euclidean distance (MED) of superimposed codewords under power constraints. While SCMA codebooks with a larger MED are expected to obtain a better BER performance, no optimal SCMA codebook in terms of MED maximization, to the authors' best knowledge, has been reported in the SCMA literature yet. In this paper, a new iterative algorithm based on alternating maximization with exact penalty is proposed for the MED maximization problem. The proposed algorithm, when supplied with appropriate initial points and parameters, achieves a set of codebooks of all users whose MED is larger than any previously reported results. A Lagrange dual problem is derived which provides an upper bound of MED of any set of codebooks. Even though there is still a nonzero gap between the achieved MED and the upper bound given by the dual problem, simulation results demonstrate clear advantages in error rate performances of the proposed set of codebooks over all existing ones not only in AWGN channels but also in some downlink scenarios that fit in 5G/NR applications, making it a good codebook candidate thereof. The proposed set of SCMA codebooks, however, are not shown to outperform existing ones in uplink channels or in the case where non-consecutive OFDMA subcarriers are used. The correctness and accuracy of error curves in the simulation results are further confirmed by the coincidences with the theoretical upper bounds of error rates derived for any given set of codebooks.

cs.IT

MIMO Speech Compression and Enhancement Based on Convolutional Denoising Autoencoder

For speech-related applications in IoT environments, identifying effective methods to handle interference noises and compress the amount of data in transmissions is essential to achieve high-quality services. In this study, we propose a novel multi-input multi-output speech compression and enhancement (MIMO-SCE) system based on a convolutional denoising autoencoder (CDAE) model to simultaneously improve speech quality and reduce the dimensions of transmission data. Compared with conventional single-channel and multi-input single-output systems, MIMO systems can be employed in applications that handle multiple acoustic signals need to be handled. We investigated two CDAE models, a fully convolutional network (FCN) and a Sinc FCN, as the core models in MIMO systems. The experimental results confirm that the proposed MIMO-SCE framework effectively improves speech quality and intelligibility while reducing the amount of recording data by a factor of 7 for transmission.

eess.AS

Interference-Precancelled Pilot Design for LMMSE Channel Estimation of GFDM

Generalized frequency division multiplexing (GFDM) is a promising candidate waveform for next-generation wireless communication systems. However, GFDM channel estimation is still challenging due to the inherent interference. In this paper, we formulate a pilot design framework with linear minimum mean square error (LMMSE) channel estimation for GFDM, and propose a novel pilot design to achieve interference precancellation during pilot generation with the fixed transmit sample values at selected frequency bins. Numerical results demonstrate that the proposed method reduces the channel estimation mean square error and the symbol error rate (SER) in high signal-to-noise ratio (SNR) regions, compared with the conventional methods.

cs.IT

Reducing Cubic Metric of Circularly Pulse-Shaped OFDM Signals Through Constellation Shaping Optimization With Performance Constraints

Circularly pulse-shaped orthogonal frequency division multiplexing (CPS-OFDM) is one of the most promising 5G waveforms that addresses two physical layer signal requirements of low out-of-subband emission (OSBE) and low peak-to-average power ratio (PAPR) with flexibility in parameter adaptation. In this paper, a constellation shaping optimization method is proposed to further reduce the cubic metric (CM) of CPS-OFDM signals for the case that demands rather high power amplifier (PA) efficiency at the transmitter. Simulation results demonstrate the superiority of the proposed scheme in CM reduction, and the corresponding benefits of spectral regrowth mitigation and spectral efficiency improvement.

cs.IT

Circularly Pulse-Shaped Precoding for OFDM: A New Waveform and Its Optimization Design for 5G New Radio

A new circularly pulse-shaped (CPS) precoding orthogonal frequency division multiplexing (OFDM) waveform, or CPS-OFDM for short, is proposed in this paper. CPS-OFDM, characterized by user-specific precoder flexibility, possesses the advantages of both low out-of-subband emission (OSBE) and low peak-to-average power ratio (PAPR), which are two major desired physical layer signal properties for various scenarios in 5G New Radio (NR), including fragmented spectrum access, new types of user equipments (UEs), and communications at high carrier frequencies. As opposed to most of existing waveform candidates using windowing or filtering techniques, CPS-OFDM prevents block extension that causes extra inter-block interference (IBI) and envelope fluctuation unfriendly to signal reception and power amplifier (PA) efficiency, respectively. An optimization problem of the prototype shaping vector built in the CPS precoder is formulated to minimize the variance of instantaneous power (VIP) with controllable OSBE power (OSBEP) and noise enhancement penalty (NEP). In order to solve the optimization problem involving a quartic objective function, the majorization-minimization (MM) algorithmic framework is exploited. By proving the convexity of the proposed problem, the globally optimal solution invariant of coming data is guaranteed to be attained via numbers of iterations. Simulation results demonstrate the advantages of the proposed scheme in terms of detection reliability and spectral efficiency for practical 5G cases such as asynchronous transmissions and mixed numerologies.

cs.IT

Frequency-Domain Decoupling for MIMO-GFDM Spatial Multiplexing

Generalized frequency division multiplexing (GFDM) is considered a non-orthogonal waveform and known to encounter difficulties when using in the spatial multiplexing mode of multiple-input-multiple-output (MIMO) scenario. In this paper, a class of GFDM prototype filters, under which the GFDM system is free from inter-subcarrier interference, is investigated, enabling frequency-domain decoupling in the processing at the GFDM receiver. An efficient MIMO-GFDM detection method based on depth-first sphere decoding is then proposed with such class of filters. Numerical results confirm a significant reduction in complexity, especially when the number of subcarriers is large, compared with existing methods presented in recent years.

cs.IT

Matrix Characterization for GFDM Systems: Low-Complexity MMSE Receivers and Optimal Prototype Filters

In this paper, a new matrix-based characterization of generalized-frequency-division-multiplexing (GFDM) transmitter matrices is proposed, as opposed to traditional vector-based characterization with prototype filters. The characterization facilitates deriving properties of GFDM (transmitter) matrices, including conditions for GFDM matrices being nonsingular and unitary, respectively. Using the new characterization, the necessary and sufficient conditions for the existence of a form of low-complexity implementation for a minimum mean square error (MMSE) receiver are derived. Such an implementation exists under multipath channels if the GFDM transmitter matrix is selected to be unitary. For cases where this implementation does not exist, a low-complexity suboptimal MMSE receiver is proposed, with its performance approximating that of an MMSE receiver. The new characterization also enables derivations of optimal prototype filters in terms of minimizing receiver mean square error (MSE). They are found to correspond to the use of unitary GFDM matrices under many scenarios. The use of such optimal filters in GFDM systems does not cause the problem of noise enhancement, thereby demonstrating the same MSE performance as orthogonal frequency division multiplexing. Moreover, we find that GFDM matrices with a size of power of two are verified to exist in the class of unitary GFDM matrices. Finally, while the out-of-band (OOB) radiation performance of systems using a unitary GFDM matrix is not optimal in general, it is shown that the OOB radiation can be satisfactorily low if parameters in the new characterization are carefully chosen.

cs.IT

Robust Beamforming Against Direction-of-Arrival Mismatch Using Subspace-Constrained Diagonal Loading

In this study, a new subspace-constrained diagonal loading (SSC-DL) method is presented for robust beamforming against the issue of a mismatched direction of arrival (DoA), based on an extension to the well known diagonal loading (DL) technique. One important difference of the proposed SSC-DL from conventional DL is that it imposes an additional constraint to restrict the optimal weight vector within a subspace whose basis vectors are determined by a number of angles neighboring to the estimated DoA. Unlike many existing methods which resort to a beamwidth expansion, the weight vector produced by SSC-DL has a relatively small beamwidth around the DoA of the target signal. Yet, the SSC-DL beamformer has a great interference suppression level, thereby achieving an improved overall SINR performance. Simulation results suggest the proposed method has a near-to-optimal SINR performance.

math.OC

Wavelet speech enhancement based on nonnegative matrix factorization

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a long time span. However, the short-time Fourier transform (STFT) that creates the spectrogram in general distorts the original signal and thereby limits the capability of the associated speech enhancement techniques. In this study, we propose a novel speech enhancement method that adopts the algorithms of discrete wavelet packet transform (DWPT) and nonnegative matrix factorization (NMF) in order to conquer the aforementioned limitation. In brief, the DWPT is first applied to split a time-domain speech signal into a series of subband signals without introducing any distortion. Then we exploit NMF to highlight the speech component for each subband. Finally, the enhanced subband signals are joined together via the inverse DWPT to reconstruct a noise-reduced signal in time domain. We evaluate the proposed DWPT-NMF based speech enhancement method on the MHINT task. Experimental results show that this new method behaves very well in prompting speech quality and intelligibility and it outperforms the convnenitional STFT-NMF based method.

cs.SD

A Cramer-Rao Bound for Semi-Blind Channel Estimation in Redundant Block Transmission Systems

A Cramer-Rao bound (CRB) for semi-blind channel estimators in redundant block transmission systems is derived. The derived CRB is valid for any system adopting a full-rank linear redundant precoder, including the popular cyclic-prefixed orthogonal frequency-division multiplexing system. Simple forms of CRBs for multiple complex parameters, either unconstrained or constrained by a holomorphic function, are also derived, which facilitate the CRB derivation of the problem of interest. The derived CRB is a lower bound on the variance of any unbiased semi-blind channel estimator, and can serve as a tractable performance metric for system design.

cs.IT

Cramer-Rao Bound for Blind Channel Estimators in Redundant Block Transmission Systems

In this paper, we derive the Cramer-Rao bound (CRB) for blind channel estimation in redundant block transmission systems, a lower bound for the mean squared error of any blind channel estimators. The derived CRB is valid for any full-rank linear redundant precoder, including both zero-padded (ZP) and cyclic-prefixed (CP) precoders. A simple form of CRBs for multiple complex parameters is also derived and presented which facilitates the CRB derivation of the problem of interest. A comparison is made between the derived CRBs and performances of existing subspace-based blind channel estimators for both CP and ZP systems. Numerical results show that there is still some room for performance improvement of blind channel estimators.

cs.IT