Source author record

Xiaojun Yuan

Xiaojun Yuan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

50works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

50 published item(s)

preprint2025arXiv

Continuous Angular Power Spectrum Recovery From Channel Covariance via Chebyshev Polynomials

This paper proposes a Chebyshev polynomial expansion framework for the recovery of a continuous angular power spectrum (APS) from channel covariance. By exploiting the orthogonality of Chebyshev polynomials in a transformed domain, we derive an exact series representation of the covariance and reformulate the inherently ill-posed APS inversion as a finite-dimensional linear regression problem via truncation. The associated approximation error is directly controlled by the tail of the APS's Chebyshev series and decays rapidly with increasing angular smoothness. Building on this representation, we derive an exact semidefinite characterization of nonnegative APS and introduce a derivative-based regularizer that promotes smoothly varying APS profiles while preserving transitions of clusters. Simulation results show that the proposed Chebyshev-based framework yields accurate APS reconstruction, and enables reliable downlink (DL) covariance prediction from uplink (UL) measurements in a frequency division duplex (FDD) setting. These findings indicate that jointly exploiting smoothness and nonnegativity in a Chebyshev domain provides an effective tool for covariance-domain processing in multi-antenna systems.

preprint2024arXiv

Hybrid Vector Message Passing for Generalized Bilinear Factorization

In this paper, we propose a new message passing algorithm that utilizes hybrid vector message passing (HVMP) to solve the generalized bilinear factorization (GBF) problem. The proposed GBF-HVMP algorithm integrates expectation propagation (EP) and variational message passing (VMP) via variational free energy minimization, yielding tractable Gaussian messages. Furthermore, GBF-HVMP enables vector/matrix variables rather than scalar ones in message passing, resulting in a loop-free Bayesian network that improves convergence. Numerical results show that GBF-HVMP significantly outperforms state-of-the-art methods in terms of NMSE performance and computational complexity.

preprint2022arXiv

Closed-Loop Data Transcription to an LDR via Minimaxing Rate Reduction

This work proposes a new computational framework for learning a structured generative model for real-world datasets. In particular, we propose to learn a closed-loop transcription between a multi-class multi-dimensional data distribution and a linear discriminative representation (LDR) in the feature space that consists of multiple independent multi-dimensional linear subspaces. In particular, we argue that the optimal encoding and decoding mappings sought can be formulated as the equilibrium point of a two-player minimax game between the encoder and decoder. A natural utility function for this game is the so-called rate reduction, a simple information-theoretic measure for distances between mixtures of subspace-like Gaussians in the feature space. Our formulation draws inspiration from closed-loop error feedback from control systems and avoids expensive evaluating and minimizing approximated distances between arbitrary distributions in either the data space or the feature space. To a large extent, this new formulation unifies the concepts and benefits of Auto-Encoding and GAN and naturally extends them to the settings of learning a both discriminative and generative representation for multi-class and multi-dimensional real-world data. Our extensive experiments on many benchmark imagery datasets demonstrate tremendous potential of this new closed-loop formulation: under fair comparison, visual quality of the learned decoder and classification performance of the encoder is competitive and often better than existing methods based on GAN, VAE, or a combination of both. Unlike existing generative models, the so learned features of the multiple classes are structured: different classes are explicitly mapped onto corresponding independent principal subspaces in the feature space. Source code can be found at https://github.com/Delay-Xili/LDR.

preprint2022arXiv

Elevation Angle-Dependent 3D Trajectory Design for Aerial RIS-aided Communication

This paper investigates an aerial reconfigurable intelligent surface (RIS)-aided communication system under the probabilistic line-of-sight (LoS) channel, where an unmanned aerial vehicle (UAV) equipped with an RIS is deployed to assist two ground nodes in their information exchange. An optimization problem with the objective of maximizing the minimum average achievable rate is formulated to jointly design the communication scheduling, the RIS's phase shift, and the three-dimensional (3D) UAV trajectory. To solve such a non-convex problem, we propose an efficient iterative algorithm to obtain its suboptimal solution. Simulation results show that our proposed design significantly outperforms the existing schemes and provides new insights into the elevation angle and distance trade-off for the UAV-borne RIS communication system.

preprint2022arXiv

Energy-Efficient UAV Communications: A Generalised Propulsion Energy Consumption Model

This paper proposes a generalised propulsion energy consumption model (PECM) for rotary-wing ummanned aerial vehicles (UAVs) under the consideration of the practical thrust-to-weight ratio (TWR) with respect to the velocity, acceleration and direction change of the UAVs. To verify the effectiveness of the proposed PECM, we consider a UAV-enabled communication system, where a rotary-wing UAV serves multiple ground users as an aerial base station. We aim to maximize the energy efficiency (EE) of the UAV by jointly optimizing the user scheduling and UAV trajectory variables. However, the formulated problem is a non-convex fractional integer programming problem, which is challenging to obtain its optimal solution. To tackle this, we propose an efficient iterative algorithm by decomposing the original problem into two sub-problems to obtain a suboptimal solution based on the successive convex approximation technique. Simulation results show that the optimized UAV trajectory by applying the proposed PECM are smoother and the corresponding EE has significant improvement as compared to other benchmark schemes.

preprint2022arXiv

Energy-Efficient UAV-Mounted RIS Assisted Mobile Edge Computing

Unmanned aerial vehicle (UAV) and reconfigurable intelligent surface (RIS) have been recently applied in the field of mobile edge computing (MEC) to improve the data exchange environment by proactively changing the wireless channels through maneuverable location deployment and intelligent signals reflection, respectively. Nevertheless, they may suffer from inherent limitations in practical scenarios. UAV-mounted RIS (U-RIS), as a promising integrated approach, can combine the advantages of UAV and RIS to break the limit. Inspired by this, we consider a novel U-RIS assisted MEC system, where a U-RIS is deployed to assist the communication between the ground users and an MEC server. The joint UAV trajectory, RIS passive beamforming and MEC resource allocation design is developed to maximize the energy efficiency (EE) of the system. To tackle the intractable non-convex problem, we divide it into two subproblems and solve them iteratively based on successive convex approximation (SCA) and the Dinkelbach method. Finally we obtain a high-performance suboptimal solution. Simulation results show that the proposed algorithm significantly improves the energy efficiency of the MEC system.

preprint2022arXiv

Frequency Reflection Modulation for Reconfigurable Intelligent Surface Aided OFDM Systems

Reconfigurable intelligent surface (RIS) based reflection modulation has been considered as a promising information delivery mechanism, and has the potential to realize passive information transfer of a RIS without consuming any additional radio frequency chain and time/frequency/energy resources. The existing on-off reflection modulation (ORM) schemes are based on manipulating the "on/off" states of RIS elements, which may lead to the degradation of RIS reflection efficiency. This paper proposes a frequency reflection modulation (FRM) method for RIS-aided OFDM systems. The FRM-OFDM scheme modulates the frequency of the incident electromagnetic waves, and the RIS information is embedded in the frequency-hoping states of RIS elements. Unlike the ORM-OFDM scheme, the FRM-OFDM scheme can achieve higher reflection efficiency, since the latter does not turn off any reflection element in reflection modulation. We propose a block coordinate descent (BCD) algorithm to maximize the user achievable rate for the FRM-OFDM system by jointly optimizing the phase shift of the RIS and the power allocation at the transmitter. Further, we design a bilinear message passing (BMP) algorithm for the bilinear recovery of both the user symbols and the RIS data. Numerical simulations have verified the efficiency of the designed BCD algorithm for system optimization and the BMP algorithm for signal detection, as well as the superiority of the proposed FRM-OFDM scheme over the ORM-OFDM scheme.

preprint2022arXiv

Full-Dimensional Rate Enhancement for UAV-Enabled Communications via Intelligent Omni-Surface

This paper investigates the achievable rate maximization problem of a downlink unmanned aerial vehicle (UAV)-enabled communication system aided by an intelligent omni-surface (IOS). Different from the state-of-the-art reconfigurable intelligent surface (RIS) that only reflects incident signals, the IOS can simultaneously reflect and transmit the signals, thereby providing full-dimensional rate enhancement. To tackle such a problem, we formulate it by jointly optimizing the IOS's phase shift and the UAV trajectory. Although it is difficult to solve it optimally due to its non-convexity, we propose an efficient iterative algorithm to obtain a high-quality suboptimal solution. Simulation results show that the IOS-assisted UAV communications can achieve more significant improvement in achievable rates than other benchmark schemes.

preprint2022arXiv

Fully Convolutional Line Parsing

We present a one-stage Fully Convolutional Line Parsing network (F-Clip) that detects line segments from images. The proposed network is very simple and flexible with variations that gracefully trade off between speed and accuracy for different applications. F-Clip detects line segments in an end-to-end fashion by predicting each line's center position, length, and angle. We further customize the design of convolution kernels of our fully convolutional network to effectively exploit the statistical priors of the distribution of line angles in real image datasets. We conduct extensive experiments and show that our method achieves a significantly better trade-off between efficiency and accuracy, resulting in a real-time line detector at up to 73 FPS on a single GPU. Such inference speed makes our method readily applicable to real-time tasks without compromising any accuracy of previous methods. Moreover, when equipped with a performance-improving backbone network, F-Clip is able to significantly outperform all state-of-the-art line detectors on accuracy at a similar or even higher frame rate. In other word, under same inference speed, F-Clip always achieving best accuracy compare with other methods. Source code https://github.com/Delay-Xili/F-Clip.

preprint2022arXiv

Hybrid Offline-Online Design for Reconfigurable Intelligent Surface Aided UAV Communication

This letter considers the reconfigurable intelligent surface (RIS)-aided unmanned aerial vehicle (UAV) communication systems in urban areas under the general Rician fading channel. A hybrid offline-online design is proposed to improve the system performance by leveraging both the statistical channel state information (S-CSI) and instantaneous channel state information (I-CSI). For the offline phase, we aim to maximize the expected average achievable rate based on the S-CSI by jointly optimizing the RIS's phase-shift and UAV trajectory. The formulated stochastic optimization problem is difficult to solve due to its non-convexity. To tackle this problem, we propose an efficient algorithm by leveraging the stochastic successive convex approximation (SSCA) techniques. For the online phase, the UAV adaptively adjusts the transmit beamforming and user scheduling according to the effective I-CSI. Numerical results verify that the proposed hybrid design performs better than various bechmark schemes, and also demonstrate a favorable trade-off between system performance and CSI overhead.

preprint2022arXiv

Intelligent Reflecting Surface Aided MIMO with Cascaded LoS Links: Channel Modelling and Full Multiplexing Region

This work studies the modelling and the optimization of intelligent reflecting surface (IRS) assisted multiple-input multiple-output (MIMO) systems through cascaded line-of-sight (LoS) links. In Part I of this work, we build up a new IRS-aided MIMO channel model, named the cascaded LoS MIMO channel. The proposed channel model consists of a transmitter (Tx) and a receiver (Rx) both equipped with uniform linear arrays, and an IRS is used to enable communications between the transmitter and the receiver through the LoS links seen by the IRS. When modeling the reflection of electromagnetic waves at the IRS, we take into account the curvature of the wavefront on different reflecting elements. Based on the established model, we study the spatial multiplexing capability of the cascaded LoS MIMO system. We introduce the notion of full multiplexing region (FMR) for the cascaded LoS MIMO channel, where the FMR is the union of Tx-IRS and IRS-Rx distance pairs that enable full multiplexing communication. Under a special passive beamforming strategy named reflective focusing, we derive an inner bound of the FMR, and provide the corresponding orientation settings of the antenna arrays that enable full multiplexing. Based on the proposed channel model and reflective focusing, the mutual information maximization problem is discussed in Part II.

preprint2022arXiv

Joint Localization and Information Transfer for Reconfigurable Intelligent Surface Aided Full-Duplex Systems

In this work, we investigate a reconfigurable intelligent surface (RIS) aided integrated sensing and communication scenario, where a base station (BS) communicates with multiple devices in a full-duplex mode, and senses the positions of these devices simultaneously. An RIS is assumed to be mounted on each device to enhance the reflected echoes. Meanwhile, the information of each device is passively transferred to the BS via reflection modulation. We aim to tackle the problem of joint localization and information retrieval at the BS. A grid based parametric model is constructed and the joint estimation problem is formulated as a compressive sensing problem. We propose a novel message-passing algorithm to solve the considered problem, and a progressive approximation method to reduce the computational complexity involved in the message passing. Moreover, an expectation-maximization (EM) algorithm is applied for tuning the grid parameters to mitigate the model mismatch problem. Finally, we analyze the efficacy of the proposed algorithm through the Bayesian Cramér-Rao bound. Numerical results demonstrate the feasibility of the proposed scheme and the superior performance of the proposed EM-based message-passing algorithm.

preprint2022arXiv

Over-the-Air Federated Multi-Task Learning Over MIMO Multiple Access Channels

With the explosive growth of data and wireless devices, federated learning (FL) over wireless medium has emerged as a promising technology for large-scale distributed intelligent systems. Yet, the urgent demand for ubiquitous intelligence will generate a large number of concurrent FL tasks, which may seriously aggravate the scarcity of communication resources. By exploiting the analog superposition of electromagnetic waves, over-the-air computation (AirComp) is an appealing solution to alleviate the burden of communication required by FL. However, sharing frequency-time resources in over-the-air computation inevitably brings about the problem of inter-task interference, which poses a new challenge that needs to be appropriately addressed. In this paper, we study over-the-air federated multi-task learning (OA-FMTL) over the multiple-input multiple-output (MIMO) multiple access (MAC) channel. We propose a novel model aggregation method for the alignment of local gradients of different devices, which alleviates the straggler problem in over-the-air computation due to the channel heterogeneity. We establish a communication-learning analysis framework for the proposed OA-FMTL scheme by considering the spatial correlation between devices, and formulate an optimization problem for the design of transceiver beamforming and device selection. To solve this problem, we develop an algorithm by using alternating optimization (AO) and fractional programming (FP), which effectively mitigates the impact of inter-task interference on the FL learning performance. We show that due to the use of the new model aggregation method, device selection is no longer essential, thereby avoiding the heavy computational burden involved in selecting active devices. Numerical results demonstrate the validity of the analysis and the superb performance of the proposed scheme.

preprint2022arXiv

Over-the-Air Federated Multi-Task Learning via Model Sparsification and Turbo Compressed Sensing

To achieve communication-efficient federated multitask learning (FMTL), we propose an over-the-air FMTL (OAFMTL) framework, where multiple learning tasks deployed on edge devices share a non-orthogonal fading channel under the coordination of an edge server (ES). In OA-FMTL, the local updates of edge devices are sparsified, compressed, and then sent over the uplink channel in a superimposed fashion. The ES employs over-the-air computation in the presence of intertask interference. More specifically, the model aggregations of all the tasks are reconstructed from the channel observations concurrently, based on a modified version of the turbo compressed sensing (Turbo-CS) algorithm (named as M-Turbo-CS). We analyze the performance of the proposed OA-FMTL framework together with the M-Turbo-CS algorithm. Furthermore, based on the analysis, we formulate a communication-learning optimization problem to improve the system performance by adjusting the power allocation among the tasks at the edge devices. Numerical simulations show that our proposed OAFMTL effectively suppresses the inter-task interference, and achieves a learning performance comparable to its counterpart with orthogonal multi-task transmission. It is also shown that the proposed inter-task power allocation optimization algorithm substantially reduces the overall communication overhead by appropriately adjusting the power allocation among the tasks.

preprint2022arXiv

Receiver Design for MIMO Unsourced Random Access with SKP Coding

In this letter, we extend the sparse Kronecker-product (SKP) coding scheme, originally designed for the additive white Gaussian noise (AWGN) channel, to multiple input multiple output (MIMO) unsourced random access (URA). With the SKP coding adopted for MIMO transmission, we develop an efficient Bayesian iterative receiver design to solve the intended challenging trilinear factorization problem. Numerical results show that the proposed design outperforms the existing counterparts, and that it performs well in all simulated settings with various antenna sizes and active-user numbers.

preprint2022arXiv

RIS-Aided Multiuser MIMO-OFDM with Linear Precoding and Iterative Detection: Analysis and Optimization

In this paper, we consider a reconfigurable intelligence surface (RIS) aided uplink multiuser multi-input multi-output (MIMO) orthogonal frequency division multiplexing (OFDM) system, where the receiver is assumed to conduct low-complexity iterative detection. We aim to minimize the total transmit power by jointly designing the precoder of the transmitter and the passive beamforming of the RIS. This problem can be tackled from the perspective of information theory. But this information-theoretic approach may involve prohibitively high complexity since the number of rate constraints that specify the capacity region of the uplink multiuser channel is exponential in the number of users. To avoid this difficulty, we formulate the design problem of the iterative receiver under the constraints of a maximal iteration number and target bit error rates of users. To tackle this challenging problem, we propose a groupwise successive interference cancellation (SIC) optimization approach, where the signals of users are decoded and cancelled in a group-by-group manner. We present a heuristic user grouping strategy, and resort to the alternating optimization technique to iteratively solve the precoding and passive beamforming sub-problems. Specifically, for the precoding sub-problem, we employ fractional programming to convert it to a convex problem; for the passive beamforming sub-problem, we adopt successive convex approximation to deal with the unit-modulus constraints of the RIS. We show that the proposed groupwise SIC approach has significant advantages in both performance and computational complexity, as compared with the counterpart approaches.

preprint2022arXiv

Sparsity Learning Based Multiuser Detection in Grant-Free Massive-Device Multiple Access

In this work, we study the multiuser detection (MUD) problem for a grant-free massive-device multiple access (MaDMA) system, where a large number of single-antenna user devices transmit sporadic data to a multi-antenna base station (BS). Specifically, we put forth two MUD schemes, termed random sparsity learning multiuser detection (RSL-MUD) and structured sparsity learning multiuser detection (SSL-MUD) for the time-slotted and non-time-slotted grant-free MaDMA systems, respectively. In the time-slotted RSL-MUD scheme, active users generate and transmit data packets with random sparsity. In the non-time-slotted SSL-MUD scheme, we introduce a sliding-window-based detection framework, and the user signals in each observation window naturally exhibit structured sparsity. We show that by exploiting the sparsity embedded in the user signals, we can recover the user activity state, the channel, and the user data in a single phase, without using pilot signals for channel estimation and/or active user identification. To this end, we develop a message-passing based statistical inference framework for the BS to blindly detect the user data without any prior knowledge of the identities and the channel state information (CSI) of the active users. Simulation results show that our RSL-MUD and SSL-MUD schemes significantly outperform their counterpart schemes in both reducing the transmission overhead and improving the error behavior of the system.

preprint2021arXiv

Bayesian User Localization and Tracking for Reconfigurable Intelligent Surface Aided MIMO Systems

In this paper, we study the user localization and tracking problem in the reconfigurable intelligent surface (RIS) aided multiple-input multiple-output (MIMO) system, where a multi-antenna base station (BS) and multiple RISs are deployed to assist the localization and tracking of a multi-antenna user. By establishing a probability transition model for user mobility, we develop a message-passing algorithm, termed the Bayesian user localization and tracking (BULT) algorithm, to estimate and track the user position and the angle-of-arrival (AoAs) at the user in an online fashion. We also derive Bayesian Cramér Rao bound (BCRB) to characterize the fundamental performance limit of the considered tracking problem. To improve the tracking performance, we optimize the beamforming design at the BS and the RISs to minimize the derived BCRB. Simulation results show that our BULT algorithm can perform close to the derived BCRB, and significantly outperforms the counterpart algorithms without exploiting the temporal correlation of the user location.

preprint2021arXiv

Deep-Learned Approximate Message Passing for Asynchronous Massive Connectivity

This paper considers the massive connectivity problem in an asynchronous grant-free random access system, where a huge number of devices sporadically transmit data to a base station (BS) with imperfect synchronization. The goal is to design algorithms for joint user activity detection, delay detection, and channel estimation. By exploiting the sparsity on both user activity and delays, we formulate a hierarchical sparse signal recovery problem in both the single-antenna and the multiple-antenna scenarios. While traditional compressed sensing algorithms can be applied to these problems, they suffer high computational complexity and often require the perfect statistical information of channel and devices. This paper solves these problems by designing the Learned Approximate Message Passing (LAMP) network, which belongs to model-driven deep learning approaches and ensures efficient performance without tremendous training data. Particularly, in the multiple-antenna scenario, we design three different LAMP structures, namely, distributed, centralized and hybrid ones, to balance the performance and complexity. Simulation results demonstrate that the proposed LAMP networks can significantly outperform the conventional AMP method thanks to their ability of parameter learning. It is also shown that LAMP has robust performance to the maximal delay spread of the asynchronous users.

preprint2021arXiv

Reconfigurable Intelligent Surface for Massive Connectivity

With the rapid development of Internet of Things (IoT), massive machine-type communication has become a promising application scenario, where a large number of devices transmit sporadically to a base station (BS). Reconfigurable intelligent surface (RIS) has been recently proposed as an innovative new technology to achieve energy efficiency and coverage enhancement by establishing favorable signal propagation environments, thereby improving data transmission in massive connectivity. Nevertheless, the BS needs to detect active devices and estimate channels to support data transmission in RIS-assisted massive access systems, which yields unique challenges. This paper shall consider an RIS-assisted uplink IoT network and aims to solve the RIS-related activity detection and channel estimation problem, where the BS detects the active devices and estimates the separated channels of the RIS-to-device link and the RIS-to-BS link. Due to limited scattering between the RIS and the BS, we model the RIS-to-BS channel as a sparse channel. As a result, by simultaneously exploiting both the sparsity of sporadic transmission in massive connectivity and the RIS-to-BS channels, we formulate the RIS-related activity detection and channel estimation problem as a sparse matrix factorization problem. Furthermore, we develop an approximate message passing (AMP) based algorithm to solve the problem based on Bayesian inference framework and reduce the computational complexity by approximating the algorithm with the central limit theorem and Taylor series arguments. Finally, extensive numerical experiments are conducted to verify the effectiveness and improvements of the proposed algorithm.

preprint2021arXiv

Semi-Blind Cascaded Channel Estimation for Reconfigurable Intelligent Surface Aided Massive MIMO

Reconfigurable intelligent surface (RIS) is envisioned to be a promising green technology to reduce the energy consumption and improve the coverage and spectral efficiency of massive multiple-input multiple-output (MIMO) wireless networks. In a RIS-aided MIMO system, the acquisition of channel state information (CSI) is important for achieving passive beamforming gains of the RIS, but is also challenging due to the cascaded property of the transmitter-RIS-receiver channel and the lack of signal processing capability of the passive RIS elements. The state-of-the-art approach for CSI acquisition in such a system is a pure training-based strategy that depends on a long sequence of pilot symbols. In this paper, we investigate semi-blind cascaded channel estimation for RIS-aided massive MIMO systems, in which the receiver simultaneously estimates the channel coefficients and the partially unknown transmit signal with a small number of pilot sequences. Specifically, we formulate the semi-blind cascaded channel estimation as a trilinear matrix factorization task. Under the Bayesian inference framework, we develop a computationally efficient iterative algorithm using the approximate message passing principle to resolve the trilinear inference problem. Meanwhile, we present an analytical framework to characterize the theoretical performance bound of the proposed approach in the large-system limit via the replica method developed in statistical physics. Extensive simulation results demonstrate the effectiveness of the proposed semi-blind cascaded channel estimation algorithm.

preprint2021arXiv

Temporal-Structure-Assisted Gradient Aggregation for Over-the-Air Federated Edge Learning

In this paper, we investigate over-the-air model aggregation in a federated edge learning (FEEL) system. We introduce a Markovian probability model to characterize the intrinsic temporal structure of the model aggregation series. With this temporal probability model, we formulate the model aggregation problem as to infer the desired aggregated update given all the past observations from a Bayesian perspective. We develop a message passing based algorithm, termed temporal-structure-assisted gradient aggregation (TSA-GA), to fulfil this estimation task with low complexity and near-optimal performance. We further establish the state evolution (SE) analysis to characterize the behaviour of the proposed TSA-GA algorithm, and derive an explicit bound of the expected loss reduction of the FEEL system under certain standard regularity conditions. In addition, we develop an expectation maximization (EM) strategy to learn the unknown parameters in the Markovian model. We show that the proposed TSAGA algorithm significantly outperforms the state-of-the-art, and is able to achieve comparable learning performance as the error-free benchmark in terms of both convergence rate and final test accuracy.

preprint2020arXiv

Denoising-based Turbo Message Passing for Compressed Video Background Subtraction

In this paper, we consider the compressed video background subtraction problem that separates the background and foreground of a video from its compressed measurements. The background of a video usually lies in a low dimensional space and the foreground is usually sparse. More importantly, each video frame is a natural image that has textural patterns. By exploiting these properties, we develop a message passing algorithm termed offline denoising-based turbo message passing (DTMP). We show that these structural properties can be efficiently handled by the existing denoising techniques under the turbo message passing framework. We further extend the DTMP algorithm to the online scenario where the video data is collected in an online manner. The extension is based on the similarity/continuity between adjacent video frames. We adopt the optical flow method to refine the estimation of the foreground. We also adopt the sliding window based background estimation to reduce complexity. By exploiting the Gaussianity of messages, we develop the state evolution to characterize the per-iteration performance of offline and online DTMP. Comparing to the existing algorithms, DTMP can work at much lower compression rates, and can subtract the background successfully with a lower mean squared error and better visual quality for both offline and online compressed video background subtraction.

preprint2020arXiv

Joint User Identification, Channel Estimation, and Signal Detection for Grant-Free NOMA

For massive machine-type communications, centralized control may incur a prohibitively high overhead. Grant-free non-orthogonal multiple access (NOMA) provides possible solutions, yet poses new challenges for efficient receiver design. In this paper, we develop a joint user identification, channel estimation, and signal detection (JUICESD) algorithm. We divide the whole detection scheme into two modules: slot-wise multi-user detection (SMD) and combined signal and channel estimation (CSCE). SMD is designed to decouple the transmissions of different users by leveraging the approximate message passing (AMP) algorithms, and CSCE is designed to deal with the nonlinear coupling of activity state, channel coefficient and transmit signal of each user separately. To address the problem that the exact calculation of the messages exchanged within CSCE and between the two modules is complicated due to phase ambiguity issues, this paper proposes a rotationally invariant Gaussian mixture (RIGM) model, and develops an efficient JUICESD-RIGM algorithm. JUICESD-RIGM achieves a performance close to JUICESD with a much lower complexity. Capitalizing on the feature of RIGM, we further analyze the performance of JUICESD-RIGM with state evolution techniques. Numerical results demonstrate that the proposed algorithms achieve a significant performance improvement over the existing alternatives, and the derived state evolution method predicts the system performance accurately.

preprint2020arXiv

Matrix-Calibration-Based Cascaded Channel Estimation for Reconfigurable Intelligent Surface Assisted Multiuser MIMO

Reconfigurable intelligent surface (RIS) is envisioned to be an essential component of the paradigm for beyond 5G networks as it can potentially provide similar or higher array gains with much lower hardware cost and energy consumption compared with the massive multiple-input multiple-output (MIMO) technology. In this paper, we focus on one of the fundamental challenges, namely the channel acquisition, in an RIS-assisted multiuser MIMO system. The state-of-the-art channel acquisition approach in such a system with fully passive RIS elements estimates the cascaded transmitter-to-RIS and RIS-to-receiver channels by adopting excessively long training sequences. To estimate the cascaded channels with an affordable training overhead, we formulate the channel estimation problem in the RIS-assisted multiuser MIMO system as a matrix-calibration based matrix factorization task. By exploiting the information on the slow-varying channel components and the hidden channel sparsity, we propose a novel message-passing based algorithm to factorize the cascaded channels. Furthermore, we present an analytical framework to characterize the theoretical performance bound of the proposed estimator in the large-system limit. Finally, we conduct simulations to verify the high accuracy and efficiency of the proposed algorithm.

preprint2020arXiv

Reconfigurable Intelligent Surface Aided Constant-Envelope Wireless Power Transfer

By reconfiguring the propagation environment of electromagnetic waves artificially, reconfigurable intelligent surfaces (RISs) have been regarded as a promising and revolutionary hardware technology to improve the energy and spectrum efficiency of wireless networks. In this paper, we study a RIS aided multiuser multiple-input single-output (MISO) wireless power transfer (WPT) system, where the transmitter is equipped with a constant-envelope analog beamformer. We formulate a novel problem to maximize the total received power of all the users by jointly optimizing the beamformer at transmitter and the phase shifts at the RISs, subject to the individual minimum received power constraints of users. We further solve the problem iteratively with a closed-form expression for each step. Numerical results show the performance gain of deploying RIS and the effectiveness of the proposed algorithm.

preprint2020arXiv

Reconfigurable Intelligent Surfaces for Energy Efficiency in D2D Communication Network

In this letter, the joint power control of D2D users and the passive beamforming of reconfigurable intelligent surfaces (RIS) for a RIS-aided device-to-device (D2D) communication network is investigated to maximize energy efficiency. This non-convex optimization problem is divided into two subproblems, which are passive beamforming and power control. The two subproblems are optimized alternately. We first decouple the passive beamforming at RIS based on the Lagrangian dual transform. This problem is solved by using fractional programming. Then we optimize the power control by using the Dinkelbach method. By iteratively solving the two subproblems, we obtain a suboptimal solution for the joint optimization problem. Numerical results have verified the effectiveness of the proposed algorithm, which can significantly improve the energy efficiency of the D2D network.

preprint2020arXiv

Reconfigurable-Intelligent-Surface Empowered Wireless Communications: Challenges and Opportunities

Reconfigurable intelligent surfaces (RISs) are regarded as a promising emerging hardware technology to improve the spectrum and energy efficiency of wireless networks by artificially reconfiguring the propagation environment of electromagnetic waves. Due to the unique advantages in enhancing wireless channel capacity, RISs have recently become a hot research topic. In this article, we focus on three fundamental physical-layer challenges for the incorporation of RISs into wireless networks, namely, channel state information acquisition, passive information transfer, and low-complexity robust system design. We summarize the state-of-the-art solutions and explore potential research directions. Furthermore, we discuss other promising research directions of RISs, including edge intelligence and physical-layer security.

preprint2020arXiv

Statistical Beamforming for FDD Downlink Massive MIMO via Spatial Information Extraction and Beam Selection

In this paper, we study the beamforming design problem in frequency-division duplexing (FDD) downlink massive MIMO systems, where instantaneous channel state information (CSI) is assumed to be unavailable at the base station (BS). We propose to extract the information of the angle-of-departures (AoDs) and the corresponding large-scale fading coefficients (a.k.a. spatial information) of the downlink channel from the uplink channel estimation procedure, based on which a novel downlink beamforming design is presented. By separating the subpaths for different users based on the spatial information and the hidden sparsity of the physical channel, we construct near-orthogonal virtual channels in the beamforming design. Furthermore, we derive a sum-rate expression and its approximations for the proposed system. Based on these closed-form rate expressions, we develop two low-complexity beam selection schemes and carry out asymptotic analysis to provide valuable insights on the system design. Numerical results demonstrate a significant performance improvement of our proposed algorithm over the state-of-the-art beamforming approach.

preprint2020arXiv

Two-Timescale Optimization for Intelligent Reflecting Surface Aided D2D Underlay Communication

The performance of a device-to-device (D2D) underlay communication system is limited by the co-channel interference between cellular users (CUs) and D2D devices. To address this challenge, an intelligent reflecting surface (IRS) aided D2D underlay system is studied in this paper. A two-timescale optimization scheme is proposed to reduce the required channel training and feedback overhead, where transmit beamforming at the base station (BS) and power control at the D2D transmitter are adapted to instantaneous effective channel state information (CSI); and the IRS phase shifts are adapted to slow-varying channel mean. Based on the two-timescale optimization scheme, we aim to maximize the D2D ergodic rate subject to a given outage probability constrained signal-to-interference-plus-noise ratio (SINR) target for the CU. The two-timescale problem is decoupled into two sub-problems, and the two sub-problems are solved iteratively with closed-form expressions. Numerical results verify that the two-timescale based optimization performs better than several baselines, and also demonstrate a favorable trade-off between system performance and CSI overhead.

preprint2019arXiv

Passive Beamforming and Information Transfer Design for Reconfigurable Intelligent Surfaces Aided Multiuser MIMO Systems

This paper investigates the passive beamforming and information transfer (PBIT) technique for the multiuser multiple-input multiple-output (MIMO) systems with the aid of a reconfigurable intelligent surface (RIS), where the RIS enhances the primary communication via passive beamforming and at the same time delivers additional information by the spatial modulation (which adjusts the on-off states of the reflecting elements). For the passive beamforming design, we propose to maximize the sum channel capacity of the RIS-aided multiuser MIMO channel and formulate the problem as a two-step stochastic program. A sample average approximation (SAA) based iterative algorithm is developed for the efficient passive beamforming design of the considered scheme. To strike a balance between complexity and performance, we then propose a simplified beamforming algorithm by approximating the stochastic program as a deterministic alternating optimization problem. For the receiver design, the signal detection at the receiver is a bilinear estimation problem since the RIS information is multiplicatively modulated onto the reflected signals of the reflecting elements. To solve this bilinear estimation problem, we develop a turbo message passing (TMP) algorithm in which the factor graph associated with the problem is divided into two modules: one for the estimation of the user signals and the other for the estimation of the RIS's on-off states. The two modules are executed iteratively to yield a near-optimal low-complexity solution. Furthermore, we extend the design of the multiuser MIMO PBIT scheme from single-RIS to multi-RIS, by leveraging the similarity between the single-RIS and multi-RIS system models. Extensive simulation results are provided to demonstrate the advantages of our passive beamforming and receiver designs.

preprint2019arXiv

Variance State Propagation for Structured Sparse Bayesian Learning

We propose a compressed sensing algorithm termed variance state propagation (VSP) for block-sparse signals, i.e., sparse signals that have nonzero coefficients occurring in clusters. The VSP algorithm is developed under the Bayesian framework. A hierarchical Gaussian prior is introduced to depict the clustered patterns in the sparse signal. Markov random field (MRF) is introduced to characterize the state of the variances of the Gaussian priors. Such a hierarchical prior has the potential to encourage clustered patterns and suppress isolated coefficients whose patterns are different from their respective neighbors. The core idea of our algorithm is to iteratively update the variances in the prior Gaussian distribution. The message passing technique is employed in the design of the algorithm. For messages that are difficult to calculate, we correspondingly design reasonable methods to achieve approximate calculations. The hyperparameters can be updated within the iteration process. Simulation results demonstrate that the VSP algorithm is able to handle a variety of block-sparse signal recovery tasks and presents a significant advantage over the existing methods.

preprint2018arXiv

TARM: A Turbo-type Algorithm for Affine Rank Minimization

The affine rank minimization (ARM) problem arises in many real-world applications. The goal is to recover a low-rank matrix from a small amount of noisy affine measurements. The original problem is NP-hard, and so directly solving the problem is computationally prohibitive. Approximate low-complexity solutions for ARM have recently attracted much research interest. In this paper, we design an iterative algorithm for ARM based on message passing principles. The proposed algorithm is termed turbo-type ARM (TARM), as inspired by the recently developed turbo compressed sensing algorithm for sparse signal recovery. We show that, when the linear operator for measurement is right-orthogonally invariant (ROIL), a scalar function called state evolution can be established to accurately predict the behaviour of the TARM algorithm. We also show that TARM converges much faster than the counterpart algorithms for low-rank matrix recovery. We further extend the TARM algorithm for matrix completion, where the measurement operator corresponds to a random selection matrix. We show that, although the state evolution is not accurate for matrix completion, the TARM algorithm with carefully tuned parameters still significantly outperforms its counterparts.

preprint2016arXiv

D-OAMP: A Denoising-based Signal Recovery Algorithm for Compressed Sensing

Approximate message passing (AMP) is an efficient iterative signal recovery algorithm for compressed sensing (CS). For sensing matrices with independent and identically distributed (i.i.d.) Gaussian entries, the behavior of AMP can be asymptotically described by a scaler recursion called state evolution. Orthogonal AMP (OAMP) is a variant of AMP that imposes a divergence-free constraint on the denoiser. In this paper, we extend OAMP to incorporate generic denoisers, hence the name D-OAMP. Our numerical results show that state evolution predicts the performance of D-OAMP well for generic denoisers when i.i.d. Gaussian or partial orthogonal sensing matrices are involved. We compare the performances of denosing-AMP (D-AMP) and D-OAMP for recovering natural images from CS measurements. Simulation results show that D-OAMP outperforms D-AMP in both convergence speed and recovery accuracy for partial orthogonal sensing matrices.

preprint2016arXiv

Fundamental Limits of Training-Based Multiuser MIMO Systems

In this paper, we endeavour to seek a fundamental understanding of the potentials and limitations of training-based multiuser multiple-input multiple-output (MIMO) systems. In a multiuser MIMO system, users are geographically separated. So, the near-far effect plays an indispensable role in channel fading. The existing optimal training design for conventional MIMO does not take the near-far effect into account, and thus is not applicable to a multiuser MIMO system. In this work, we use the majorization theory as a basic tool to study the tradeoff between the channel estimation quality and the information throughput. We establish tight upper and lower bounds of the throughput, and prove that the derived lower bound is asymptotically optimal for throughput maximization at high signal-to-noise ratio. Our analysis shows that the optimal training sequences for throughput maximization in a multiuser MIMO system are in general not orthogonal to each other. Furthermore, due to the near-far effect, the optimal training design for throughput maximization is to deactivate a portion of users with the weakest channels in transmission. These observations shed light on the practical design of training-based multiuser MIMO systems.

preprint2016arXiv

Locally Orthogonal Training Design for Cloud-RANs Based on Graph Coloring

We consider training-based channel estimation for a cloud radio access network (CRAN), in which a large amount of remote radio heads (RRHs) and users are randomly scattered over the service area. In this model, assigning orthogonal training sequences to all users will incur a substantial overhead to the overall network, and is even impossible when the number of users is large. Therefore, in this paper, we introduce the notion of local orthogonality, under which the training sequence of a user is orthogonal to those of the other users in its neighborhood. We model the design of locally orthogonal training sequences as a graph coloring problem. Then, based on the theory of random geometric graph, we show that the minimum training length scales in the order of $\ln K$, where $K$ is the number of users covered by a CRAN. This indicates that the proposed training design yields a scalable solution to sustain the need of large-scale cooperation in CRANs. Numerical results show that the proposed scheme outperforms other reference schemes.

preprint2016arXiv

MIMO Multiway Distributed-Relay Channel with Full Data Exchange: An Achievable Rate Perspective

We consider efficient communications over the multiple-input multiple-output (MIMO) multiway distributed relay channel (MDRC) with full data exchange, where each user, equipped with multiple antennas, broadcasts its message to all the other users via the help of a number of distributive relays. We propose a physical-layer network coding (PNC) based scheme involving linear precoding for channel alignment nested lattice coding for PNC, and lattice-based precoding for interference mitigation, We show that, with the proposed scheme, distributed relaying achieves the same sum-rate as cooperative relaying in the high SNR regime. We also show that the proposed scheme achieve the asymptotic sum capacity of the MIMO MDRC within a constant gap at high SNR. Numerical results demonstrate that the proposed scheme considerably outperforms the existing schemes including decode-and-forward and amplify-and-forward.

preprint2016arXiv

Optimal Degrees of Freedom Region for the Asymmetric MIMO Y Channel

This letter studies the optimal degrees of freedom (DoF) region for the asymmetric three-user MIMO Y channel with antenna configuration $(M_1,M_2,M_3,N)$, where $M_i$ is the number of antennas at user $i$ and $N$ is the number of antennas at the relay node. The converse is proved by using the cut-set theorem and the genie-message approach. To prove the achievability, we divide the DoF tuples in the outer bound into two cases. For each case, we show that the DoF tuples are achievable by collectively utilizing antenna deactivation, pairwise signal alignment and cyclic signal alignment techniques. This work not only offers a complete characterization of DoF region for the considered channel model, but also provides a new and elegant achievability proof.

preprint2015arXiv

Compute-Compress-and-Forward: Exploiting Asymmetry of Wireless Relay Networks

Compute-and-forward (CF) harnesses interference in a wireless networkby allowing relays to compute combinations of source messages. The computed message combinations at relays are correlated, and so directly forwarding these combinations to a destination generally incurs information redundancy and spectrum inefficiency. To address this issue, we propose a novel relay strategy, termed compute-compress-and-forward (CCF). In CCF, source messages are encoded using nested lattice codes constructed on a chain of nested coding and shaping lattices. A key difference of CCF from CF is an extra compressing stage inserted in between the computing and forwarding stages of a relay, so as to reduce the forwarding information rate of the relay. The compressing stage at each relay consists of two operations: first to quantize the computed message combination on an appropriately chosen lattice (referred to as a quantization lattice), and then to take modulo on another lattice (referred to as a modulo lattice). We study the design of the quantization and modulo lattices and propose successive recovering algorithms to ensure the recoverability of source messages at destination. Based on that, we formulate a sum-rate maximization problem that is in general an NP-hard mixed integer program. A low-complexity algorithm is proposed to give a suboptimal solution. Numerical results are presented to demonstrate the superiority of CCF over the existing CF schemes.

preprint2015arXiv

On the Performance of Turbo Signal Recovery with Partial DFT Sensing Matrices

This letter is on the performance of the turbo signal recovery (TSR) algorithm for partial discrete Fourier transform (DFT) matrices based compressed sensing. Based on state evolution analysis, we prove that TSR with a partial DFT sensing matrix outperforms the well-known approximate message passing (AMP) algorithm with an independent identically distributed (IID) sensing matrix.

preprint2014arXiv

Dynamic Nested Clustering for Parallel PHY-Layer Processing in Cloud-RANs

Featured by centralized processing and cloud based infrastructure, Cloud Radio Access Network (C-RAN) is a promising solution to achieve an unprecedented system capacity in future wireless cellular networks. The huge capacity gain mainly comes from the centralized and coordinated signal processing at the cloud server. However, full-scale coordination in a large-scale C-RAN requires the processing of very large channel matrices, leading to high computational complexity and channel estimation overhead. To resolve this challenge, we exploit the near-sparsity of large C-RAN channel matrices, and derive a unified theoretical framework for clustering and parallel processing. Based on the framework, we propose a dynamic nested clustering (DNC) algorithm that not only greatly improves the system scalability in terms of baseband-processing and channel-estimation complexity, but also is amenable to various parallel processing strategies for different data center architectures. With the proposed algorithm, we show that the computation time for the optimal linear detector is greatly reduced from $O(N^3)$ to no higher than $O(N^{\frac{42}{23}})$, where $N$ is the number of RRHs in C-RAN.

preprint2014arXiv

MIMO Multiway Relaying with Clustered Full Data Exchange: Signal Space Alignment and Degrees of Freedom

We investigate achievable degrees of freedom (DoF) for a multiple-input multiple-output (MIMO) multiway relay channel (mRC) with $L$ clusters and $K$ users per cluster. Each user is equipped with $M$ antennas and the relay with $N$ antennas. We assume a new data exchange model, termed \emph{clustered full data exchange}, i.e., each user in a cluster wants to learn the messages of all the other users in the same cluster. Novel signal alignment techniques are developed to systematically construct the beamforming matrices at the users and the relay for efficient physical-layer network coding. Based on that, we derive an achievable DoF of the MIMO mRC with an arbitrary network configuration of $L$ and $K$, as well as with an arbitrary antenna configuration of $M$ and $N$. We show that our proposed scheme achieves the DoF capacity when $\frac{M}{N} \leq \frac{1}{LK-1}$ and $\frac{M}{N} \geq \frac{(K-1)L+1}{KL}$.

preprint2014arXiv

MIMO Multiway Relaying with Pairwise Data Exchange: A Degrees of Freedom Perspective

In this paper, we study achievable degrees of freedom (DoF) of a multiple-input multiple-output (MIMO) multiway relay channel (mRC) where $K$ users, each equipped with $M$ antennas, exchange messages in a pairwise manner via a common $N$-antenna relay node. % A novel and systematic way of joint beamforming design at the users and at the relay is proposed to align signals for efficient implementation of physical-layer network coding (PNC). It is shown that, when the user number $K=3$, the proposed beamforming design can achieve the DoF capacity of the considered mRC for any $(M,N)$ setups. % For the scenarios with $K>3$, we show that the proposed signaling scheme can be improved by disabling a portion of relay antennas so as to align signals more efficiently. Our analysis reveals that the obtained achievable DoF is always piecewise linear, and is bounded either by the number of user antennas $M$ or by the number of relay antennas $N$. Further, we show that the DoF capacity can be achieved for $\frac{M}{N} \in \left(0,\frac{K-1}{K(K-2)} \right]$ and $\frac{M}{N} \in \left[\frac{1}{K(K-1)}+\frac{1}{2},\infty \right)$, which provides a broader range of the DoF capacity than the existing results. Asymptotic DoF as $K\rightarrow \infty$ is also derived based on the proposed signaling scheme.

preprint2014arXiv

Multiple-Input Multiple-Output Two-Way Relaying: A Space-Division Approach

We propose a novel space-division based network-coding scheme for multiple-input multiple-output (MIMO) two-way relay channels (TWRCs), in which two multi-antenna users exchange information via a multi-antenna relay. In the proposed scheme, the overall signal space at the relay is divided into two subspaces. In one subspace, the spatial streams of the two users have nearly orthogonal directions, and are completely decoded at the relay. In the other subspace, the signal directions of the two users are nearly parallel, and linear functions of the spatial streams are computed at the relay, following the principle of physical-layer network coding (PNC). Based on the recovered messages and message-functions, the relay generates and forwards network-coded messages to the two users. We show that, at high signal-to-noise ratio (SNR), the proposed scheme achieves the asymptotic sum rate capacity of MIMO TWRCs within 1/2log(5/4) = 0.161 bits per user-antenna for any antenna configuration and channel realization. We perform large-system analysis to derive the average sum-rate of the proposed scheme over Rayleigh-fading MIMO TWRCs. We show that the average asymptotic sum rate gap to the capacity upper bound is at most 0.053 bits per relay-antenna. It is demonstrated that the proposed scheme significantly outperforms the existing schemes.

preprint2014arXiv

Towards the Asymptotic Sum Capacity of the MIMO Cellular Two-Way Relay Channel

In this paper, we consider the transceiver and relay design for multiple-input multiple-output (MIMO) cellular two-way relay channel (cTWRC), where a multi-antenna base station (BS) exchanges information with multiple multi-antenna mobile stations via a multi-antenna relay station (RS). We propose a novel two-way relaying scheme to approach the sum capacity of the MIMO cTWRC.

preprint2014arXiv

Turbo Compressed Sensing with Partial DFT Sensing Matrix

In this letter, we propose a turbo compressed sensing algorithm with partial discrete Fourier transform (DFT) sensing matrices. Interestingly, the state evolution of the proposed algorithm is shown to be consistent with that derived using the replica method. Numerical results demonstrate that the proposed algorithm outperforms the well-known approximate message passing (AMP) algorithm when a partial DFT sensing matrix is involved.

preprint2013arXiv

Wireless MIMO Switching: Weighted Sum Mean Square Error and Sum Rate Optimization

This paper addresses joint transceiver and relay design for a wireless multiple-input-multiple-output (MIMO) switching scheme that enables data exchange among multiple users. Here, a multi-antenna relay linearly precodes the received (uplink) signals from multiple users before forwarding the signal in the downlink, where the purpose of precoding is to let each user receive its desired signal with interference from other users suppressed. The problem of optimizing the precoder based on various design criteria is typically non-convex and difficult to solve. The main contribution of this paper is a unified approach to solve the weighted sum mean square error (MSE) minimization and weighted sum rate maximization problems in MIMO switching. Specifically, an iterative algorithm is proposed for jointly optimizing the relay's precoder and the users' receive filters to minimize the weighted sum MSE. It is also shown that the weighted sum rate maximization problem can be reformulated as an iterated weighted sum MSE minimization problem and can therefore be solved similarly to the case of weighted sum MSE minimization. With properly chosen initial values, the proposed iterative algorithms are asymptotically optimal in both high and low signal-to-noise ratio (SNR) regimes for MIMO switching, either with or without self-interference cancellation (a.k.a., physical-layer network coding). Numerical results show that the optimized MIMO switching scheme based on the proposed algorithms significantly outperforms existing approaches in the literature.

preprint2012arXiv

Achievable Rates of MIMO Systems with Linear Precoding and Iterative LMMSE Detection

We establish area theorems for iterative detection over coded linear systems (including multiple-input multipleoutput (MIMO) channels, inter-symbol-interference (ISI) channels, and orthogonal frequency-division multiplexing (OFDM) systems). We propose a linear precoding technique that asymptotically ensures the Gaussianness of the messages passed in iterative detection, as the transmission block length tends to infinity. We show that the proposed linear precoding scheme with iterative linear minimum mean-square error (LMMSE) detection is potentially information lossless, under various assumptions on the availability of channel state information at the transmitter (CSIT). Numerical results are provided to verify our analysis.

preprint2012arXiv

Eigen-Direction Alignment Based Physical-Layer Network Coding for MIMO Two-Way Relay Channels

In this paper, we propose a novel communication strategy which incorporates physical-layer network coding (PNC) into multiple-input multiple output (MIMO) two-way relay channels (TWRCs). At the heart of the proposed scheme lies a new key technique referred to as eigen-direction alignment (EDA) precoding. The EDA precoding efficiently aligns the two-user's eigen-modes into the same directions. Based on that, we carry out multi-stream PNC over the aligned eigen-modes. We derive an achievable rate of the proposed EDA-PNC scheme, based on nested lattice codes, over a MIMO TWRC. Asymptotic analysis shows that the proposed EDA-PNC scheme approaches the capacity upper bound as the number of user antennas increases towards infinity. For a finite number of user antennas, we formulate the design criterion of the optimal EDA precoder and present solutions. Numerical results show that there is only a marginal gap between the achievable rate of the proposed EDA-PNC scheme and the capacity upper bound of the MIMO TWRC, in the median-to-large SNR region. We also show that the proposed EDA-PNC scheme significantly outperforms existing amplify-and-forward and decode-and-forward based schemes for MIMO TWRCs.