Source author record

Mohammad M. Mansour

Mohammad M. Mansour appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

11works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

11 published item(s)

preprint2021arXiv

Deep Learning Based Frequency-Selective Channel Estimation for Hybrid mmWave MIMO Systems

Millimeter wave (mmWave) massive multiple-input multiple-output (MIMO) systems typically employ hybrid mixed signal processing to avoid expensive hardware and high training overheads. {However, the lack of fully digital beamforming at mmWave bands imposes additional challenges in channel estimation. Prior art on hybrid architectures has mainly focused on greedy optimization algorithms to estimate frequency-flat narrowband mmWave channels, despite the fact that in practice, the large bandwidth associated with mmWave channels results in frequency-selective channels. In this paper, we consider a frequency-selective wideband mmWave system and propose two deep learning (DL) compressive sensing (CS) based algorithms for channel estimation.} The proposed algorithms learn critical apriori information from training data to provide highly accurate channel estimates with low training overhead. In the first approach, a DL-CS based algorithm simultaneously estimates the channel supports in the frequency domain, which are then used for channel reconstruction. The second approach exploits the estimated supports to apply a low-complexity multi-resolution fine-tuning method to further enhance the estimation performance. Simulation results demonstrate that the proposed DL-based schemes significantly outperform conventional orthogonal matching pursuit (OMP) techniques in terms of the normalized mean-squared error (NMSE), computational complexity, and spectral efficiency, particularly in the low signal-to-noise ratio regime. When compared to OMP approaches that achieve an NMSE gap of \$\unit[\{4-10\}]{dB}\$ with respect to the Cramer Rao Lower Bound (CRLB), the proposed algorithms reduce the CRLB gap to only \$\unit[\{1-1.5\}]{dB}\$, while significantly reducing complexity by two orders of magnitude.

preprint2021arXiv

Terahertz-Band MIMO-NOMA: Adaptive Superposition Coding and Subspace Detection

We consider the problem of efficient ultra-massive multiple-input multiple-output (UM-MIMO) data detection in terahertz (THz)-band non-orthogonal multiple access (NOMA) systems. We argue that the most common THz NOMA configuration is power-domain superposition coding over quasi-optical doubly-massive MIMO channels. We propose spatial tuning techniques that modify antenna subarray arrangements to enhance channel conditions. Towards recovering the superposed data at the receiver side, we propose a family of data detectors based on low-complexity channel matrix puncturing, in which higher-order detectors are dynamically formed from lower-order component detectors. We first detail the proposed solutions for the case of superposition coding of multiple streams in point-to-point THz MIMO links. We then extend the study to multi-user NOMA, in which randomly distributed users get grouped into narrow cell sectors and are allocated different power levels depending on their proximity to the base station. We show that successive interference cancellation is carried with minimal performance and complexity costs under spatial tuning. We derive approximate bit error rate (BER) equations, and we propose an architectural design to illustrate complexity reductions. Under typical THz conditions, channel puncturing introduces more than an order of magnitude reduction in BER at high signal-to-noise ratios while reducing complexity by approximately 90%.

preprint2020arXiv

Efficient Angle-Domain Processing for FDD-based Cell-free Massive MIMO Systems

Cell-free massive MIMO communications is an emerging network technology for 5G wireless communications wherein distributed multi-antenna access points (APs) serve many users simultaneously. Most prior work on cell-free massive MIMO systems assume time-division duplexing mode, although frequency-division duplexing (FDD) systems dominate current wireless standards. The key challenges in FDD massive MIMO systems are channel-state information (CSI) acquisition and feedback overhead. To address these challenges, we exploit the so-called angle reciprocity of multipath components in the uplink and downlink, so that the required CSI acquisition overhead scales only with the number of served users, and not the number of AP antennas nor APs. We propose a low complexity multipath component estimation technique and present linear angle-of-arrival (AoA)-based beamforming/combining schemes for FDD-based cell-free massive MIMO systems. We analyze the performance of these schemes by deriving closed-form expressions for the mean-square-error of the estimated multipath components, as well as expressions for the uplink and downlink spectral efficiency. Using semi-definite programming, we solve a max-min power allocation problem that maximizes the minimum user rate under per-user power constraints. Furthermore, we present a user-centric (UC) AP selection scheme in which each user chooses a subset of APs to improve the overall energy efficiency of the system. Simulation results demonstrate that the proposed multipath component estimation technique outperforms conventional subspace-based and gradient-descent based techniques. We also show that the proposed beamforming and combining techniques along with the proposed power control scheme substantially enhance the spectral and energy efficiencies with an adequate number of antennas at the APs.

preprint2020arXiv

Optimal Augmented-Channel Puncturing for Low-Complexity Soft-Output MIMO Detectors

We propose a computationally-efficient soft-output detector for multiple-input multiple-output channels based on augmented channel puncturing in order to reduce tree processing complexity. The proposed detector, dubbed augmented WL detector (AWLD), employs a punctured channel with a special structure derived by triangulizing the original channel in augmented form, followed by Gaussian elimination. We prove that these punctured channels are optimal in maximizing the lower-bound on the achievable information rate (AIR) based on a newly proposed mismatched detection model. We show that the AWLD decomposes into a minimum mean-square error (MMSE) prefilter and channel-gain compensation stages, followed by a regular unaugmented WL detector (WLD). It attains the same performance as the existing AIR partial marginalization (AIR-PM) detector, but with much simpler processing.

preprint2016arXiv

Comments on "A Square-Root-Free Matrix Decomposition Method for Energy-Efficient Least Square Computation on Embedded Systems"

A square-root-free matrix QR decomposition (QRD) scheme was rederived in [1] based on [2] to simplify computations when solving least-squares (LS) problems on embedded systems. The scheme of [1] aims at eliminating both the square-root and division operations in the QRD normalization and backward substitution steps in the LS computations. It is claimed in [1] that the LS solution only requires finding the directions of the orthogonal basis of the matrix in question, regardless of the normalization of their Euclidean norms. MIMO detection problems have been named as potential applications that benefit from this. While this is true for unconstrained LS problems, we conversely show here that constrained LS problems such as MIMO detection still require computing the norms of the orthogonal basis to produce the correct result.

preprint2016arXiv

Modulation Classification via Subspace Detection in MIMO Systems

The problem of efficient modulation classification (MC) in multiple-input multiple-output (MIMO) systems is considered. Per-layer likelihood-based MC is proposed by employing subspace decomposition to partially decouple the transmitted streams. When detecting the modulation type of the stream of interest, a dense constellation is assumed on all remaining streams. The proposed classifier outperforms existing MC schemes at a lower complexity cost, and can be efficiently implemented in the context of joint MC and subspace data detection.

preprint2015arXiv

A Low-Complexity Detection Algorithm for the Primary Synchronization Signal in LTE

One of the challenging tasks in LTE baseband receiver design is synchronization, which determines the symbol boundary and transmitted frame start-time, and performs cell identification. Conventional algorithms are based on correlation methods that involve a large number of multiplications and thus lead to high receiver hardware complexity and power consumption. In this paper, a hardware-efficient synchronization algorithm for frame timing based on K-means clustering schemes is proposed. The algorithm reduces the complexity of the primary synchronization signal for LTE from 24 complex-multiplications, currently best known in the literature, to just 8. Simulation results demonstrate that the proposed algorithm has negligible performance degradation with reduced complexity relative to conventional techniques.

preprint2015arXiv

Inter-Frame Coding For Broadcast Communication

A novel inter-frame coding approach to the problem of varying channel-state conditions in broadcast wireless communication is developed in this paper; this problem causes the appropriate code-rate to vary across different transmitted frames and different receivers as well. The main aspect of the proposed approach is that it incorporates an iterative rate-matching process into the decoding of the received set of frames, such that: throughout inter-frame decoding, the code-rate of each frame is progressively lowered to or below the appropriate value, prior to applying or re-applying conventional physical-layer channel decoding on it. This iterative rate-matching process is asymptotically analyzed in this paper. It is shown to be optimal, in the sense defined in the paper. Consequently, the data-rates achievable by the proposed scheme are derived. Overall, it is concluded that, compared to the existing solutions, inter-frame coding presents a better complexity versus data-rate tradeoff. In terms of complexity, the overhead of inter-frame decoding includes operations that are similar in type and scheduling to those employed in the relatively- simple iterative erasure decoding. In terms of data-rates, compared to the state-of-the-art two-stage scheme involving both error-correcting and erasure coding, inter-frame coding increases the data-rate by a factor that reaches up to 1.55x.

preprint2015arXiv

Multi-User MIMO Receivers With Partial State Information

We consider a multi-user multiple-input multiple-output (MU-MIMO) system that uses orthogonal frequency division multiplexing (OFDM). Several receivers are developed for data detection of MU-MIMO transmissions where two users share the same OFDM time and frequency resources. The receivers have partial state information about the MU-MIMO transmission with each receiver having knowledge of the MU-MIMO channel, however the modulation constellation of the co-scheduled user is unknown. We propose a joint maximum likelihood (ML) modulation classification of the co-scheduled user and data detection receiver using the max-log-MAP approximation. It is shown that the decision metric for the modulation classification is an accumulation over a set of tones of Euclidean distance computations that are also used by the max-log-MAP detector for bit log-likelihood ratio (LLR) soft decision generation. An efficient hardware implementation emerges that exploits this commonality between the classification and detection steps and results in sharing of the hardware resources. Comparisons of the link performance of the proposed receiver to several linear receivers is demonstrated through computer simulations. It is shown that the proposed receiver offers \unit[1.5]{dB} improvement in signal-to-noise ratio (SNR) over the nulling projection receiver at $1\%$ block error rate (BLER) for $64$-QAM with turbo code rate of $1/2$ in the case of zero transmit and receiver antenna correlations. However, in the case of high antenna correlation, the linear receiver approaches suffer significant loss relative to the optimal receiver.

preprint2015arXiv

Optimized Configurable Architectures for Scalable Soft-Input Soft-Output MIMO Detectors with 256-QAM

This paper presents an optimized low-complexity and high-throughput multiple-input multiple-output (MIMO) signal detector core for detecting spatially-multiplexed data streams. The core architecture supports various layer configurations up to 4, while achieving near-optimal performance, as well as configurable modulation constellations up to 256-QAM on each layer. The core is capable of operating as a soft-input soft-output log-likelihood ratio (LLR) MIMO detector which can be used in the context of iterative detection and decoding. High area-efficiency is achieved via algorithmic and architectural optimizations performed at two levels. First, distance computations and slicing operations for an optimal 2-layer maximum a posteriori (MAP) MIMO detector are optimized to eliminate the use of multipliers and reduce the overhead of slicing in the presence of soft-input LLRs. We show that distances can be easily computed using elementary addition operations, while optimal slicing is done via efficient comparisons with soft decision boundaries, resulting in a simple feed-forward pipelined architecture. Second, to support more layers, an efficient channel decomposition scheme is presented that reduces the detection of multiple layers into multiple 2-layer detection subproblems, which map onto the 2-layer core with a slight modification using a distance accumulation stage and a post-LLR processing stage. Various architectures are accordingly developed to achieve a desired detection throughput and run-time reconfigurability by time-multiplexing of one or more component cores. The proposed core is applied as well to design an optimal multi-user MIMO detector for LTE. The core occupies an area of 1.58MGE and achieves a throughput of 733 Mbps for 256-QAM when synthesized in 90 nm CMOS.

preprint2014arXiv

Pruned Bit-Reversal Permutations: Mathematical Characterization, Fast Algorithms and Architectures

A mathematical characterization of serially-pruned permutations (SPPs) employed in variable-length permuters and their associated fast pruning algorithms and architectures are proposed. Permuters are used in many signal processing systems for shuffling data and in communication systems as an adjunct to coding for error correction. Typically only a small set of discrete permuter lengths are supported. Serial pruning is a simple technique to alter the length of a permutation to support a wider range of lengths, but results in a serial processing bottleneck. In this paper, parallelizing SPPs is formulated in terms of recursively computing sums involving integer floor and related functions using integer operations, in a fashion analogous to evaluating Dedekind sums. A mathematical treatment for bit-reversal permutations (BRPs) is presented, and closed-form expressions for BRP statistics are derived. It is shown that BRP sequences have weak correlation properties. A new statistic called permutation inliers that characterizes the pruning gap of pruned interleavers is proposed. Using this statistic, a recursive algorithm that computes the minimum inliers count of a pruned BR interleaver (PBRI) in logarithmic time complexity is presented. This algorithm enables parallelizing a serial PBRI algorithm by any desired parallelism factor by computing the pruning gap in lookahead rather than a serial fashion, resulting in significant reduction in interleaving latency and memory overhead. Extensions to 2-D block and stream interleavers, as well as applications to pruned fast Fourier transforms and LTE turbo interleavers, are also presented. Moreover, hardware-efficient architectures for the proposed algorithms are developed. Simulation results demonstrate 3 to 4 orders of magnitude improvement in interleaving time compared to existing approaches.