Source author record

Qiuting Huang

Qiuting Huang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2022arXiv

EC-GSM-IoT Network Synchronization with Support for Large Frequency Offsets

EDGE-based EC-GSM-IoT is a promising candidate for the billion-device cellular IoT (cIoT), providing similar coverage and battery life as NB-IoT. The goal of 20 dB coverage extension compared to EDGE poses significant challenges for the initial network synchronization, which has to be performed well below the thermal noise floor, down to an SNR of -8.5 dB. We present a low-complexity synchronization algorithm supporting up to 50 kHz initial frequency offset, thus enabling the use of a low-cost +/-25 ppm oscillator. The proposed algorithm does not only fulfill the 3GPP requirements, but surpasses them by 3 dB, enabling communication with an SNR of -11.5 dB or a maximum coupling loss of up to 170.5 dB.

preprint2020arXiv

Energy-Efficient Hardware-Accelerated Synchronization for Shared-L1-Memory Multiprocessor Clusters

The steeply growing performance demands for highly power- and energy-constrained processing systems such as end-nodes of the internet-of-things (IoT) have led to parallel near-threshold computing (NTC), joining the energy-efficiency benefits of low-voltage operation with the performance typical of parallel systems. Shared-L1-memory multiprocessor clusters are a promising architecture, delivering performance in the order of GOPS and over 100 GOPS/W of energy-efficiency. However, this level of computational efficiency can only be reached by maximizing the effective utilization of the processing elements (PEs) available in the clusters. Along with this effort, the optimization of PE-to-PE synchronization and communication is a critical factor for performance. In this work, we describe a light-weight hardware-accelerated synchronization and communication unit (SCU) for tightly-coupled clusters of processors. We detail the architecture, which enables fine-grain per-PE power management, and its integration into an eight-core cluster of RISC-V processors. To validate the effectiveness of the proposed solution, we implemented the eight-core cluster in advanced 22nm FDX technology and evaluated performance and energy-efficiency with tunable microbenchmarks and a set of real-life applications and kernels. The proposed solution allows synchronization-free regions as small as 42 cycles, over 41 times smaller than the baseline implementation based on fast test-and-set access to L1 memory when constraining the microbenchmarks to 10% synchronization overhead. When evaluated on the real-life DSP-applications, the proposed SCU improves performance by up to 92% and 23% on average and energy efficiency by up to 98% and 39% on average.

preprint2016arXiv

Maximum-Likelihood Detection for Energy-Efficient Timing Acquisition in NB-IoT

Initial timing acquisition in narrow-band IoT (NB-IoT) devices is done by detecting a periodically transmitted known sequence. The detection has to be done at lowest possible latency, because the RF-transceiver, which dominates downlink power consumption of an NB-IoT modem, has to be turned on throughout this time. Auto-correlation detectors show low computational complexity from a signal processing point of view at the price of a higher detection latency. In contrast a maximum likelihood cross-correlation detector achieves low latency at a higher complexity as shown in this paper. We present a hardware implementation of the maximum likelihood cross-correlation detection. The detector achieves an average detection latency which is a factor of two below that of an auto-correlation method and is able to reduce the required energy per timing acquisition by up to 34%.

preprint2015arXiv

A Low-complexity Channel Shortening Receiver with Diversity Support for Evolved 2G Device

The second generation (2G) cellular networks are the current workhorse for machine-to-machine (M2M) communications. Diversity in 2G devices can be present both in form of multiple receive branches and blind repetitions. In presence of diversity, intersymbol interference (ISI) equalization and co-channel interference (CCI) suppression are usually very complex. In this paper, we consider the improvements for 2G devices with receive diversity. We derive a low-complexity receiver based on a channel shortening filter, which allows to sum up all diversity branches to a single stream after filtering while keeping the full diversity gain. The summed up stream is subsequently processed by a single stream Max-log-MAP (MLM) equalizer. The channel shortening filter is designed to maximize the mutual information lower bound (MILB) with the Ungerboeck detection model. Its filter coefficients can be obtained mainly by means of discrete-Fourier transforms (DFTs). Compared with the state-of-art homomorphic (HOM) filtering based channel shortener which cooperates with a delayed-decision feedback MLM (DDF-MLM) equalizer, the proposed MILB channel shortener has superior performance. Moreover, the equalization complexity, in terms of real-valued multiplications, is decreased by a factor that equals the number of diversity branches.

preprint2014arXiv

A Signal Processor for Gaussian Message Passing

In this paper, we present a novel signal processing unit built upon the theory of factor graphs, which is able to address a wide range of signal processing algorithms. More specifically, the demonstrated factor graph processor (FGP) is tailored to Gaussian message passing algorithms. We show how to use a highly configurable systolic array to solve the message update equations of nodes in a factor graph efficiently. A proper instruction set and compilation procedure is presented. In a recursive least squares channel estimation example we show that the FGP can compute a message update faster than a state-ofthe- art DSP. The results demonstrate the usabilty of the FGP architecture as a flexible HW accelerator for signal-processing and communication systems.