Source author record

Kan Wu

Kan Wu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

9works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

9 published item(s)

preprint2022arXiv

Free-space point-to-multiplepoint optical frequency transfer with lens assisted integrated beam steering

We report on the realization of high-performance silica integrated two-dimensional lens assisted beam-steering (LABS) arrays along with the first-of-their-kind point-to-multiplepoint optical frequency transfer. {The LABS equips with $N$ antennas} and has the capability to produce arbitrary number of output beams with different output angles with the simple control complexity. We demonstrate that the LABS has 16 scanning angles, which can support {the access capability for the maximum of simultaneous 16 user nodes.} The coaxial configuration for transmitting and receiving the light as a monolithic transceiver allows us to reduce the out-of-loop phase noise significantly. Finally, the LABS-based non-blocking point-to-multiplepoint in-door free-space optical frequency transfer links with 24 m and 50 m free-space links are shown. After being compensated for the free-space link up to 50 m, the fractional frequency instability of $4.5\times10^{-17}$ and $7.7\times10^{-20}$ at the averaging time of 1 s and 20,000 s, respectively, can be achieved. The present work proves the potential application of the 2D LABS in free-space optical time-frequency transfer and provides a guidance for developing a chip-scale optical time-frequency transfer system.

preprint2022arXiv

MiniViT: Compressing Vision Transformers with Weight Multiplexing

Vision Transformer (ViT) models have recently drawn much attention in computer vision due to their high model capability. However, ViT models suffer from huge number of parameters, restricting their applicability on devices with limited memory. To alleviate this problem, we propose MiniViT, a new compression framework, which achieves parameter reduction in vision transformers while retaining the same performance. The central idea of MiniViT is to multiplex the weights of consecutive transformer blocks. More specifically, we make the weights shared across layers, while imposing a transformation on the weights to increase diversity. Weight distillation over self-attention is also applied to transfer knowledge from large-scale ViT models to weight-multiplexed compact models. Comprehensive experiments demonstrate the efficacy of MiniViT, showing that it can reduce the size of the pre-trained Swin-B transformer by 48\%, while achieving an increase of 1.0\% in Top-1 accuracy on ImageNet. Moreover, using a single-layer of parameters, MiniViT is able to compress DeiT-B by 9.7 times from 86M to 9M parameters, without seriously compromising the performance. Finally, we verify the transferability of MiniViT by reporting its performance on downstream benchmarks. Code and models are available at here.

preprint2022arXiv

TinyViT: Fast Pretraining Distillation for Small Vision Transformers

Vision transformer (ViT) recently has drawn great attention in computer vision due to its remarkable model capability. However, most prevailing ViT models suffer from huge number of parameters, restricting their applicability on devices with limited resources. To alleviate this issue, we propose TinyViT, a new family of tiny and efficient small vision transformers pretrained on large-scale datasets with our proposed fast distillation framework. The central idea is to transfer knowledge from large pretrained models to small ones, while enabling small models to get the dividends of massive pretraining data. More specifically, we apply distillation during pretraining for knowledge transfer. The logits of large teacher models are sparsified and stored in disk in advance to save the memory cost and computation overheads. The tiny student transformers are automatically scaled down from a large pretrained model with computation and parameter constraints. Comprehensive experiments demonstrate the efficacy of TinyViT. It achieves a top-1 accuracy of 84.8% on ImageNet-1k with only 21M parameters, being comparable to Swin-B pretrained on ImageNet-21k while using 4.2 times fewer parameters. Moreover, increasing image resolutions, TinyViT can reach 86.5% accuracy, being slightly better than Swin-L while using only 11% parameters. Last but not the least, we demonstrate a good transfer ability of TinyViT on various downstream tasks. Code and models are available at https://github.com/microsoft/Cream/tree/main/TinyViT.

preprint2016arXiv

Analysis and Approximation of Dual Tandem Queues with Finite Buffer Capacity

Tandem queues with finite buffer capacity commonly exist in practical applications. By viewing a tandem queue as an integrated system, an innovative approach has been developed to analyze its performance through the insight from reduction method. In our approach, the starvation at the bottleneck caused by service time randomness is modeled and captured by interruptions. Fundamental properties of tandem queues with finite buffer capacity are examined. We show that in general system service rate of a dual tandem queue with finite buffer capacity is equal or smaller than its bottleneck service rate, and virtual interruptions, which are the extra idle period at the bottleneck caused by the non-bottlenecks, depend on arrival rates. Hence, system service rate is a function of arrival rate when the buffer capacity of a tandem queue is finite. Approximation for the mean queue time of a dual tandem queue has been developed through the concept of virtual interruptions.

preprint2014arXiv

A Unified View on Planning, Scheduling and Dispatching in Production Systems

Planning, scheduling and dispatching play critical roles in the operations of a supply chain. Their definitions are clearly given through a unified view in this paper. The distinction between planning and scheduling is analyzed from the view point of microeconomics and queueing theory. The distinction between scheduling and dispatching is analyzed from the view point of computational complexity and hierarchical decomposition. Based on the elasticity of price and capacity, planning can be separated into demand planning or capacity planning. Scheduling period is the time horizon where price and average production cost are insensitive to the production rate. The critical roles of the master production schedule and move targets in job scheduling have been explained through the concept of hierarchical decomposition. Dispatching is the last layer of job scheduling in the hierarchical decomposition. The advantage of pull and push systems has been compared and analyzed systematically.

preprint2014arXiv

Effects of carbon nanotubes and graphene oxide absorbers on the noise of mode-locked fiber lasers

Phase noise is very important for the ultrafast pulse application in telecommunication, ultrafast diagnose, material science, and biology. In this paper, two types of carbon nano-materials, single-wall carbon nanotube and graphene oxide, are investigated for noise suppression in ultrafast photonics. Various properties of the wall-paper SAs, such as saturable intensity, optical absorption and degree of purity, are found to be key factors determining the phase noise of the ultrafast pulses. A reduced-noise femtosecond fiber laser is experimentally demonstrated by optimizing the above parameters of carbon material based SAs. The phase noise reduction more than 10 dB at 10 kHz can be obtained in the experiments. To our knowledge, this is the first time that the relationship between different carbon material based SAs and the phase noise of mode-locked lasers has been investigated. This work will pave the way to get a high-quality ultrashort pulse in passively mode-locked fiber lasers.

preprint2014arXiv

Measurement of a topological edge invariant in a microwave network

We report on the measurement of topological invariants in an electromagnetic topological insulator analog formed by a microwave network, consisting of the winding numbers of scattering matrix eigenvalues. The experiment can be regarded as a variant of a topological pump, with non-zero winding implying the existence of topological edge states. In microwave networks, unlike most other systems exhibiting topological insulator physics, the winding can be directly observed. The effects of loss on the experimental results, and on the topological edge states, is discussed.

preprint2013arXiv

Computing matrix inversion with optical networks

With this paper we bring about a discussion on the computing potential of complex optical networks and provide experimental demonstration that an optical fiber network can be used as an analog processor to calculate matrix inversion. A 3x3 matrix is inverted as a proof-of-concept demonstration using a fiber network containing three nodes and operating at telecomm wavelength. For an NxN matrix, the overall solving time (including setting time of the matrix elements and calculation time of inversion) scales as O(N^2), whereas matrix inversion by most advanced computer algorithms requires ~O(N^2.37) computational time. For well-conditioned matrices, the error of the inversion performed optically is found to be less than 3%, limited by the accuracy of measurement equipment.

preprint2010arXiv

Phase Noise and Intensity Noise of the Pulse Train Generated from Mode-locked Lasers in the Demodulation Measurement

The phase noise and intensity noise of a pulse train are theoretically analyzed in the demodulation measurement. The effect of pulse asymmetry is discussed for the first time using Fourier series. Experimentally, photodetectors with different bandwidth and incident power levels are compared to achieve minimum pulse distortion.