Catalog footprint

What is connected

151works
62topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

151 published item(s)

preprint2026arXiv

AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents

GPU kernel optimization is increasingly critical for efficient deep learning systems, but writing high-performance kernels still requires substantial low-level expertise. Recent AI coding agents can iteratively read code, invoke compilers and profilers, and refine implementations, yet existing kernel benchmarks evaluate single LLM calls rather than full agent workflows, and none include both kernel-to-kernel optimization and unseen-configuration generalization testing. We present AgentKernelArena, an open-source benchmark for measuring AI coding agents on GPU kernel optimization. The benchmark contains 196 tasks spanning HIP-to-HIP optimization, Triton-to-Triton optimization, and PyTorch-to-HIP translation, and evaluates complete agent workflows in isolated workspaces using gated compilation, correctness, and performance checks, centralized scoring and an unseen-configuration generalization protocol that tests whether optimizations transfer to input configurations the agent never observed. Across production agents including Cursor Agent, Claude Code, and Codex Agent, we find near-perfect compilation and high correctness rates on most task categories, with the strongest configurations achieving mean speedups of up to 6.89x on PyTorch-to-HIP, 6.69x on HIP-to-HIP, and 2.13x on Triton-to-Triton tasks. Our unseen-configuration evaluation shows that HIP-to-HIP and Triton-to-Triton optimizations largely transfer to unseen input shapes, while PyTorch-to-HIP exhibits substantial correctness drops, indicating that agents generating kernels from scratch frequently hardcode shape-specific assumptions. AgentKernelArena is designed as a modular, extensible framework for rigorous evaluation of agentic GPU kernel optimization across agents, tasks, and hardware targets.

preprint2026arXiv

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However, existing methods usually follow an output-as-input latent paradigm and yield unstable gains. We identify evidence for a feature-space mismatch that can contribute to this instability: dominant visual-latent models build on pre-norm MLLMs and reuse decoder hidden states as predicted latent inputs, even though these states occupy a substantially different norm regime from the input embeddings the model was trained to consume~\citep{xie2025mhc,li2026siamesenorm,team2026attention}. This mismatch can make direct latent feedback unreliable. Motivated by this diagnosis, we propose \textbf{GAP}, a \textbf{G}ranular \textbf{A}lignment \textbf{P}aradigm for visual latent modeling. GAP aligns visual latent reasoning at three levels: feature-level alignment maps decoder outputs into input-compatible visual latents through a lightweight PCA-aligned latent head; context-level alignment grounds latent targets with inspectable auxiliary visual supervision; and capacity-guided alignment assigns latent supervision selectively to examples where the base MLLM struggles. On Qwen2.5-VL 7B, the resulting model achieves the best mean aggregate perception and reasoning performance among our supervised variants. Inference-time intervention probing further suggests that generated latents provide task-relevant visual signal beyond merely adding token slots.

preprint2025arXiv

FDP: A Frequency-Decomposition Preprocessing Pipeline for Unsupervised Anomaly Detection in Brain MRI

Due to the diversity of brain anatomy and the scarcity of annotated data, supervised anomaly detection for brain MRI remains challenging, driving the development of unsupervised anomaly detection (UAD) approaches. Current UAD methods typically utilize artificially generated noise perturbations on healthy MRIs to train generative models for normal anatomy reconstruction, enabling anomaly detection via residual maps. However, such simulated anomalies lack the biophysical fidelity and morphological complexity characteristic of true clinical lesions. To advance UAD in brain MRI, we conduct the first systematic frequency-domain analysis of pathological signatures, revealing two key properties: (1) anomalies exhibit unique frequency patterns distinguishable from normal anatomy, and (2) low-frequency signals maintain consistent representations across healthy scans. These insights motivate our Frequency-Decomposition Preprocessing (FDP) framework, the first UAD method to leverage frequency-domain reconstruction for simultaneous pathology suppression and anatomical preservation. FDP can integrate seamlessly with existing anomaly simulation techniques, consistently enhancing detection performance across diverse architectures while maintaining diagnostic fidelity. Experimental results demonstrate that FDP consistently improves anomaly detection performance when integrated with existing methods. Notably, FDP achieves a 17.63% increase in DICE score with LDM while maintaining robust improvements across multiple baselines. The code is available at https://github.com/ls1rius/MRI_FDP.

preprint2024arXiv

3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction.

preprint2024arXiv

Deep peak property learning for efficient chiral molecules ECD spectra prediction

Chiral molecule assignation is crucial for asymmetric catalysis, functional materials, and the drug industry. The conventional approach requires theoretical calculations of electronic circular dichroism (ECD) spectra, which is time-consuming and costly. To speed up this process, we have incorporated deep learning techniques for the ECD prediction. We first set up a large-scale dataset of Chiral Molecular ECD spectra (CMCDS) with calculated ECD spectra. We further develop the ECDFormer model, a Transformer-based model to learn the chiral molecular representations and predict corresponding ECD spectra with improved efficiency and accuracy. Unlike other models for spectrum prediction, our ECDFormer creatively focused on peak properties rather than the whole spectrum sequence for prediction, inspired by the scenario of chiral molecule assignation. Specifically, ECDFormer predicts the peak properties, including number, position, and symbol, then renders the ECD spectra from these peak properties, which significantly outperforms other models in ECD prediction, Our ECDFormer reduces the time of acquiring ECD spectra from 1-100 hours per molecule to 1.5s.

preprint2024arXiv

Jet Schemes, Quantum Dilogarithm and Feigin-Stoyanovsky's Principal Subspaces

We analyze the structure of Feigin-Stoyanovsky's principal subspaces of affine Lie algebra from the jet algebra viewpoint. For type $A$ level one principal subspaces, we show that their shifted multi-graded Hilbert series can be expressed either using the quantum dilogarithm or as certain generating functions ``counting" finite-dimensional representations of $A$-type quivers. This notably results in novel fermionic character formulas for these principal subspaces. Moreover, our result implies that all level one principal subspaces of type $A$ are ``classically free" as vertex algebras. We also analyze infinite jet algebras associated to principal subspaces of affine vertex algebras $L_{1}(\mathfrak{so}_5)$, $L_{1}(\mathfrak{so}_8)$ and $L_1(\frak{g}_2)$. We derive a new character formula for the principal subspace of $L_1(\mathfrak{so}_5)$, proving that it is classically free, and present evidence that the principal subspaces of $L_1(\mathfrak{so}_8)$ and of $L_1(\frak{g}_2)$ are also classically free.

preprint2023arXiv

Improving photon number resolvability of a superconducting nanowire detector array using a level comparator circuit

Photon number resolving (PNR) capability is very important in many optical applications, including quantum information processing, fluorescence detection, and few-photon-level ranging and imaging. Superconducting nanowire single-photon detectors (SNSPDs) with a multipixel interleaved architecture give the array an excellent spatial PNR capability. However, the signal-to-noise ratio (SNR) of the photon number resolution (SNRPNR) of the array will be degraded with increasing the element number due to the electronic noise in the readout circuit, which limits the PNR resolution as well as the maximum PNR number. In this study, a 16-element interleaved SNSPD array was fabricated, and the PNR capability of the array was investigated and analyzed. By introducing a level comparator circuit (LCC), the SNRPNR of the detector array was improved over a factor of four. In addition, we performed a statistical analysis of the photon number on this SNSPD array with LCC, showing that the LCC method effectively enhances the PNR resolution. Besides, the system timing jitter of the detector was reduced from 90 ps to 72 ps due to the improved electrical SNR.

preprint2023arXiv

OccluMix: Towards De-Occlusion Virtual Try-on by Semantically-Guided Mixup

Image Virtual try-on aims at replacing the cloth on a personal image with a garment image (in-shop clothes), which has attracted increasing attention from the multimedia and computer vision communities. Prior methods successfully preserve the character of clothing images, however, occlusion remains a pernicious effect for realistic virtual try-on. In this work, we first present a comprehensive analysis of the occlusions and categorize them into two aspects: i) Inherent-Occlusion: the ghost of the former cloth still exists in the try-on image; ii) Acquired-Occlusion: the target cloth warps to the unreasonable body part. Based on the in-depth analysis, we find that the occlusions can be simulated by a novel semantically-guided mixup module, which can generate semantic-specific occluded images that work together with the try-on images to facilitate training a de-occlusion try-on (DOC-VTON) framework. Specifically, DOC-VTON first conducts a sharpened semantic parsing on the try-on person. Aided by semantics guidance and pose prior, various complexities of texture are selectively blending with human parts in a copy-and-paste manner. Then, the Generative Module (GM) is utilized to take charge of synthesizing the final try-on image and learning to de-occlusion jointly. In comparison to the state-of-the-art methods, DOC-VTON achieves better perceptual quality by reducing occlusion effects.

preprint2022arXiv

A photon counting reconstructive spectrometer combining metasurfaces and superconducting nanowire single-photon detectors

Faint light spectroscopy has many important applications such as fluorescence spectroscopy, lidar and astronomical observations. However, long measurement time limit its application on real-time measurement. In this work, a photon counting reconstructive spectrometer combining metasurfaces and superconducting nanowire single photon detectors (SNSPDs) was proposed. A prototype device was fabricated on a silicon on isolator (SOI) substrate, and its performance was characterized. Experiment results show that this device support spectral reconstruction of mono-color lights with a resolution of 2 nm in the wavelength region of 1500 nm ~ 1600 nm. The detection efficiency of this device is 1.4% ~ 3.2% in this wavelength region. The measurement time required by this photon counting reconstructive spectrometer was also investigated experimentally, showing its potential to be applied in the scenarios requiring real-time measurement.

preprint2022arXiv

An Efficient Training Approach for Very Large Scale Face Recognition

Face recognition has achieved significant progress in deep learning era due to the ultra-large-scale and welllabeled datasets. However, training on the outsize datasets is time-consuming and takes up a lot of hardware resource. Therefore, designing an efficient training approach is indispensable. The heavy computational and memory costs mainly result from the million-level dimensionality of thefully connected (FC) layer. To this end, we propose a novel training approach, termed Faster Face Classification (F2C), to alleviate time and cost without sacrificing the performance. This method adopts Dynamic Class Pool (DCP) for storing and updating the identities features dynamically, which could be regarded as a substitute for the FC layer. DCP is efficiently time-saving and cost-saving, as its smaller size with the independence from the whole face identities together. We further validate the proposed F2C method across several face benchmarks and private datasets, and display comparable results, meanwhile the speed is faster than state-of-the-art FC-based methods in terms of recognition accuracy and hardware costs. Moreover, our method is further improved by a well-designed dual data loader including indentity-based and instancebased loaders, which makes it more efficient for the updating DCP parameters.

preprint2022arXiv

An Empirical Study of Yanked Releases in the Rust Package Registry

Cargo, the software packaging manager of Rust, provides a yank mechanism to support release-level deprecation, which can prevent packages from depending on yanked releases. Most prior studies focused on code-level (i.e., deprecated APIs) and package-level deprecation (i.e., deprecated packages). However, few studies have focused on release-level deprecation. In this study, we investigate how often and how the yank mechanism is used, the rationales behind its usage, and the adoption of yanked releases in the Cargo ecosystem. Our study shows that 9.6% of the packages in Cargo have at least one yanked release, and the proportion of yanked releases kept increasing from 2014 to 2020. Package owners yank releases for other reasons than withdrawing a defective release, such as fixing a release that does not follow semantic versioning or indicating a package is removed or replaced. In addition, we found that 46% of the packages directly adopted at least one yanked release and the yanked releases propagated through the dependency network, which leads to 1.4% of the releases in the ecosystem having unresolved dependencies.

preprint2022arXiv

An Empirical Study on Distribution Shift Robustness From the Perspective of Pre-Training and Data Augmentation

The performance of machine learning models under distribution shift has been the focus of the community in recent years. Most of current methods have been proposed to improve the robustness to distribution shift from the algorithmic perspective, i.e., designing better training algorithms to help the generalization in shifted test distributions. This paper studies the distribution shift problem from the perspective of pre-training and data augmentation, two important factors in the practice of deep learning that have not been systematically investigated by existing work. By evaluating seven pre-trained models, including ResNets and ViT's with self-supervision and supervision mode, on five important distribution-shift datasets, from WILDS and DomainBed benchmarks, with five different learning algorithms, we provide the first comprehensive empirical study focusing on pre-training and data augmentation. With our empirical result obtained from 1,330 models, we provide the following main observations: 1) ERM combined with data augmentation can achieve state-of-the-art performance if we choose a proper pre-trained model respecting the data property; 2) specialized algorithms further improve the robustness on top of ERM when handling a specific type of distribution shift, e.g., GroupDRO for spurious correlation and CORAL for large-scale out-of-distribution data; 3) Comparing different pre-training modes, architectures and data sizes, we provide novel observations about pre-training on distribution shift, which sheds light on designing or selecting pre-training strategy for different kinds of distribution shifts. In summary, our empirical study provides a comprehensive baseline for a wide range of pre-training models fine-tuned with data augmentation, which potentially inspires research exploiting the power of pre-training and data augmentation in the future of distribution shift study.

preprint2022arXiv

Antiferromagnetic structure and magnetic properties of Dy2O2Te: An isostructural analog of the rare-earth superconductors R2O2Bi

The rare-earth compounds R2O2Bi (R=Tb, Dy, Er, Lu, Y) are newly discovered superconductors in the vicinity of a rare-earth magnetic long-range order. In this work, we determine the magnetic order of the parent compound Dy2O2Te by neutron scattering as the A-type antiferromagnetic structure below the Néel temperature TN=9.7K. The large staggered magnetic moment 9.4(1) μB per Dy at T=3.5K lies in the basal ab plane. In a magnetic field, anomalous magnetic properties including the bifurcation between zero-field- and field-cooling magnetization, a butterfly-shaped magnetic hysteresis, and slow magnetic relaxation emerge, which are related to the field-induced metamagnetic transitions in Dy2O2Te. Our experimental findings could stimulate further research on the relation between antiferromagnetism and superconductivity in these rare-earth compounds.

preprint2022arXiv

Cats: Complementary CNN and Transformer Encoders for Segmentation

Recently, deep learning methods have achieved state-of-the-art performance in many medical image segmentation tasks. Many of these are based on convolutional neural networks (CNNs). For such methods, the encoder is the key part for global and local information extraction from input images; the extracted features are then passed to the decoder for predicting the segmentations. In contrast, several recent works show a superior performance with the use of transformers, which can better model long-range spatial dependencies and capture low-level details. However, transformer as sole encoder underperforms for some tasks where it cannot efficiently replace the convolution based encoder. In this paper, we propose a model with double encoders for 3D biomedical image segmentation. Our model is a U-shaped CNN augmented with an independent transformer encoder. We fuse the information from the convolutional encoder and the transformer, and pass it to the decoder to obtain the results. We evaluate our methods on three public datasets from three different challenges: BTCV, MoDA and Decathlon. Compared to the state-of-the-art models with and without transformers on each task, our proposed method obtains higher Dice scores across the board.

preprint2022arXiv

CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation

Unsupervised domain adaptation (UDA) aims to transfer knowledge learned from a labeled source domain to a different unlabeled target domain. Most existing UDA methods focus on learning domain-invariant feature representation, either from the domain level or category level, using convolution neural networks (CNNs)-based frameworks. One fundamental problem for the category level based UDA is the production of pseudo labels for samples in target domain, which are usually too noisy for accurate domain alignment, inevitably compromising the UDA performance. With the success of Transformer in various tasks, we find that the cross-attention in Transformer is robust to the noisy input pairs for better feature alignment, thus in this paper Transformer is adopted for the challenging UDA task. Specifically, to generate accurate input pairs, we design a two-way center-aware labeling algorithm to produce pseudo labels for target samples. Along with the pseudo labels, a weight-sharing triple-branch transformer framework is proposed to apply self-attention and cross-attention for source/target feature learning and source-target domain alignment, respectively. Such design explicitly enforces the framework to learn discriminative domain-specific and domain-invariant representations simultaneously. The proposed method is dubbed CDTrans (cross-domain transformer), and it provides one of the first attempts to solve UDA tasks with a pure transformer solution. Experiments show that our proposed method achieves the best performance on public UDA datasets, e.g. VisDA-2017 and DomainNet. Code and models are available at https://github.com/CDTrans/CDTrans.

preprint2022arXiv

CGAR: Critic Guided Action Redistribution in Reinforcement Leaning

Training a game-playing reinforcement learning agent requires multiple interactions with the environment. Ignorant random exploration may cause a waste of time and resources. It's essential to alleviate such waste. As discussed in this paper, under the settings of the off-policy actor critic algorithms, we demonstrate that the critic can bring more expected discounted rewards than or at least equal to the actor. Thus, the Q value predicted by the critic is a better signal to redistribute the action originally sampled from the policy distribution predicted by the actor. This paper introduces the novel Critic Guided Action Redistribution (CGAR) algorithm and tests it on the OpenAI MuJoCo tasks. The experimental results demonstrate that our method improves the sample efficiency and achieves state-of-the-art performance. Our code can be found at https://github.com/tairanhuang/CGAR.

preprint2022arXiv

Chip-scale Spontaneous Quasi-Phase-Matched Micro-Racetrack Resonator

Due to their capacity for non-classical light generation, high-efficiency second-order nonlinear parametric processes play an important role in quantum photonic technology, and chip-scale realization of these processes is recognized as the key to building efficient light sources for integrated quantum photonic circuits. To achieve ultra-high nonlinear conversion efficiency, traditional method uses quasi-phase matching (QPM) technology. However, QPM requires electric field poling, which is incompatible with the CMOS fabrication process, and this hinders the wafer-scale production of integrated photonic circuits. In this paper, we demonstrate efficient spontaneous quasi-phase matched (SQPM) frequency conversion in a micro-racetrack resonator. Our approach does not involve poling, but exploits the anisotropy of the ferroelectric crystals to allow the phase-matching condition to be fulfilled spontaneously as the TE-polarized light circulates in a specifically designed racetrack resonator. SQPM second harmonic generation is observed with a normalized intracavity conversion efficiency of 0.85%/W, corresponding to the 111st-order QPM. This could theoretically reach 186,000%/W by first-order QPM. In this case such high intracavity conversion efficiency can be implemented in practice with an optimized outward coupling. Our configurable SQPM approach will benefit the application of nonlinear frequency conversion in chip-scale integrated photonics with CMOS-compatible fabrication processes, and is applicable to other on-chip nonlinear processes such as quantum frequency conversion or frequency-comb generation.

preprint2022arXiv

Correlating exciton coherence length, localization, and its optical lineshape. I. a finite temperature solution of the Davydov soliton model

The lineshape of spectroscopic transitions offer windows into the local environment of a system. Here, we present a novel approach for connecting the lineshape of a molecular exciton to finite-temperature lattice vibrations within the context of the Davydov soliton model (A. S. Davydov and N. I. Kislukha, Phys. Stat. Sol. {\bf 59},465(1973)). Our results are based upon a numerically exact, self-consistent treatment of the model in which thermal effects are introduced as fluctuations about the zero-temperature localized soliton state. We find that both the energy fluctuations and the localization can be described in terms of a parameter-free, reduced description by introducing a critical temperature below which exciton self-trapping is expected to be stable. Above this temperature, the self-consistent ansatz relating the lattice distortion to the exciton wavefunction breaks down. Our theoretical model coorelates well with both experimental observations on molecular J-aggregate and resolves one of the critical issues concerning the finite temperture stability of soliton states in alpha-helices and protein peptide chains.

preprint2022arXiv

Criteria Comparative Learning for Real-scene Image Super-Resolution

Real-scene image super-resolution aims to restore real-world low-resolution images into their high-quality versions. A typical RealSR framework usually includes the optimization of multiple criteria which are designed for different image properties, by making the implicit assumption that the ground-truth images can provide a good trade-off between different criteria. However, this assumption could be easily violated in practice due to the inherent contrastive relationship between different image properties. Contrastive learning (CL) provides a promising recipe to relieve this problem by learning discriminative features using the triplet contrastive losses. Though CL has achieved significant success in many computer vision tasks, it is non-trivial to introduce CL to RealSR due to the difficulty in defining valid positive image pairs in this case. Inspired by the observation that the contrastive relationship could also exist between the criteria, in this work, we propose a novel training paradigm for RealSR, named Criteria Comparative Learning (Cria-CL), by developing contrastive losses defined on criteria instead of image patches. In addition, a spatial projector is proposed to obtain a good view for Cria-CL in RealSR. Our experiments demonstrate that compared with the typical weighted regression strategy, our method achieves a significant improvement under similar parameter settings.

preprint2022arXiv

Cross Vision-RF Gait Re-identification with Low-cost RGB-D Cameras and mmWave Radars

Human identification is a key requirement for many applications in everyday life, such as personalized services, automatic surveillance, continuous authentication, and contact tracing during pandemics, etc. This work studies the problem of cross-modal human re-identification (ReID), in response to the regular human movements across camera-allowed regions (e.g., streets) and camera-restricted regions (e.g., offices) deployed with heterogeneous sensors. By leveraging the emerging low-cost RGB-D cameras and mmWave radars, we propose the first-of-its-kind vision-RF system for cross-modal multi-person ReID at the same time. Firstly, to address the fundamental inter-modality discrepancy, we propose a novel signature synthesis algorithm based on the observed specular reflection model of a human body. Secondly, an effective cross-modal deep metric learning model is introduced to deal with interference caused by unsynchronized data across radars and cameras. Through extensive experiments in both indoor and outdoor environments, we demonstrate that our proposed system is able to achieve ~92.5% top-1 accuracy and ~97.5% top-5 accuracy out of 56 volunteers. We also show that our proposed system is able to robustly reidentify subjects even when multiple subjects are present in the sensors' field of view.

preprint2022arXiv

Detection of Flare-induced Plasma Flows in the Corona of EV Lac with X-ray Spectroscopy

Stellar flares are characterized by sudden enhancement of electromagnetic radiation from the atmospheres of stars. Compared to their solar counterparts, our knowledge on the coronal plasma dynamics of stellar flares and their connection to coronal mass ejections (CMEs) remains very limited. With time-resolved high-resolution spectroscopic observations from the \textit{Chandra} X-ray observatory, we detected noticeable coronal plasma flows during several stellar flares on a nearby dMe star EV Lac. In the observed spectra of O~{\sc{viii}} (3 MK), Fe~{\sc{xvii}} (6 MK), Mg~{\sc{xii}} (10 MK), and Si~{\sc{xiv}} (16 MK) lines, these flare-induced upflows/downflows appear as significant Doppler shifts of several tens to \speed{130}, and the upflow velocity generally increases with temperature. Variable line ratios of the Si~{\sc{xiii}} triplet reveal that these plasma flows in most flares are accompanied by an increase of the coronal plasma density and temperature. We interpret these results as X-ray evidences for chromospheric evaporation on EV Lac. In two successive flares, the plasma flow pattern and a sharp increase of the measured coronal density are highly suggestive of explosive evaporation. The transition from redshifts to blueshifts in such an explosive evaporation occurs at a temperature of at least 10 MK, much higher than that observed in solar flares ($\sim$1 MK). However, in one flare the cool and warm upflows appear to be accompanied by a decreasing plasma density, which might be explained by a stellar filament/prominence eruption coupled to this flare. These results provide important clues to understand the coronal plasma dynamics during flares on M dwarfs.

preprint2022arXiv

DLME: Deep Local-flatness Manifold Embedding

Manifold learning (ML) aims to seek low-dimensional embedding from high-dimensional data. The problem is challenging on real-world datasets, especially with under-sampling data, and we find that previous methods perform poorly in this case. Generally, ML methods first transform input data into a low-dimensional embedding space to maintain the data's geometric structure and subsequently perform downstream tasks therein. The poor local connectivity of under-sampling data in the former step and inappropriate optimization objectives in the latter step leads to two problems: structural distortion and underconstrained embedding. This paper proposes a novel ML framework named Deep Local-flatness Manifold Embedding (DLME) to solve these problems. The proposed DLME constructs semantic manifolds by data augmentation and overcomes the structural distortion problem using a smoothness constrained based on a local flatness assumption about the manifold. To overcome the underconstrained embedding problem, we design a loss and theoretically demonstrate that it leads to a more suitable embedding based on the local flatness. Experiments on three types of datasets (toy, biological, and image) for various downstream tasks (classification, clustering, and visualization) show that our proposed DLME outperforms state-of-the-art ML and contrastive learning methods.

preprint2022arXiv

DnSwin: Toward Real-World Denoising via Continuous Wavelet Sliding-Transformer

Real-world image denoising is a practical image restoration problem that aims to obtain clean images from in-the-wild noisy inputs. Recently, the Vision Transformer (ViT) has exhibited a strong ability to capture long-range dependencies, and many researchers have attempted to apply the ViT to image denoising tasks. However, a real-world image is an isolated frame that makes the ViT build long-range dependencies based on the internal patches, which divides images into patches, disarranges noise patterns and damages gradient continuity. In this article, we propose to resolve this issue by using a continuous Wavelet Sliding-Transformer that builds frequency correspondences under real-world scenes, called DnSwin. Specifically, we first extract the bottom features from noisy input images by using a convolutional neural network (CNN) encoder. The key to DnSwin is to extract high-frequency and low-frequency information from the observed features and build frequency dependencies. To this end, we propose a Wavelet Sliding-Window Transformer (WSWT) that utilizes the discrete wavelet transform (DWT), self-attention and the inverse DWT (IDWT) to extract deep features. Finally, we reconstruct the deep features into denoised images using a CNN decoder. Both quantitative and qualitative evaluations conducted on real-world denoising benchmarks demonstrate that the proposed DnSwin performs favorably against the state-of-the-art methods.

preprint2022arXiv

Dynamic Gradient Reactivation for Backward Compatible Person Re-identification

We study the backward compatible problem for person re-identification (Re-ID), which aims to constrain the features of an updated new model to be comparable with the existing features from the old model in galleries. Most of the existing works adopt distillation-based methods, which focus on pushing new features to imitate the distribution of the old ones. However, the distillation-based methods are intrinsically sub-optimal since it forces the new feature space to imitate the inferior old feature space. To address this issue, we propose the Ranking-based Backward Compatible Learning (RBCL), which directly optimizes the ranking metric between new features and old features. Different from previous methods, RBCL only pushes the new features to find best-ranking positions in the old feature space instead of strictly alignment, and is in line with the ultimate goal of backward retrieval. However, the sharp sigmoid function used to make the ranking metric differentiable also incurs the gradient vanish issue, therefore stems the ranking refinement during the later period of training. To address this issue, we propose the Dynamic Gradient Reactivation (DGR), which can reactivate the suppressed gradients by adding dynamic computed constant during forward step. To further help targeting the best-ranking positions, we include the Neighbor Context Agents (NCAs) to approximate the entire old feature space during training. Unlike previous works which only test on the in-domain settings, we make the first attempt to introduce the cross-domain settings (including both supervised and unsupervised), which are more meaningful and difficult. The experimental results on all five settings show that the proposed RBCL outperforms previous state-of-the-art methods by large margins under all settings.

preprint2022arXiv

Efficient force field and energy emulation through partition of permutationally equivalent atoms

Gaussian process (GP) emulator has been used as a surrogate model for predicting force field and molecular potential, to overcome the computational bottleneck of molecular dynamics simulation. Integrating both atomic force and energy in predictions was found to be more accurate than using energy alone, yet it requires $O((NM)^3)$ computational operations for computing the likelihood function and making predictions, where $N$ is the number of atoms and $M$ is the number of simulated configurations in the training sample, due to the inversion of a large covariance matrix. The large computational need limits its applications to emulating simulation of small molecules. The computational challenge of using both gradient information and function values in GPs was recently noticed in statistics and machine learning communities, where conventional approximation methods, such as the low rank decomposition or sparse approximation, may not work well. Here we introduce a new approach, the atomized force field (AFF) model, that integrates both force and energy in the emulator with many fewer computational operations. The drastic reduction on computation is achieved by utilizing the naturally sparse structure of the covariance satisfying the constraints of the energy conservation and permutation symmetry of atoms. The efficient machine learning algorithm extends the limits of its applications on larger molecules under the same computational budget, with nearly no loss of predictive accuracy. Furthermore, our approach contains uncertainty assessment of predictions of atomic forces and potentials, useful for developing a sequential design over the chemical input space, with almost no increase in computational cost.

preprint2022arXiv

ELUCID VII: Using Constrained Hydro Simulations to Explore the Gas Component of the Cosmic Web

Using reconstructed initial conditions in the SDSS survey volume, we carry out constrained hydrodynamic simulations in three regions representing different types of the cosmic web: the Coma cluster of galaxies; the SDSS great wall; and a large low-density region at $z\sim 0.05$. These simulations, which include star formation and stellar feedback but no AGN formation and feedback, are used to investigate the properties and evolution of intergalactic and intra-cluster media. About half of the warm-hot intergalactic gas is associated with filaments in the local cosmic web. Gas in the outskirts of massive filaments and halos can be heated significantly by accretion shocks generated by mergers of filaments and halos, respectively, and there is a tight correlation between gas temperature and the strength of the local tidal field. The simulations also predict some discontinuities associated with shock fronts and contact edges, which can be tested using observations of the thermal SZ effect and X-rays. A large fraction of the sky is covered by Ly$α$ and OVI absorption systems, and most of the OVI systems and low-column density HI systems are associated with filaments in the cosmic web. The constrained simulations, which follow the formation and heating history of the observed cosmic web, provide an important avenue to interpret observational data. With full information about the origin and location of the cosmic gas to be observed, such simulations can also be used to develop observational strategies.

preprint2022arXiv

EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose Estimation

Locating 3D objects from a single RGB image via Perspective-n-Points (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, so that 2D-3D point correspondences can be partly learned by backpropagating the gradient w.r.t. object pose. Yet, learning the entire set of unrestricted 2D-3D points from scratch fails to converge with existing approaches, since the deterministic pose is inherently non-differentiable. In this paper, we propose the EPro-PnP, a probabilistic PnP layer for general end-to-end pose estimation, which outputs a distribution of pose on the SE(3) manifold, essentially bringing categorical Softmax to the continuous domain. The 2D-3D coordinates and corresponding weights are treated as intermediate variables learned by minimizing the KL divergence between the predicted and target pose distribution. The underlying principle unifies the existing approaches and resembles the attention mechanism. EPro-PnP significantly outperforms competitive baselines, closing the gap between PnP-based method and the task-specific leaders on the LineMOD 6DoF pose estimation and nuScenes 3D object detection benchmarks.

preprint2022arXiv

GiraffeDet: A Heavy-Neck Paradigm for Object Detection

In conventional object detection frameworks, a backbone body inherited from image recognition models extracts deep latent features and then a neck module fuses these latent features to capture information at different scales. As the resolution in object detection is much larger than in image recognition, the computational cost of the backbone often dominates the total inference cost. This heavy-backbone design paradigm is mostly due to the historical legacy when transferring image recognition models to object detection rather than an end-to-end optimized design for object detection. In this work, we show that such paradigm indeed leads to sub-optimal object detection models. To this end, we propose a novel heavy-neck paradigm, GiraffeDet, a giraffe-like network for efficient object detection. The GiraffeDet uses an extremely lightweight backbone and a very deep and large neck module which encourages dense information exchange among different spatial scales as well as different levels of latent semantics simultaneously. This design paradigm allows detectors to process the high-level semantic information and low-level spatial information at the same priority even in the early stage of the network, making it more effective in detection tasks. Numerical evaluations on multiple popular object detection benchmarks show that GiraffeDet consistently outperforms previous SOTA models across a wide spectrum of resource constraints. The source code is available at https://github.com/jyqi/GiraffeDet.

preprint2022arXiv

Global spherically symmetric solutions to degenerate compressible Navier-Stokes equations with large data and far field vacuum

We consider the initial-boundary value problem (IBVP) for the isentropic compressible Navier-Stokes equations (\textbf{CNS}) in the domain exterior to a ball in $\mathbb R^d$ $(d=2\ \text{or} \ 3)$. When viscosity coefficients are given as a constant multiple of the mass density $ρ$, based on some analysis of the nonlinear structure of this system, we prove the global existence of the unique spherically symmetric classical solution for (large) initial data with spherical symmetry and far field vacuum in some inhomogeneous Sobolev spaces. Moreover, the solutions we obtained have the conserved total mass and finite total energy. $ρ$ keeps positive in the domain considered but decays to zero in the far field, which is consistent with the facts that the total mass is conserved, and \textbf{CNS} is a model of non-dilute fluids where $ρ$ is bounded away from the vacuum. To prove the existence, on the one hand, we consider a well-designed reformulated structure by introducing some new variables, which, actually, can transfer the degeneracies of the time evolution and the viscosity to the possible singularity of some special source terms. On the other hand, it is observed that, for the spherically symmetric flow, the radial projection of the so-called effective velocity $\boldsymbol{v} =U+\nabla φ(ρ)$ ($U$ is the velocity of the fluid, and $φ(ρ)$ is a function of $ρ$ defined via the shear viscosity coefficient $μ(ρ)$: $φ'(ρ)=2μ(ρ)/ρ^2$), verifies a damped transport equation which provides the possibility to obtain its upper bound. Then combined with the BD entropy estimates, one can obtain the required uniform a priori estimates of the solution. It is worth pointing out that the frame work on the well-posedness theory established here can be applied to the shallow water equations.

preprint2022arXiv

Graph Convolution for Re-ranking in Person Re-identification

Nowadays, deep learning is widely applied to extract features for similarity computation in person re-identification (re-ID) and have achieved great success. However, due to the non-overlapping between training and testing IDs, the difference between the data used for model training and the testing data makes the performance of learned feature degraded during testing. Hence, re-ranking is proposed to mitigate this issue and various algorithms have been developed. However, most of existing re-ranking methods focus on replacing the Euclidean distance with sophisticated distance metrics, which are not friendly to downstream tasks and hard to be used for fast retrieval of massive data in real applications. In this work, we propose a graph-based re-ranking method to improve learned features while still keeping Euclidean distance as the similarity metric. Inspired by graph convolution networks, we develop an operator to propagate features over an appropriate graph. Since graph is the essential key for the propagation, two important criteria are considered for designing the graph, and three different graphs are explored accordingly. Furthermore, a simple yet effective method is proposed to generate a profile vector for each tracklet in videos, which helps extend our method to video re-ID. Extensive experiments on three benchmark data sets, e.g., Market-1501, Duke, and MARS, demonstrate the effectiveness of our proposed approach.

preprint2022arXiv

Graph Neural Networks for Double-Strand DNA Breaks Prediction

Double-strand DNA breaks (DSBs) are a form of DNA damage that can cause abnormal chromosomal rearrangements. Recent technologies based on high-throughput experiments have obvious high costs and technical challenges.Therefore, we design a graph neural network based method to predict DSBs (GraphDSB), using DNA sequence features and chromosome structure information. In order to improve the expression ability of the model, we introduce Jumping Knowledge architecture and several effective structural encoding methods. The contribution of structural information to the prediction of DSBs is verified by the experiments on datasets from normal human epidermal keratinocytes (NHEK) and chronic myeloid leukemia cell line (K562), and the ablation studies further demonstrate the effectiveness of the designed components in the proposed GraphDSB framework. Finally, we use GNNExplainer to analyze the contribution of node features and topology to DSBs prediction, and proved the high contribution of 5-mer DNA sequence features and two chromatin interaction modes.

preprint2022arXiv

Image-to-Video Re-Identification via Mutual Discriminative Knowledge Transfer

The gap in representations between image and video makes Image-to-Video Re-identification (I2V Re-ID) challenging, and recent works formulate this problem as a knowledge distillation (KD) process. In this paper, we propose a mutual discriminative knowledge distillation framework to transfer a video-based richer representation to an image based representation more effectively. Specifically, we propose the triplet contrast loss (TCL), a novel loss designed for KD. During the KD process, the TCL loss transfers the local structure, exploits the higher order information, and mitigates the misalignment of the heterogeneous output of teacher and student networks. Compared with other losses for KD, the proposed TCL loss selectively transfers the local discriminative features from teacher to student, making it effective in the ReID. Besides the TCL loss, we adopt mutual learning to regularize both the teacher and student networks training. Extensive experiments demonstrate the effectiveness of our method on the MARS, DukeMTMC-VideoReID and VeRi-776 benchmarks.

preprint2022arXiv

Improved Fine-Tuning by Better Leveraging Pre-Training Data

As a dominant paradigm, fine-tuning a pre-trained model on the target data is widely used in many deep learning applications, especially for small data sets. However, recent studies have empirically shown that training from scratch has the final performance that is no worse than this pre-training strategy once the number of training samples is increased in some vision tasks. In this work, we revisit this phenomenon from the perspective of generalization analysis by using excess risk bound which is popular in learning theory. The result reveals that the excess risk bound may have a weak dependency on the pre-trained model. The observation inspires us to leverage pre-training data for fine-tuning, since this data is also available for fine-tuning. The generalization result of using pre-training data shows that the excess risk bound on a target task can be improved when the appropriate pre-training data is included in fine-tuning. With the theoretical motivation, we propose a novel selection strategy to select a subset from pre-training data to help improve the generalization on the target task. Extensive experimental results for image classification tasks on 8 benchmark data sets verify the effectiveness of the proposed data selection based fine-tuning pipeline.

preprint2022arXiv

Improved Knowledge Distillation via Full Kernel Matrix Transfer

Knowledge distillation is an effective way for model compression in deep learning. Given a large model (i.e., teacher model), it aims to improve the performance of a compact model (i.e., student model) by transferring the information from the teacher. Various information for distillation has been studied. Recently, a number of works propose to transfer the pairwise similarity between examples to distill relative information. However, most of efforts are devoted to developing different similarity measurements, while only a small matrix consisting of examples within a mini-batch is transferred at each iteration that can be inefficient for optimizing the pairwise similarity over the whole data set. In this work, we aim to transfer the full similarity matrix effectively. The main challenge is from the size of the full matrix that is quadratic to the number of examples. To address the challenge, we decompose the original full matrix with Nystr{ö}m method. By selecting appropriate landmark points, our theoretical analysis indicates that the loss for transfer can be further simplified. Concretely, we find that the difference between the original full kernel matrices between teacher and student can be well bounded by that of the corresponding partial matrices, which only consists of similarities between original examples and landmark points. Compared with the full matrix, the size of the partial matrix is linear in the number of examples, which improves the efficiency of optimization significantly. The empirical study on benchmark data sets demonstrates the effectiveness of the proposed algorithm. Code is available at \url{https://github.com/idstcv/KDA}.

preprint2022arXiv

Influences of the dissipative topological edge state on quantized transport in MnBi2Te4

The beauty of quantum Hall (QH) effect is the metrological precision of Hall resistance quantization that originates from the topological edge states. Understanding the factors that lead to quantization breakdown not only provides important insights on the nature of the topological protection of these edge states, but is beneficial for device applications involving such quantized transport. In this work, we combine conventional transport and real space conductivity mapping to investigate whether the quantization breakdown is tied to the disappearance of edge state in the hotly studied MnBi2Te4 system. Our experimental results unambiguously show that topological edge state does exist when quantization breakdown occurs. Such edge state is dissipative in nature and could lead to a quantization breakdown due to its diffusive character causing overlapping with bulk and other edge states in real devices. Our findings bring attentions to issues that are generally inaccessible in the transport study of QH, but can play important roles in practical measurements and device applications.

preprint2022arXiv

Joint learning of object graph and relation graph for visual question answering

Modeling visual question answering(VQA) through scene graphs can significantly improve the reasoning accuracy and interpretability. However, existing models answer poorly for complex reasoning questions with attributes or relations, which causes false attribute selection or missing relation in Figure 1(a). It is because these models cannot balance all kinds of information in scene graphs, neglecting relation and attribute information. In this paper, we introduce a novel Dual Message-passing enhanced Graph Neural Network (DM-GNN), which can obtain a balanced representation by properly encoding multi-scale scene graph information. Specifically, we (i)transform the scene graph into two graphs with diversified focuses on objects and relations; Then we design a dual structure to encode them, which increases the weights from relations (ii)fuse the encoder output with attribute features, which increases the weights from attributes; (iii)propose a message-passing mechanism to enhance the information transfer between objects, relations and attributes. We conduct extensive experiments on datasets including GQA, VG, motif-VG and achieve new state of the art.

preprint2022arXiv

Large Quality Factor Enhancement Based on Cascaded Uniform Lithium Niobate Bichromatic Photonic Crystal Cavities

In this paper, by cascading several bichromatic photonic crystals we demonstrate that the quality factor can be much larger compared with that in an isolated cavity without increasing the total size of the device. We take lithium niobate photonic crystal as an example to illustrate that the simulated quality factor of the cascaded cavity can attain 10^5 with a 70° slant angle, which is an order of magnitude larger than that in isolated cavity. The device can be fabricated easily by current etching technique for lithium niobate. We have fabricated the proposed device experimentally including holes with 70° slant angle. This work is expected to provide guidance to the design of photonic crystal cavity with high-quality factor.

preprint2022arXiv

Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combine the motion and speech audio of the speaker using a motion-audio cross attention transformer. Furthermore, we enable non-deterministic prediction by learning a discrete latent representation of realistic listener motion with a novel motion-encoding VQ-VAE. Our method organically captures the multimodal and non-deterministic nature of nonverbal dyadic interactions. Moreover, it produces realistic 3D listener facial motion synchronous with the speaker (see video). We demonstrate that our method outperforms baselines qualitatively and quantitatively via a rich suite of experiments. To facilitate this line of research, we introduce a novel and large in-the-wild dataset of dyadic conversations. Code, data, and videos available at https://evonneng.github.io/learning2listen/.

preprint2022arXiv

Location reference recognition from texts: A survey and comparison

A vast amount of location information exists in unstructured texts, such as social media posts, news stories, scientific articles, web pages, travel blogs, and historical archives. Geoparsing refers to the process of recognizing location references from texts and identifying their geospatial representations. While geoparsing can benefit many domains, a summary of the specific applications is still missing. Further, there lacks a comprehensive review and comparison of existing approaches for location reference recognition, which is the first and a core step of geoparsing. To fill these research gaps, this review first summarizes seven typical application domains of geoparsing: geographic information retrieval, disaster management, disease surveillance, traffic management, spatial humanities, tourism management, and crime management. We then review existing approaches for location reference recognition by categorizing these approaches into four groups based on their underlying functional principle: rule-based, gazetteer matching-based, statistical learning-based, and hybrid approaches. Next, we thoroughly evaluate the correctness and computational efficiency of the 27 most widely used approaches for location reference recognition based on 26 public datasets with different types of texts (e.g., social media posts and news stories) containing 39,736 location references across the world. Results from this thorough evaluation can help inform future methodological developments for location reference recognition, and can help guide the selection of proper approaches based on application needs.

preprint2022arXiv

MAE-DET: Revisiting Maximum Entropy Principle in Zero-Shot NAS for Efficient Object Detection

In object detection, the detection backbone consumes more than half of the overall inference cost. Recent researches attempt to reduce this cost by optimizing the backbone architecture with the help of Neural Architecture Search (NAS). However, existing NAS methods for object detection require hundreds to thousands of GPU hours of searching, making them impractical in fast-paced research and development. In this work, we propose a novel zero-shot NAS method to address this issue. The proposed method, named MAE-DET, automatically designs efficient detection backbones via the Maximum Entropy Principle without training network parameters, reducing the architecture design cost to nearly zero yet delivering the state-of-the-art (SOTA) performance. Under the hood, MAE-DET maximizes the differential entropy of detection backbones, leading to a better feature extractor for object detection under the same computational budgets. After merely one GPU day of fully automatic design, MAE-DET innovates SOTA detection backbones on multiple detection benchmark datasets with little human intervention. Comparing to ResNet-50 backbone, MAE-DET is $+2.0\%$ better in mAP when using the same amount of FLOPs/parameters, and is $1.54$ times faster on NVIDIA V100 at the same mAP. Code and pre-trained models are available at https://github.com/alibaba/lightweight-neuralarchitecture-search.

preprint2022arXiv

Massive Star-Forming Galaxies Have Converted Most of Their Halo Gas into Stars

In the local Universe, the efficiency for converting baryonic gas into stars is very low. In dark matter halos where galaxies form and evolve, the average efficiency varies with galaxy stellar mass and has a maximum of about twenty percent for Milky-Way-like galaxies. The low efficiency at higher mass is believed to be produced by some quenching processes, such as the feedback from active galactic nuclei. We perform an analysis of weak lensing and satellite kinematics for SDSS central galaxies. Our results reveal that the efficiency is much higher, more than sixty percent, for a large population of massive star-forming galaxies around $10^{11}M_{\odot}$. This suggests that these galaxies acquired most of the gas in their halos and converted it into stars without being affected significantly by quenching processes. This population of galaxies is not reproduced in current galaxy formation models, indicating that our understanding of galaxy formation is incomplete. The implications of our results on circumgalactic media, star formation quenching and disc galaxy rotation curves are discussed. We also examine systematic uncertainties in halo-mass and stellar-mass measurements that might influence our results.

preprint2022arXiv

ModDrop++: A Dynamic Filter Network with Intra-subject Co-training for Multiple Sclerosis Lesion Segmentation with Missing Modalities

Multiple Sclerosis (MS) is a chronic neuroinflammatory disease and multi-modality MRIs are routinely used to monitor MS lesions. Many automatic MS lesion segmentation models have been developed and have reached human-level performance. However, most established methods assume the MRI modalities used during training are also available during testing, which is not guaranteed in clinical practice. Previously, a training strategy termed Modality Dropout (ModDrop) has been applied to MS lesion segmentation to achieve the state-of-the-art performance with missing modality. In this paper, we present a novel method dubbed ModDrop++ to train a unified network adaptive to an arbitrary number of input MRI sequences. ModDrop++ upgrades the main idea of ModDrop in two key ways. First, we devise a plug-and-play dynamic head and adopt a filter scaling strategy to improve the expressiveness of the network. Second, we design a co-training strategy to leverage the intra-subject relation between full modality and missing modality. Specifically, the intra-subject co-training strategy aims to guide the dynamic head to generate similar feature representations between the full- and missing-modality data from the same subject. We use two public MS datasets to show the superiority of ModDrop++. Source code and trained models are available at https://github.com/han-liu/ModDropPlusPlus.

preprint2022arXiv

MogFace: Towards a Deeper Appreciation on Face Detection

Benefiting from the pioneering design of generic object detectors, significant achievements have been made in the field of face detection. Typically, the architectures of the backbone, feature pyramid layer, and detection head module within the face detector all assimilate the excellent experience from general object detectors. However, several effective methods, including label assignment and scale-level data augmentation strategy, fail to maintain consistent superiority when applying on the face detector directly. Concretely, the former strategy involves a vast body of hyper-parameters and the latter one suffers from the challenge of scale distribution bias between different detection tasks, which both limit their generalization abilities. Furthermore, in order to provide accurate face bounding boxes for facial down-stream tasks, the face detector imperatively requires the elimination of false alarms. As a result, practical solutions on label assignment, scale-level data augmentation, and reducing false alarms are necessary for advancing face detectors. In this paper, we focus on resolving three aforementioned challenges that exiting methods are difficult to finish off and present a novel face detector, termed MogFace. In our Mogface, three key components, Adaptive Online Incremental Anchor Mining Strategy, Selective Scale Enhancement Strategy and Hierarchical Context-Aware Module, are separately proposed to boost the performance of face detectors. Finally, to the best of our knowledge, our MogFace is the best face detector on the Wider Face leader-board, achieving all champions across different testing scenarios. The code is available at \url{https://github.com/damo-cv/MogFace}.

preprint2022arXiv

Multi-view Point Cloud Registration based on Evolutionary Multitasking with Bi-Channel Knowledge Sharing Mechanism

Multi-view point cloud registration is fundamental in 3D reconstruction. Since there are close connections between point clouds captured from different viewpoints, registration performance can be enhanced if these connections be harnessed properly. Therefore, this paper models the registration problem as multi-task optimization, and proposes a novel bi-channel knowledge sharing mechanism for effective and efficient problem solving. The modeling of multi-view point cloud registration as multi-task optimization are twofold. By simultaneously considering the local accuracy of two point clouds as well as the global consistency posed by all the point clouds involved, a fitness function with an adaptive threshold is derived. Also a framework of the co-evolutionary search process is defined for the concurrent optimization of multiple fitness functions belonging to related tasks. To enhance solution quality and convergence speed, the proposed bi-channel knowledge sharing mechanism plays its role. The intra-task knowledge sharing introduces aiding tasks that are much simpler to solve, and useful information is shared across aiding tasks and the original tasks, accelerating the search process. The inter-task knowledge sharing explores commonalities buried among the original tasks, aiming to prevent tasks from getting stuck to local optima. Comprehensive experiments conducted on model object as well as scene point clouds show the efficacy of the proposed method.

preprint2022arXiv

Nonlinear manipulation of orbital angular momentum spectra with second- and third- harmonic generation in a quasi-periodically poled crystal

Optical orbital angular momentum (OAM), as an important degree of freedom of light, has been attracted extensive attention, due to its intrinsic feature of natural discrete infinite dimension. Manipulation of OAM spectra is crucial for many impressive applications from classical to quantum realms, in particular, nonlinear manipulation of OAM spectra. Here we realized the nonlinear manipulation of OAM spectra by using the simultaneous second- and third-harmonic generation in a single nonlinear crystal of quasi-periodically poled potassium titanyl phosphate, for fundamental waves with a variety of OAM spectra, especially for customized OAM spectra of the second and third harmonics. The experimental results confirmed the theoretical predictions. Our approach not only provides a novel way to manipulate OAM spectra at new shorter wavelengths that are hard to be directly generated, but also may find new applications towards multiplexing in classical optics and high-dimensional information processing in quantum optics.

preprint2022arXiv

NSSIA: A New Self-Sovereign Identity Scheme with Accountability

Self-Sovereign Identity (SSI) is a new distributed method for identity management, commonly used to address the problem that users are lack of control over their identities. However, the excessive pursuit of self-sovereignty in the most existing SSI schemes hinders sanctions against attackers. To deal with the malicious behavior, a few SSI schemes introduce accountability mechanisms, but they sacrifice users' privacy. What's more, the digital identities (static strings or updatable chains) in the existing SSI schemes are as inputs to a third-party executable program (mobile app, smart contract, etc.) to achieve identity reading, storing and proving, users' self-sovereignty are weakened. To solve the above problems, we present a new self-sovereign identity scheme to strike a balance between privacy and accountability and get rid of the dependence on the third-party program. In our scheme, one and only individual-specific executable code is generated as a digital avatar-i for each human to interact with others in cyberspace without a third-party program, in which the embedding of biometrics enhances uniqueness and user control over their identity. In addition, a joint accountability mechanism, which is based on the shamir (t, n) threshold algorithm and a consortium blockchain, is designed to restrict the power of each regulatory authority and protect users' privacy. Finally, we analyze the security, SSI properties and conduct detailed experiments in term of the cost of computation, storage and blockchain gas. The analysis results indicate that our scheme resists the known attacks and fulfills all the six SSI properties. Compared with the state-of-the-art schemes, the extensive experiment results show that the cost is larger in server storage, blockchain storage and blockchain gas, but is still low enough for practical situations.

preprint2022arXiv

NTIRE 2022 Challenge on Efficient Super-Resolution: Methods and Results

This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-resolve an input image with a magnification factor of $\times$4 based on pairs of low and corresponding high resolution images. The aim was to design a network for single image super-resolution that achieved improvement of efficiency measured according to several metrics including runtime, parameters, FLOPs, activations, and memory consumption while at least maintaining the PSNR of 29.00dB on DIV2K validation set. IMDN is set as the baseline for efficiency measurement. The challenge had 3 tracks including the main track (runtime), sub-track one (model complexity), and sub-track two (overall performance). In the main track, the practical runtime performance of the submissions was evaluated. The rank of the teams were determined directly by the absolute value of the average runtime on the validation set and test set. In sub-track one, the number of parameters and FLOPs were considered. And the individual rankings of the two metrics were summed up to determine a final ranking in this track. In sub-track two, all of the five metrics mentioned in the description of the challenge including runtime, parameter count, FLOPs, activations, and memory consumption were considered. Similar to sub-track one, the rankings of five metrics were summed up to determine a final ranking. The challenge had 303 registered participants, and 43 teams made valid submissions. They gauge the state-of-the-art in efficient single image super-resolution.

preprint2022arXiv

Pancreas segmentation with probabilistic map guided bi-directional recurrent UNet

Pancreas segmentation in medical imaging data is of great significance for clinical pancreas diagnostics and treatment. However, the large population variations in the pancreas shape and volume cause enormous segmentation difficulties, even for state-of-the-art algorithms utilizing fully-convolutional neural networks (FCNs). Specifically, pancreas segmentation suffers from the loss of spatial information in 2D methods, and the high computational cost of 3D methods. To alleviate these problems, we propose a probabilistic-map-guided bi-directional recurrent UNet (PBR-UNet) architecture, which fuses intra-slice information and inter-slice probabilistic maps into a local 3D hybrid regularization scheme, which is followed by bi-directional recurrent network optimization. The PBR-UNet method consists of an initial estimation module for efficiently extracting pixel-level probabilistic maps and a primary segmentation module for propagating hybrid information through a 2.5D U-Net architecture. Specifically, local 3D information is inferred by combining an input image with the probabilistic maps of the adjacent slices into multichannel hybrid data, and then hierarchically aggregating the hybrid information of the entire segmentation network. Besides, a bi-directional recurrent optimization mechanism is developed to update the hybrid information in both the forward and the backward directions. This allows the proposed network to make full and optimal use of the local context information. Quantitative and qualitative evaluation was performed on the NIH Pancreas-CT dataset, and our proposed PBR-UNet method achieved better segmentation results with less computational cost compared to other state-of-the-art methods.

preprint2022arXiv

PMAL: Open Set Recognition via Robust Prototype Mining

Open Set Recognition (OSR) has been an emerging topic. Besides recognizing predefined classes, the system needs to reject the unknowns. Prototype learning is a potential manner to handle the problem, as its ability to improve intra-class compactness of representations is much needed in discrimination between the known and the unknowns. In this work, we propose a novel Prototype Mining And Learning (PMAL) framework. It has a prototype mining mechanism before the phase of optimizing embedding space, explicitly considering two crucial properties, namely high-quality and diversity of the prototype set. Concretely, a set of high-quality candidates are firstly extracted from training samples based on data uncertainty learning, avoiding the interference from unexpected noise. Considering the multifarious appearance of objects even in a single category, a diversity-based strategy for prototype set filtering is proposed. Accordingly, the embedding space can be better optimized to discriminate therein the predefined classes and between known and unknowns. Extensive experiments verify the two good characteristics (i.e., high-quality and diversity) embraced in prototype mining, and show the remarkable performance of the proposed framework compared to state-of-the-arts.

preprint2022arXiv

Point RCNN: An Angle-Free Framework for Rotated Object Detection

Rotated object detection in aerial images is still challenging due to arbitrary orientations, large scale and aspect ratio variations, and extreme density of objects. Existing state-of-the-art rotated object detection methods mainly rely on angle-based detectors. However, angle regression can easily suffer from the long-standing boundary problem. To tackle this problem, we propose a purely angle-free framework for rotated object detection, called Point RCNN, which mainly consists of PointRPN and PointReg. In particular, PointRPN generates accurate rotated RoIs (RRoIs) by converting the learned representative points with a coarse-to-fine manner, which is motivated by RepPoints. Based on the learned RRoIs, PointReg performs corner points refinement for more accurate detection. In addition, aerial images are often severely unbalanced in categories, and existing methods almost ignore this issue. In this paper, we also experimentally verify that re-sampling the images of the rare categories will stabilize training and further improve the detection performance. Experiments demonstrate that our Point RCNN achieves the new state-of-the-art detection performance on commonly used aerial datasets, including DOTA-v1.0, DOTA-v1.5, and HRSC2016.

preprint2022arXiv

Real-World Image Super-Resolution by Exclusionary Dual-Learning

Real-world image super-resolution is a practical image restoration problem that aims to obtain high-quality images from in-the-wild input, has recently received considerable attention with regard to its tremendous application potentials. Although deep learning-based methods have achieved promising restoration quality on real-world image super-resolution datasets, they ignore the relationship between L1- and perceptual- minimization and roughly adopt auxiliary large-scale datasets for pre-training. In this paper, we discuss the image types within a corrupted image and the property of perceptual- and Euclidean- based evaluation protocols. Then we propose a method, Real-World image Super-Resolution by Exclusionary Dual-Learning (RWSR-EDL) to address the feature diversity in perceptual- and L1- based cooperative learning. Moreover, a noise-guidance data collection strategy is developed to address the training time consumption in multiple datasets optimization. When an auxiliary dataset is incorporated, RWSR-EDL achieves promising results and repulses any training time increment by adopting the noise-guidance data collection strategy. Extensive experiments show that RWSR-EDL achieves competitive performance over state-of-the-art methods on four in-the-wild image super-resolution datasets.

preprint2022arXiv

Region-Based Evidential Deep Learning to Quantify Uncertainty and Improve Robustness of Brain Tumor Segmentation

Despite recent advances in the accuracy of brain tumor segmentation, the results still suffer from low reliability and robustness. Uncertainty estimation is an efficient solution to this problem, as it provides a measure of confidence in the segmentation results. The current uncertainty estimation methods based on quantile regression, Bayesian neural network, ensemble, and Monte Carlo dropout are limited by their high computational cost and inconsistency. In order to overcome these challenges, Evidential Deep Learning (EDL) was developed in recent work but primarily for natural image classification. In this paper, we proposed a region-based EDL segmentation framework that can generate reliable uncertainty maps and robust segmentation results. We used the Theory of Evidence to interpret the output of a neural network as evidence values gathered from input features. Following Subjective Logic, evidence was parameterized as a Dirichlet distribution, and predicted probabilities were treated as subjective opinions. To evaluate the performance of our model on segmentation and uncertainty estimation, we conducted quantitative and qualitative experiments on the BraTS 2020 dataset. The results demonstrated the top performance of the proposed method in quantifying segmentation uncertainty and robustly segmenting tumors. Furthermore, our proposed new framework maintained the advantages of low computational cost and easy implementation and showed the potential for clinical application.

preprint2022arXiv

SBPF: Sensitiveness Based Pruning Framework For Convolutional Neural Network On Image Classification

Pruning techniques are used comprehensively to compress convolutional neural networks (CNNs) on image classification. However, the majority of pruning methods require a well pre-trained model to provide useful supporting parameters, such as C1-norm, BatchNorm value and gradient information, which may lead to inconsistency of filter evaluation if the parameters of the pre-trained model are not well optimized. Therefore, we propose a sensitiveness based method to evaluate the importance of each layer from the perspective of inference accuracy by adding extra damage for the original model. Because the performance of the accuracy is determined by the distribution of parameters across all layers rather than individual parameter, the sensitiveness based method will be robust to update of parameters. Namely, we can obtain similar importance evaluation of each convolutional layer between the imperfect-trained and fully trained models. For VGG-16 on CIFAR-10, even when the original model is only trained with 50 epochs, we can get same evaluation of layer importance as the results when the model is trained fully. Then we will remove filters proportional from each layer by the quantified sensitiveness. Our sensitiveness based pruning framework is verified efficiently on VGG-16, a customized Conv-4 and ResNet-18 with CIFAR-10, MNIST and CIFAR-100, respectively.

preprint2022arXiv

Scaled ReLU Matters for Training Vision Transformers

Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty are empirically analysed in ~\cite{xiao2021early}, and the authors conjecture that the issue lies with the \textit{patchify-stem} of ViT models and propose that early convolutions help transformers see better. In this paper, we further investigate this problem and extend the above conclusion: only early convolutions do not help for stable training, but the scaled ReLU operation in the \textit{convolutional stem} (\textit{conv-stem}) matters. We verify, both theoretically and empirically, that scaled ReLU in \textit{conv-stem} not only improves training stabilization, but also increases the diversity of patch tokens, thus boosting peak performance with a large margin via adding few parameters and flops. In addition, extensive experiments are conducted to demonstrate that previous ViTs are far from being well trained, further showing that ViTs have great potential to be a better substitute of CNNs.

preprint2022arXiv

Semantic Data Augmentation based Distance Metric Learning for Domain Generalization

Domain generalization (DG) aims to learn a model on one or more different but related source domains that could be generalized into an unseen target domain. Existing DG methods try to prompt the diversity of source domains for the model's generalization ability, while they may have to introduce auxiliary networks or striking computational costs. On the contrary, this work applies the implicit semantic augmentation in feature space to capture the diversity of source domains. Concretely, an additional loss function of distance metric learning (DML) is included to optimize the local geometry of data distribution. Besides, the logits from cross entropy loss with infinite augmentations is adopted as input features for the DML loss in lieu of the deep features. We also provide a theoretical analysis to show that the logits can approximate the distances defined on original features well. Further, we provide an in-depth analysis of the mechanism and rational behind our approach, which gives us a better understanding of why leverage logits in lieu of features can help domain generalization. The proposed DML loss with the implicit augmentation is incorporated into a recent DG method, that is, Fourier Augmented Co-Teacher framework (FACT). Meanwhile, our method also can be easily plugged into various DG methods. Extensive experiments on three benchmarks (Digits-DG, PACS and Office-Home) have demonstrated that the proposed method is able to achieve the state-of-the-art performance.

preprint2022arXiv

Task Adaptive Parameter Sharing for Multi-Task Learning

Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple downstream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a general method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing resources used and competition between tasks. TAPS solves a joint optimization problem which determines which layers to share with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers encourages weight sharing with the base model. Compared to other methods, TAPS retains high accuracy on downstream tasks while introducing few task-specific parameters. Moreover, TAPS is agnostic to the model architecture and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement.

preprint2022arXiv

Ultrafast coherent interlayer phonon dynamics in atomically thin layers of MnBi2Te4

The atomically thin MnBi2Te4 crystal is a novel magnetic topological insulator, exhibiting exotic quantum physics. Here we report a systematic investigation of ultrafast carrier dynamics and coherent interlayer phonons in few-layer MnBi2Te4 as a function of layer number using time-resolved pump-probe reflectivity spectroscopy. Pronounced coherent phonon oscillations from the interlayer breathing mode are directly observed in the time domain. We find that the coherent oscillation frequency, the photocarrier and coherent phonon decay rates all depend sensitively on the sample thickness. The time-resolved measurements are complemented by ultralow-frequency Raman spectroscopy measurements, which both confirm the interlayer breathing mode and additionally enable observation of the interlayer shear mode. The layer dependence of these modes allows us to extract both the out-of-plane and in-plane interlayer force constants. Our studies not only reveal the interlayer van der Waals coupling strengths, but also shed light on the ultrafast optical properties of this novel two-dimensional material.

preprint2022arXiv

Unsupervised Cross-Modality Domain Adaptation for Segmenting Vestibular Schwannoma and Cochlea with Data Augmentation and Model Ensemble

Magnetic resonance images (MRIs) are widely used to quantify vestibular schwannoma and the cochlea. Recently, deep learning methods have shown state-of-the-art performance for segmenting these structures. However, training segmentation models may require manual labels in target domain, which is expensive and time-consuming. To overcome this problem, domain adaptation is an effective way to leverage information from source domain to obtain accurate segmentations without requiring manual labels in target domain. In this paper, we propose an unsupervised learning framework to segment the VS and cochlea. Our framework leverages information from contrast-enhanced T1-weighted (ceT1-w) MRIs and its labels, and produces segmentations for T2-weighted MRIs without any labels in the target domain. We first applied a generator to achieve image-to-image translation. Next, we ensemble outputs from an ensemble of different models to obtain final segmentations. To cope with MRIs from different sites/scanners, we applied various 'online' augmentations during training to better capture the geometric variability and the variability in image appearance and quality. Our method is easy to build and produces promising segmentations, with a mean Dice score of 0.7930 and 0.7432 for VS and cochlea respectively in the validation set.

preprint2022arXiv

Unsupervised Visual Representation Learning by Online Constrained K-Means

Cluster discrimination is an effective pretext task for unsupervised representation learning, which often consists of two phases: clustering and discrimination. Clustering is to assign each instance a pseudo label that will be used to learn representations in discrimination. The main challenge resides in clustering since prevalent clustering methods (e.g., k-means) have to run in a batch mode. Besides, there can be a trivial solution consisting of a dominating cluster. To address these challenges, we first investigate the objective of clustering-based representation learning. Based on this, we propose a novel clustering-based pretext task with online \textbf{Co}nstrained \textbf{K}-m\textbf{e}ans (\textbf{CoKe}). Compared with the balanced clustering that each cluster has exactly the same size, we only constrain the minimal size of each cluster to flexibly capture the inherent data structure. More importantly, our online assignment method has a theoretical guarantee to approach the global optimum. By decoupling clustering and discrimination, CoKe can achieve competitive performance when optimizing with only a single view from each instance. Extensive experiments on ImageNet and other benchmark data sets verify both the efficacy and efficiency of our proposal. Code is available at \url{https://github.com/idstcv/CoKe}.

preprint2021arXiv

1st Place Solution to ECCV-TAO-2020: Detect and Represent Any Object for Tracking

We extend the classical tracking-by-detection paradigm to this tracking-any-object task. Solid detection results are first extracted from TAO dataset. Some state-of-the-art techniques like \textbf{BA}lanced-\textbf{G}roup \textbf{S}oftmax (\textbf{BAGS}\cite{li2020overcoming}) and DetectoRS\cite{qiao2020detectors} are integrated during detection. Then we learned appearance features to represent any object by training feature learning networks. We ensemble several models for improving detection and feature representation. Simple linking strategies with most similar appearance features and tracklet-level post association module are finally applied to generate final tracking results. Our method is submitted as \textbf{AOA} on the challenge website. Code is available at https://github.com/feiaxyt/Winner_ECCV20_TAO.

preprint2021arXiv

A linearized framework and a new benchmark for model selection for fine-tuning

Fine-tuning from a collection of models pre-trained on different domains (a "model zoo") is emerging as a technique to improve test accuracy in the low-data regime. However, model selection, i.e. how to pre-select the right model to fine-tune from a model zoo without performing any training, remains an open topic. We use a linearized framework to approximate fine-tuning, and introduce two new baselines for model selection -- Label-Gradient and Label-Feature Correlation. Since all model selection algorithms in the literature have been tested on different use-cases and never compared directly, we introduce a new comprehensive benchmark for model selection comprising of: i) A model zoo of single and multi-domain models, and ii) Many target tasks. Our benchmark highlights accuracy gain with model zoo compared to fine-tuning Imagenet models. We show our model selection baseline can select optimal models to fine-tune in few selections and has the highest ranking correlation to fine-tuning accuracy compared to existing algorithms.

preprint2021arXiv

AsymptoticNG: A regularized natural gradient optimization algorithm with look-ahead strategy

Optimizers that further adjust the scale of gradient, such as Adam, Natural Gradient (NG), etc., despite widely concerned and used by the community, are often found poor generalization performance, compared with Stochastic Gradient Descent (SGD). They tend to converge excellently at the beginning of training but are weak at the end. An immediate idea is to complement the strengths of these algorithms with SGD. However, a truncated replacement of optimizer often leads to a crash of the update pattern, and new algorithms often require many iterations to stabilize their search direction. Driven by this idea and to address this problem, we design and present a regularized natural gradient optimization algorithm with look-ahead strategy, named asymptotic natural gradient (ANG). According to the total iteration step, ANG dynamic assembles NG and Euclidean gradient, and updates parameters along the new direction using the intensity of NG. Validation experiments on CIFAR10 and CIFAR100 data sets show that ANG can update smoothly and stably at the second-order speed, and achieve better generalization performance.

preprint2021arXiv

DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech

With the number of smart devices increasing, the demand for on-device text-to-speech (TTS) increases rapidly. In recent years, many prominent End-to-End TTS methods have been proposed, and have greatly improved the quality of synthesized speech. However, to ensure the qualified speech, most TTS systems depend on large and complex neural network models, and it's hard to deploy these TTS systems on-device. In this paper, a small-footprint, fast, stable network for on-device TTS is proposed, named as DeviceTTS. DeviceTTS makes use of a duration predictor as a bridge between encoder and decoder so as to avoid the problem of words skipping and repeating in Tacotron. As we all know, model size is a key factor for on-device TTS. For DeviceTTS, Deep Feedforward Sequential Memory Network (DFSMN) is used as the basic component. Moreover, to speed up inference, mix-resolution decoder is proposed for balance the inference speed and speech quality. Experiences are done with WORLD and LPCNet vocoder. Finally, with only 1.4 million model parameters and 0.099 GFLOPS, DeviceTTS achieves comparable performance with Tacotron and FastSpeech. As far as we know, the DeviceTTS can meet the needs of most of the devices in practical application.

preprint2021arXiv

Edge, Structure and Texture Refinement for Retrospective High Quality MRI Restoration using Deep Learning

22. Shortening acquisition time and reducing the motion-artifact are two of the most critical issues in MRI. As a promising solution, high-quality MRI image restoration provides a new approach to achieve higher resolution without costing additional acquisition time or modification on the pulse sequences. Recently, as to the rise of deep learning, convolutional neural networks have been proposed for super-resolution (SR) image generation and motion-artifact reduction (MAR) for MRI. Recent studies suggest that using perceptual feature space loss and k space loss to capture the perceptual information and high-frequency information of images, respectively. However, the quality of reconstructed SR and MAR MR images is limited because the most important details of the informative area in the MR image, the edges and the structure, cannot be very well restored. Besides, lots of the SR approaches are trained by using low-resolution images generated by applying bicubic or blur-downscale degradation, which cannot represent the real process of MRI measurement. Such inconsistencies lead to performance degradation in the reconstruction of SR images as well. This study reveals that using the L1 loss of SSIM and gradient map edge quality loss could force the deep learning model to focus on studying the features of edges and structure details of MR images, thus generating SR images with more accurate, fruitful information and reduced motion-artifact. We employed a state-of-the-art model, RCAN, as the network framework in both SR and MAR tasks, trained the model by using low-resolution images and motion-artifact affected images which were generated by emulating how they are measured in the real MRI measurement to ensure the model can be easily applied in the practical clinic environment, and verified the trained model could work fairly well.

preprint2021arXiv

Experimental Side-Channel-Free Quantum Key Distribution

Quantum key distribution can provide unconditionally secure key exchange for remote users in theory. In practice, however, in most quantum key distribution systems, quantum hackers might steal the secure keys by listening to the side channels in the source, such as the photon frequency spectrum, emission time, propagation direction, spatial angular momentum, and so on. It is hard to prevent such kinds of attacks because side channels may exist in any of the encoding space whether the designers take care of or not. Here we report an experimental realization of a side-channel-free quantum key distribution protocol which is not only measurement-device-independent, but also immune to all side-channel attacks in the source. We achieve a secure key rate of 4.80e-7 per pulse through 50 km fiber spools.

preprint2021arXiv

Frenkel biexcitons in hybrid HJ photophysical aggregates

Frenkel excitons, the primary photoexcitations in organic semiconductors that are unequivocally responsible for the optical properties of this materials class, are predicted to form \emph{bound} exciton pairs, i.e., biexcitons. These are key intermediates, ubiquitous in many relevant photophysical processes; for example, they determine the exciton bimolecular annihilation dynamics in such systems. Deciphering the details of biexciton correlations is, thus, of utmost importance to understand the optical processes in these semiconductors. To date, however, due to their spectral ambiguity, there has been only scant direct evidence of bound biexcitons, limiting the insights that can be gained. Moreover, a quantum-mechanical basis describing biexciton correlation/stability has so far been lacking. By employing nonlinear coherent spectroscopy, we identify here bound biexcitons in a model polymeric semiconductor. We find, unexpectedly, that excitons with \emph{interchain} vibronic dispersion reveal \emph{intrachain} biexciton correlations and vice versa. Moreover, using a Frenkel exciton model, we can relate the biexciton binding energy to molecular parameters quantified by quantum chemistry, including the magnitude and sign of the exciton-exciton interaction the inter-site hopping energies. Therefore, our work promises a window towards general insights into the many-body electronic structure in polymeric semiconductors and beyond; e.g., other excitonic systems such as organic semiconductor crystals, molecular aggregates, photosynthetic light-harvesting complexes, or DNA.

preprint2021arXiv

High-performance quantum entanglement generation via cascaded second-order nonlinear processes

In this paper, we demonstrate the generation of high-performance entangled photon-pairs in different degrees of freedom from a single piece of fiber pigtailed periodically poled LiNbO$_3$ (PPLN) waveguide. We utilize cascaded second-order nonlinear optical processes, i.e. second-harmonic generation (SHG) and spontaneous parametric down conversion (SPDC), to generate photon-pairs. Previously, the performance of the photon pairs is contaminated by Raman noise photons from the fiber pigtails. Here by integrating the PPLN waveguide with noise rejecting filters, we obtain a coincidence-to-accidental ratio (CAR) higher than 52,600 with photon-pair generation and detection rate of 52.3 kHz and 3.5 kHz, respectively. Energy-time, frequency-bin and time-bin entanglement is prepared by coherently superposing correlated two-photon states in these degrees of freedom, respectively. The energy-time entangled two-photon states achieve the maximum value of CHSH-Bell inequality of S=2.708$\pm$0.024 with a two-photon interference visibility of 95.74$\pm$0.86%. The frequency-bin entangled two-photon states achieve fidelity of 97.56$\pm$1.79% with a spatial quantum beating visibility of 96.85$\pm$2.46%. The time-bin entangled two-photon states achieve the maximum value of CHSH-Bell inequality of S=2.595$\pm$0.037 and quantum tomographic fidelity of 89.07$\pm$4.35%. Our results provide a potential candidate for quantum light source in quantum photonics.

preprint2021arXiv

Magnetization-tuned topological quantum phase transition in MnBi2Te4 devices

Recently, the intrinsic magnetic topological insulator MnBi2Te4 has attracted enormous research interest due to the great success in realizing exotic topological quantum states, such as the quantum anomalous Hall effect (QAHE), axion insulator state, high-Chern-number and high-temperature Chern insulator states. One key issue in this field is to effectively manipulate these states and control topological phase transitions. Here, by systematic angle-dependent transport measurements, we reveal a magnetization-tuned topological quantum phase transition from Chern insulator to magnetic insulator with gapped Dirac surface states in MnBi2Te4 devices. Specifically, as the magnetic field is tilted away from the out-of-plane direction by around 40-60 degrees, the Hall resistance deviates from the quantization value and a colossal, anisotropic magnetoresistance is detected. The theoretical analyses based on modified Landauer-Buttiker formalism show that the field-tilt-driven switching from ferromagnetic state to canted antiferromagnetic state induces a topological quantum phase transition from Chern insulator to magnetic insulator with gapped Dirac surface states in MnBi2Te4 devices. Our work provides an efficient means for modulating topological quantum states and topological quantum phase transitions.

preprint2021arXiv

Metrological characterisation of non-Gaussian entangled states of superconducting qubits

Multipartite entangled states are significant resources for both quantum information processing and quantum metrology. In particular, non-Gaussian entangled states are predicted to achieve a higher sensitivity of precision measurements than Gaussian states. On the basis of metrological sensitivity, the conventional linear Ramsey squeezing parameter (RSP) efficiently characterises the Gaussian entangled atomic states but fails for much wider classes of highly sensitive non-Gaussian states. These complex non-Gaussian entangled states can be classified by the nonlinear squeezing parameter (NLSP), as a generalisation of the RSP with respect to nonlinear observables, and identified via the Fisher information. However, the NLSP has never been measured experimentally. Using a 19-qubit programmable superconducting processor, here we report the characterisation of multiparticle entangled states generated during its nonlinear dynamics. First, selecting 10 qubits, we measure the RSP and the NLSP by single-shot readouts of collective spin operators in several different directions. Then, by extracting the Fisher information of the time-evolved state of all 19 qubits, we observe a large metrological gain of 9.89$^{+0.28}_{-0.29}$ dB over the standard quantum limit, indicating a high level of multiparticle entanglement for quantum-enhanced phase sensitivity. Benefiting from high-fidelity full controls and addressable single-shot readouts, the superconducting processor with interconnected qubits provides an ideal platform for engineering and benchmarking non-Gaussian entangled states that are useful for quantum-enhanced metrology.

preprint2021arXiv

Object Detection Made Simpler by Eliminating Heuristic NMS

We show a simple NMS-free, end-to-end object detection framework, of which the network is a minimal modification to a one-stage object detector such as the FCOS detection model [Tian et al. 2019]. We attain on par or even improved detection accuracy compared with the original one-stage detector. It performs detection at almost the same inference speed, while being even simpler in that now the post-processing NMS (non-maximum suppression) is eliminated during inference. If the network is capable of identifying only one positive sample for prediction for each ground-truth object instance in an image, then NMS would become unnecessary. This is made possible by attaching a compact PSS head for automatic selection of the single positive sample for each instance (see Fig. 1). As the learning objective involves both one-to-many and one-to-one label assignments, there is a conflict in the labels of some training examples, making the learning challenging. We show that by employing a stop-gradient operation, we can successfully tackle this issue and train the detector. On the COCO dataset, our simple design achieves superior performance compared to both the FCOS baseline detector with NMS post-processing and the recent end-to-end NMS-free detectors. Our extensive ablation studies justify the rationale of the design choices.

preprint2021arXiv

Position-Aware Tagging for Aspect Sentiment Triplet Extraction

Aspect Sentiment Triplet Extraction (ASTE) is the task of extracting the triplets of target entities, their associated sentiment, and opinion spans explaining the reason for the sentiment. Existing research efforts mostly solve this problem using pipeline approaches, which break the triplet extraction process into several stages. Our observation is that the three elements within a triplet are highly related to each other, and this motivates us to build a joint model to extract such triplets using a sequence tagging approach. However, how to effectively design a tagging approach to extract the triplets that can capture the rich interactions among the elements is a challenging research question. In this work, we propose the first end-to-end model with a novel position-aware tagging scheme that is capable of jointly extracting the triplets. Our experimental results on several existing datasets show that jointly capturing elements in the triplet using our approach leads to improved performance over the existing approaches. We also conducted extensive experiments to investigate the model effectiveness and robustness.

preprint2021arXiv

Quantum key distribution over 658 km fiber with distributed vibration sensing

Twin-field quantum key distribution (TF-QKD) promises ultra-long secure key distribution which surpasses the rate distance limit and can reduce the number of the trusted nodes in long-haul quantum network. Tremendous efforts have been made towards implementation of TF-QKD, among which, the secure key with finite size analysis can distribute more than 500 km in the lab and in the field. Here, we demonstrate the sending-or-not-sending TF-QKD experimentally, achieving a secure key distribution with finite size analysis over 658 km ultra-low-loss optical fiber, improve the secure distance record by around 100 km. Meanwhile, in a TF-QKD system, any phase fluctuation due to temperature variation and ambient variation during the channel must be recorded and compensated, and all these phase information can then be utilized to sense the channel vibration perturbations. With our QKD system, we recovered the external vibrational perturbations on the fiber generated by an artificial vibroseis and successfully located the perturbation position with a resolution better than 1 km. Our results not only set a new distance record of QKD, but also demonstrate that the redundant information of TF-QKD can be used for remote sensing of the channel vibration, which can find applications in earthquake detection and landslide monitoring besides secure communication.

preprint2021arXiv

Spectrally multiplexed heralded single photon source at telecom-band

Heralded single photon source (HSPS) is an important way in generating genuine single photon, having advantages of experimental simplicity and versatility. However, HSPS intrinsically suffers from the trade-off between the heralded single photon rate and the single photon purity. To overcome this, one can apply multiplexing technology in different degrees of freedom to enhance the performance of HSPS. Here, by employing spectral multiplexing and active feed-forward spectral manipulating, we demonstrate a HSPS at 1.5 μm telecom-band. Our experimental results show that the spectral multiplexing effectively erases the frequency correlation of pair source and significantly improves the heralded single photon rate while keeping the g{^(^2^)}(0) as low as 0.0006{\pm}0.0001. The Hong-Ou-Mandel interference between the heralded single photons and photons from an independent weak coherent source indicates a high indistinguishability. Our results pave a way for scalable HSPS by spectral multiplexing towards deterministic single photon emission.

preprint2020arXiv

A 16-channel fiber array-coupled superconducting single-photon detector array with average system detection efficiency over 60% at telecom wavelength

We report a compact, scalable, and high-performance superconducting nanowire single-photon detector (SNSPD) array by using a multichannel optical fiber array-coupled configuration. For single pixels with an active area of 18 um in diameter and illuminated at the telecom wavelength of 1550 nm, we achieved a pixel yield of 13/16 on one chip, an average system detection efficiency of 69% at a dark count rate of 160 cps, a minimum timing jitter of 74 ps, and a maximum count rate of ~40 Mcps. The optical crosstalk coefficient between adjacent channels is better than -60 dB. The performance of the fiber array-coupled detectors is comparable with a standalone detector coupled to a single fiber. Our method is promising for the development of scalable, high-performance, and high-yield SNSPDs.

preprint2020arXiv

A Federated Multi-View Deep Learning Framework for Privacy-Preserving Recommendations

Privacy-preserving recommendations are recently gaining momentum, since the decentralized user data is increasingly harder to collect, by recommendation service providers, due to the serious concerns over user privacy and data security. This situation is further exacerbated by the strict government regulations such as Europe's General Data Privacy Regulations(GDPR). Federated Learning(FL) is a newly developed privacy-preserving machine learning paradigm to bridge data repositories without compromising data security and privacy. Thus many federated recommendation(FedRec) algorithms have been proposed to realize personalized privacy-preserving recommendations. However, existing FedRec algorithms, mostly extended from traditional collaborative filtering(CF) method, cannot address cold-start problem well. In addition, their performance overhead w.r.t. model accuracy, trained in a federated setting, is often non-negligible comparing to centralized recommendations. This paper studies this issue and presents FL-MV-DSSM, a generic content-based federated multi-view recommendation framework that not only addresses the cold-start problem, but also significantly boosts the recommendation performance by learning a federated model from multiple data source for capturing richer user-level features. The new federated multi-view setting, proposed by FL-MV-DSSM, opens new usage models and brings in new security challenges to FL in recommendation scenarios. We prove the security guarantees of \xxx, and empirical evaluations on FL-MV-DSSM and its variations with public datasets demonstrate its effectiveness. Our codes will be released if this paper is accepted.

preprint2020arXiv

An entanglement-based quantum network based on symmetric dispersive optics quantum key distribution

Quantum key distribution (QKD) is a crucial technology for information security in the future. Developing simple and efficient ways to establish QKD among multiple users are important to extend the applications of QKD in communication networks. Herein, we proposed a scheme of symmetric dispersive optics QKD (DO-QKD) and demonstrated an entanglement-based quantum network based on it. In the experiment, a broadband entanglement photon pair source was shared by end users via wavelength and space division multiplexing. The wide spectrum of generated entangled photon pairs was divided into 16 combinations of frequency-conjugate channels. Photon pairs in each channel combination supported a fully-connected subnet with 8 users by a passive beam splitter. Eventually, it showed that an entanglement-based QKD network over 100 users could be supported by one entangled photon pair source in this architecture. It has great potential on applications of local quantum networks with large user number.

preprint2020arXiv

ARCH: Animatable Reconstruction of Clothed Humans

In this paper, we propose ARCH (Animatable Reconstruction of Clothed Humans), a novel end-to-end framework for accurate reconstruction of animation-ready 3D clothed humans from a monocular image. Existing approaches to digitize 3D humans struggle to handle pose variations and recover details. Also, they do not produce models that are animation ready. In contrast, ARCH is a learned pose-aware model that produces detailed 3D rigged full-body human avatars from a single unconstrained RGB image. A Semantic Space and a Semantic Deformation Field are created using a parametric 3D body estimator. They allow the transformation of 2D/3D clothed humans into a canonical space, reducing ambiguities in geometry caused by pose variations and occlusions in training data. Detailed surface geometry and appearance are learned using an implicit function representation with spatial local features. Furthermore, we propose additional per-pixel supervision on the 3D reconstruction using opacity-aware differentiable rendering. Our experiments indicate that ARCH increases the fidelity of the reconstructed humans. We obtain more than 50% lower reconstruction errors for standard metrics compared to state-of-the-art methods on public datasets. We also show numerous qualitative examples of animated, high-quality reconstructed avatars unseen in the literature so far.

preprint2020arXiv

Attribute Mix: Semantic Data Augmentation for Fine Grained Recognition

Collecting fine-grained labels usually requires expert-level domain knowledge and is prohibitive to scale up. In this paper, we propose Attribute Mix, a data augmentation strategy at attribute level to expand the fine-grained samples. The principle lies in that attribute features are shared among fine-grained sub-categories, and can be seamlessly transferred among images. Toward this goal, we propose an automatic attribute mining approach to discover attributes that belong to the same super-category, and Attribute Mix is operated by mixing semantically meaningful attribute features from two images. Attribute Mix is a simple but effective data augmentation strategy that can significantly improve the recognition performance without increasing the inference budgets. Furthermore, since attributes can be shared among images from the same super-category, we further enrich the training samples with attribute level labels using images from the generic domain. Experiments on widely used fine-grained benchmarks demonstrate the effectiveness of our proposed method.

preprint2020arXiv

CSRN: Collaborative Sequential Recommendation Networks for News Retrieval

Nowadays, news apps have taken over the popularity of paper-based media, providing a great opportunity for personalization. Recurrent Neural Network (RNN)-based sequential recommendation is a popular approach that utilizes users' recent browsing history to predict future items. This approach is limited that it does not consider the societal influences of news consumption, i.e., users may follow popular topics that are constantly changing, while certain hot topics might be spreading only among specific groups of people. Such societal impact is difficult to predict given only users' own reading histories. On the other hand, the traditional User-based Collaborative Filtering (UserCF) makes recommendations based on the interests of the "neighbors", which provides the possibility to supplement the weaknesses of RNN-based methods. However, conventional UserCF only uses a single similarity metric to model the relationships between users, which is too coarse-grained and thus limits the performance. In this paper, we propose a framework of deep neural networks to integrate the RNN-based sequential recommendations and the key ideas from UserCF, to develop Collaborative Sequential Recommendation Networks (CSRNs). Firstly, we build a directed co-reading network of users, to capture the fine-grained topic-specific similarities between users in a vector space. Then, the CSRN model encodes users with RNNs, and learns to attend to neighbors and summarize what news they are reading at the moment. Finally, news articles are recommended according to both the user's own state and the summarized state of the neighbors. Experiments on two public datasets show that the proposed model outperforms the state-of-the-art approaches significantly.

preprint2020arXiv

Demonstration of Planar Ultrananocrystalline Diamond Field Emission Source Operating in SRF Injector at 2 Kelvin

Reported here is the first demonstration of electron beam generation in an SRF TESLA 1.3 GHz gun equipped with field emission cathode when operated at 2 Kelvin. The cathode is submicron film of nitrogen-incorporated ultrananocrystalline diamond [(N)UNCD] deposited atop a Nb RRR300 cathode plug. The output current was measured to increase exponentially as a function of the cavity gradient. Our results demonstrate a feasible path toward simplified fully cryogenic SRF injector technology. One important finding is that the electron emitter made of (N)UNCD, a material long been known as a highly efficient field emission material, demonstrated a record low turn-on gradient of 0.6 MV/m. A hypothesis explaining this behavior is proposed.

preprint2020arXiv

DR Loss: Improving Object Detection by Distributional Ranking

Most of object detection algorithms can be categorized into two classes: two-stage detectors and one-stage detectors. Recently, many efforts have been devoted to one-stage detectors for the simple yet effective architecture. Different from two-stage detectors, one-stage detectors aim to identify foreground objects from all candidates in a single stage. This architecture is efficient but can suffer from the imbalance issue with respect to two aspects: the inter-class imbalance between the number of candidates from foreground and background classes and the intra-class imbalance in the hardness of background candidates, where only a few candidates are hard to be identified. In this work, we propose a novel distributional ranking (DR) loss to handle the challenge. For each image, we convert the classification problem to a ranking problem, which considers pairs of candidates within the image, to address the inter-class imbalance problem. Then, we push the distributions of confidence scores for foreground and background towards the decision boundary. After that, we optimize the rank of the expectations of derived distributions in lieu of original pairs. Our method not only mitigates the intra-class imbalance issue in background candidates but also improves the efficiency for the ranking algorithm. By merely replacing the focal loss in RetinaNet with the developed DR loss and applying ResNet-101 as the backbone, mAP of the single-scale test on COCO can be improved from 39.1% to 41.7% without bells and whistles, which demonstrates the effectiveness of the proposed loss function. Code is available at \url{https://github.com/idstcv/DR_loss}.

preprint2020arXiv

Effect of dispersion on indistinguishability between single-photon wave-packets

With propagating through a dispersive medium, the temporal-spectral profile of laser pulses should be inevitably modified. Although such dispersion effect has been well studied in classical optics, its effect on a single-photon wave-packet, i.e., the matter wave of a single-photon, has not yet been entirely revealed. In this paper, we investigate the effect of dispersion on indistinguishability of single-photon wave-packets through the Hong-Ou-Mandel (HOM) interference. By dispersively manipulating two indistinguishable single-photon wave-packets before interfering with each other, we observe that the difference of the second-order dispersion between two optical paths of the HOM interferometer can be mapped to the interference curve, indicating that (1) with the same amount of dispersion effect in both paths, the HOM interference curve must be only determined by the intrinsic indistinguishability between the wave-packets, i.e., dispersion cancellation due to the indistinguishability between Feynman paths; (2) unbalanced dispersion effect in two paths cannot be cancelled and will broaden the interference curve thus providing a way to measure the second-order dispersion coefficient. Our results suggest a more comprehensive understanding of the single-photon wave-packet and pave ways to explore further applications of the HOM interference.

preprint2020arXiv

Efficient Graph-Based Active Learning with Probit Likelihood via Gaussian Approximations

We present a novel adaptation of active learning to graph-based semi-supervised learning (SSL) under non-Gaussian Bayesian models. We present an approximation of non-Gaussian distributions to adapt previously Gaussian-based acquisition functions to these more general cases. We develop an efficient rank-one update for applying "look-ahead" based methods as well as model retraining. We also introduce a novel "model change" acquisition function based on these approximations that further expands the available collection of active learning acquisition functions for such methods.

preprint2020arXiv

Electronic states and magnetic response of MnBi2Te4 by scanning tunneling microscopy and spectroscopy

Exotic quantum phenomena have been demonstrated in recently discovered intrinsic magnetic topological insulator MnBi2Te4. At its two-dimensional limit, quantum anomalous Hall (QAH) effect and axion insulator state are observed in odd and even layers of MnBi2Te4, respectively. The measured band structures exhibit intriguing and complex properties. Here we employ low-temperature scanning tunneling microscopy to study its surface states and magnetic response. The quasiparticle interference patterns indicate that the electronic structures on the topmost layer of MnBi2Te4 is different from that of the expected out-of-plane A-type antiferromagnetic phase. The topological surface states may be embedded in deeper layers beneath the topmost surface. Such novel electronic structure presumably related to the modification of crystalline structure during sample cleaving and re-orientation of magnetic moment of Mn atoms near the surface. Mn dopants substituted at the Bi site on the second atomic layer are observed. The ratio of Mn/Bi substitutions is 5%. The electronic structures are fluctuating at atomic scale on the surface, which can affect the magnetism of MnBi2Te4. Our findings shed new lights on the magnetic property of MnBi2Te4 and thus the design of magnetic topological insulators.

preprint2020arXiv

Generative Tweening: Long-term Inbetweening of 3D Human Motions

The ability to generate complex and realistic human body animations at scale, while following specific artistic constraints, has been a fundamental goal for the game and animation industry for decades. Popular techniques include key-framing, physics-based simulation, and database methods via motion graphs. Recently, motion generators based on deep learning have been introduced. Although these learning models can automatically generate highly intricate stylized motions of arbitrary length, they still lack user control. To this end, we introduce the problem of long-term inbetweening, which involves automatically synthesizing complex motions over a long time interval given very sparse keyframes by users. We identify a number of challenges related to this problem, including maintaining biomechanical and keyframe constraints, preserving natural motions, and designing the entire motion sequence holistically while considering all constraints. We introduce a biomechanically constrained generative adversarial network that performs long-term inbetweening of human motions, conditioned on keyframe constraints. This network uses a novel two-stage approach where it first predicts local motion in the form of joint angles, and then predicts global motion, i.e. the global path that the character follows. Since there are typically a number of possible motions that could satisfy the given user constraints, we also enable our network to generate a variety of outputs with a scheme that we call Motion DNA. This approach allows the user to manipulate and influence the output content by feeding seed motions (DNA) to the network. Trained with 79 classes of captured motion data, our network performs robustly on a variety of highly complex motion styles.

preprint2020arXiv

Hierarchically Robust Representation Learning

With the tremendous success of deep learning in visual tasks, the representations extracted from intermediate layers of learned models, that is, deep features, attract much attention of researchers. Previous empirical analysis shows that those features can contain appropriate semantic information. Therefore, with a model trained on a large-scale benchmark data set (e.g., ImageNet), the extracted features can work well on other tasks. In this work, we investigate this phenomenon and demonstrate that deep features can be suboptimal due to the fact that they are learned by minimizing the empirical risk. When the data distribution of the target task is different from that of the benchmark data set, the performance of deep features can degrade. Hence, we propose a hierarchically robust optimization method to learn more generic features. Considering the example-level and concept-level robustness simultaneously, we formulate the problem as a distributionally robust optimization problem with Wasserstein ambiguity set constraints, and an efficient algorithm with the conventional training pipeline is proposed. Experiments on benchmark data sets demonstrate the effectiveness of the robust deep representations.

preprint2020arXiv

High-Chern-Number and High-Temperature Quantum Hall Effect without Landau Levels

The quantum Hall effect (QHE) with quantized Hall resistance of h/νe2 starts the research on topological quantum states and lays the foundation of topology in physics. Afterwards, Haldane proposed the QHE without Landau levels, showing nonzero Chern number |C|=1, which has been experimentally observed at relatively low temperatures. For emerging physics and low-power-consumption electronics, the key issues are how to increase the working temperature and realize high Chern numbers (C>1). Here, we report the experimental discovery of high-Chern-number QHE (C=2) without Landau levels and C=1 Chern insulator state displaying nearly quantized Hall resistance plateau above the Néel temperature in MnBi2Te4 devices. Our observations provide a new perspective on topological matter and open new avenues for exploration of exotic topological quantum states and topological phase transitions at higher temperatures.

preprint2020arXiv

Hybrid Differentially Private Federated Learning on Vertically Partitioned Data

We present HDP-VFL, the first hybrid differentially private (DP) framework for vertical federated learning (VFL) to demonstrate that it is possible to jointly learn a generalized linear model (GLM) from vertically partitioned data with only a negligible cost, w.r.t. training time, accuracy, etc., comparing to idealized non-private VFL. Our work builds on the recent advances in VFL-based collaborative training among different organizations which rely on protocols like Homomorphic Encryption (HE) and Secure Multi-Party Computation (MPC) to secure computation and training. In particular, we analyze how VFL's intermediate result (IR) can leak private information of the training data during communication and design a DP-based privacy-preserving algorithm to ensure the data confidentiality of VFL participants. We mathematically prove that our algorithm not only provides utility guarantees for VFL, but also offers multi-level privacy, i.e. DP w.r.t. IR and joint differential privacy (JDP) w.r.t. model weights. Experimental results demonstrate that our work, under adequate privacy budgets, is quantitatively and qualitatively similar to GLMs, learned in idealized non-private VFL setting, rather than the increased cost in memory and processing time in most prior works based on HE or MPC. Our codes will be released if this paper is accepted.

preprint2020arXiv

Intuitive, Interactive Beard and Hair Synthesis with Generative Models

We present an interactive approach to synthesizing realistic variations in facial hair in images, ranging from subtle edits to existing hair to the addition of complex and challenging hair in images of clean-shaven subjects. To circumvent the tedious and computationally expensive tasks of modeling, rendering and compositing the 3D geometry of the target hairstyle using the traditional graphics pipeline, we employ a neural network pipeline that synthesizes realistic and detailed images of facial hair directly in the target image in under one second. The synthesis is controlled by simple and sparse guide strokes from the user defining the general structural and color properties of the target hairstyle. We qualitatively and quantitatively evaluate our chosen method compared to several alternative approaches. We show compelling interactive editing results with a prototype user interface that allows novice users to progressively refine the generated image to match their desired hairstyle, and demonstrate that our approach also allows for flexible and high-fidelity scalp hair synthesis.

preprint2020arXiv

Learning Formation of Physically-Based Face Attributes

Based on a combined data set of 4000 high resolution facial scans, we introduce a non-linear morphable face model, capable of producing multifarious face geometry of pore-level resolution, coupled with material attributes for use in physically-based rendering. We aim to maximize the variety of face identities, while increasing the robustness of correspondence between unique components, including middle-frequency geometry, albedo maps, specular intensity maps and high-frequency displacement details. Our deep learning based generative model learns to correlate albedo and geometry, which ensures the anatomical correctness of the generated assets. We demonstrate potential use of our generative model for novel identity generation, model fitting, interpolation, animation, high fidelity data visualization, and low-to-high resolution data domain transferring. We hope the release of this generative model will encourage further cooperation between all graphics, vision, and data focused professionals while demonstrating the cumulative value of every individual's complete biometric profile.

preprint2020arXiv

Learning to Generate Diverse Dance Motions with Transformer

With the ongoing pandemic, virtual concerts and live events using digitized performances of musicians are getting traction on massive multiplayer online worlds. However, well choreographed dance movements are extremely complex to animate and would involve an expensive and tedious production process. In addition to the use of complex motion capture systems, it typically requires a collaborative effort between animators, dancers, and choreographers. We introduce a complete system for dance motion synthesis, which can generate complex and highly diverse dance sequences given an input music sequence. As motion capture data is limited for the range of dance motions and styles, we introduce a massive dance motion data set that is created from YouTube videos. We also present a novel two-stream motion transformer generative model, which can generate motion sequences with high flexibility. We also introduce new evaluation metrics for the quality of synthesized dance motions, and demonstrate that our system can outperform state-of-the-art methods. Our system provides high-quality animations suitable for large crowds for virtual concerts and can also be used as reference for professional animation pipelines. Most importantly, we show that vast online videos can be effective in training dance motion models.

preprint2020arXiv

Lightweight Mask R-CNN for Long-Range Wireless Power Transfer Systems

Resonant Beam Charging (RBC) is a wireless charging technology which supports multi-watt power transfer over meter-level distance. The features of safety, mobility and simultaneous charging capability enable RBC to charge multiple mobile devices safely at the same time. To detect the devices that need to be charged, a Mask R-CNN based dection model is proposed in previous work. However, considering the constraints of the RBC system, it's not easy to apply Mask R-CNN in lightweight hardware-embedded devices because of its heavy model and huge computation. Thus, we propose a machine learning detection approach which provides a lighter and faster model based on traditional Mask R-CNN. The proposed approach makes the object detection much easier to be transplanted on mobile devices and reduce the burden of hardware computation. By adjusting the structure of the backbone and the head part of Mask R-CNN, we reduce the average detection time from $1.02\mbox{s}$ per image to $0.6132\mbox{s}$, and reduce the model size from $245\mbox{MB}$ to $47.1\mbox{MB}$. The improved model is much more suitable for the application in the RBC system.

preprint2020arXiv

Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model

End-to-end (E2E) systems have played a more and more important role in automatic speech recognition (ASR) and achieved great performance. However, E2E systems recognize output word sequences directly with the input acoustic feature, which can only be trained on limited acoustic data. The extra text data is widely used to improve the results of traditional artificial neural network-hidden Markov model (ANN-HMM) hybrid systems. The involving of extra text data to standard E2E ASR systems may break the E2E property during decoding. In this paper, a novel modular E2E ASR system is proposed. The modular E2E ASR system consists of two parts: an acoustic-to-phoneme (A2P) model and a phoneme-to-word (P2W) model. The A2P model is trained on acoustic data, while extra data including large scale text data can be used to train the P2W model. This additional data enables the modular E2E ASR system to model not only the acoustic part but also the language part. During the decoding phase, the two models will be integrated and act as a standard acoustic-to-word (A2W) model. In other words, the proposed modular E2E ASR system can be easily trained with extra text data and decoded in the same way as a standard E2E ASR system. Experimental results on the Switchboard corpus show that the modular E2E model achieves better word error rate (WER) than standard A2W models.

preprint2020arXiv

Monocular Real-Time Volumetric Performance Capture

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template model. Our system reconstructs a fully textured 3D human from each frame by leveraging Pixel-Aligned Implicit Function (PIFu). While PIFu achieves high-resolution reconstruction in a memory-efficient manner, its computationally expensive inference prevents us from deploying such a system for real-time applications. To this end, we propose a novel hierarchical surface localization algorithm and a direct rendering method without explicitly extracting surface meshes. By culling unnecessary regions for evaluation in a coarse-to-fine manner, we successfully accelerate the reconstruction by two orders of magnitude from the baseline without compromising the quality. Furthermore, we introduce an Online Hard Example Mining (OHEM) technique that effectively suppresses failure modes due to the rare occurrence of challenging examples. We adaptively update the sampling probability of the training data based on the current reconstruction accuracy, which effectively alleviates reconstruction artifacts. Our experiments and evaluations demonstrate the robustness of our system to various challenging angles, illuminations, poses, and clothing styles. We also show that our approach compares favorably with the state-of-the-art monocular performance capture. Our proposed approach removes the need for multi-view studio settings and enables a consumer-accessible solution for volumetric capture.

preprint2020arXiv

Multi-Domain Learning and Identity Mining for Vehicle Re-Identification

This paper introduces our solution for the Track2 in AI City Challenge 2020 (AICITY20). The Track2 is a vehicle re-identification (ReID) task with both the real-world data and synthetic data. Our solution is based on a strong baseline with bag of tricks (BoT-BS) proposed in person ReID. At first, we propose a multi-domain learning method to joint the real-world and synthetic data to train the model. Then, we propose the Identity Mining method to automatically generate pseudo labels for a part of the testing data, which is better than the k-means clustering. The tracklet-level re-ranking strategy with weighted features is also used to post-process the results. Finally, with multiple-model ensemble, our method achieves 0.7322 in the mAP score which yields third place in the competition. The codes are available at https://github.com/heshuting555/AICITY2020_DMT_VehicleReID.

preprint2020arXiv

Neural Architecture Design for GPU-Efficient Networks

Many mission-critical systems are based on GPU for inference. It requires not only high recognition accuracy but also low latency in responding time. Although many studies are devoted to optimizing the structure of deep models for efficient inference, most of them do not leverage the architecture of \textbf{modern GPU} for fast inference, leading to suboptimal performance. To address this issue, we propose a general principle for designing GPU-efficient networks based on extensive empirical studies. This design principle enables us to search for GPU-efficient network structures effectively by a simple and lightweight method as opposed to most Neural Architecture Search (NAS) methods that are complicated and computationally expensive. Based on the proposed framework, we design a family of GPU-Efficient Networks, or GENets in short. We did extensive evaluations on multiple GPU platforms and inference engines. While achieving $\geq 81.3\%$ top-1 accuracy on ImageNet, GENet is up to $6.4$ times faster than EfficienNet on GPU. It also outperforms most state-of-the-art models that are more efficient than EfficientNet in high precision regimes. Our source code and pre-trained models are available from \url{https://github.com/idstcv/GPU-Efficient-Networks}.

preprint2020arXiv

Non-equilibrium states of a plasmonic Dicke model with coherent and dissipative surface plasmon-quantum emitter interactions

Hybrid photonic-plasmonic nanostructures allow one to engineer coupling of quantum emitters and cavity modes accounting for the direct coherent and environment mediated dissipative pathways. Using generalized plasmonic Dicke model, we explore the non-equilibrium phase diagram with respect to these interactions. The analysis shows that their interplay results in the extension of the superradiant and regular lasing states to the dissipative coupling regime and an emergent lasing phase without population inversion having boundary with the superradiant and normal states. Calculated photon emission spectra are demonstrated to carry distinct signatures of these phases.

preprint2020arXiv

NTIRE 2020 Challenge on Real-World Image Super-Resolution: Methods and Results

This paper reviews the NTIRE 2020 challenge on real world super-resolution. It focuses on the participating methods and final results. The challenge addresses the real world setting, where paired true high and low-resolution images are unavailable. For training, only one set of source input images is therefore provided along with a set of unpaired high-quality target images. In Track 1: Image Processing artifacts, the aim is to super-resolve images with synthetically generated image processing artifacts. This allows for quantitative benchmarking of the approaches \wrt a ground-truth image. In Track 2: Smartphone Images, real low-quality smart phone images have to be super-resolved. In both tracks, the ultimate goal is to achieve the best perceptual quality, evaluated using a human study. This is the second challenge on the subject, following AIM 2019, targeting to advance the state-of-the-art in super-resolution. To measure the performance we use the benchmark protocol from AIM 2019. In total 22 teams competed in the final testing phase, demonstrating new and innovative solutions to the problem.

preprint2020arXiv

Objects detection for remote sensing images based on polar coordinates

Arbitrary-oriented object detection is an important task in the field of remote sensing object detection. Existing studies have shown that the polar coordinate system has obvious advantages in dealing with the problem of rotating object modeling, that is, using fewer parameters to achieve more accurate rotating object detection. However, present state-of-the-art detectors based on deep learning are all modeled in Cartesian coordinates. In this article, we introduce the polar coordinate system to the deep learning detector for the first time, and propose an anchor free Polar Remote Sensing Object Detector (P-RSDet), which can achieve competitive detection accuracy via uses simpler object representation model and less regression parameters. In P-RSDet method, arbitrary-oriented object detection can be achieved by predicting the center point and regressing one polar radius and two polar angles. Besides, in order to express the geometric constraint relationship between the polar radius and the polar angle, a Polar Ring Area Loss function is proposed to improve the prediction accuracy of the corner position. Experiments on DOTA, UCAS-AOD and NWPU VHR-10 datasets show that our P-RSDet achieves state-of-the-art performances with simpler model and less regression parameters.

preprint2020arXiv

On equivariant oriented cohomology of Bott-Samelson varieties

For any Bott-Samelson resolution $q_{I}:\hat{X_{I}}\rightarrow G/B$ of the flag variety $G/B$, and any torus equivariant oriented cohomology $h_T$, we compute the restriction formula of certain basis $η_L$ of $h_T(\hat{X_{I}})$ determined by the projective bundle formula. As an application, we show that $h_T(\hat{X_{I}})$ embeds into the equivariant oriented cohomology of $T$-fixed points, and the image can be characterized by using the Goresky-Kottwitz-MacPherson (GKM) description. Furthermore, we compute the push-forward of the basis $η_L$ onto $h_T(G/B)$, and their restriction formula.

preprint2020arXiv

On the Continuity of Rotation Representations in Neural Networks

In neural networks, it is often desirable to work with various representations of the same space. For example, 3D rotations can be represented with quaternions or Euler angles. In this paper, we advance a definition of a continuous representation, which can be helpful for training deep neural networks. We relate this to topological concepts such as homeomorphism and embedding. We then investigate what are continuous and discontinuous representations for 2D, 3D, and n-dimensional rotations. We demonstrate that for 3D rotations, all representations are discontinuous in the real Euclidean spaces of four or fewer dimensions. Thus, widely used representations such as quaternions and Euler angles are discontinuous and difficult for neural networks to learn. We show that the 3D rotations have continuous representations in 5D and 6D, which are more suitable for learning. We also present continuous representations for the general case of the n-dimensional rotation group SO(n). While our main focus is on rotations, we also show that our constructions apply to other groups such as the orthogonal group and similarity transforms. We finally present empirical results, which show that our continuous rotation representations outperform discontinuous ones for several practical problems in graphics and vision, including a simple autoencoder sanity test, a rotation estimator for 3D point clouds, and an inverse kinematics solver for 3D human poses.

preprint2020arXiv

One-Shot Identity-Preserving Portrait Reenactment

We present a deep learning-based framework for portrait reenactment from a single picture of a target (one-shot) and a video of a driving subject. Existing facial reenactment methods suffer from identity mismatch and produce inconsistent identities when a target and a driving subject are different (cross-subject), especially in one-shot settings. In this work, we aim to address identity preservation in cross-subject portrait reenactment from a single picture. We introduce a novel technique that can disentangle identity from expressions and poses, allowing identity preserving portrait reenactment even when the driver's identity is very different from that of the target. This is achieved by a novel landmark disentanglement network (LD-Net), which predicts personalized facial landmarks that combine the identity of the target with expressions and poses from a different subject. To handle portrait reenactment from unseen subjects, we also introduce a feature dictionary-based generative adversarial network (FD-GAN), which locally translates 2D landmarks into a personalized portrait, enabling one-shot portrait reenactment under large pose and expression variations. We validate the effectiveness of our identity disentangling capabilities via an extensive ablation study, and our method produces consistent identities for cross-subject portrait reenactment. Our comprehensive experiments show that our method significantly outperforms the state-of-the-art single-image facial reenactment methods. We will release our code and models for academic use.

preprint2020arXiv

Oriented Objects as pairs of Middle Lines

The detection of oriented objects is frequently appeared in the field of natural scene text detection as well as object detection in aerial images. Traditional detectors for oriented objects are common to rotate anchors on the basis of the RCNN frameworks, which will multiple the number of anchors with a variety of angles, coupled with rotating NMS algorithm, the computational complexities of these models are greatly increased. In this paper, we propose a novel model named Oriented Objects Detection Network O^2-DNet to detect oriented objects by predicting a pair of middle lines inside each target. O^2-DNet is an one-stage, anchor-free and NMS-free model. The target line segments of our model are defined as two corresponding middle lines of original rotating bounding box annotations which can be transformed directly instead of additional manual tagging. Experiments show that our O^2-DNet achieves excellent performance on ICDAR 2015 and DOTA datasets. It is noteworthy that the objects in COCO can be regard as a special form of oriented objects with an angle of 90 degrees. O^2-DNet can still achieve competitive results in these general natural object detection datasets.

preprint2020arXiv

Prediction of Ternary Fluorooxoborates with Coplanar Triangle Units [BOxF3-x]x- From First-Principles

Ten new ternary fluorooxoborate structures were obtained from first-principles prediction. Coplanar aligned triangle structure units [BO2F]2- and [BOF2]- like [BO3]3- in borates were found from the computational simulation. We identified new covalent coordination patterns of the F atom connected with the B atoms which are located in the bridging site, -B--F--B-. Besides, one molecular crystal with [B4O4F4] molecular unit was attached.

preprint2020arXiv

Probing exciton/exciton interactions with entangled photons: theory

Quantum entangled photons provide a sensitive probe of many-body interactions and offer an unique experimental portal for quantifying many-body correlations in a material system. In this paper, we present a theoretical demonstration of how photon-photon entanglement can be generated via interactions between coupled qubits. Here we develop a model for the scattering of an entangled pair of photons from a molecular dimer. We develop a diagrammatic theory for the scattering matrix and show that one can correlate the von Neumann entropy of the outgoing bi-photon wave function to exciton exchange and repulsion interactions. We conclude by discussing possible experimental scenarios for realizing these ideas.

preprint2020arXiv

Reciprocal Collision Avoidance for General Nonlinear Agents using Reinforcement Learning

Finding feasible and collision-free paths for multiple nonlinear agents is challenging in the decentralized scenarios due to limited available information of other agents and complex dynamics constraints. In this paper, we propose a fast multi-agent collision avoidance algorithm for general nonlinear agents with continuous action space, where each agent observes only positions and velocities of nearby agents. To reduce online computation, we first decompose the multi-agent scenario and solve a two agents collision avoidance problem using reinforcement learning (RL). When extending the trained policy to a multi-agent problem, safety is ensured by introducing the optimal reciprocal collision avoidance (ORCA) as linear constraints and the overall collision avoidance action could be found through simple convex optimization. Most existing RL-based multi-agent collision avoidance algorithms rely on the direct control of agent velocities. In sharp contrasts, our approach is applicable to general nonlinear agents. Realistic simulations based on nonlinear bicycle agent models are performed with various challenging scenarios, indicating a competitive performance of the proposed method in avoiding collisions, congestion and deadlock with smooth trajectories.

preprint2020arXiv

Rethinking the Hyperparameters for Fine-tuning

Fine-tuning from pre-trained ImageNet models has become the de-facto standard for various computer vision tasks. Current practices for fine-tuning typically involve selecting an ad-hoc choice of hyperparameters and keeping them fixed to values normally used for training from scratch. This paper re-examines several common practices of setting hyperparameters for fine-tuning. Our findings are based on extensive empirical evaluation for fine-tuning on various transfer learning benchmarks. (1) While prior works have thoroughly investigated learning rate and batch size, momentum for fine-tuning is a relatively unexplored parameter. We find that the value of momentum also affects fine-tuning performance and connect it with previous theoretical findings. (2) Optimal hyperparameters for fine-tuning, in particular, the effective learning rate, are not only dataset dependent but also sensitive to the similarity between the source domain and target domain. This is in contrast to hyperparameters for training from scratch. (3) Reference-based regularization that keeps models close to the initial model does not necessarily apply for "dissimilar" datasets. Our findings challenge common practices of fine-tuning and encourages deep learning practitioners to rethink the hyperparameters for fine-tuning.

preprint2020arXiv

Revisiting the Continuity of Rotation Representations in Neural Networks

In this paper, we provide some careful analysis of certain pathological behavior of Euler angles and unit quaternions encountered in previous works related to rotation representation in neural networks. In particular, we show that for certain problems, these two representations will provably produce completely wrong results for some inputs, and that this behavior is inherent in the topological property of the problem itself and is not caused by unsuitable network architectures or training procedures. We further show that previously proposed embeddings of $\mathrm{SO}(3)$ into higher dimensional Euclidean spaces aimed at fixing this behavior are not universally effective, due to possible symmetry in the input causing changes to the topology of the input space. We propose an ensemble trick as an alternative solution.

preprint2020arXiv

Robust and Precise Vehicle Localization based on Multi-sensor Fusion in Diverse City Scenes

We present a robust and precise localization system that achieves centimeter-level localization accuracy in disparate city scenes. Our system adaptively uses information from complementary sensors such as GNSS, LiDAR, and IMU to achieve high localization accuracy and resilience in challenging scenes, such as urban downtown, highways, and tunnels. Rather than relying only on LiDAR intensity or 3D geometry, we make innovative use of LiDAR intensity and altitude cues to significantly improve localization system accuracy and robustness. Our GNSS RTK module utilizes the help of the multi-sensor fusion framework and achieves a better ambiguity resolution success rate. An error-state Kalman filter is applied to fuse the localization measurements from different sources with novel uncertainty estimation. We validate, in detail, the effectiveness of our approaches, achieving 5-10cm RMS accuracy and outperforming previous state-of-the-art systems. Importantly, our system, while deployed in a large autonomous driving fleet, made our vehicles fully autonomous in crowded city streets despite road construction that occurred from time to time. A dataset including more than 60 km real traffic driving in various urban roads is used to comprehensively test our system.

preprint2020arXiv

Robust axion insulator and Chern insulator phases in a two-dimensional antiferromagnetic topological insulator

The intricate interplay between nontrivial topology and magnetism in two-dimensional (2D) materials has led to the emergence of many novel phenomena and functionalities. An outstanding example is the quantum anomalous Hall (QAH) effect, which was realized in magnetically doped topological insulators (TIs) in the absence of magnetic field. Recently, the layered van der Waals compound MnBi2Te4 has been theoretically predicted and experimentally verified to be a TI with interlayer antiferromagnetic (AFM) order. It is a rare stoichiometric material with coexisting topology and magnetism, thus represents a perfect building block for complex topological-magnetic structures. Here we investigate the quantum transport behaviors of both bulk crystal and exfoliated MnBi2Te4 flakes in a field effect transistor geometry. In the 6 septuple layers (SLs) device tuned into the insulating regime, we observe a large longitudinal resistance and zero Hall plateau, which are characteristic of the axion insulator state. The robust axion insulator state occurs in zero magnetic field, over a wide magnetic field range, and at relatively high temperatures. Moreover, a moderate magnetic field drives a quantum phase transition from the axion insulator phase to a Chern insulator phase with zero longitudinal resistance and quantized Hall resistance h/e2 (h is the Plank constant and e is the elemental charge). These results pave the road for using even-number-SL MnBi2Te4 to realize the quantized topological magnetoelectric effect and axion electrodynamics in condensed matter systems.

preprint2020arXiv

Semi-Anchored Detector for One-Stage Object Detection

A standard one-stage detector is comprised of two tasks: classification and regression. Anchors of different shapes are introduced for each location in the feature map to mitigate the challenge of regression for multi-scale objects. However, the performance of classification can degrade due to the highly class-imbalanced problem in anchors. Recently, many anchor-free algorithms have been proposed to classify locations directly. The anchor-free strategy benefits the classification task but can lead to sup-optimum for the regression task due to the lack of prior bounding boxes. In this work, we propose a semi-anchored framework. Concretely, we identify positive locations in classification, and associate multiple anchors to the positive locations in regression. With ResNet-101 as the backbone, the proposed semi-anchored detector achieves 43.6% mAP on COCO data set, which demonstrates the state-of-art performance among one-stage detectors.

preprint2020arXiv

SoftTriple Loss: Deep Metric Learning Without Triplet Sampling

Distance metric learning (DML) is to learn the embeddings where examples from the same class are closer than examples from different classes. It can be cast as an optimization problem with triplet constraints. Due to the vast number of triplet constraints, a sampling strategy is essential for DML. With the tremendous success of deep learning in classifications, it has been applied for DML. When learning embeddings with deep neural networks (DNNs), only a mini-batch of data is available at each iteration. The set of triplet constraints has to be sampled within the mini-batch. Since a mini-batch cannot capture the neighbors in the original set well, it makes the learned embeddings sub-optimal. On the contrary, optimizing SoftMax loss, which is a classification loss, with DNN shows a superior performance in certain DML tasks. It inspires us to investigate the formulation of SoftMax. Our analysis shows that SoftMax loss is equivalent to a smoothed triplet loss where each class has a single center. In real-world data, one class can contain several local clusters rather than a single one, e.g., birds of different poses. Therefore, we propose the SoftTriple loss to extend the SoftMax loss with multiple centers for each class. Compared with conventional deep metric learning algorithms, optimizing SoftTriple loss can learn the embeddings without the sampling phase by mildly increasing the size of the last fully connected layer. Experiments on the benchmark fine-grained data sets demonstrate the effectiveness of the proposed loss function. Code is available at https://github.com/idstcv/SoftTriple

preprint2020arXiv

Some remarks on associated varieties of vertex operator superalgebras

We study several families of vertex operator superalgebras from a jet (super)scheme point of view. We provide new examples of vertex algebras which are "chiralizations" of their Zhu's Poisson algebras $R_V$. Our examples come from affine $C_\ell^{(1)}$-series vertex algebras ($\ell \geq 1$), certain $N=1$ superconformal vertex algebras, Feigin-Stoyanovsky principal subspaces, Feigin-Stoyanovsky type subspaces, graph vertex algebras $W_Γ$, and extended Virasoro vertex algebra. We also give a counterexample to the chiralization property for the $N=2$ superconformal vertex algebra of central charge $1$.

preprint2020arXiv

Towards human-interpretable, automated learning of feedback control for the mixing layer

We propose an automated analysis of the flow control behaviour from an ensemble of control laws and associated time-resolved flow snapshots. The input may be the rich data base of machine learning control (MLC) optimizing a feedback law for a cost function in the plant. The proposed methodology provides (1) insights into control landscape which maps control laws to performance including extrema and ridge-lines, (2) a catalogue of representative flow states and their contribution to cost function for investigated control laws and (3) a visualization of the dynamics. Key enablers are classification and feature extraction methods of machine learning. The analysis is successfully applied to the stabilization of a mixing layer with sensor-based feedback driving an upstream actuator. The fluctuation energy is reduced by 26\%. The control replaces unforced Kelvin-Helmholtz vortices with subsequent vortex pairing by higher-frequency Kelvin-Helmholtz structures of lower energy. These efforts target a human interpretable, fully automated analysis of MLC identifying qualitatively different actuation regimes, distilling corresponding coherent structures, and developing a digital twin of the plant.

preprint2020arXiv

Ultra-sensitive nanometric flat pigment for binocular stereoscopic image

Two-dimensional (2D) transition metal dichalcogenides (TMDs) with tantalizing layer-dependent electronic and optical properties have emerged as a new paradigm for integrated flat opto-electronic devices. However, daunting challenges remain in deterministic fabrication of TMD layers with demanded shapes and thicknesses as well as light field manipulation in such atomic-thick layers with vanishingly small thicknesses compared to the wavelength. Here, we demonstrate ultra-sensitive light field manipulation in full visible ranges based on laser exfoliating MoS2 layers with nanometric precisions. The nontrivial interfacial phase shifts stemming from the unique dispersion of MoS2 layers integrated on the metallic substrate empower an ultra-sensitive resonance manipulation up to 12.8 nm per MoS2 layer across the entire visible bands, which is more than five times larger than their counterparts. The interlayer van der Waals interactions endow a laser exfoliation method for on-demand patterning MoS2 with atomic thickness precisions and subwavelength feature sizes in a facile and lithography-free fashion. With this, nanometric flat color prints and further binocular stereoscopic views by multi-perspective diffractive images can be realized. Our results with demonstrated practicality unlock full potentials and pave the way for widespread applications of emerging 2D flat optics.

preprint2019arXiv

An Experimentally Verified Approach to non-Entanglement-Breaking Channel Certification

Ensuring the non-entanglement-breaking (non-EB) property of quantum channels is crucial for the effective distribution and storage of quantum states. However, a practical method for direct and accurate certification of the non-EB feature is highly desirable. Here, we propose and verify a realistic source based measurement device independent certification of non-EB channels. Our method is resilient to repercussions on the certification from experimental conditions, such as multiphotons and imperfect state preparation, and can be implemented with information incomplete set. We achieve good agreement between experimental outcomes and theoretical predictions, which is validated by the expected results of the ideal semi-quantum signaling game, and accurately certify the non-EB channels. Furthermore, our approach is highly robust to effects from noise. Therefore, the proposed approach can be expected to play a significant role in the design and evaluation of realistic quantum channels.

preprint2019arXiv

Baryon -- anti-Baryon Photoproduction

Baryon - anti-baryon photoproduction off a proton target is being studied in detail at Jefferson Lab with the GlueX Experiment. We observe $p\bar{p}$ and, for the first time, $Λ\barΛ$ photoproduction (with $Λ\rightarrow π^- p$, $\barΛ \rightarrow π^+ \bar{p}$) from thresholds up to $E_γ = 11.4$ GeV. Preliminary spectra from data accumulated during the GlueX Phase-I period are shown. Angular distributions of the photoproduced hyperons indicate more than one production mechanism in the reaction channel $γp \rightarrow Λ\barΛp$. A Monte Carlo simulation with four mechanisms, tested through comparison between simulation and experimental data, is presented.

preprint2019arXiv

High-speed measurement-device-independent quantum key distribution with integrated silicon photonics

Measurement-device-independent quantum key distribution (MDI-QKD) removes all detector side channels and enables secure QKD with an untrusted relay. It is suitable for building a star-type quantum access network, where the complicated and expensive measurement devices are placed in the central untrusted relay and each user requires only a low-cost transmitter, such as an integrated photonic chip. Here, we experimentally demonstrate a 1.25 GHz silicon photonic chip-based MDI-QKD system using polarization encoding. The photonic chip transmitters integrate the necessary encoding components for a standard QKD source. We implement random modulations of polarization states and decoy intensities, and demonstrate a finite-key secret rate of 31 bps over 36 dB channel loss (or 180 km standard fiber). This key rate is higher than state-of-the-art MDI-QKD experiments. The results show that silicon photonic chip-based MDI-QKD, benefiting from miniaturization, low-cost manufacture and compatibility with CMOS microelectronics, is a promising solution for future quantum secure networks.

preprint2019arXiv

Sending-or-Not-Sending with Independent Lasers: Secure Twin-Field Quantum Key Distribution Over 509 km

Twin field quantum key distribution promises high key rates at long distance to beat the rate distance limit. Here, applying the sending or not sending TF QKD protocol, we experimentally demonstrate a secure key distribution breaking the absolute key rate limit of repeaterless QKD over 509 km, 408 km ultra-low loss optical fibre and 350 km standard optical fibre. Two independent lasers are used as the source with remote frequency locking technique over 500 km fiber distance; Practical optical fibers are used as the optical path with appropriate noise filtering; And finite key effects are considered in the key rate analysis. The secure key rates obtained at different distances are more than 5 times higher than the conditional limit of repeaterless QKD, a bound value assuming the same detection loss in the comparison. The achieved secure key rate is also higher than that a traditional QKD protocol running with a perfect repeaterless QKD device and even if an infinite number of sent pulses. Our result shows that the protocol and technologies applied in this experiment enable TF QKD to achieve high secure key rate at long distribution distance, and hence practically useful for field implementation of intercity QKD.

preprint2017arXiv

Locating any two vertices on Hamiltonian cycles

In this paper we give a proof of Enomoto's conjecture for graphs of sufficiently large order. Enomoto's conjecture states that, if $G$ is a graph of order $n$ with minimum degree $δ(G)\geq \frac{n}{2}+1$, then for any pair of vertices $x$, $y$ in $G$, there is a Hamiltonian cycle $C$ of $G$ such that $d_C(x,y)=\lfloor \frac{n}{2}\rfloor$. The main tools of our proof are Regularity Lemma of Szemerédi and Blow-up Lemma of Komlós et al.

preprint2016arXiv

An adaptive algorithm based on the shifted inverse iteration for the Steklov eigenvalue problem

This paper proposes and analyzes an a posteriori error estimator for the finite element multi-scale discretization approximation of the Steklov eigenvalue problem. Based on the a posteriori error estimates, an adaptive algorithm of shifted inverse iteration type is designed. Finally, numerical experiments comparing the performances of three kinds of different adaptive algorithms are provided, which illustrate the efficiency of the adaptive algorithm proposed here.

preprint2016arXiv

Capturing Dynamic Textured Surfaces of Moving Targets

We present an end-to-end system for reconstructing complete watertight and textured models of moving subjects such as clothed humans and animals, using only three or four handheld sensors. The heart of our framework is a new pairwise registration algorithm that minimizes, using a particle swarm strategy, an alignment error metric based on mutual visibility and occlusion. We show that this algorithm reliably registers partial scans with as little as 15% overlap without requiring any initial correspondences, and outperforms alternative global registration algorithms. This registration algorithm allows us to reconstruct moving subjects from free-viewpoint video produced by consumer-grade sensors, without extensive sensor calibration, constrained capture volume, expensive arrays of cameras, or templates of the subject geometry.

preprint2016arXiv

Deep CTR Prediction in Display Advertising

Click through rate (CTR) prediction of image ads is the core task of online display advertising systems, and logistic regression (LR) has been frequently applied as the prediction model. However, LR model lacks the ability of extracting complex and intrinsic nonlinear features from handcrafted high-dimensional image features, which limits its effectiveness. To solve this issue, in this paper, we introduce a novel deep neural network (DNN) based model that directly predicts the CTR of an image ad based on raw image pixels and other basic features in one step. The DNN model employs convolution layers to automatically extract representative visual features from images, and nonlinear CTR features are then learned from visual features and other contextual features by using fully-connected layers. Empirical evaluations on a real world dataset with over 50 million records demonstrate the effectiveness and efficiency of this method.

preprint2016arXiv

Dense Human Body Correspondences Using Convolutional Networks

We propose a deep learning approach for finding dense correspondences between 3D scans of people. Our method requires only partial geometric information in the form of two depth maps or partial reconstructed surfaces, works for humans in arbitrary poses and wearing any clothing, does not require the two people to be scanned from similar viewpoints, and runs in real time. We use a deep convolutional neural network to train a feature descriptor on depth map pixels, but crucially, rather than training the network to solve the shape correspondence problem directly, we train it to solve a body region classification problem, modified to increase the smoothness of the learned descriptors near region boundaries. This approach ensures that nearby points on the human body are nearby in feature space, and vice versa, rendering the feature descriptor suitable for computing dense correspondences between the scans. We validate our method on real and synthetic data for both clothed and unclothed humans, and show that our correspondences are more robust than is possible with state-of-the-art unsupervised methods, and more accurate than those found using methods that require full watertight 3D geometry.

preprint2016arXiv

Excited-State Structure Modifications Due to Molecular Substituents and Exciton Scattering in Conjugated Molecules

Attachment of chemical substituents (such as polar moieties) constitutes an efficient and convenient way to modify physical and chemical properties of conjugated polymers and oligomers. Associated modifications in the molecular electronic states can be comprehensively described by examining scattering of excitons in the polymer's backbone at the scattering center representing the chemical substituent. Here, we implement effective tight-binding models as a tool to examine the analytical properties of the exciton scattering matrices in semi-infinite polymer chains with substitutions. We demonstrate that chemical interactions between the substitution and attached polymer are adequately described by the analytical properties of the scattering matrices. In particular, resonant and bound electronic excitations are expressed via the positions of zeros and poles of the scattering amplitude, analytically continued to complex values of exciton quasi-momenta. We exemplify the formulated concepts by analyzing excited states in conjugated phenylacetylenes substituted by perylene.

preprint2016arXiv

Photorealistic Facial Texture Inference Using Deep Neural Networks

We present a data-driven inference method that can synthesize a photorealistic texture map of a complete 3D face model given a partial 2D view of a person in the wild. After an initial estimation of shape and low-frequency albedo, we compute a high-frequency partial texture map, without the shading component, of the visible face area. To extract the fine appearance details from this incomplete input, we introduce a multi-scale detail analysis technique based on mid-layer feature correlations extracted from a deep convolutional neural network. We demonstrate that fitting a convex combination of feature correlations from a high-resolution face database can yield a semantically plausible facial detail description of the entire face. A complete and photorealistic texture map can then be synthesized by iteratively optimizing for the reconstructed feature correlations. Using these high-resolution textures and a commercial rendering framework, we can produce high-fidelity 3D renderings that are visually comparable to those obtained with state-of-the-art multi-view face capture systems. We demonstrate successful face reconstructions from a wide range of low resolution input images, including those of historical figures. In addition to extensive evaluations, we validate the realism of our results using a crowdsourced user study.

preprint2016arXiv

Probing polaron excitation spectra in organic semiconductors by photoinduced-absorption-detected two-dimensional coherent spectroscopy

We report a theoretical description and experimental implementation of a novel two-dimensional coherent excitation spectroscopy based on quasi-steady-state photoinduced absorption measurement of a long-lived nonlinear population. We have studied a semiconductor-polymer:fullerene-derivative distributed heterostructure by measuring the 2D excitation spectrum by means of photoluminescence, photocurrent and photoinduced absorption from metastable polaronic products. We conclude that the photoinduced absorption probe is a viable and valuable probe in this family of 2D coherent spectroscopies.

preprint2016arXiv

Real-Time Facial Segmentation and Performance Capture from RGB Input

We introduce the concept of unconstrained real-time 3D facial performance capture through explicit semantic segmentation in the RGB input. To ensure robustness, cutting edge supervised learning approaches rely on large training datasets of face images captured in the wild. While impressive tracking quality has been demonstrated for faces that are largely visible, any occlusion due to hair, accessories, or hand-to-face gestures would result in significant visual artifacts and loss of tracking accuracy. The modeling of occlusions has been mostly avoided due to its immense space of appearance variability. To address this curse of high dimensionality, we perform tracking in unconstrained images assuming non-face regions can be fully masked out. Along with recent breakthroughs in deep learning, we demonstrate that pixel-level facial segmentation is possible in real-time by repurposing convolutional neural networks designed originally for general semantic segmentation. We develop an efficient architecture based on a two-stream deconvolution network with complementary characteristics, and introduce carefully designed training samples and data augmentation strategies for improved segmentation accuracy and robustness. We adopt a state-of-the-art regression-based facial tracking framework with segmented face images as training, and demonstrate accurate and uninterrupted facial performance capture in the presence of extreme occlusion and even side views. Furthermore, the resulting segmentation can be directly used to composite partial 3D face models on the input images and enable seamless facial manipulation tasks, such as virtual make-up or face replacement.

preprint2016arXiv

Superconducting nanowire single photon detector at 532 nm and demonstration in satellite laser ranging

Superconducting nanowire single-photon detectors (SNSPDs) at a wavelength of 532 nm were designed and fabricated aiming to satellite laser ranging (SLR) applications. The NbN SNSPDs were fabricated on one-dimensional photonic crystals with a sensitive-area diameter of 42 um. The devices were coupled with multimode fiber (phi=50um) and exhibited a maximum system detection efficiency of 75% at an extremely low dark count rate of <0.1 Hz. An SLR experiment using an SNSPD at a wavelength of 532 nm was successfully demonstrated. The results showed a depth ranging with a precision of ~8.0 mm for the target satellite LARES, which is ~3,000 km away from the ground ranging station at the Sheshan Observatory.

preprint2016arXiv

Transverse beam size measurement system using visible synchrotron radiation at HLS II

An interferometer system and an imaging system using visible synchrotron radiation (SR) have been installed in HLS II storage ring. Simulations of these two systems are given using Synchrotron Radiation Workshop(SRW) code. With these two systems, the beam energy spread and the beam emittance can be measured. A detailed description of these two systems and the measurement method is given in this paper. The measurement results of beam size, emittance and energy spread are given at the end.

preprint2015arXiv

Beam size and position measurement based on logarithm processing algorithm in HLS II

A logarithm processing algorithm to measure beam transverse size and position is proposed and preliminary experimental results in Hefei Light Source II (HLS II) are given. The algorithm is based on only 4 successive channels of 16 anode channels of multianode photomultiplier tube (MAPMT) R5900U-00-L16 which has typical rise time of 0.6 ns and effective area of 0.8x16 mm for a single anode channel. In the paper, we firstly elaborate the simulation results of the algorithm with and without channel inconsistency. Then we calibrate the channel inconsistency and verify the algorithm using general current signal processor Libera Photon in low-speed scheme. Finally we get turn-by-turn beam size and position and calculate the vertical tune in high-speed scheme. The experimental results show that measured values fit well with simulation results after channel differences are calibrated and the fractional part of the tune in vertical direction is 0.3628 which is very close to the nominal value 0.3621.

preprint2015arXiv

Breaking the Barriers to True Augmented Reality

In recent years, Augmented Reality (AR) and Virtual Reality (VR) have gained considerable commercial traction, with Facebook acquiring Oculus VR for \$2 billion, Magic Leap attracting more than \$500 million of funding, and Microsoft announcing their HoloLens head-worn computer. Where is humanity headed: a brave new dystopia-or a paradise come true? In this article, we present discussions, which started at the symposium "Making Augmented Reality Real", held at Nara Institute of Science and Technology in August 2014. Ten scientists were invited to this three-day event, which started with a full day of public presentations and panel discussions (video recordings are available at the event web page), followed by two days of roundtable discussions addressing the future of AR and VR.

preprint2015arXiv

Dark counts of superconducting nanowire single-photon detector under illumination

An abnormal increase in the SDE was observed for superconducting nanowire single-photon detectors (SNSPDs) when the bias current (Ib) was close to the switching current (Isw). By introducing the time-correlated single-photon counting technique, we investigated the temporal histogram of the detection counts of an SNSPD under illumination. The temporal information helps us to distinguish photon counts from dark counts in the time domain. In this manner, the dark count rate (DCR) under illumination and the accurate SDE can be determined. The DCR under moderate illumination may be significantly larger than the conventional DCR measured without illumination under a high Ib, which causes the abnormal increase in the SDE. The increased DCR may be explained by the suppression of Isw under illumination.

preprint2015arXiv

Nonlinear differentiation equation and analytic function spaces

In this paper we consider the nonlinear complex differential equation $$(f^{(k)})^{n_{k}}+A_{k-1}(z)(f^{(k-1)})^{n_{k-1}}+\cdot\cdot\cdot+A_{1}(z)(f')^{n_{1}}+A_{0}(z)f^{n_{0}}=0, $$where $ A_{j}(z)$, $ j=0, \cdots, k-1 $, are analytic in the unit disk $ \mathbb{D} $, $ n_{j}\in R^{+} $ for all $ j=0, \cdots, k $. We investigate this nonlinear differential equation from two aspects. On one hand, we provide some sufficient conditions on coefficients such that all solutions of this equation belong to a class of Möbius invariant function space, the so-called $Q_K$ space. On the other hand, we find some growth estimates for the analytic solutions of this equation if the coefficients belong to some analytic function spaces.

preprint2015arXiv

The lower bound property of the Morley element eigenvalues

In this paper, we prove that the Morley element eigenvalues approximate the exact ones from below on regular meshes, including adaptive local refined meshes, for the fourth-order elliptic eigenvalue problems with the clamped boundary condition in any dimension. And we implement the adaptive computation to obtain lower bounds of the Morley element eigenvalues for the vibration problem of clamped plate under tension.

preprint2015arXiv

Two-dimensional coherent photocurrent excitation spectroscopy in a polymer solar cell

In high-performance solar cells based on polymeric semiconductors, the mechanism of photocarrier generation on $<100$-fs timescales is yet to be unravelled. In particular the dynamics of early-time electronic coupling between excitons on polymer chains and charge-transfer states need to be investigated in order to develop a detailed picture of ultrafast processes involved in photocurrent production. In this proceeding, we report preliminary measurements using a novel spectroscopy that can measure such correlations: two-dimensional coherent photocurrent excitation spectroscopy. This nonlinear technique measures off-diagonal spectral correlations in a two-dimensional photocurrent excitation spectrum. We interpret these sectroscopic measurements in light of recent theoretical predictions.

preprint2014arXiv

Algorithms to test open set condition for self-similar set related to P.V. numbers

Fix a P.V. number $λ^{-1}>1.$ Given $\mathbf{p}=(p_{1},\cdots,p_{m})\in \mathbb{N}^{m}$, $\mathbf{b}=(b_{1},\cdots,b_{m})\in \mathbb{Q^{m}$, for the self-similar set $E_{\mathbf{p},\mathbf{b}}=\cup_{i=1}^{m}(λ^{p_{i}}E_{\mathbf{p},\mathbf{b}}+b_{i})$ we find an efficient algorithm to test whether $E_{\mathbf{p},\mathbf{b}}$ satisfies the open set condition (strong separation condition) or not.

preprint2014arXiv

Application of Artificial Neural Networks in Predicting Abrasion Resistance of Solution Polymerized Styrene-Butadiene Rubber Based Composites

Abrasion resistance of solution polymerized styrene-butadiene rubber (SSBR) based composites is a typical and crucial property in practical applications. Previous studies show that the abrasion resistance can be calculated by the multiple linear regression model. In our study, considering this relationship can also be described into the non-linear conditions, a Multilayer Feed-forward Neural Networks model with 3 nodes (MLFN-3) was successfully established to describe the relationship between the abrasion resistance and other properties, using 23 groups of data, with the RMS error 0.07. Our studies have proved that Artificial Neural Networks (ANN) model can be used to predict the SSBR-based composites, which is an accurate and robust process.

preprint2014arXiv

Application of Multilayer Feedforward Neural Networks in Predicting Tree Height and Forest Stock Volume of Chinese Fir

Wood increment is critical information in forestry management. Previous studies used mathematics models to describe complex growing pattern of forest stand, in order to determine the dynamic status of growing forest stand in multiple conditions. In our research, we aimed at studying non-linear relationships to establish precise and robust Artificial Neural Networks (ANN) models to predict the precise values of tree height and forest stock volume based on data of Chinese fir. Results show that Multilayer Feedforward Neural Networks with 4 nodes (MLFN-4) can predict the tree height with the lowest RMS error (1.77); Multilayer Feedforward Neural Networks with 7 nodes (MLFN-7) can predict the forest stock volume with the lowest RMS error (4.95). The training and testing process have proved that our models are precise and robust.

preprint2014arXiv

Determination of Boiling Range of Xylene Mixed in PX Device Using Artificial Neural Networks

Determination of boiling range of xylene mixed in PX device is currently a crucial topic in the practical applications because of the recent disputes of PX project in China. In our study, instead of determining the boiling range of xylene mixed by traditional approach in laboratory or industry, we successfully established two Artificial Neural Networks (ANNs) models to determine the initial boiling point and final boiling point respectively. Results show that the Multilayer Feedforward Neural Networks (MLFN) model with 7 nodes (MLFN-7) is the best model to determine the initial boiling point of xylene mixed, with the RMS error 0.18; while the MLFN model with 4 nodes (MLFN-4) is the best model to determine the final boiling point of xylene mixed, with the RMS error 0.75. The training and testing processes both indicate that the models we developed are robust and precise. Our research can effectively avoid the damage of the PX device to human body and environment.

preprint2014arXiv

Nonideal optical cavity structure of superconducting nanowire single photon detector

Optical cavity structure has been proven to be a crucial factor for obtaining high detection efficiency in superconducting nanowire single photon detector (SNSPD). Practically, complicated fabrication processes may result in a non-ideal optical cavity structure. The cross-sectional transmission electron microscope (TEM) image of SNSPD fabricated in this study shows unexpected arc-shaped optical cavities which could have originated due to the over-etching of SiO2 layer while defining NbN nanowire. The effects of the arc-shaped optical cavity structure, such as the wavelength dependence of the optical absorption efficiency for different polarization, were analyzed by performing optical simulations using finite-difference time-domain method. The central wavelength of the device is found to exhibit a blue shift owing to the arced cavity structure. This effect is equivalent to the flat cavity with a reduced height. The results may give interesting reference for SNSPD design and fabrication.

preprint2014arXiv

Superconducting nanowire single photon detector with on-chip bandpass filter

Dark count rate is one of the key parameters limiting the performance of the superconducting nanowire single photon detector (SNSPD). We have designed a multi-layer film bandpass filter that can be integrated onto the SNSPD to suppress the dark counts contributed by the stray light and blackbody radiation of the fiber. The bandpass filter is composed of 16 SiO2/Si bilayers deposited onto the backside of a thermally oxidized Si substrate. The substrate shows an excellent bandpass filter effect and provides a high transmittance of 88% at the central wavelength of the pass band, which is the same as that of the bare substrate. The SNSPDs fabricated on the substrate integrated with the bandpass filter show conspicuous wavelength-sensitive detection efficiency. The background dark count rate is reduced by two orders of magnitude to sub-Hz compared with the conventional SNSPD (a few tens of Hz). The detector exhibits a system detection efficiency of 56% at DCR of 1 Hz, with the measured minimal noise equivalent power reaching 2e-19 w/Hz1/2.

preprint2013arXiv

Millimeter-scale and large-angle self-collimation in a photonic crystal composed of silicon nanorods

We report the observation of a large-angle self-collimation phenomenon occurring in photonic crystals composed of nanorods. Electromagnetic waves incident onto such photonic crystals from directions covering a wide-range of incident angles become highly localized along a single array of rods, which results in narrow-beam propagation without divergence. A propagation length of 0.4 mm is experimentally observed over the wavelength range of 1540 nm to 1570 nm, even in the large incident angle case, which is a very considerable length scale for on-chip optical interconnection.

preprint2010arXiv

Interaction induced ferro-electricity in the rotational states of polar molecules

We show that a ferro-electric quantum phase transition can be driven by the dipolar interaction of polar molecules in the presence a micro-wave field. The obtained ferro-electricity crucially depends on the harmonic confinement potential, and the resulting dipole moment persists even when the external field is turned off adiabatically. The transition is shown to be second order for fermions and for bosons of a smaller permanent dipole moment, but is first order for bosons of a larger moment. Our results suggest the possibility of manipulating the microscopic rotational state of polar molecules by tuning the trap's aspect ratio (and other mesoscopic parameters), even though the later's energy scale is smaller than the former's by six orders of magnitude.

preprint2010arXiv

Large-scale Graphitic Thin Films Synthesized on Ni and Transferred to Insulators: Structural and Electronic Properties

We present a comprehensive study of the structural and electronic properties of ultrathin films containing graphene layers synthesized by chemical vapor deposition (CVD) based surface segregation on polycrystalline Ni foils then transferred onto insulating SiO2/Si substrates. Films of size up to several mm's have been synthesized. Structural characterizations by atomic force microscopy (AFM), scanning tunneling microscopy (STM), cross-sectional transmission electron microscopy (XTEM) and Raman spectroscopy confirm that such large scale graphitic thin films (GTF) contain both thick graphite regions and thin regions of few layer graphene. The films also contain many wrinkles, with sharply-bent tips and dislocations revealed by XTEM, yielding insights on the growth and buckling processes of the GTF. Measurements on mm-scale back-gated transistor devices fabricated from the transferred GTF show ambipolar field effect with resistance modulation ~50% and carrier mobilities reaching ~2000 cm^2/Vs. We also demonstrate quantum transport of carriers with phase coherence length over 0.2 $μ$m from the observation of 2D weak localization in low temperature magneto-transport measurements. Our results show that despite the non-uniformity and surface roughness, such large-scale, flexible thin films can have electronic properties promising for device applications.

preprint2001arXiv

Fast Tree Search for Enumeration of a Lattice Model of Protein Folding

Using a fast tree-searching algorithm and a Pentium cluster, we enumerated all the sequences and compact conformations (structures) for a protein folding model on a cubic lattice of size $4\times3\times3$. We used two types of amino acids -- hydrophobic (H) and polar (P) -- to make up the sequences, so there were $2^{36} \approx 6.87 \times 10^{10}$ different sequences. The total number of distinct structures was 84,731,192. We made use of a simple solvation model in which the energy of a sequence folded into a structure is minus the number of hydrophobic amino acids in the ``core'' of the structure. For every sequence, we found its ground state or ground states, i.e., the structure or structures for which its energy is lowest. About 0.3% of the sequences have a unique ground state. The number of structures that are unique ground states of at least one sequence is 2,662,050, about 3% of the total number of structures. However, these ``designable'' structures differ drastically in their designability, defined as the number of sequences whose unique ground state is that structure. To understand this variation in designability, we studied the distribution of structures in a high dimensional space in which each structure is represented by a string of 1's and 0's, denoting core and surface sites, respectively.

preprint1998arXiv

Dynamics and stress in gravity driven granular flow

We study, using simulations, the steady-state flow of dry sand driven by gravity in two-dimensions. An investigation of the microscopic grain dynamics reveals that grains remain separated but with a power-law distribution of distances and times between collisions. While there are large random grain velocities, many of these fluctuations are correlated across the system and local rearrangements are very slow. Stresses in the system are almost entirely transfered by collisions and the structure of the stress tensor comes almost entirely from a bias in the directions in which collisions occur.