Catalog footprint

What is connected

213works
51topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

213 published item(s)

preprint2026arXiv

An Investigation into the Applicability of Friction Velocity Estimation Methods for the Channel with A Deposit Body

This study investigates how the deposit body influences friction characteristics by altering local flow fields, which is closely related to bed shear stress. Using generalized flume experiments, the study assesses the applicability of classical uniform flow friction models in deposit body river sections, revealing frictional changes induced by flow field non-uniformity. Initially, based on \( \frac{u_*}{\sqrt{\overline{w'^2}}} = 0.85 \sim 1.15 \) under uniform flow conditions as the judgment basis, the reliability of the classical model is verified. Four models are then applied to estimate near-bed friction velocity in deposit body sections. Results show a significant alignment between the longitudinal velocity gradient and the peak friction velocity derived from the turbulent kinetic energy method (TKE). Dimensional analysis of friction indicators reveals that: (a) friction velocity is primarily influenced by turbulence intensity, with constricted and narrowed sections resembling uniform flow, while the expansion section forms a peak; (b) models incorporating flow field fluctuations (TKE, Vertical Turbulence Kinetic Energy (TKE w'), Reynolds shear stress method (RSS)) effectively capture the impact of non-uniform flow fields on friction characteristics; (c) when energy states are low or when deposit body proportions are large, the deposit body's resistance ratio increases, and peak friction velocity rises. This study provides theoretical insights into friction estimation and sediment transport in non-uniform flow fields of deposit bodies.

preprint2026arXiv

Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models

Reinforcement learning from human feedback (RLHF) effectively promotes preference alignment of text-to-image (T2I) diffusion models. To improve computational efficiency, direct preference optimization (DPO), which avoids explicit reward modeling, has been widely studied. However, its reliance on binary feedback limits it to coarse-grained modeling on chosen-rejected pairs, resulting in suboptimal optimization. In this paper, we propose ArenaPO, which leverages Arena scores as offline rewards to provide refined feedback, thus achieving efficient and fine-grained optimization without a reward model. This enables ArenaPO to benefit from both the rich rewards of traditional RLHF and the efficiency of DPO. Specifically, we first construct a model Arena in which each model's capability is represented as a Gaussian distribution, and infer these capabilities by traversing the annotated pairwise preferences. Each output image is treated as a sample from the corresponding capability distribution. Then, for a image pair, conditioned on the two capability distributions and the observed pairwise preference, the absolute quality gap is estimated using latent-variable inference based on truncated normal distribution, which serves as fine-grained feedback during training. It does not require a reward model and can be computed offline, thus introducing no additional training overhead. We conduct ArenaPO training on Pick-a-Pic v2 and HPD v3 datasets, showing that ArenaPO consistently outperforms existing baselines.

preprint2026arXiv

Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization

Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transferability, casting doubt on the feasibility of transferable multimodal jailbreaks. We revisit this conclusion under a strictly untargeted threat model without enforcing a fixed prefix or response pattern. Our preliminary experiment reveals that refusal behavior concentrates at high-entropy tokens during autoregressive decoding, and non-refusal tokens already carry substantial probability mass among the top-ranked candidates before attack. Motivated by this finding, we propose Untargeted Jailbreak via Entropy Maximization(UJEM)-KL, a lightweight attack that maximizes entropy at these decision tokens to flip refusal outcomes, while stabilizing the remaining low-entropy positions to preserve output quality. Across three VLMs and two safety benchmarks, UJEM-KL achieves competitive white-box attack success rates and consistently improves transferability, while remaining effective under representative defenses. Our experimental results indicate that the limited transferability primarily stems from overly constrained optimization objectives.

preprint2026arXiv

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer accuracy. We introduce DocScope, a benchmark that formulates long-document QA as a structured reasoning trajectory prediction problem: given a complete PDF document and a question, the model outputs evidence pages, supporting evidence regions, relevant factual statements, and a final answer. We design a four-stage evaluation protocol -- Page Localization, Region Grounding, Fact Extraction, and Answer Verification -- that audits each level of the trajectory independently through inter-stage decoupling, with all judges selected and calibrated via human alignment studies. DocScope comprises 1,124 questions derived from 273 documents, with all hierarchical evidence annotations completed by human annotators. We benchmark 6 proprietary models, 12 open-weight models, and several domain-specific systems. Our experiments reveal that answer accuracy cannot substitute for trajectory-level evaluation: even among correct answers, the highest observed rate of complete evidence chains is only 29\%. Across all models, region grounding remains the weakest trajectory stage. Furthermore, the primary difficulty stems from aggregating evidence dispersed across long distances and multiple document clusters, while an oracle study identifies faithful perception and fact extraction as the dominant capability bottleneck. Cross-architecture comparisons further suggest that activated parameter count matters more than total scale. The benchmark and code will be publicly released at https://github.com/MiliLab/DocScope.

preprint2026arXiv

Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization but limited reasoning, whereas multimodal large language models (MLLMs) provide flexible reasoning but weak grounding in specialized medical domains. We present Echo-α, an agentic multimodal reasoning model for ultrasound interpretation that unifies these strengths within an invoke-and-reason framework. Echo-α is trained to coordinate organ-specific detector outputs, integrate them with global visual context, and convert the resulting evidence into grounded diagnostic decisions beyond detector-only inference. This behavior is established through a nine-task supervised curriculum and then refined by sequential reinforcement learning under different reward trade-offs, yielding Echo-α-Grounding for lesion anchoring and Echo-α-Diagnosis for final diagnosis. On multi-center renal and breast ultrasound benchmarks, Echo-α outperforms competitive baselines on both grounding and diagnosis. In particular, on cross-center test sets, Echo-α-Grounding attains 56.73%/43.78% F1@0.5 and Echo- α-Diagnosis reaches 74.90%/49.20% overall accuracy on renal/breast ultrasound. These results suggest that agentic multimodal reasoning can turn specialized detectors into verifiable clinical evidence, offering a practical route toward ultrasound AI systems that are more accurate, interpretable, and transferable. The repository is at https://github.com/MiliLab/Echo-Alpha.

preprint2026arXiv

EmoKGEdit: Training-free Affective Injection via Visual Cue Transformation

Existing image emotion editing methods struggle to disentangle emotional cues from latent content representations, often yielding weak emotional expression and distorted visual structures. To bridge this gap, we propose EmoKGEdit, a novel training-free framework for precise and structure-preserving image emotion editing. Specifically, we construct a Multimodal Sentiment Association Knowledge Graph (MSA-KG) to disentangle the intricate relationships among objects, scenes, attributes, visual clues and emotion. MSA-KG explicitly encode the causal chain among object-attribute-emotion, and as external knowledge to support chain of thought reasoning, guiding the multimodal large model to infer plausible emotion-related visual cues and generate coherent instructions. In addition, based on MSA-KG, we design a disentangled structure-emotion editing module that explicitly separates emotional attributes from layout features within the latent space, which ensures that the target emotion is effectively injected while strictly maintaining visual spatial coherence. Extensive experiments demonstrate that EmoKGEdit achieves excellent performance in both emotion fidelity and content preservation, and outperforms the state-of-the-art methods.

preprint2026arXiv

EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, an emotion semantic graph is constructed to capture the relational structure among emotions, objects, and visual attributes. To enhance the discriminability and transferability of emotion representations, we employ adversarial regularization, aligning the latent emotion distributions across modalities. Building upon EmoLat, a cross-modal sentiment transfer framework is proposed to manipulate image sentiment via joint embedding of text and EmoLat features. The network is optimized using a multi-objective loss incorporating semantic consistency, emotion alignment, and adversarial regularization. To support effective modeling, we construct EmoSpace Set, a large-scale benchmark dataset comprising images with dense annotations on emotions, object semantics, and visual attributes. Extensive experiments on EmoSpace Set demonstrate that our approach significantly outperforms existing state-of-the-art methods in both quantitative metrics and qualitative transfer fidelity, establishing a new paradigm for controllable image sentiment editing guided by textual input. The EmoSpace Set and all the code are available at http://github.com/JingVIPLab/EmoLat.

preprint2026arXiv

From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models

Latent actions serve as an intermediate representation that enables consistent modeling of vision-language-action (VLA) models across heterogeneous datasets. However, approaches to supervising VLAs with latent actions are fragmented and lack a systematic comparison. This work structures the study of latent action supervision from two perspectives: (i) regularizing the trajectory via image-based latent actions, and (ii) unifying the target space with action-based latent actions. Under a unified VLA baseline, we instantiate and compare four representative integration strategies. Our results reveal a formulation-task correspondence: image-based latent actions benefit long-horizon reasoning and scene-level generalization, whereas action-based latent actions excel at complex motor coordination. Furthermore, we find that directly supervising the VLM with discrete latent action tokens yields the most effective performance. Finally, our experiments offer initial insights into the benefits of latent action supervision in mixed-data, suggesting a promising direction for VLA training. Code is available at https://github.com/RUCKBReasoning/From_Pixels_to_Tokens.

preprint2026arXiv

GP-GS: Gaussian Processes Densification for 3D Gaussian Splatting

3D Gaussian Splatting (3DGS) enables photorealistic rendering but suffers from artefacts due to sparse Structure-from-Motion (SfM) initialisation. To address this limitation, we propose GP-GS, a Gaussian Process (GP) based densification framework for 3DGS optimisation. GP-GS formulates point cloud densification as a continuous regression problem, where a GP learns a local mapping from 2D pixel coordinates to 3D position and colour attributes. An adaptive neighbourhood-based sampling strategy generates candidate pixels for inference, while GP-predicted uncertainty is used to filter unreliable predictions, reducing noise and preserving geometric structure. Extensive experiments on synthetic and real-world benchmarks demonstrate that GP-GS consistently improves reconstruction quality and rendering fidelity, achieving up to 1.12 dB PSNR improvement over strong baselines.

preprint2026arXiv

High-Ti induced planar-fault transformation toward superlattice extrinsic stacking faults and microtwins in crept CoNi-based superalloys

Controlling planar fault shearing mechanisms is key for improving the high-temperature creep performance of gamma prime-strengthened high-temperature superalloys. This work examines how the Ti concentration in L12-strengthened CoNi-based alloys affects planar fault formation during creep. Interrupted compressive creep tests were conducted at 1223 K under air with a constant load stress of 241 MPa. We found, for the first time, that high Ti additions shift the dominant gamma prime shearing mode from antiphase boundaries (APBs) in Ti-free and low-Ti alloys to superlattice extrinsic stacking faults (SESFs). Systematic ab initio calculations show that in high-Ti alloys, the elevated APB energy renders APB-shearing mode unfavorable. Nevertheless, the SESF energy decreases relative to that in low-Ti compositions, and an increased ratio of complex intrinsic stacking fault (CISF) to SESF energy promote the transformation of high-energy CISFs into lower-energy SESFs. Chemical analysis using scanning transmission electron microscopy combined with energy-dispersive X-ray spectroscopy further reveals that, SESFs in high-Ti alloys are enriched in Ti, Mo and W, yet no grid-like ordering is observed. Together with the ab initio calculations, Mo and W additions in high Ti alloys could facilitate the transformation from L12 structure to low-energy D024 structure, indicating Mo and W segregation along SESFs is energetically favourable. Furthermore, the successive SESF thickening facilitates microtwinning in the absence of D024 ordering along SESFs, as an additional big carrier for creep strain. These new findings clarify the role of Ti in controlling planar fault shearing mechanisms, providing new insights for optimizing the creep performance of next-generation CoNi-based superalloys.

preprint2026arXiv

InkDiffuser: High-Fidelity One-shot Chinese Calligraphy via Differentiable Morphological Optimization

Current Chinese calligraphy generation methods suffer from poor stroke rendering and unrealistic ink morphology, resulting in outputs with limited visual fidelity and artistic fluidity. To address this problem, we propose \textbf{InkDiffuser}, a diffusion-based generative framework for one-shot Chinese calligraphy synthesis. To guarantee high-fidelity rendering, we introduce two core contributions: a high-frequency enhancement mechanism and a Differentiable Ink Structure (DIS) loss that explicitly regularizes ink morphology. Inspired by the observation that high-frequency information in individual samples typically carries contour details, we enhance content extraction by explicitly fusing high-frequency representations for more accurate font structure. Furthermore, we propose a differentiable ink structure loss that integrates differentiable morphological operations into the diffusion process. By allowing the model to learn an explicit decomposition of ink-trace structures, DIS facilitates fine-grained refinement of stroke contours and delivers significantly improved visual realism in the generated calligraphy. Extensive experiments on various calligraphic styles and complex characters demonstrate that InkDiffuser can generate superior calligraphy fonts with realistic ink rendering effects from only a single reference glyph and outperform existing few-shot font generation approaches in structural consistency, detail fidelity, and visual authenticity. The code is available at the following address: https://github.com/JingVIPLab/InkDiffuser.

preprint2026arXiv

InpaintHuman: Reconstructing Occluded Humans with Multi-Scale UV Mapping and Identity-Preserving Diffusion Inpainting

Reconstructing complete and animatable 3D human avatars from monocular videos remains challenging, particularly under severe occlusions. While 3D Gaussian Splatting has enabled photorealistic human rendering, existing methods struggle with incomplete observations, often producing corrupted geometry and temporal inconsistencies. We present InpaintHuman, a novel method for generating high-fidelity, complete, and animatable avatars from occluded monocular videos. Our approach introduces two key innovations: (i) a multi-scale UV-parameterized representation with hierarchical coarse-to-fine feature interpolation, enabling robust reconstruction of occluded regions while preserving geometric details; and (ii) an identity-preserving diffusion inpainting module that integrates textual inversion with semantic-conditioned guidance for subject-specific, temporally coherent completion. Unlike SDS-based methods, our approach employs direct pixel-level supervision to ensure identity fidelity. Experiments on synthetic benchmarks (PeopleSnapshot, ZJU-MoCap) and real-world scenarios (OcMotion) demonstrate competitive performance with consistent improvements in reconstruction quality across diverse poses and viewpoints.

preprint2026arXiv

Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions

The exponential growth of Large Language Models (LLMs) continues to highlight the need for efficient strategies to meet ever-expanding computational and data demands. This survey provides a comprehensive analysis of two complementary paradigms: Knowledge Distillation (KD) and Dataset Distillation (DD), both aimed at compressing LLMs while preserving their advanced reasoning capabilities and linguistic diversity. We first examine key methodologies in KD, such as task-specific alignment, rationale-based training, and multi-teacher frameworks, alongside DD techniques that synthesize compact, high-impact datasets through optimization-based gradient matching, latent space regularization, and generative synthesis. Building on these foundations, we explore how integrating KD and DD can produce more effective and scalable compression strategies. Together, these approaches address persistent challenges in model scalability, architectural heterogeneity, and the preservation of emergent LLM abilities. We further highlight applications across domains such as healthcare and education, where distillation enables efficient deployment without sacrificing performance. Despite substantial progress, open challenges remain in preserving emergent reasoning and linguistic diversity, enabling efficient adaptation to continually evolving teacher models and datasets, and establishing comprehensive evaluation protocols. By synthesizing methodological innovations, theoretical foundations, and practical insights, our survey charts a path toward sustainable, resource-efficient LLMs through the tighter integration of KD and DD principles.

preprint2026arXiv

Learning Cross-Atlas Consistent Brain Disorder Representations via Disentangled Multi-Atlas Functional Connectivity Learning

Functional connectivity (FC) derived from resting-state fMRI is widely used to characterize large-scale brain network alterations in neurological and psychiatric disorders. However, FC construction critically depends on the choice of brain atlas, and different parcellations may emphasize distinct organizational features, leading to heterogeneous and sometimes inconsistent representations. Existing multi-atlas approaches partially alleviate this issue but often fuse atlas-derived features or predictions at a relatively shallow level, while single-atlas disentanglement methods do not explicitly address cross-atlas heterogeneity. We propose Multi-Atlas Disentangled Connectivity LEarning (MADCLE), a multi-branch representation learning framework that jointly encodes FC matrices derived from different brain atlases. Rather than introducing a single explicitly shared latent variable across parcellations, MADCLE learns atlas-wise disease-related representations and encourages them to be cross-atlas consistent through distributional alignment. Meanwhile, covariate-related and atlas-dependent residual factors are modeled separately using covariate similarity supervision, atlas-specific reconstruction, and decorrelation constraints, thereby reducing the leakage of non-disease and parcellation-dependent information into the disease-related embeddings. Experiments on the ADNI and ADHD-200 datasets suggest that MADCLE achieves competitive or improved performance compared with single-atlas baselines, multi-atlas GNN/Transformer models, and recent multi-atlas consistency frameworks. These results support the potential value of structured disentanglement for FC-based disorder identification under heterogeneous parcellation schemes.

preprint2026arXiv

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

Large Language Models (LLMs) have demonstrated remarkable capabilities. However, their massive parameter scale leads to significant resource consumption and latency during inference. Post-training weight-only quantization offers a promising solution by reducing model size and accelerating token generation through alleviating the memory-bound issue. Nevertheless, the presence of inherent systematic outliers in weights continues to be a major obstacle. While existing methods, such as scaling and rotation, attempt to address this issue, the performance remains unsatisfactory. In this paper, we propose Outlier Self-Absorption Quantization (OSAQ), which performs additive weight suppression guided by the second-order low-rank property for low-bit weight-only quantization of LLMs. Specifically, we observe that the Hessian exhibits low-rank consistency across different inputs, with certain directions consistently showing vanishing curvature. Leveraging this property, we identify a stable null space of the Hessian and then construct an additive weight transformation by linearly combining the vectors within this null space, thereby suppressing weight outliers without affecting the task loss. This additive transformation can be absorbed into the weights offline, requiring no inter-layer transformations and introducing no inference overhead. Moreover, the construction is efficiently achieved by a closed-form solution, without resource-intensive training or iterative procedures. Extensive experiments demonstrate that OSAQ effectively suppresses outliers and enhances low-bit quantization performance. For instance, in 2-bit quantization, OSAQ, when integrated with GPTQ, achieves over 40% lower perplexity compared to vanilla GPTQ.

preprint2026arXiv

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition

Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referring cues. The task is challenging because the relevance of different modalities varies across referring expressions and scenes, while existing methods typically treat multimodal cues as homogeneous inputs for fusion, prompting, or reasoning, making them vulnerable to irrelevant or misleading modalities. To address this problem, we propose PRIMED, inspired by the biased competition theory in cognitive neuroscience, which explicitly models both visual perception and language-driven prior modulation, and enables more accurate Ref-AVS by adaptive modality suppression. Specifically, a Modality Prior Decoder first estimates whether the referring expression relies primarily on audio, vision, or their joint interaction, generating a modality prior to adaptively guide high-level attention. A Token Distiller further extracts compact global visual tokens from high-level features and shares them across Competition-aware Cross-modal Fusion modules to provide hierarchical global context. Additionally, we introduce a Spatial-Aware Semantic Alignment loss to further enhance foreground-background discrimination through contrastive learning. Extensive experiments on the Ref-AVS benchmark demonstrate that PRIMED achieves state-of-the-art overall performance.

preprint2026arXiv

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation

Long-horizon robotic manipulation requires dense feedback that reflects how a task advances through its procedural stages, not merely whether the final outcome is successful. Existing reward models often rely on trajectory-level success labels or time-based interpolation, which can conflate elapsed time with true task progress and therefore fail to capture unfinished steps, stagnation, and failure states. We present ProcVLM, a progress-aware vision-language model that learns procedure-grounded progress as a dense reward signal for manipulation. Rather than deriving progress from terminal outcomes or temporal proxies, ProcVLM grounds progress estimation in procedural structure and intra-stage visual change, and further adopts a reasoning-before-estimation paradigm that infers the remaining atomic actions before estimating task progress. Specifically, we construct this supervision by synthesizing frame-level subtask-semantic annotations, assigning progress budgets according to subtask structure, and distributing each budget based on intra-subtask visual change. To train ProcVLM at scale, we build a standardized procedural supervision synthesis pipeline and construct ProcCorpus-60M from 30 embodied datasets with 60M annotated frames, from which we derive ProcVQA for procedure-aware pretraining, with progress estimation as the central task alongside action segmentation and future planning. Experiments on ProcVQA and reward-model benchmarks show that ProcVLM improves embodied procedural reasoning and yields more discriminative trajectory-internal progress estimates than representative baselines, supporting its use as a dense reward model for downstream reward-guided policy optimization. Project page: https://procvlm.github.io/

preprint2026arXiv

SAMe: A Semantic Anatomy Mapping Engine for Robotic Ultrasound

Robotic ultrasound has advanced local image-driven control, contact regulation, and view optimization, yet current systems lack the anatomical understanding needed to determine what to scan, where to begin, and how to adapt to individual patient anatomy. These gaps make systems still reliant on expert intervention to initiate scanning. Here we present SAMe, a semantic anatomy mapping engine that provides robotic ultrasound with an explicit anatomical prior layer. SAMe addresses scan initiation as a target-to-anatomy-to-action process: it grounds under-specified clinical complaints into structured target organs, instantiates a patient-specific anatomical representation for the grounded targets from a single external body image, and translates this representation into control-facing 6-DoF probe initialization states without any additional registration using preoperative CT or MRI. The anatomical representation maintained by SAMe is explicit, lightweight (single-organ inference in 0.08s), and compatible with downstream control by design. Across semantic grounding, anatomical instantiation, and real-robot evaluation, SAMe shows strong performance across the full initialization pipeline. In real-robot experiments, centroid-based SAMe initialization outperformed the body-keypoint-based heuristic baseline under a budget-matched single-target setting for both liver (86.7% versus 46.7%) and kidney (80.0% versus 73.3%) initialization. Furthermore, The trial-level organ-hit rate reached 97.3% for liver and 83.3% for kidney when multiple candidate targets were available. These results establish an explicit anatomical prior layer that addresses scan initialization and is designed to support broader downstream autonomous scanning pipelines, providing the anatomical foundation for complaint-driven, anatomically informed robotic ultrasonography.

preprint2026arXiv

Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning

We present Seirênes, a self-play RL framework that transforms contextual interference from a failure mode of LLM reasoning into an internal training signal for co-evolving more resilient reasoners. While RL with verifiable rewards has significantly advanced reasoning capabilities, models can still exhibit fragility when encountering non-idealized contexts: scenarios characterized by superfluous information, tangential instructions, or incidental correlations that differ from the clean distributions typical of standard benchmarks. Seirênes harnesses this vulnerability through a parameter-shared and adversarial self-play loop. Within this framework, a single model is trained to both construct plausible yet distracting contexts that expose its own reasoning blind spots, and solve problems by discerning the essential task from these perturbations to recover the core underlying logic. By pitting these competing objectives against each other, Seirênes compels the model to move beyond superficial pattern matching and anchors its capabilities in robust underlying reasoning. This continuous interaction sustains an informative co-evolutionary curriculum as the model improves. Across seven mathematical reasoning benchmarks and model scales from 4B to 30B, Seirênes achieves average gains of +10.2, +9.1, and +7.2 points. Besides, distracting contexts produced by the 4B Seirênes model reduce the accuracy of top-tier closed-source models (GPT and Gemini) by roughly 4--5 points, revealing Seirênes' general ability to uncover reasoning models' blind spots.

preprint2026arXiv

Structural Energy Guidance for View-Consistent Text-to-3D Generation

Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work identifies viewpoint bias in 2D diffusion priors as the main cause and proposes Structural Energy-Guided Sampling (SEGS), a training-free and plug-and-play framework to improve multi-view consistency. SEGS constructs a structural energy in the PCA subspace of U-Net features and injects its gradient into the denoising process. It can be easily integrated into SDS/VSD pipelines without retraining. Experiments show that SEGS reduces the Janus Rate by about 10% on average and improves View-CS scores across multiple baselines, including DreamFusion, Magic3D, and LucidDreamer. This method effectively alleviates viewpoint artifacts while preserving appearance fidelity, providing a flexible solution for high-quality text-to-3D content generation.

preprint2026arXiv

TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL

In Text-to-SQL tasks, existing LLM-based methods often include extensive database schemas in prompts, leading to long context lengths and increased prefilling latency. While user queries typically focus on recurrent table sets-offering an opportunity for KV cache sharing across queries-current inference engines, such as SGLang and vLLM, generate redundant prefix cache copies when processing user queries with varying table orders. To address this inefficiency, we propose precomputing table representations as KV caches offline and querying the required ones online. A key aspect of our approach is the computation of table caches while preserving primary foreign key relationships between tables. Additionally, we construct a Table Trie structure to facilitate efficient KV cache lookups during inference. To enhance cache performance, we introduce a cache management system with a query reranking strategy to improve cache hit rates and a computation loading pipeline for parallelizing model inference and cache loading. Experimental results show that our proposed TableCache achieves up to a 3.62x speedup in Time to First Token (TTFT) with negligible performance degradation.

preprint2026arXiv

UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs

Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-scale scene context and micro-scale targets. We refer to this empirical gap as a "resolution illusion": higher input resolution provides the appearance of richer visual detail, but does not necessarily yield reliable perception of spatially small, task-relevant evidence. To benchmark this challenge, we introduce UHR-Micro, a benchmark comprising 11,253 instructions grounded in 1,212 UHR images, designed to evaluate VLMs at the spatial limits of native Earth observation imagery. UHR-Micro spans diverse micro-target scales, context requirements, task families, and visual conditions, and provides diagnostic annotations that support controlled evaluation and fine-grained error attribution. Experiments with representative high-resolution VLMs show substantial failures in spatial grounding and evidence parsing, despite access to high-resolution inputs. Further analysis suggests that these failures are not fully resolved by increasing model capacity, but are closely tied to insufficient guidance in locating and using task-relevant micro-evidence. Motivated by this finding, we propose Micro-evidence Active Perception (MAP), a reference agent that decomposes queries into evidence-seeking steps, actively inspects candidate regions, and grounds its answers in localized observations. MAP-Agent improves micro-level perception by making high-resolution reasoning evidence-centered rather than image-centered. Together, UHR-Micro and MAP-Agent provide a diagnostic platform for evaluating, understanding, and advancing high-resolution reasoning in Earth observation VLMs. Datasets and source code were released at https://github.com/MiliLab/UHR-Micro.

preprint2026arXiv

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence. However, this task remains challenging as it requires bridging discrete code representations with continuous visual dynamics. Existing optimization-based methods often destroy topological consistency, while general-purpose LLMs rely on rigid CSS/SMIL transformations, failing to model geometry-level non-rigid deformations. To address these limitations, we present VAnim, the first LLM-based framework for open-domain text-to-SVG animation. We reconceptualize animation not as sequence generation, but as Sparse State Updates (SSU) on a persistent SVG DOM tree. This paradigm compresses sequence length by over 9.8x while preserving the SVG DOM structure and non-participating elements by construction. To enable precise control, we propose an Identification-First Motion Planning mechanism that grounds textual instructions in explicit visual entities. Furthermore, to overcome the non-differentiable nature of SVG rendering, we employ Rendering-Aware Reinforcement Learning via Group Relative Policy Optimization (GRPO). By leveraging a hybrid reward from a state-of-the-art video perception encoder, we align discrete code updates with high-fidelity visual feedback. We also introduce SVGAnim-134k, the first benchmark for vector animation. Extensive experiments demonstrate that VAnim significantly outperforms state-of-the-art baselines in semantic alignment and structural validity, with additional appendix metrics further validating motion quality and identity preservation.

preprint2026arXiv

VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despite the strong multimodal video understanding capabilities of recent Video-LLMs, their performance on existing Video TextVQA benchmarks remains limited. To better understand this gap, we conduct an upper-bound analysis through frame-wise question answering, counting a sample as correct if any frame yields the right answer, which significantly outperforms direct video-based inference and reveals a substantial performance gap. The results suggest that the primary bottleneck lies in the localization of key question-relevant evidence, rather than in reasoning capacity itself. Building on this insight, we propose a question-guided agent framework that explicitly anchors the relevant keyframes before answering. The approach operates effectively in a training-free setting and consistently surpasses direct video inference. With additional supervised fine-tuning (SFT) and reinforcement learning (RL), it achieves an average improvement of +12.12 in accuracy and +11.15 in ANLS across benchmarks, establishing new state-of-the-art results. Our study underscores the critical role of explicit keyframe anchoring for advancing Video TextVQA. The code will be publicly released.

preprint2026arXiv

XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression

Learning-based 3D visual geometry models have benefited substantially from large-scale transformers. Among these, StreamVGGT leverages frame-wise causal attention for strong streaming reconstruction, but suffers from unbounded KV cache growth, leading to escalating memory consumption and inference latency as input frames accumulate. We propose XStreamVGGT, a tuning-free approach that systematically compresses the KV cache through joint pruning and quantization, enabling extremely memory-efficient streaming inference. Specifically, redundant KVs originating from multi-view inputs are pruned through efficient token importance identification, enabling a fixed memory budget. Leveraging the unique distribution of KV tensors, we incorporate KV quantization to further reduce memory consumption. Extensive evaluations show that XStreamVGGT achieves mostly negligible performance degradation while substantially reducing memory usage by 4.42$\times$ and accelerating inference by 5.48$\times$, enabling scalable and practical streaming 3D applications. The code is available at https://github.com/ywh187/XStreamVGGT/.

preprint2025arXiv

GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking

Video text spotting (VTS) extends image text spotting (ITS) by adding text tracking, significantly increasing task complexity. Despite progress in VTS, existing methods still fall short of the performance seen in ITS. This paper identifies a key limitation in current video text spotters: limited recognition capability, even after extensive end-to-end training. To address this, we propose GoMatching++, a parameter- and data-efficient method that transforms an off-the-shelf image text spotter into a video specialist. The core idea lies in freezing the image text spotter and introducing a lightweight, trainable tracker, which can be optimized efficiently with minimal training data. Our approach includes two key components: (1) a rescoring mechanism to bridge the domain gap between image and video data, and (2) the LST-Matcher, which enhances the frozen image text spotter's ability to handle video text. We explore various architectures for LST-Matcher to ensure efficiency in both parameters and training data. As a result, GoMatching++ sets new performance records on challenging benchmarks such as ICDAR15-video, DSText, and BOVText, while significantly reducing training costs. To address the lack of curved text datasets in VTS, we introduce ArTVideo, a new benchmark featuring over 30% curved text with detailed annotations. We also provide a comprehensive statistical analysis and experimental results for ArTVideo. We believe that GoMatching++ and the ArTVideo benchmark will drive future advancements in video text spotting. The source code, models and dataset are publicly available at https://github.com/Hxyz-123/GoMatching.

preprint2024arXiv

Automated Detection of Myopic Maculopathy in MMAC 2023: Achievements in Classification, Segmentation, and Spherical Equivalent Prediction

Myopic macular degeneration is the most common complication of myopia and the primary cause of vision loss in individuals with pathological myopia. Early detection and prompt treatment are crucial in preventing vision impairment due to myopic maculopathy. This was the focus of the Myopic Maculopathy Analysis Challenge (MMAC), in which we participated. In task 1, classification of myopic maculopathy, we employed the contrastive learning framework, specifically SimCLR, to enhance classification accuracy by effectively capturing enriched features from unlabeled data. This approach not only improved the intrinsic understanding of the data but also elevated the performance of our classification model. For Task 2 (segmentation of myopic maculopathy plus lesions), we have developed independent segmentation models tailored for different lesion segmentation tasks and implemented a test-time augmentation strategy to further enhance the model's performance. As for Task 3 (prediction of spherical equivalent), we have designed a deep regression model based on the data distribution of the dataset and employed an integration strategy to enhance the model's prediction accuracy. The results we obtained are promising and have allowed us to position ourselves in the Top 6 of the classification task, the Top 2 of the segmentation task, and the Top 1 of the prediction task. The code is available at \url{https://github.com/liyihao76/MMAC_LaTIM_Solution}.

preprint2024arXiv

Enhancing RBF-FD Efficiency for Highly Non-Uniform Node Distributions via Adaptivity

Radial basis function generated finite-difference (RBF-FD) methods have recently gained popularity due to their flexibility with irregular node distributions. However, the convergence theories in the literature, when applied to nonuniform node distributions, require shrinking fill distance and do not take advantage of areas with high data density. Non-adaptive approach using same stencil size and degree of appended polynomial will have higher local accuracy at high density region, but has no effect on the overall order of convergence and could be a waste of computational power. This work proposes an adaptive RBF-FD method that utilizes the local data density to achieve a desirable order accuracy. By performing polynomial refinement and using adaptive stencil size based on data density, the adaptive RBF-FD method yields differentiation matrices with higher sparsity while achieving the same user-specified convergence order for nonuniform point distributions. This allows the method to better leverage regions with higher node density, maintaining both accuracy and efficiency compared to standard non-adaptive RBF-FD methods.

preprint2024arXiv

Minor Issues Escalated to Critical Levels in Large Samples: A Permutation-Based Fix

In the big data era, the need to reevaluate traditional statistical methods is paramount due to the challenges posed by vast datasets. While larger samples theoretically enhance accuracy and hypothesis testing power without increasing false positives, practical concerns about inflated Type-I errors persist. The prevalent belief is that larger samples can uncover subtle effects, necessitating dual consideration of p-value and effect size. Yet, the reliability of p-values from large samples remains debated. This paper warns that larger samples can exacerbate minor issues into significant errors, leading to false conclusions. Through our simulation study, we demonstrate how growing sample sizes amplify issues arising from two commonly encountered violations of model assumptions in real-world data and lead to incorrect decisions. This underscores the need for vigilant analytical approaches in the era of big data. In response, we introduce a permutation-based test to counterbalance the effects of sample size and assumption discrepancies by neutralizing them between actual and permuted data. We demonstrate that this approach effectively stabilizes nominal Type I error rates across various sample sizes, thereby ensuring robust statistical inferences even amidst breached conventional assumptions in big data. For reproducibility, our R codes are publicly available at: \url{https://github.com/ubcxzhang/bigDataIssue}.

preprint2024arXiv

Robust single-particle cryo-EM image denoising and restoration

Cryo-electron microscopy (cryo-EM) has achieved near-atomic level resolution of biomolecules by reconstructing 2D micrographs. However, the resolution and accuracy of the reconstructed particles are significantly reduced due to the extremely low signal-to-noise ratio (SNR) and complex noise structure of cryo-EM images. In this paper, we introduce a diffusion model with post-processing framework to effectively denoise and restore single particle cryo-EM images. Our method outperforms the state-of-the-art (SOTA) denoising methods by effectively removing structural noise that has not been addressed before. Additionally, more accurate and high-resolution three-dimensional reconstruction structures can be obtained from denoised cryo-EM images.

preprint2024arXiv

SimDistill: Simulated Multi-modal Distillation for BEV 3D Object Detection

Multi-view camera-based 3D object detection has become popular due to its low cost, but accurately inferring 3D geometry solely from camera data remains challenging and may lead to inferior performance. Although distilling precise 3D geometry knowledge from LiDAR data could help tackle this challenge, the benefits of LiDAR information could be greatly hindered by the significant modality gap between different sensory modalities. To address this issue, we propose a Simulated multi-modal Distillation (SimDistill) method by carefully crafting the model architecture and distillation strategy. Specifically, we devise multi-modal architectures for both teacher and student models, including a LiDAR-camera fusion-based teacher and a simulated fusion-based student. Owing to the ``identical'' architecture design, the student can mimic the teacher to generate multi-modal features with merely multi-view images as input, where a geometry compensation module is introduced to bridge the modality gap. Furthermore, we propose a comprehensive multi-modal distillation scheme that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space. Incorporating them together, our SimDistill can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment. Extensive experiments validate the effectiveness and superiority of SimDistill over state-of-the-art methods, achieving an improvement of 4.8\% mAP and 4.1\% NDS over the baseline detector. The source code will be released at https://github.com/ViTAE-Transformer/SimDistill.

preprint2023arXiv

Multi-modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation

3D point cloud semantic segmentation has a wide range of applications. Recently, weakly supervised point cloud segmentation methods have been proposed, aiming to alleviate the expensive and laborious manual annotation process by leveraging scene-level labels. However, these methods have not effectively exploited the rich geometric information (such as shape and scale) and appearance information (such as color and texture) present in RGB-D scans. Furthermore, current approaches fail to fully leverage the point affinity that can be inferred from the feature extraction network, which is crucial for learning from weak scene-level labels. Additionally, previous work overlooks the detrimental effects of the long-tailed distribution of point cloud data in weakly supervised 3D semantic segmentation. To this end, this paper proposes a simple yet effective scene-level weakly supervised point cloud segmentation method with a newly introduced multi-modality point affinity inference module. The point affinity proposed in this paper is characterized by features from multiple modalities (e.g., point cloud and RGB), and is further refined by normalizing the classifier weights to alleviate the detrimental effects of long-tailed distribution without the need of the prior of category distribution. Extensive experiments on the ScanNet and S3DIS benchmarks verify the effectiveness of our proposed method, which outperforms the state-of-the-art by ~4% to ~6% mIoU. Codes are released at https://github.com/Sunny599/AAAI24-3DWSSG-MMA.

preprint2023arXiv

Towards Deeper Understanding of Camouflaged Object Detection

Preys in the wild evolve to be camouflaged to avoid being recognized by predators. In this way, camouflage acts as a key defence mechanism across species that is critical to survival. To detect and segment the whole scope of a camouflaged object, camouflaged object detection (COD) is introduced as a binary segmentation task, with the binary ground truth camouflage map indicating the exact regions of the camouflaged objects. In this paper, we revisit this task and argue that the binary segmentation setting fails to fully understand the concept of camouflage. We find that explicitly modeling the conspicuousness of camouflaged objects against their particular backgrounds can not only lead to a better understanding about camouflage, but also provide guidance to designing more sophisticated camouflage techniques. Furthermore, we observe that it is some specific parts of camouflaged objects that make them detectable by predators. With the above understanding about camouflaged objects, we present the first triple-task learning framework to simultaneously localize, segment, and rank camouflaged objects, indicating the conspicuousness level of camouflage. As no corresponding datasets exist for either the localization model or the ranking model, we generate localization maps with an eye tracker, which are then processed according to the instance level labels to generate our ranking-based training and testing dataset. We also contribute the largest COD testing set to comprehensively analyse performance of the COD models. Experimental results show that our triple-task learning framework achieves new state-of-the-art, leading to a more explainable COD network. Our code, data, and results are available at: \url{https://github.com/JingZhang617/COD-Rank-Localize-and-Segment}.

preprint2022arXiv

A Multi-Strategy based Pre-Training Method for Cold-Start Recommendation

Cold-start problem is a fundamental challenge for recommendation tasks. The recent self-supervised learning (SSL) on Graph Neural Networks (GNNs) model, PT-GNN, pre-trains the GNN model to reconstruct the cold-start embeddings and has shown great potential for cold-start recommendation. However, due to the over-smoothing problem, PT-GNN can only capture up to 3-order relation, which can not provide much useful auxiliary information to depict the target cold-start user or item. Besides, the embedding reconstruction task only considers the intra-correlations within the subgraph of users and items, while ignoring the inter-correlations across different subgraphs. To solve the above challenges, we propose a multi-strategy based pre-training method for cold-start recommendation (MPT), which extends PT-GNN from the perspective of model architecture and pretext tasks to improve the cold-start recommendation performance. Specifically, in terms of the model architecture, in addition to the short-range dependencies of users and items captured by the GNN encoder, we introduce a Transformer encoder to capture long-range dependencies. In terms of the pretext task, in addition to considering the intra-correlations of users and items by the embedding reconstruction task, we add embedding contrastive learning task to capture inter-correlations of users and items. We train the GNN and Transformer encoders on these pretext tasks under the meta-learning setting to simulate the real cold-start scenario, making the model easily and rapidly being adapted to new cold-start users and items. Experiments on three public recommendation datasets show the superiority of the proposed MPT model against the vanilla GNN models, the pre-training GNN model on user/item embedding inference and the recommendation task.

preprint2022arXiv

A Roadmap for Big Model

With the rapid development of deep learning, training Big Models (BMs) for multiple downstream tasks becomes a popular paradigm. Researchers have achieved various outcomes in the construction of BMs and the BM application in many fields. At present, there is a lack of research work that sorts out the overall progress of BMs and guides the follow-up research. In this paper, we cover not only the BM technologies themselves but also the prerequisites for BM training and applications with BMs, dividing the BM review into four parts: Resource, Models, Key Technologies and Application. We introduce 16 specific BM-related topics in those four parts, they are Data, Knowledge, Computing System, Parallel Training System, Language Model, Vision Model, Multi-modal Model, Theory&Interpretability, Commonsense Reasoning, Reliability&Security, Governance, Evaluation, Machine Translation, Text Generation, Dialogue and Protein Research. In each topic, we summarize clearly the current studies and propose some future research directions. At the end of this paper, we conclude the further development of BMs in a more general view.

preprint2022arXiv

A State Transition Model for Mobile Notifications via Survival Analysis

Mobile notifications have become a major communication channel for social networking services to keep users informed and engaged. As more mobile applications push notifications to users, they constantly face decisions on what to send, when and how. A lack of research and methodology commonly leads to heuristic decision making. Many notifications arrive at an inappropriate moment or introduce too many interruptions, failing to provide value to users and spurring users' complaints. In this paper we explore unique features of interactions between mobile notifications and user engagement. We propose a state transition framework to quantitatively evaluate the effectiveness of notifications. Within this framework, we develop a survival model for badging notifications assuming a log-linear structure and a Weibull distribution. Our results show that this model achieves more flexibility for applications and superior prediction accuracy than a logistic regression model. In particular, we provide an online use case on notification delivery time optimization to show how we make better decisions, drive more user engagement, and provide more value to users.

preprint2022arXiv

AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation

Recently, medical report generation, which aims to automatically generate a long and coherent descriptive paragraph of a given medical image, has received growing research interests. Different from the general image captioning tasks, medical report generation is more challenging for data-driven neural models. This is mainly due to 1) the serious data bias: the normal visual regions dominate the dataset over the abnormal visual regions, and 2) the very long sequence. To alleviate above two problems, we propose an AlignTransformer framework, which includes the Align Hierarchical Attention (AHA) and the Multi-Grained Transformer (MGT) modules: 1) AHA module first predicts the disease tags from the input image and then learns the multi-grained visual features by hierarchically aligning the visual regions and disease tags. The acquired disease-grounded visual features can better represent the abnormal regions of the input image, which could alleviate data bias problem; 2) MGT module effectively uses the multi-grained features and Transformer framework to generate the long medical report. The experiments on the public IU-Xray and MIMIC-CXR datasets show that the AlignTransformer can achieve results competitive with state-of-the-art methods on the two datasets. Moreover, the human evaluation conducted by professional radiologists further proves the effectiveness of our approach.

preprint2022arXiv

Approximating diamond principles on products at an inaccessible cardinal

We isolate \emph{the approximating diamond principles}, which are consequences of the diamond principle at an inaccessible cardinal. We use these principles to find new methods for negating the diamond principle at large cardinals. Most notably, we demonstrate, using Gitik's overlapping extenders forcing, a new method to get the consistency of the failure of the diamond principle at a large cardinal $θ$ without changing cofinalities or adding fast clubs to $θ$. In addition, we show that the approximating diamond principles necessarily hold at a weakly compact cardinal. This result, combined with the fact that in all known models where the diamond principle fails the approximating diamond principles also fail at an inaccessible cardinal, exhibits essential combinatorial obstacles to make the diamond principle fail at a weakly compact cardinal.

preprint2022arXiv

Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images

Semantic change detection (SCD) extends the multi-class change detection (MCD) task to provide not only the change locations but also the detailed land-cover/land-use (LCLU) categories before and after the observation intervals. This fine-grained semantic change information is very useful in many applications. Recent studies indicate that the SCD can be modeled through a triple-branch Convolutional Neural Network (CNN), which contains two temporal branches and a change branch. However, in this architecture, the communications between the temporal branches and the change branch are insufficient. To overcome the limitations in existing methods, we propose a novel CNN architecture for the SCD, where the semantic temporal features are merged in a deep CD unit. Furthermore, we elaborate on this architecture to reason the bi-temporal semantic correlations. The resulting Bi-temporal Semantic Reasoning Network (Bi-SRNet) contains two types of semantic reasoning blocks to reason both single-temporal and cross-temporal semantic correlations, as well as a novel loss function to improve the semantic consistency of change detection results. Experimental results on a benchmark dataset show that the proposed architecture obtains significant accuracy improvements over the existing approaches, while the added designs in the Bi-SRNet further improves the segmentation of both semantic categories and the changed areas. The codes in this paper are accessible at: github.com/ggsDing/Bi-SRNet.

preprint2022arXiv

Blockchain-assisted Undisclosed IIoT Vulnerabilities Trusted Sharing Protection with Dynamic Token

With the large-scale deployment of industrial internet of things (IIoT) devices, the number of vulnerabilities that threaten IIoT security is also growing dramatically, including a mass of undisclosed IIoT vulnerabilities that lack mitigation measures. Coordination Vulnerabilities Disclosure (CVD) is one of the most popular vulnerabilities sharing solutions, in which some security workers (SWs) can develop undisclosed vulnerabilities patches together. However, CVD assumes that sharing participants (SWs) are all honest, and thus offering chances for dishonest SWs to leak undisclosed IIoT vulnerabilities. To combat such threats, we propose an Undisclosed IIoT Vulnerabilities Trusted Sharing Protection (UIV-TSP) scheme with dynamic token. In this article, a dynamic token is an implicit access credential for an SW to acquire an undisclosed vulnerability information, which is only held by the system and constantly updated as the SW access. Meanwhile, the latest updated token can be stealthily sneaked into the acquired information as the traceability token. Once the undisclosed vulnerability information leaves the SW host, the embedded self-destruct program will be automatically triggered to prevent leaks since the destination MAC address in the traceability token has changed. To quickly distinguish dishonest SWs, trust mechanism is adopted to evaluate the trust value of SWs. Moreover, we design a blockchain-assisted continuous logs storage method to achieve the tamper-proofing of dynamic token and the transparency of undisclosed IIoT vulnerabilities sharing. The simulation results indicate that our proposed scheme is resilient to suppress dishonest SWs and protect the IoT undisclosed vulnerabilities effectively.

preprint2022arXiv

BMD: A General Class-balanced Multicentric Dynamic Prototype Strategy for Source-free Domain Adaptation

Source-free Domain Adaptation (SFDA) aims to adapt a pre-trained source model to the unlabeled target domain without accessing the well-labeled source data, which is a much more practical setting due to the data privacy, security, and transmission issues. To make up for the absence of source data, most existing methods introduced feature prototype based pseudo-labeling strategies to realize self-training model adaptation. However, feature prototypes are obtained by instance-level predictions based feature clustering, which is category-biased and tends to result in noisy labels since the visual domain gaps between source and target are usually different between categories. In addition, we found that a monocentric feature prototype may be ineffective to represent each category and introduce negative transfer, especially for those hard-transfer data. To address these issues, we propose a general class-Balanced Multicentric Dynamic prototype (BMD) strategy for the SFDA task. Specifically, for each target category, we first introduce a global inter-class balanced sampling strategy to aggregate potential representative target samples. Then, we design an intra-class multicentric clustering strategy to achieve more robust and representative prototypes generation. In contrast to existing strategies that update the pseudo label at a fixed training period, we further introduce a dynamic pseudo labeling strategy to incorporate network update information during model adaptation. Extensive experiments show that the proposed model-agnostic BMD strategy significantly improves representative SFDA methods to yield new state-of-the-art results. The code is available at https://github.com/ispc-lab/BMD.

preprint2022arXiv

CODE: Contrastive Pre-training with Adversarial Fine-tuning for Zero-shot Expert Linking

Expert finding, a popular service provided by many online websites such as Expertise Finder, LinkedIn, and AMiner, is beneficial to seeking candidate qualifications, consultants, and collaborators. However, its quality is suffered from lack of ample sources of expert information. This paper employs AMiner as the basis with an aim at linking any external experts to the counterparts on AMiner. As it is infeasible to acquire sufficient linkages from arbitrary external sources, we explore the problem of zero-shot expert linking. In this paper, we propose CODE, which first pre-trains an expert linking model by contrastive learning on AMiner such that it can capture the representation and matching patterns of experts without supervised signals, then it is fine-tuned between AMiner and external sources to enhance the models transferability in an adversarial manner. For evaluation, we first design two intrinsic tasks, author identification and paper clustering, to validate the representation and matching capability endowed by contrastive learning. Then the final external expert linking performance on two genres of external sources also implies the superiority of the adversarial fine-tuning method. Additionally, we show the online deployment of CODE, and continuously improve its online performance via active learning.

preprint2022arXiv

Compactness and Guessing Principles in the Radin Extensions

We investigate the interaction between compactness principles and guessing principles in the Radin forcing extensions. In particular, we show that in any Radin forcing extension with respect to a measure sequence on $κ$, if $κ$ is weakly compact, then $\diamondsuit(κ)$ holds. This provides contrast with a well-known theorem of Woodin, who showed that in a certain Radin extension over a suitably prepared ground model relative to the existence of large cardinals, the diamond principle fails at a strongly inaccessible Mahlo cardinal. Refining the analysis of the Radin extensions, we consistently demonstrate a scenario where a compactness principle, stronger than the diagonal stationary reflection principle, holds yet the diamond principle fails at a strongly inaccessible cardinal, improving a result from \cite{BN19}.

preprint2022arXiv

DearKD: Data-Efficient Early Knowledge Distillation for Vision Transformers

Transformers are successfully applied to computer vision due to their powerful modeling capacity with self-attention. However, the excellent performance of transformers heavily depends on enormous training images. Thus, a data-efficient transformer solution is urgently needed. In this work, we propose an early knowledge distillation framework, which is termed as DearKD, to improve the data efficiency required by transformers. Our DearKD is a two-stage framework that first distills the inductive biases from the early intermediate layers of a CNN and then gives the transformer full play by training without distillation. Further, our DearKD can be readily applied to the extreme data-free case where no real images are available. In this case, we propose a boundary-preserving intra-divergence loss based on DeepInversion to further close the performance gap against the full-data counterpart. Extensive experiments on ImageNet, partial ImageNet, data-free setting and other downstream tasks prove the superiority of DearKD over its baselines and state-of-the-art methods.

preprint2022arXiv

Direct Visualization and Manipulation of Tunable Quantum Well State in Semiconducting Nb2SiTe4

Quantum well states (QWSs) can form at the surface or interfaces of materials with confinement potential. They have broad applications in electronic and optical devices such as high mobility electron transistor, photodetector and quantum well laser. The properties of the QWSs are usually the key factors for the performance of the devices. However, direct visualization and manipulation of such states are in general challenging. In this work, by using angle-resolved photoemission spectroscopy (ARPES) and scanning tunneling microscopy/spectroscopy (STM/STS), we directly probe the QWSs generated on the vacuum interface of a narrow band gap semiconductor Nb2SiTe4. Interestingly, the position and splitting of QWSs could be easily manipulated via potassium (K) dosage onto the sample surface. Our results suggest Nb2SiTe4 to be an intriguing semiconductor system to study and engineer the QWSs, which has great potential in device applications.

preprint2022arXiv

Durable and Recoverable Hydrophilicity of Polyethylene Terephthalate Fabric Prepared with Plasma Selective Etching

Durable delustered PET (PET-TiO2) fabrics super hydrophilic surface has been obtained by plasma selecting etching. The aging effect of their hydrophilicity after plasma treatment has been investigated with storage time. After Ar/O2 radio frequency (RF) plasma treatment for only 7 min, PET-TiO2 fabric showed water contact angle of 0o. After 10 month storage time, it keeps its water contact angle below 75.7o. Further more, with Xenon light irradiation for 10 min, it is firstly found that it has well-recovered water contact angle to 5°. While the contact angle of PET fabric for 7 min returns to 123.0° and its hydrophilicity disappeared almost completely and showed no response to Xenon light irradiation. The water absorption rate of 7 min plasma treated PET-TiO2 fabric increased by 57.54%. By field emission scanning electron microscopy (FE-SEM), X-ray photoelectron spectroscopy (XPS) and X-ray diffraction analysis(XRD) measurement, waviness structure of humps and ridges with irregular particles or pits were found on the plasma treated PET-TiO2 fabric surface and increased Ti atomic percentage was observed. It is verified that TiO2 particles inside PET-TiO2 fiber have been exposed to its surface by plasma selective etching of its organic component. It suppresses the aging effect and is characterized with durable and recoverable hydrophilicity. This one step, quick, green and cost-resonable manufacture method has a pratical application for durable superhydrophilic surfaces.

preprint2022arXiv

DUT: Learning Video Stabilization by Simply Watching Unstable Videos

Previous deep learning-based video stabilizers require a large scale of paired unstable and stable videos for training, which are difficult to collect. Traditional trajectory-based stabilizers, on the other hand, divide the task into several sub-tasks and tackle them subsequently, which are fragile in textureless and occluded regions regarding the usage of hand-crafted features. In this paper, we attempt to tackle the video stabilization problem in a deep unsupervised learning manner, which borrows the divide-and-conquer idea from traditional stabilizers while leveraging the representation power of DNNs to handle the challenges in real-world scenarios. Technically, DUT is composed of a trajectory estimation stage and a trajectory smoothing stage. In the trajectory estimation stage, we first estimate the motion of keypoints, initialize and refine the motion of grids via a novel multi-homography estimation strategy and a motion refinement network, respectively, and get the grid-based trajectories via temporal association. In the trajectory smoothing stage, we devise a novel network to predict dynamic smoothing kernels for trajectory smoothing, which can well adapt to trajectories with different dynamic patterns. We exploit the spatial and temporal coherence of keypoints and grid vertices to formulate the training objectives, resulting in an unsupervised training scheme. Experiment results on public benchmarks show that DUT outperforms state-of-the-art methods both qualitatively and quantitatively. The source code is available at https://github.com/Annbless/DUTCode.

preprint2022arXiv

Energy-Based Generative Cooperative Saliency Prediction

Conventional saliency prediction models typically learn a deterministic mapping from an image to its saliency map, and thus fail to explain the subjective nature of human attention. In this paper, to model the uncertainty of visual saliency, we study the saliency prediction problem from the perspective of generative models by learning a conditional probability distribution over the saliency map given an input image, and treating the saliency prediction as a sampling process from the learned distribution. Specifically, we propose a generative cooperative saliency prediction framework, where a conditional latent variable model (LVM) and a conditional energy-based model (EBM) are jointly trained to predict salient objects in a cooperative manner. The LVM serves as a fast but coarse predictor to efficiently produce an initial saliency map, which is then refined by the iterative Langevin revision of the EBM that serves as a slow but fine predictor. Such a coarse-to-fine cooperative saliency prediction strategy offers the best of both worlds. Moreover, we propose a "cooperative learning while recovering" strategy and apply it to weakly supervised saliency prediction, where saliency annotations of training images are partially observed. Lastly, we find that the learned energy function in the EBM can serve as a refinement module that can refine the results of other pre-trained saliency prediction models. Experimental results show that our model can produce a set of diverse and plausible saliency maps of an image, and obtain state-of-the-art performance in both fully supervised and weakly supervised saliency prediction tasks.

preprint2022arXiv

Exploring Depth Contribution for Camouflaged Object Detection

Camouflaged object detection (COD) aims to segment camouflaged objects hiding in the environment, which is challenging due to the similar appearance of camouflaged objects and their surroundings. Research in biology suggests depth can provide useful object localization cues for camouflaged object discovery. In this paper, we study the depth contribution for camouflaged object detection, where the depth maps are generated with existing monocular depth estimation (MDE) methods. Due to the domain gap between the MDE dataset and our COD dataset, the generated depth maps are not accurate enough to be directly used. We then introduce two solutions to avoid the noisy depth maps from dominating the training process. Firstly, we present an auxiliary depth estimation branch ("ADE"), aiming to regress the depth maps. We find that "ADE" is especially necessary for our "generated depth" scenario. Secondly, we introduce a multi-modal confidence-aware loss function via a generative adversarial network to weigh the contribution of depth for camouflaged object detection. Our extensive experiments on various camouflaged object detection datasets explain that the existing "sensor depth" based RGB-D segmentation techniques work poorly with "generated depth", and our proposed two solutions work cooperatively, achieving effective depth contribution exploration for camouflaged object detection.

preprint2022arXiv

Exploring Sequence Feature Alignment for Domain Adaptive Detection Transformers

Detection transformers have recently shown promising object detection results and attracted increasing attention. However, how to develop effective domain adaptation techniques to improve its cross-domain performance remains unexplored and unclear. In this paper, we delve into this topic and empirically find that direct feature distribution alignment on the CNN backbone only brings limited improvements, as it does not guarantee domain-invariant sequence features in the transformer for prediction. To address this issue, we propose a novel Sequence Feature Alignment (SFA) method that is specially designed for the adaptation of detection transformers. Technically, SFA consists of a domain query-based feature alignment (DQFA) module and a token-wise feature alignment (TDA) module. In DQFA, a novel domain query is used to aggregate and align global context from the token sequence of both domains. DQFA reduces the domain discrepancy in global feature representations and object relations when deploying in the transformer encoder and decoder, respectively. Meanwhile, TDA aligns token features in the sequence from both domains, which reduces the domain gaps in local and instance-level feature representations in the transformer encoder and decoder, respectively. Besides, a novel bipartite matching consistency loss is proposed to enhance the feature discriminability for robust object detection. Experiments on three challenging benchmarks show that SFA outperforms state-of-the-art domain adaptive object detection methods. Code has been made available at: https://github.com/encounter1997/SFA.

preprint2022arXiv

FakeCLR: Exploring Contrastive Learning for Solving Latent Discontinuity in Data-Efficient GANs

Data-Efficient GANs (DE-GANs), which aim to learn generative models with a limited amount of training data, encounter several challenges for generating high-quality samples. Since data augmentation strategies have largely alleviated the training instability, how to further improve the generative performance of DE-GANs becomes a hotspot. Recently, contrastive learning has shown the great potential of increasing the synthesis quality of DE-GANs, yet related principles are not well explored. In this paper, we revisit and compare different contrastive learning strategies in DE-GANs, and identify (i) the current bottleneck of generative performance is the discontinuity of latent space; (ii) compared to other contrastive learning strategies, Instance-perturbation works towards latent space continuity, which brings the major improvement to DE-GANs. Based on these observations, we propose FakeCLR, which only applies contrastive learning on perturbed fake samples, and devises three related training techniques: Noise-related Latent Augmentation, Diversity-aware Queue, and Forgetting Factor of Queue. Our experimental results manifest the new state of the arts on both few-shot generation and limited-data generation. On multiple datasets, FakeCLR acquires more than 15% FID improvement compared to existing DE-GANs. Code is available at https://github.com/iceli1007/FakeCLR.

preprint2022arXiv

FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis

In recent years, the security of AI systems has drawn increasing research attention, especially in the medical imaging realm. To develop a secure medical image analysis (MIA) system, it is a must to study possible backdoor attacks (BAs), which can embed hidden malicious behaviors into the system. However, designing a unified BA method that can be applied to various MIA systems is challenging due to the diversity of imaging modalities (e.g., X-Ray, CT, and MRI) and analysis tasks (e.g., classification, detection, and segmentation). Most existing BA methods are designed to attack natural image classification models, which apply spatial triggers to training images and inevitably corrupt the semantics of poisoned pixels, leading to the failures of attacking dense prediction models. To address this issue, we propose a novel Frequency-Injection based Backdoor Attack method (FIBA) that is capable of delivering attacks in various MIA tasks. Specifically, FIBA leverages a trigger function in the frequency domain that can inject the low-frequency information of a trigger image into the poisoned image by linearly combining the spectral amplitude of both images. Since it preserves the semantics of the poisoned image pixels, FIBA can perform attacks on both classification and dense prediction models. Experiments on three benchmarks in MIA (i.e., ISIC-2019 for skin lesion classification, KiTS-19 for kidney tumor segmentation, and EAD-2019 for endoscopic artifact detection), validate the effectiveness of FIBA and its superiority over state-of-the-art methods in attacking MIA models as well as bypassing backdoor defense. Source code will be available at https://github.com/HazardFY/FIBA.

preprint2022arXiv

Generative Transformer for Accurate and Reliable Salient Object Detection

Transformer, which originates from machine translation, is particularly powerful at modeling long-range dependencies. Currently, the transformer is making revolutionary progress in various vision tasks, leading to significant performance improvements compared with the convolutional neural network (CNN) based frameworks. In this paper, we conduct extensive research on exploiting the contributions of transformers for accurate and reliable salient object detection. For the former, we apply transformer to a deterministic model, and explain that the effective structure modeling and global context modeling abilities lead to its superior performance compared with the CNN based frameworks. For the latter, we observe that both CNN and transformer based frameworks suffer greatly from the over-confidence issue, where the models tend to generate wrong predictions with high confidence. To estimate the reliability degree of both CNN- and transformer-based frameworks, we further present a latent variable model, namely inferential generative adversarial network (iGAN), based on the generative adversarial network (GAN). The stochastic attribute of the latent variable makes it convenient to estimate the predictive uncertainty, serving as an auxiliary output to evaluate the reliability of model prediction. Different from the conventional GAN, which defines the distribution of the latent variable as fixed standard normal distribution $\mathcal{N}(0,\mathbf{I})$, the proposed iGAN infers the latent variable by gradient-based Markov Chain Monte Carlo (MCMC), namely Langevin dynamics, leading to an input-dependent latent variable model. We apply our proposed iGAN to both fully and weakly supervised salient object detection, and explain that iGAN within the transformer framework leads to both accurate and reliable salient object detection.

preprint2022arXiv

GETAM: Gradient-weighted Element-wise Transformer Attention Map for Weakly-supervised Semantic segmentation

Weakly Supervised Semantic Segmentation (WSSS) is challenging, particularly when image-level labels are used to supervise pixel level prediction. To bridge their gap, a Class Activation Map (CAM) is usually generated to provide pixel level pseudo labels. CAMs in Convolutional Neural Networks suffer from partial activation ie, only the most discriminative regions are activated. Transformer based methods, on the other hand, are highly effective at exploring global context with long range dependency modeling, potentially alleviating the "partial activation" issue. In this paper, we propose the first transformer based WSSS approach, and introduce the Gradient weighted Element wise Transformer Attention Map (GETAM). GETAM shows fine scale activation for all feature map elements, revealing different parts of the object across transformer layers. Further, we propose an activation aware label completion module to generate high quality pseudo labels. Finally, we incorporate our methods into an end to end framework for WSSS using double backward propagation. Extensive experiments on PASCAL VOC and COCO demonstrate that our results beat the state-of-the-art end-to-end approaches by a significant margin, and outperform most multi-stage methods.m most multi-stage methods.

preprint2022arXiv

GMFlow: Learning Optical Flow via Global Matching

Learning-based optical flow estimation has been dominated with the pipeline of cost volume with convolutions for flow regression, which is inherently limited to local correlations and thus is hard to address the long-standing challenge of large displacements. To alleviate this, the state-of-the-art framework RAFT gradually improves its prediction quality by using a large number of iterative refinements, achieving remarkable performance but introducing linearly increasing inference time. To enable both high accuracy and efficiency, we completely revamp the dominant flow regression pipeline by reformulating optical flow as a global matching problem, which identifies the correspondences by directly comparing feature similarities. Specifically, we propose a GMFlow framework, which consists of three main components: a customized Transformer for feature enhancement, a correlation and softmax layer for global feature matching, and a self-attention layer for flow propagation. We further introduce a refinement step that reuses GMFlow at higher feature resolution for residual flow prediction. Our new framework outperforms 31-refinements RAFT on the challenging Sintel benchmark, while using only one refinement and running faster, suggesting a new paradigm for accurate and efficient optical flow estimation. Code is available at https://github.com/haofeixu/gmflow.

preprint2022arXiv

Graph Contrastive Learning for Anomaly Detection

Graph-based anomaly detection has been widely used for detecting malicious activities in real-world applications. Existing attempts to address this problem have thus far focused on structural feature engineering or learning in the binary classification regime. In this work, we propose to leverage graph contrastive coding and present the supervised GraphCAD model for contrasting abnormal nodes with normal ones in terms of their distances to the global context (e.g., the average of all nodes). To handle scenarios with scarce labels, we further enable GraphCAD as a self-supervised framework by designing a graph corrupting strategy for generating synthetic node labels. To achieve the contrastive objective, we design a graph neural network encoder that can infer and further remove suspicious links during message passing, as well as learn the global context of the input graph. We conduct extensive experiments on four public datasets, demonstrating that 1) GraphCAD significantly and consistently outperforms various advanced baselines and 2) its self-supervised version without fine-tuning can achieve comparable performance with its fully supervised version.

preprint2022arXiv

Halogenation induced transition of superconductor-to-semiconductor in MXene-like MOene with direct band gap and long carrier lifetime

Traditional MXenes with intriguing mechanical and electronic properties, together with the fertilities of elemental compositions and chemical decorations have aroused much attentions. However, the semiconducting traits with direc band gap are extremetely rare among reported MXenes. Thus, broadening the family of MXene beyond carbides and nitrides with unique behaviors is still an extraordinary and fascinating field.

preprint2022arXiv

I3CL:Intra- and Inter-Instance Collaborative Learning for Arbitrary-shaped Scene Text Detection

Existing methods for arbitrary-shaped text detection in natural scenes face two critical issues, i.e., 1) fracture detections at the gaps in a text instance; and 2) inaccurate detections of arbitrary-shaped text instances with diverse background context. To address these issues, we propose a novel method named Intra- and Inter-Instance Collaborative Learning (I3CL). Specifically, to address the first issue, we design an effective convolutional module with multiple receptive fields, which is able to collaboratively learn better character and gap feature representations at local and long ranges inside a text instance. To address the second issue, we devise an instance-based transformer module to exploit the dependencies between different text instances and a global context module to exploit the semantic context from the shared background, which are able to collaboratively learn more discriminative text feature representation. In this way, I3CL can effectively exploit the intra- and inter-instance dependencies together in a unified end-to-end trainable framework. Besides, to make full use of the unlabeled data, we design an effective semi-supervised learning method to leverage the pseudo labels via an ensemble strategy. Without bells and whistles, experimental results show that the proposed I3CL sets new state-of-the-art results on three challenging public benchmarks, i.e., an F-measure of 77.5% on ICDAR2019-ArT, 86.9% on Total-Text, and 86.4% on CTW-1500. Notably, our I3CL with the ResNeSt-101 backbone ranked 1st place on the ICDAR2019-ArT leaderboard. The source code will be available at https://github.com/ViTAE-Transformer/ViTAE-Transformer-Scene-Text-Detection.

preprint2022arXiv

Improving RGB-D Point Cloud Registration by Learning Multi-scale Local Linear Transformation

Point cloud registration aims at estimating the geometric transformation between two point cloud scans, in which point-wise correspondence estimation is the key to its success. In addition to previous methods that seek correspondences by hand-crafted or learnt geometric features, recent point cloud registration methods have tried to apply RGB-D data to achieve more accurate correspondence. However, it is not trivial to effectively fuse the geometric and visual information from these two distinctive modalities, especially for the registration problem. In this work, we propose a new Geometry-Aware Visual Feature Extractor (GAVE) that employs multi-scale local linear transformation to progressively fuse these two modalities, where the geometric features from the depth data act as the geometry-dependent convolution kernels to transform the visual features from the RGB data. The resultant visual-geometric features are in canonical feature spaces with alleviated visual dissimilarity caused by geometric changes, by which more reliable correspondence can be achieved. The proposed GAVE module can be readily plugged into recent RGB-D point cloud registration framework. Extensive experiments on 3D Match and ScanNet demonstrate that our method outperforms the state-of-the-art point cloud registration methods even without correspondence or pose supervision. The code is available at: https://github.com/514DNA/LLT.

preprint2022arXiv

Information-Theoretic Odometry Learning

In this paper, we propose a unified information theoretic framework for learning-motivated methods aimed at odometry estimation, a crucial component of many robotics and vision tasks such as navigation and virtual reality where relative camera poses are required in real time. We formulate this problem as optimizing a variational information bottleneck objective function, which eliminates pose-irrelevant information from the latent representation. The proposed framework provides an elegant tool for performance evaluation and understanding in information-theoretic language. Specifically, we bound the generalization errors of the deep information bottleneck framework and the predictability of the latent representation. These provide not only a performance guarantee but also practical guidance for model design, sample collection, and sensor selection. Furthermore, the stochastic latent representation provides a natural uncertainty measure without the needs for extra structures or computations. Experiments on two well-known odometry datasets demonstrate the effectiveness of our method.

preprint2022arXiv

Injecting Numerical Reasoning Skills into Knowledge Base Question Answering Models

Embedding-based methods are popular for Knowledge Base Question Answering (KBQA), but few current models have numerical reasoning skills and thus struggle to answer ordinal constrained questions. This paper proposes a new embedding-based KBQA framework which particularly takes numerical reasoning into account. We present NumericalTransformer on top of NSM, a state-of-the-art embedding-based KBQA model, to create NT-NSM. To enable better training, we propose two pre-training tasks with explicit numerical-oriented loss functions on two generated training datasets and a template-based data augmentation method for enriching ordinal constrained QA dataset. Extensive experiments on KBQA benchmarks demonstrate that with the help of our training algorithm, NT-NSM is empowered with numerical reasoning skills and substantially outperforms the baselines in answering ordinal constrained questions.

preprint2022arXiv

Interference of the scattered vector light fields from two optically levitated nanoparticles

We experimentally study the interference of dipole scattered light from two optically levitated nanoparticles in vacuum, which present an environment free of particle-substrate interactions. We illuminate the two trapped nanoparticles with a linearly polarized probe beam orthogonal to the propagation of the trapping laser beams. The scattered light from the nanoparticles are collected by a high numerical aperture (NA) objective lens and imaged. The interference fringes from the scattered vector light for the different dipole orientations in image and Fourier space are observed. Especially, the interference fringes of two scattered light fields with polarization vortex show the π shift of the interference fringes between inside and outside the center region of the two nanoparticles in the image space. As far as we know, this is the first experimental observation of the interference of scattered vector light fields from two dipoles in free space. This work also provides a simple and direct method to determine the spatial scales between optically levitated nanoparticles by the interference fringes.

preprint2022arXiv

JPerceiver: Joint Perception Network for Depth, Pose and Layout Estimation in Driving Scenes

Depth estimation, visual odometry (VO), and bird's-eye-view (BEV) scene layout estimation present three critical tasks for driving scene perception, which is fundamental for motion planning and navigation in autonomous driving. Though they are complementary to each other, prior works usually focus on each individual task and rarely deal with all three tasks together. A naive way is to accomplish them independently in a sequential or parallel manner, but there are many drawbacks, i.e., 1) the depth and VO results suffer from the inherent scale ambiguity issue; 2) the BEV layout is directly predicted from the front-view image without using any depth-related information, although the depth map contains useful geometry clues for inferring scene layouts. In this paper, we address these issues by proposing a novel joint perception framework named JPerceiver, which can simultaneously estimate scale-aware depth and VO as well as BEV layout from a monocular video sequence. It exploits the cross-view geometric transformation (CGT) to propagate the absolute scale from the road layout to depth and VO based on a carefully-designed scale loss. Meanwhile, a cross-view and cross-modal transfer (CCT) module is devised to leverage the depth clues for reasoning road and vehicle layout through an attention mechanism. JPerceiver can be trained in an end-to-end multi-task learning way, where the CGT scale loss and CCT module promote inter-task knowledge transfer to benefit feature learning of each task. Experiments on Argoverse, Nuscenes and KITTI show the superiority of JPerceiver over existing methods on all the above three tasks in terms of accuracy, model size, and inference speed. The code and models are available at~\href{https://github.com/sunnyHelen/JPerceiver}{https://github.com/sunnyHelen/JPerceiver}.

preprint2022arXiv

Knowledge Learning with Crowdsourcing: A Brief Review and Systematic Perspective

Big data have the characteristics of enormous volume, high velocity, diversity, value-sparsity, and uncertainty, which lead the knowledge learning from them full of challenges. With the emergence of crowdsourcing, versatile information can be obtained on-demand so that the wisdom of crowds is easily involved to facilitate the knowledge learning process. During the past thirteen years, researchers in the AI community made great efforts to remove the obstacles in the field of learning from crowds. This concentrated survey paper comprehensively reviews the technical progress in crowdsourcing learning from a systematic perspective that includes three dimensions of data, models, and learning processes. In addition to reviewing existing important work, the paper places a particular emphasis on providing some promising blueprints on each dimension as well as discussing the lessons learned from our past research work, which will light up the way for new researchers and encourage them to pursue new contributions.

preprint2022arXiv

Learning Affordance Grounding from Exocentric Images

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. Human has the ability that transform the various exocentric interactions to invariant egocentric affordance so as to counter the impact of interactive diversity. To empower an agent with such ability, this paper proposes a task of affordance grounding from exocentric view, i.e., given exocentric human-object interaction and egocentric object images, learning the affordance knowledge of the object and transferring it to the egocentric image using only the affordance label as supervision. To this end, we devise a cross-view knowledge transfer framework that extracts affordance-specific features from exocentric interactions and enhances the perception of affordance regions by preserving affordance correlation. Specifically, an Affordance Invariance Mining module is devised to extract specific clues by minimizing the intra-class differences originated from interaction habits in exocentric images. Besides, an Affordance Co-relation Preserving strategy is presented to perceive and localize affordance by aligning the co-relation matrix of predicted results between the two views. Particularly, an affordance grounding dataset named AGD20K is constructed by collecting and labeling over 20K images from 36 affordance categories. Experimental results demonstrate that our method outperforms the representative models in terms of objective metrics and visual quality. Code: github.com/lhc1224/Cross-View-AG.

preprint2022arXiv

Local Partial Zero-Forcing Combining for Cell-Free Massive MIMO Systems

Cell-free massive multiple-input multiple-output (MIMO) provides more uniform spectral efficiency (SE) for users (UEs) than cellular technology. The main challenge to achieve the benefits of cell-free massive MIMO is to realize signal processing in a scalable way. In this paper, we consider scalable fullpilot zero-forcing (FZF), partial FZF (PFZF), protective weak PFZF (PWPFZF), and local regularized ZF (LRZF) combining by exploiting channel statistics. We derive closed-form expressions of the uplink SE for FZF, PFZF, and PWPFZF combining with large-scale fading decoding over independent Rayleigh fading channels, taking channel estimation errors and pilot contamination into account. Moreover, we investigate the impact of the number of pilot sequences, antennas per AP, and APs on the performance. Numerical results show that LRZF provides the highest SE. However, PWPFZF is preferable when the number of pilot sequences is large and the number of antennas per AP is small. The reason is that PWPFZF has lower computational complexity and the SE expression can be computed in closed-form. Furthermore, we investigate the performance of PWPFZF combining with fractional power control and the numerical results show that it improves the performance of weak UEs and realizes uniformly good service for all UEs in a scalable fashion.

preprint2022arXiv

Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images

Long-range contextual information is crucial for the semantic segmentation of High-Resolution (HR) Remote Sensing Images (RSIs). However, image cropping operations, commonly used for training neural networks, limit the perception of long-range contexts in large RSIs. To overcome this limitation, we propose a Wide-Context Network (WiCoNet) for the semantic segmentation of HR RSIs. Apart from extracting local features with a conventional CNN, the WiCoNet has an extra context branch to aggregate information from a larger image area. Moreover, we introduce a Context Transformer to embed contextual information from the context branch and selectively project it onto the local features. The Context Transformer extends the Vision Transformer, an emerging kind of neural network, to model the dual-branch semantic correlations. It overcomes the locality limitation of CNNs and enables the WiCoNet to see the bigger picture before segmenting the land-cover/land-use (LCLU) classes. Ablation studies and comparative experiments conducted on several benchmark datasets demonstrate the effectiveness of the proposed method. In addition, we present a new Beijing Land-Use (BLU) dataset. This is a large-scale HR satellite dataset with high-quality and fine-grained reference labels, which can facilitate future studies in this field.

preprint2022arXiv

Low-Complexity Block Coordinate Descend Based Multiuser Detection for Uplink Grant-Free NOMA

Grant-free non-orthogonal multiple access (NOMA) scheme is considered as a promising candidate for the enabling of massive connectivity and reduced signalling overhead for Internet of Things (IoT) applications in massive machine-type communication (mMTC) networks. Exploiting the inherent nature of sporadic transmissions in the grant-free NOMA systems, compressed sensing based multiuser detection (CS-MUD) has been deemed as a powerful solution to user activity detection (UAD) and data detection (DD). In this paper, block coordinate descend (BCD) method is employed in CS-MUD to reduce the computational complexity. We propose two modified BCD based algorithms, called enhanced BCD (EBCD) and complexity reduction enhanced BCD (CR-EBCD), respectively. To be specific, by incorporating a novel candidate set pruning mechanism into the original BCD framework, our proposed EBCD algorithm achieves remarkable CS-MUD performance improvement. In addition, the proposed CR-EBCD algorithm further ameliorates the proposed EBCD by eliminating the redundant matrix multiplications during the iteration process. As a consequence, compared with the proposed EBCD algorithm, our proposed CR-EBCD algorithm enjoys two orders of magnitude complexity saving without any CS-MUD performance degradation, rendering it a viable solution for future mMTC scenarios. Extensive simulation results demonstrate the bound-approaching performance as well as ultra-low computational complexity.

preprint2022arXiv

MeshMAE: Masked Autoencoders for 3D Mesh Data Analysis

Recently, self-supervised pre-training has advanced Vision Transformers on various tasks w.r.t. different data modalities, e.g., image and 3D point cloud data. In this paper, we explore this learning paradigm for 3D mesh data analysis based on Transformers. Since applying Transformer architectures to new modalities is usually non-trivial, we first adapt Vision Transformer to 3D mesh data processing, i.e., Mesh Transformer. In specific, we divide a mesh into several non-overlapping local patches with each containing the same number of faces and use the 3D position of each patch's center point to form positional embeddings. Inspired by MAE, we explore how pre-training on 3D mesh data with the Transformer-based structure benefits downstream 3D mesh analysis tasks. We first randomly mask some patches of the mesh and feed the corrupted mesh into Mesh Transformers. Then, through reconstructing the information of masked patches, the network is capable of learning discriminative representations for mesh data. Therefore, we name our method MeshMAE, which can yield state-of-the-art or comparable performance on mesh analysis tasks, i.e., classification and segmentation. In addition, we also conduct comprehensive ablation studies to show the effectiveness of key designs in our method.

preprint2022arXiv

MetaCVR: Conversion Rate Prediction via Meta Learning in Small-Scale Recommendation Scenarios

Different from large-scale platforms such as Taobao and Amazon, CVR modeling in small-scale recommendation scenarios is more challenging due to the severe Data Distribution Fluctuation (DDF) issue. DDF prevents existing CVR models from being effective since 1) several months of data are needed to train CVR models sufficiently in small scenarios, leading to considerable distribution discrepancy between training and online serving; and 2) e-commerce promotions have significant impacts on small scenarios, leading to distribution uncertainty of the upcoming time period. In this work, we propose a novel CVR method named MetaCVR from a perspective of meta learning to address the DDF issue. Firstly, a base CVR model which consists of a Feature Representation Network (FRN) and output layers is designed and trained sufficiently with samples across months. Then we treat time periods with different data distributions as different occasions and obtain positive and negative prototypes for each occasion using the corresponding samples and the pre-trained FRN. Subsequently, a Distance Metric Network (DMN) is devised to calculate the distance metrics between each sample and all prototypes to facilitate mitigating the distribution uncertainty. At last, we develop an Ensemble Prediction Network (EPN) which incorporates the output of FRN and DMN to make the final CVR prediction. In this stage, we freeze the FRN and train the DMN and EPN with samples from recent time period, therefore effectively easing the distribution discrepancy. To the best of our knowledge, this is the first study of CVR prediction targeting the DDF issue in small-scale recommendation scenarios. Experimental results on real-world datasets validate the superiority of our MetaCVR and online A/B test also shows our model achieves impressive gains of 11.92% on PCVR and 8.64% on GMV.

preprint2022arXiv

Microscopic nuclear equation of state at finite temperature and stellar stability

A microscopic nuclear equation of state compatible with all current astrophysical constraints constructed within the Brueckner-Hartree-Fock formalism is presented and extended in a consistent way to finite temperature. The effects of finite temperature on the properties of neutron stars are studied in detail and a universal relation regarding stellar stability is proposed.

preprint2022arXiv

Model-Driven Deep Learning-Based MIMO-OFDM Detector: Design, Simulation, and Experimental Results

Multiple-input multiple-output orthogonal frequency division multiplexing (MIMO-OFDM), a fundamental transmission scheme, promises high throughput and robustness against multipath fading. However, these benefits rely on the efficient detection strategy at the receiver and come at the expense of the extra bandwidth consumed by the cyclic prefix (CP). We use the iterative orthogonal approximate message passing (OAMP) algorithm in this paper as the prototype of the detector because of its remarkable potential for interference suppression. However, OAMP is computationally expensive for the matrix inversion per iteration. We replace the matrix inversion with the conjugate gradient (CG) method to reduce the complexity of OAMP. We further unfold the CG-based OAMP algorithm into a network and tune the critical parameters through deep learning (DL) to enhance detection performance. Simulation results and complexity analysis show that the proposed scheme has significant gain over other iterative detection methods and exhibits comparable performance to the state-of-the-art DL-based detector at a reduced computational cost. Furthermore, we design a highly efficient CP-free MIMO-OFDM receiver architecture to remove the CP overhead. This architecture first eliminates the intersymbol interference by buffering the previously recovered data and then detects the signal using the proposed detector. Numerical experiments demonstrate that the designed receiver offers a higher spectral efficiency than traditional receivers. Finally, over-the-air tests verify the effectiveness and robustness of the proposed scheme in realistic environments.

preprint2022arXiv

Observation of Dimension-Crossover of a Tunable 1D Dirac Fermion in Topological Semimetal NbSi$_x$Te$_2$

Condensed matter systems in low dimensions exhibit emergent physics that does not exist in three dimensions. When electrons are confined to one dimension (1D), some significant electronic states appear, such as charge density wave, spin-charge separations and Su-Schrieffer-Heeger (SSH) topological state. However, a clear understanding of how the 1D electronic properties connects with topology is currently lacking. Here we systematically investigated the characteristic 1D Dirac fermion electronic structure originated from the metallic NbTe$_2$ chains on the surface of the composition-tunable layered compound NbSi$_x$Te$_2$ ($x$ = 0.40 and 0.43) using angle-resolved photoemission spectroscopy. We found the Dirac fermion forms a Dirac nodal line structure protected by the combined $\widetilde{\mathcal{M}}{\rm_y}$ and time-reversal symmetry T and proves the NbSi$_x$Te$_2$ system as a topological semimetal, in consistent with the ab-initio calculations. As $x$ decreases, the interaction between adjacent NbTe2 chains increases and Dirac fermion goes through a dimension-crossover from 1D to 2D, as evidenced by the variation of its Fermi surface and Fermi velocity across the Brillouin zone in consistence with a Dirac SSH model. Our findings demonstrate a tunable 1D Dirac electron system, which offers a versatile platform for the exploration of intriguing 1D physics and device applications.

preprint2022arXiv

On-demand assembly of optically-levitated nanoparticle arrays in vacuum

Realizing a large-scale fully controllable quantum system is a challenging task in current physical research and has broad applications. Ultracold atom and molecule arrays in optical tweezers in vacuum have been used for quantum simulation, quantum metrology and quantum computing. Recently, quantum ground state cooling of the center-of-mass motion of a single optically levitated nanoparticle in vacuum was demonstrated, providing unprecedented opportunities for studying macroscopic quantum mechanics and precision measurements. In this work, we create a reconfigurable optically-levitated nanoparticle array in vacuum. Our optically-levitated nanoparticle array allows full control of individual nanoparticles to form an arbitrary pattern and detect their motion. As a concrete example, we choose two nanoparticles without rotation signals from an array to synthesize a nanodumbbell in-situ by merging them into one trap. The nanodumbbell synthesized in-situ can rotate beyond 1 GHz. Our work provides a new platform for studying macroscopic many-body physics.

preprint2022arXiv

Output Feedback Control of Radially-Dependent Reaction-Diffusion PDEs on Balls of Arbitrary Dimensions

Recently, the problem of boundary stabilization and estimation for unstable linear constant-coefficient reaction-diffusion equation on n-balls (in particular, disks and spheres) has been solved by means of the backstepping method. However, the extension of this result to spatially-varying coefficients is far from trivial. Some early success has been achieved under simplifying conditions, such as radially-varying reaction coefficients under revolution symmetry, on a disk or a sphere. These particular cases notwithstanding, the problem remains open. The main issue is that the equations become singular in the radius; when applying the backstepping method, the same type of singularity appears in the kernel equations. Traditionally, well-posedness of these equations has been proved by transforming them into integral equations and then applying the method of successive approximations. In this case, with the resulting integral equation becoming singular, successive approximations do not easily apply. This paper takes a different route and directly addresses the kernel equations via a power series approach, finding in the process the required conditions for the radially-varying reaction (namely, analyticity and evenness) and showing the existence and convergence of the series solution. This approach provides a direct numerical method that can be readily applied, despite singularities, to both control and observer boundary design problems.

preprint2022arXiv

Quantum Dynamics of Cold Atomic Gas with $SU(1,1)$ Symmetry

Motivated by recent advances in quantum dynamics, we investigate the dynamics of the system with $SU(1,1)$ symmetry. Instead of performing the time-ordered integral for the evolution operator of the time-dependent Hamiltonian, we show that the time evolution operator can be expressed as an $SU(1,1)$ group element. Since the $SU(1,1)$ group describes the "rotation" on a hyperbolic surface, the dynamics can be visualized on a Poincaré disk, a stereographic projection of the upper hyperboloid. As an example, we present the trajectory of the revival of Bose-Einstein condensation and that of the scale-invariant Fermi gas on the Poincaré disk. Further considering the quantum gas in the oscillating lattice, we also study the dynamics of the system with time-dependent single-particle dispersion. Our results are hopefully to be checked in current experiments.

preprint2022arXiv

Re-weighting Negative Samples for Model-Agnostic Matching

Recommender Systems (RS), as an efficient tool to discover users' interested items from a very large corpus, has attracted more and more attention from academia and industry. As the initial stage of RS, large-scale matching is fundamental yet challenging. A typical recipe is to learn user and item representations with a two-tower architecture and then calculate the similarity score between both representation vectors, which however still struggles in how to properly deal with negative samples. In this paper, we find that the common practice that randomly sampling negative samples from the entire space and treating them equally is not an optimal choice, since the negative samples from different sub-spaces at different stages have different importance to a matching model. To address this issue, we propose a novel method named Unbiased Model-Agnostic Matching Approach (UMA$^2$). It consists of two basic modules including 1) General Matching Model (GMM), which is model-agnostic and can be implemented as any embedding-based two-tower models; and 2) Negative Samples Debias Network (NSDN), which discriminates negative samples by borrowing the idea of Inverse Propensity Weighting (IPW) and re-weighs the loss in GMM. UMA$^2$ seamlessly integrates these two modules in an end-to-end multi-task learning framework. Extensive experiments on both real-world offline dataset and online A/B test demonstrate its superiority over state-of-the-art methods.

preprint2022arXiv

ReAct: Temporal Action Detection with Relational Queries

This work aims at advancing temporal action detection (TAD) using an encoder-decoder framework with action queries, similar to DETR, which has shown great success in object detection. However, the framework suffers from several problems if directly applied to TAD: the insufficient exploration of inter-query relation in the decoder, the inadequate classification training due to a limited number of training samples, and the unreliable classification scores at inference. To this end, we first propose a relational attention mechanism in the decoder, which guides the attention among queries based on their relations. Moreover, we propose two losses to facilitate and stabilize the training of action classification. Lastly, we propose to predict the localization quality of each action query at inference in order to distinguish high-quality queries. The proposed method, named ReAct, achieves the state-of-the-art performance on THUMOS14, with much lower computational costs than previous methods. Besides, extensive ablation studies are conducted to verify the effectiveness of each proposed component. The code is available at https://github.com/sssste/React.

preprint2022arXiv

Recurrent Glimpse-based Decoder for Detection with Transformer

Although detection with Transformer (DETR) is increasingly popular, its global attention modeling requires an extremely long training period to optimize and achieve promising detection performance. Alternative to existing studies that mainly develop advanced feature or embedding designs to tackle the training issue, we point out that the Region-of-Interest (RoI) based detection refinement can easily help mitigate the difficulty of training for DETR methods. Based on this, we introduce a novel REcurrent Glimpse-based decOder (REGO) in this paper. In particular, the REGO employs a multi-stage recurrent processing structure to help the attention of DETR gradually focus on foreground objects more accurately. In each processing stage, visual features are extracted as glimpse features from RoIs with enlarged bounding box areas of detection results from the previous stage. Then, a glimpse-based decoder is introduced to provide refined detection results based on both the glimpse features and the attention modeling outputs of the previous stage. In practice, REGO can be easily embedded in representative DETR variants while maintaining their fully end-to-end training and inference pipelines. In particular, REGO helps Deformable DETR achieve 44.8 AP on the MSCOCO dataset with only 36 training epochs, compared with the first DETR and the Deformable DETR that require 500 and 50 epochs to achieve comparable performance, respectively. Experiments also show that REGO consistently boosts the performance of different DETR detectors by up to 7% relative gain at the same setting of 50 training epochs. Code is available via https://github.com/zhechen/Deformable-DETR-REGO.

preprint2022arXiv

RGB-D Saliency Detection via Cascaded Mutual Information Minimization

Existing RGB-D saliency detection models do not explicitly encourage RGB and depth to achieve effective multi-modal learning. In this paper, we introduce a novel multi-stage cascaded learning framework via mutual information minimization to "explicitly" model the multi-modal information between RGB image and depth data. Specifically, we first map the feature of each mode to a lower dimensional feature vector, and adopt mutual information minimization as a regularizer to reduce the redundancy between appearance features from RGB and geometric features from depth. We then perform multi-stage cascaded learning to impose the mutual information minimization constraint at every stage of the network. Extensive experiments on benchmark RGB-D saliency datasets illustrate the effectiveness of our framework. Further, to prosper the development of this field, we contribute the largest (7x larger than NJU2K) dataset, which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations. Based on these rich labels, we additionally construct four new benchmarks with strong baselines and observe some interesting phenomena, which can motivate future model design. Source code and dataset are available at "https://github.com/JingZhang617/cascaded_rgbd_sod".

preprint2022arXiv

Robust control problems of BSDEs coupled with value functions

A robust control problem is considered in this paper, where the controlled stochastic differential equations (SDEs) include ambiguity parameters and their coefficients satisfy non-Lipschitz continuous and non-linear growth conditions, the objective function is expressed as a backward stochastic differential equation (BSDE) with the generator depending on the value function. We establish the existence and uniqueness of the value function in a proper space and provide a verification theorem. Moreover, we apply the results to solve two typical optimal investment problems in the market with ambiguity, one of which is with Heston stochastic volatility model. In particular, we establish some sharp estimations for Heston model with ambiguity parameters.

preprint2022arXiv

Salient Object Detection via Bounding-box Supervision

The success of fully supervised saliency detection models depends on a large number of pixel-wise labeling. In this paper, we work on bounding-box based weakly-supervised saliency detection to relieve the labeling effort. Given the bounding box annotation, we observe that pixels inside the bounding box may contain extensive labeling noise. However, as a large amount of background is excluded, the foreground bounding box region contains a less complex background, making it possible to perform handcrafted features-based saliency detection with only the cropped foreground region. As the conventional handcrafted features are not representative enough, leading to noisy saliency maps, we further introduce structure-aware self-supervised loss to regularize the structure of the prediction. Further, we claim that pixels outside the bounding box should be background, thus partial cross-entropy loss function can be used to accurately localize the accurate background region. Experimental results on six benchmark RGB saliency datasets illustrate the effectiveness of our model.

preprint2022arXiv

SASA: Semantics-Augmented Set Abstraction for Point-based 3D Object Detection

Although point-based networks are demonstrated to be accurate for 3D point cloud modeling, they are still falling behind their voxel-based competitors in 3D detection. We observe that the prevailing set abstraction design for down-sampling points may maintain too much unimportant background information that can affect feature learning for detecting objects. To tackle this issue, we propose a novel set abstraction method named Semantics-Augmented Set Abstraction (SASA). Technically, we first add a binary segmentation module as the side output to help identify foreground points. Based on the estimated point-wise foreground scores, we then propose a semantics-guided point sampling algorithm to help retain more important foreground points during down-sampling. In practice, SASA shows to be effective in identifying valuable points related to foreground objects and improving feature learning for point-based 3D detection. Additionally, it is an easy-to-plug-in module and able to boost various point-based detectors, including single-stage and two-stage ones. Extensive experiments on the popular KITTI and nuScenes datasets validate the superiority of SASA, lifting point-based detection models to reach comparable performance to state-of-the-art voxel-based methods.

preprint2022arXiv

SerialTrack: ScalE and Rotation Invariant Augmented Lagrangian Particle Tracking

We present a new particle tracking algorithm to accurately resolve large deformation and rotational motion fields, which takes advantage of both local and global particle tracking algorithms. We call this method the ScalE and Rotation Invariant Augmented Lagrangian Particle Tracking (SerialTrack). This method builds an iterative scale and rotation invariant topology-based feature for each particle within a multi-scale tracking algorithm. The global kinematic compatibility condition is applied as a global augmented Lagrangian constraint to enhance the tracking accuracy. An open source software package implementing this numerical approach to track both 2D and 3D, incremental and cumulative deformation fields is provided.

preprint2022arXiv

Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question Answering

Recent works on knowledge base question answering (KBQA) retrieve subgraphs for easier reasoning. A desired subgraph is crucial as a small one may exclude the answer but a large one might introduce more noises. However, the existing retrieval is either heuristic or interwoven with the reasoning, causing reasoning on the partial subgraphs, which increases the reasoning bias when the intermediate supervision is missing. This paper proposes a trainable subgraph retriever (SR) decoupled from the subsequent reasoning process, which enables a plug-and-play framework to enhance any subgraph-oriented KBQA model. Extensive experiments demonstrate SR achieves significantly better retrieval and QA performance than existing retrieval methods. Via weakly supervised pre-training as well as the end-to-end fine-tuning, SRl achieves new state-of-the-art performance when combined with NSM, a subgraph-oriented reasoner, for embedding-based KBQA methods.

preprint2022arXiv

Subtype-Former: a deep learning approach for cancer subtype discovery with multi-omics data

Motivation: Cancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer subtyping. However, precise cancer subtyping remains challenging due to the large amount and high dimensionality of omics data. Results: This study proposed Subtype-Former, a deep learning method based on MLP and Transformer Block, to extract the low-dimensional representation of the multi-omics data. K-means and Consensus Clustering are also used to achieve accurate subtyping results. We compared Subtype-Former with the other state-of-the-art subtyping methods across the TCGA 10 cancer types. We found that Subtype-Former can perform better on the benchmark datasets of more than 5000 tumors based on the survival analysis. In addition, Subtype-Former also achieved outstanding results in pan-cancer subtyping, which can help analyze the commonalities and differences across various cancer types at the molecular level. Finally, we applied Subtype-Former to the TCGA 10 types of cancers. We identified 50 essential biomarkers, which can be used to study targeted cancer drugs and promote the development of cancer treatments in the era of precision medicine.

preprint2022arXiv

Towards Data-Efficient Detection Transformers

Detection Transformers have achieved competitive performance on the sample-rich COCO dataset. However, we show most of them suffer from significant performance drops on small-size datasets, like Cityscapes. In other words, the detection transformers are generally data-hungry. To tackle this problem, we empirically analyze the factors that affect data efficiency, through a step-by-step transition from a data-efficient RCNN variant to the representative DETR. The empirical results suggest that sparse feature sampling from local image areas holds the key. Based on this observation, we alleviate the data-hungry issue of existing detection transformers by simply alternating how key and value sequences are constructed in the cross-attention layer, with minimum modifications to the original models. Besides, we introduce a simple yet effective label augmentation method to provide richer supervision and improve data efficiency. Experiments show that our method can be readily applied to different detection transformers and improve their performance on both small-size and sample-rich datasets. Code will be made publicly available at \url{https://github.com/encounter1997/DE-DETRs}.

preprint2022arXiv

Towards Scale Consistent Monocular Visual Odometry by Learning from the Virtual World

Monocular visual odometry (VO) has attracted extensive research attention by providing real-time vehicle motion from cost-effective camera images. However, state-of-the-art optimization-based monocular VO methods suffer from the scale inconsistency problem for long-term predictions. Deep learning has recently been introduced to address this issue by leveraging stereo sequences or ground-truth motions in the training dataset. However, it comes at an additional cost for data collection, and such training data may not be available in all datasets. In this work, we propose VRVO, a novel framework for retrieving the absolute scale from virtual data that can be easily obtained from modern simulation environments, whereas in the real domain no stereo or ground-truth data are required in either the training or inference phases. Specifically, we first train a scale-aware disparity network using both monocular real images and stereo virtual data. The virtual-to-real domain gap is bridged by using an adversarial training strategy to map images from both domains into a shared feature space. The resulting scale-consistent disparities are then integrated with a direct VO system by constructing a virtual stereo objective that ensures the scale consistency over long trajectories. Additionally, to address the suboptimality issue caused by the separate optimization backend and the learning process, we further propose a mutual reinforcement pipeline that allows bidirectional information flow between learning and optimization, which boosts the robustness and accuracy of each other. We demonstrate the effectiveness of our framework on the KITTI and vKITTI2 datasets.

preprint2022arXiv

Towards Scale-Aware, Robust, and Generalizable Unsupervised Monocular Depth Estimation by Integrating IMU Motion Dynamics

Unsupervised monocular depth and ego-motion estimation has drawn extensive research attention in recent years. Although current methods have reached a high up-to-scale accuracy, they usually fail to learn the true scale metric due to the inherent scale ambiguity from training with monocular sequences. In this work, we tackle this problem and propose DynaDepth, a novel scale-aware framework that integrates information from vision and IMU motion dynamics. Specifically, we first propose an IMU photometric loss and a cross-sensor photometric consistency loss to provide dense supervision and absolute scales. To fully exploit the complementary information from both sensors, we further drive a differentiable camera-centric extended Kalman filter (EKF) to update the IMU preintegrated motions when observing visual measurements. In addition, the EKF formulation enables learning an ego-motion uncertainty measure, which is non-trivial for unsupervised methods. By leveraging IMU during training, DynaDepth not only learns an absolute scale, but also provides a better generalization ability and robustness against vision degradation such as illumination change and moving objects. We validate the effectiveness of DynaDepth by conducting extensive experiments and simulations on the KITTI and Make3D datasets.

preprint2022arXiv

Transformer Networks for Predictive Group Elevator Control

We propose a Predictive Group Elevator Scheduler by using predictive information of passengers arrivals from a Transformer based destination predictor and a linear regression model that predicts remaining time to destinations. Through extensive empirical evaluation, we find that the savings of Average Waiting Time (AWT) could be as high as above 50% for light arrival streams and around 15% for medium arrival streams in afternoon down-peak traffic regimes. Such results can be obtained after carefully setting the Predicted Probability of Going to Elevator (PPGE) threshold, thus avoiding a majority of false predictions for people heading to the elevator, while achieving as high as 80% of true predictive elevator landings as early as after having seen only 60% of the whole trajectory of a passenger.

preprint2022arXiv

Tuning of Quantum Paraelectricity of M-type Hexaferrite BaFe12O19 by External Parameters

M-type hexaferrite BaFe12O19 was recently reported to be a new type of quantum paraelectrics with triangular lattice by showing a low temperature dielectric plateau due to quantum fluctuation. It has also been proposed to have a possible quantum-dipole liquid ground state. To suppress its quantum fluctuations and reach a possible quantum critical point, we have tuned its quantum paraelectricity in three ways: (i) 57Fe isotope replacement; (ii) in-plane compressive strain; and (iii) hydrostatic pressure. It is found that 95% 57Fe replacement and the in-plane strain are more effective to drive its ground state closer to a critical region by inducing a peak feature in the temperature dependence of dielectric constant. In contrast, the application of hydrostatic pressure pushed the system away from the quantum critical point by gradually suppressing the plateau feature in dielectric constant. Our combined efforts reveal the potential of the M-type hexaferrites for studying the quantum critical behaviors.

preprint2021arXiv

Atomic-scale investigation of the irradiation-resistant effect of symmetric tilt grain boundaries of Fe-Ni-Cr alloy

In this paper, the Fe-20Ni-25Cr alloy that is used for fuel cladding or pressure vessels with various grain boundaries (GBs) was investigated by employing molecular dynamics simulations. The bi-crystals comprised of Σ3(111), Σ3(112), Σ9(114), Σ11(113), Σ19(116), and Σ17(223) types GBs were considered to systematically examine the interplay between irradiation defects, irradiation microstructure evolution under stress, and irradiation mechanical properties with irradiation intensity, coincidence site lattice parameter, tilt angle, and GB thickness. It is found that irradiated vacancies and interstitials are annihilated by competitive GB absorption and recombination. Bias absorption of interstitials is observed for most bi-crystals except Σ3(111) and Σ11(113) at 15 keV incident energy, and results in abundant residual vacancies clusters in grain interior. In addition, different GBs exhibit quite diverse irradiation defect sink ability, and the number of residual vacancies is inversely related to the GB thickness, where Σ3(111) and Σ11(113) GBs with narrow GB thickness are weak in defect absorption and the others are strong. Furthermore, uniaxial tensile simulations perpendicular to the GB reveal that all of the mechanical performance of bi-crystals deteriorates after irradiation, which originates from dislocation propagation facilitated by irradiation defect clusters. In particular, regardless of whether the irradiation is applied, the maximum tensile strain, toughness, and Youngs modulus are monotonically correlated with GB tilt angle, while the ultimate tensile strength is stable for larger GB CSL parameter. Finally, on the basis of the evolution of the irradiation defects, microstructures, and mechanical performances, we proposed guidelines of rational design of irradiation-resistant Fe-Ni-Cr alloy.

preprint2021arXiv

Capsule Graph Neural Networks with EM Routing

To effectively classify graph instances, graph neural networks need to have the capability to capture the part-whole relationship existing in a graph. A capsule is a group of neurons representing complicated properties of entities, which has shown its advantages in traditional convolutional neural networks. This paper proposed novel Capsule Graph Neural Networks that use the EM routing mechanism (CapsGNNEM) to generate high-quality graph embeddings. Experimental results on a number of real-world graph datasets demonstrate that the proposed CapsGNNEM outperforms nine state-of-the-art models in graph classification tasks.

preprint2021arXiv

Learning structure-aware semantic segmentation with image-level supervision

Compared with expensive pixel-wise annotations, image-level labels make it possible to learn semantic segmentation in a weakly-supervised manner. Within this pipeline, the class activation map (CAM) is obtained and further processed to serve as a pseudo label to train the semantic segmentation model in a fully-supervised manner. In this paper, we argue that the lost structure information in CAM limits its application in downstream semantic segmentation, leading to deteriorated predictions. Furthermore, the inconsistent class activation scores inside the same object contradicts the common sense that each region of the same object should belong to the same semantic category. To produce sharp prediction with structure information, we introduce an auxiliary semantic boundary detection module, which penalizes the deteriorated predictions. Furthermore, we adopt smoothness loss to encourage prediction inside the object to be consistent. Experimental results on the PASCAL-VOC dataset illustrate the effectiveness of the proposed solution.

preprint2021arXiv

LineaRE: Simple but Powerful Knowledge Graph Embedding for Link Prediction

The task of link prediction for knowledge graphs is to predict missing relationships between entities. Knowledge graph embedding, which aims to represent entities and relations of a knowledge graph as low dimensional vectors in a continuous vector space, has achieved promising predictive performance. If an embedding model can cover different types of connectivity patterns and mapping properties of relations as many as possible, it will potentially bring more benefits for link prediction tasks. In this paper, we propose a novel embedding model, namely LineaRE, which is capable of modeling four connectivity patterns (i.e., symmetry, antisymmetry, inversion, and composition) and four mapping properties (i.e., one-to-one, one-to-many, many-to-one, and many-to-many) of relations. Specifically, we regard knowledge graph embedding as a simple linear regression task, where a relation is modeled as a linear function of two low-dimensional vector-presented entities with two weight vectors and a bias vector. Since the vectors are defined in a real number space and the scoring function of the model is linear, our model is simple and scalable to large knowledge graphs. Experimental results on multiple widely used real-world datasets show that the proposed LineaRE model significantly outperforms existing state-of-the-art models for link prediction tasks.

preprint2021arXiv

Multi-Stage Transmission Line Flow Control Using Centralized and Decentralized Reinforcement Learning Agents

Planning future operational scenarios of bulk power systems that meet security and economic constraints typically requires intensive labor efforts in performing massive simulations. To automate this process and relieve engineers' burden, a novel multi-stage control approach is presented in this paper to train centralized and decentralized reinforcement learning agents that can automatically adjust grid controllers for regulating transmission line flows at normal condition and under contingencies. The power grid flow control problem is formulated as Markov Decision Process (MDP). At stage one, centralized soft actor-critic (SAC) agent is trained to control generator active power outputs in a wide area to control transmission line flows against specified security limits. If line overloading issues remain unresolved, stage two is used to train decentralized SAC agent via load throw-over at local substations. The effectiveness of the proposed approach is verified on a series of actual planning cases used for operating the power grid of SGCC Zhejiang Electric Power Company.

preprint2021arXiv

Recent Advances on π-Conjugated Polymers as Active Elements in High Performance Organic Field-Effect Transistors

As high-performance organic semiconductors, π-conjugated polymers have attracted much attention due to their charming advantages including low-cost, solution processability, mechanical flexibility, and tunable optoelectronic properties. During the past several decades, the great advances have been made in polymers-based OFETs with p-type, n-type or even ambipolar characterics. Through chemical modification and alignment optimization, lots of conjugated polymers exhibited superior mobilities, and some mobilities are even larger than 10 cm2 V-1 s-1 in OFETs, which makes them very promising for the applications in organic electronic devices. This review describes the recent progress of the high performance polymers used in OFETs from the aspects of molecular design and assembly strategy. Furthermore, the current challenges and outlook in the design and development of conjugated polymers are also mentioned.

preprint2021arXiv

Stagewise Unsupervised Domain Adaptation with Adversarial Self-Training for Road Segmentation of Remote Sensing Images

Road segmentation from remote sensing images is a challenging task with wide ranges of application potentials. Deep neural networks have advanced this field by leveraging the power of large-scale labeled data, which, however, are extremely expensive and time-consuming to acquire. One solution is to use cheap available data to train a model and deploy it to directly process the data from a specific application domain. Nevertheless, the well-known domain shift (DS) issue prevents the trained model from generalizing well on the target domain. In this paper, we propose a novel stagewise domain adaptation model called RoadDA to address the DS issue in this field. In the first stage, RoadDA adapts the target domain features to align with the source ones via generative adversarial networks (GAN) based inter-domain adaptation. Specifically, a feature pyramid fusion module is devised to avoid information loss of long and thin roads and learn discriminative and robust features. Besides, to address the intra-domain discrepancy in the target domain, in the second stage, we propose an adversarial self-training method. We generate the pseudo labels of target domain using the trained generator and divide it to labeled easy split and unlabeled hard split based on the road confidence scores. The features of hard split are adapted to align with the easy ones using adversarial learning and the intra-domain adaptation process is repeated to progressively improve the segmentation performance. Experiment results on two benchmarks demonstrate that RoadDA can efficiently reduce the domain gap and outperforms state-of-the-art methods.

preprint2021arXiv

Tunable Flux through a Synthetic Hall Tube of Neutral Fermions

Hall tube with a tunable flux is an important geometry for studying quantum Hall physics, but its experimental realization in real space is still challenging. Here, we propose to realize a synthetic Hall tube with tunable flux in a one-dimensional optical lattice with the synthetic ring dimension defined by atomic hyperfine states. We investigate the effects of the flux on the system topology and study its quench dynamics. Utilizing the tunable flux, we show how to realize topological charge pumping, where interesting charge flow and transport are observed in rotated spin basis. Finally, we show that the recently observed quench dynamics in a synthetic Hall tube can be explained by the random flux existing in the experiment.

preprint2021arXiv

Understanding WeChat User Preferences and "Wow" Diffusion

WeChat is the largest social instant messaging platform in China, with 1.1 billion monthly active users. "Top Stories" is a novel friend-enhanced recommendation engine in WeChat, in which users can read articles based on preferences of both their own and their friends. Specifically, when a user reads an article by opening it, the "click" behavior is private. Moreover, if the user clicks the "wow" button, (only) her/his direct connections will be aware of this action/preference. Based on the unique WeChat data, we aim to understand user preferences and "wow" diffusion in Top Stories at different levels. We have made some interesting discoveries. For instance, the "wow" probability of one user is negatively correlated with the number of connected components that are formed by her/his active friends, but the click probability is the opposite. We further study to what extent users' "wow" and click behavior can be predicted from their social connections. To address this problem, we present a hierarchical graph representation learning based model DiffuseGNN, which is capable of capturing the structure-based social observations discovered above. Our experiments show that the proposed method can significantly improve the prediction performance compared with alternative methods.

preprint2020arXiv

A regression-based method for detecting publication bias in multivariate meta-analysis

Publication bias occurs when the publication of research results depends not only on the quality of the research but also on its nature and direction. The consequence is that published studies may not be truly representative of all valid studies undertaken, and this bias may threaten the validity of systematic reviews and meta-analyses - on which evidence-based medicine increasingly relies. Multivariate meta-analysis has recently received increasing attention for its ability reducing potential bias and improving statistical efficiency by borrowing information across outcomes. However, detecting and accounting for publication bias are more challenging in multivariate meta-analysis setting because some studies may be completely unpublished whereas some studies may selectively report part of multiple outcomes. In this paper, we propose a score test for jointly testing publication bias for multiple outcomes, which is novel to the multivariate setting. The proposed test is a natural multivariate extension of the univariate Egger's test, and can handle the above mentioned scenarios simultaneously, It accounts for correlations among multivariate outcomes, while allowing different types of outcomes, and can borrow information across outcomes. The proposed test is shown to be more powerful than the Egger's test, Begg's test and Trim and Fill method through simulation studies. Two data analyses are given to illustrate the performance of the proposed test in practice.

preprint2020arXiv

Condensing Two-stage Detection with Automatic Object Key Part Discovery

Modern two-stage object detectors generally require excessively large models for their detection heads to achieve high accuracy. To address this problem, we propose that the model parameters of two-stage detection heads can be condensed and reduced by concentrating on object key parts. To this end, we first introduce an automatic object key part discovery task to make neural networks discover representative sub-parts in each foreground object. With these discovered key parts, we then decompose the object appearance modeling into a key part modeling process and a global modeling process for detection. Key part modeling encodes fine and detailed features from the discovered key parts, and global modeling encodes rough and holistic object characteristics. In practice, such decomposition allows us to significantly abridge model parameters without sacrificing much detection accuracy. Experiments on popular datasets illustrate that our proposed technique consistently maintains original performance while waiving around 50% of the model parameters of common two-stage detection heads, with the performance only deteriorating by 1.5% when waiving around 96% of the original model parameters. Codes are released on: https://github.com/zhechen/Condensing2stageDetection.

preprint2020arXiv

CONNA: Addressing Name Disambiguation on The Fly

Name disambiguation is a key and also a very tough problem in many online systems such as social search and academic search. Despite considerable research, a critical issue that has not been systematically studied is disambiguation on the fly -- to complete the disambiguation in the real-time. This is very challenging, as the disambiguation algorithm must be accurate, efficient, and error tolerance. In this paper, we propose a novel framework -- CONNA -- to train a matching component and a decision component jointly via reinforcement learning. The matching component is responsible for finding the top matched candidate for the given paper, and the decision component is responsible for deciding on assigning the top matched person or creating a new person. The two components are intertwined and can be bootstrapped via jointly training. Empirically, we evaluate CONNA on two name disambiguation datasets. Experimental results show that the proposed framework can achieve a 1.21%-19.84% improvement on F1-score using joint training of the matching and the decision components. The proposed CONNA has been successfully deployed on AMiner -- a large online academic search system.

preprint2020arXiv

Constraining the nuclear symmetry energy and properties of neutron star from GW170817 by Bayesian analysis

Based on the distribution of tidal deformabilities and component masses of binary neutron star merger GW170817, the parametric equation of states (EOS) are employed to probe the nuclear symmetry energy and the properties of neutron star. To obtain a proper distribution of the parameters of the EOS that is consistent with the observation, Bayesian analysis is used and the constraints of causality and maximum mass are considered. From this analysis, it is found that the symmetry energy at twice the saturation density of nuclear matter can be constrained within $E_{sym}(2{ρ_{0}})$ = $34.5^{+20.5}_{-2.3}$ MeV at 90\% credible level. Moreover, the constraints on the radii and dimensionless tidal deformabilities of canonical neutron stars are also demonstrated through this analysis, and the corresponding constraints are 10.80 km $< R_{1.4} <$ 13.20 km and $133 < Λ_{1.4} < 686$ at 90\% credible level, with the most probable value of $\bar{R}_{1.4}$ = 12.60 km and $\barΛ_{1.4}$ = 500, respectively. With respect to the prior, our result (posterior result) prefers a softer EOS, corresponding to a lower expected value of symmetry energy, a smaller radius and a smaller tidal deformability.

preprint2020arXiv

Cox regression analysis for distorted covariates with an unknown distortion function

We study inference for censored survival data where some covariates are distorted by some unknown functions of an observable confounding variable in a multiplicative form. Example of this kind of data in medical studies is the common practice to normalizing some important observed exposure variables by patients' body mass index (BMI), weight or age. Such phenomenon also appears frequently in environmental studies where ambient measure is used for normalization, and in genomic studies where library size needs to be normalized for next generation sequencing data. We propose a new covariate-adjusted Cox proportional hazards regression model and utilize the kernel smoothing method to estimate the distorting function, then employ an estimated maximum likelihood method to derive estimator for the regression parameters. We establish the large sample properties of the proposed estimator. Extensive simulation studies demonstrate that the proposed estimator performs well in correcting the bias arising from distortion. A real data set from the National Wilms' Tumor Study (NWTS) is used to illustrate the proposed approach.

preprint2020arXiv

Ensemble emotion recognizing with multiple modal physiological signals

Physiological signals that provide the objective repression of human affective states are attracted increasing attention in the emotion recognition field. However, the single signal is difficult to obtain completely and accurately description for emotion. Multiple physiological signals fusing models, building the uniform classification model by means of consistent and complementary information from different emotions to improve recognition performance. Original fusing models usually choose the particular classification method to recognition, which is ignoring different distribution of multiple signals. Aiming above problems, in this work, we propose an emotion classification model through multiple modal physiological signals for different emotions. Features are extracted from EEG, EMG, EOG signals for characterizing emotional state on valence and arousal levels. For characterization, four bands filtering theta, beta, alpha, gamma for signal preprocessing are adopted and three Hjorth parameters are computing as features. To improve classification performance, an ensemble classifier is built. Experiments are conducted on the benchmark DEAP datasets. For the two-class task, the best result on arousal is 94.42\%, the best result on valence is 94.02\%, respectively. For the four-class task, the highest average classification accuracy is 90.74, and it shows good stability. The influence of different peripheral physiological signals for results is also analyzed in this paper.

preprint2020arXiv

Entire Space Multi-Task Modeling via Post-Click Behavior Decomposition for Conversion Rate Prediction

Recommender system, as an essential part of modern e-commerce, consists of two fundamental modules, namely Click-Through Rate (CTR) and Conversion Rate (CVR) prediction. While CVR has a direct impact on the purchasing volume, its prediction is well-known challenging due to the Sample Selection Bias (SSB) and Data Sparsity (DS) issues. Although existing methods, typically built on the user sequential behavior path ``impression$\to$click$\to$purchase'', is effective for dealing with SSB issue, they still struggle to address the DS issue due to rare purchase training samples. Observing that users always take several purchase-related actions after clicking, we propose a novel idea of post-click behavior decomposition. Specifically, disjoint purchase-related Deterministic Action (DAction) and Other Action (OAction) are inserted between click and purchase in parallel, forming a novel user sequential behavior graph ``impression$\to$click$\to$D(O)Action$\to$purchase''. Defining model on this graph enables to leverage all the impression samples over the entire space and extra abundant supervised signals from D(O)Action, which will effectively address the SSB and DS issues together. To this end, we devise a novel deep recommendation model named Elaborated Entire Space Supervised Multi-task Model ($ESM^{2}$). According to the conditional probability rule defined on the graph, it employs multi-task learning to predict some decomposed sub-targets in parallel and compose them sequentially to formulate the final CVR. Extensive experiments on both offline and online environments demonstrate the superiority of $ESM^{2}$ over state-of-the-art models. The source code and dataset will be released.

preprint2020arXiv

Experimental realization of spin-tensor momentum coupling in ultracold Fermi gases

We experimentally realize the spin-tensor momentum coupling (STMC) using the three ground Zeeman states coupled by three Raman laser beams in ultracold atomic system of $^{40}$K Fermi atoms. This new type of STMC consists of two bright-state bands as a regular spin-orbit coupled spin-1/2 system and one dark-state middle band. Using radio-frequency spin-injection spectroscopy, we investigate the energy band of STMC. It is demonstrated that the middle state is a dark state in the STMC system. The realized energy band of STMC may open the door for further exploring exotic quantum matters.

preprint2020arXiv

GCC: Graph Contrastive Coding for Graph Neural Network Pre-Training

Graph representation learning has emerged as a powerful technique for addressing real-world problems. Various downstream graph learning tasks have benefited from its recent developments, such as node classification, similarity search, and graph classification. However, prior arts on graph representation learning focus on domain specific problems and train a dedicated model for each graph dataset, which is usually non-transferable to out-of-domain data. Inspired by the recent advances in pre-training from natural language processing and computer vision, we design Graph Contrastive Coding (GCC) -- a self-supervised graph neural network pre-training framework -- to capture the universal network topological properties across multiple networks. We design GCC's pre-training task as subgraph instance discrimination in and across networks and leverage contrastive learning to empower graph neural networks to learn the intrinsic and transferable structural representations. We conduct extensive experiments on three graph learning tasks and ten graph datasets. The results show that GCC pre-trained on a collection of diverse datasets can achieve competitive or better performance to its task-specific and trained-from-scratch counterparts. This suggests that the pre-training and fine-tuning paradigm presents great potential for graph representation learning.

preprint2020arXiv

Generative networks as inverse problems with fractional wavelet scattering networks

Deep learning is a hot research topic in the field of machine learning methods and applications. Generative Adversarial Networks (GANs) and Variational Auto-Encoders (VAEs) provide impressive image generations from Gaussian white noise, but both of them are difficult to train since they need to train the generator (or encoder) and the discriminator (or decoder) simultaneously, which is easy to cause unstable training. In order to solve or alleviate the synchronous training difficult problems of GANs and VAEs, recently, researchers propose Generative Scattering Networks (GSNs), which use wavelet scattering networks (ScatNets) as the encoder to obtain the features (or ScatNet embeddings) and convolutional neural networks (CNNs) as the decoder to generate the image. The advantage of GSNs is the parameters of ScatNets are not needed to learn, and the disadvantage of GSNs is that the expression ability of ScatNets is slightly weaker than CNNs and the dimensional reduction method of Principal Component Analysis (PCA) is easy to lead overfitting in the training of GSNs, and therefore affect the generated quality in the testing process. In order to further improve the quality of generated images while keep the advantages of GSNs, this paper proposes Generative Fractional Scattering Networks (GFRSNs), which use more expressive fractional wavelet scattering networks (FrScatNets) instead of ScatNets as the encoder to obtain the features (or FrScatNet embeddings) and use the similar CNNs of GSNs as the decoder to generate the image. Additionally, this paper develops a new dimensional reduction method named Feature-Map Fusion (FMF) instead of PCA for better keeping the information of FrScatNets and the effect of image fusion on the quality of image generation is also discussed.

preprint2020arXiv

Gleason Score Prediction using Deep Learning in Tissue Microarray Image

Prostate cancer (PCa) is one of the most common cancers in men around the world. The most accurate method to evaluate lesion levels of PCa is microscopic inspection of stained biopsy tissue and estimate the Gleason score of tissue microarray (TMA) image by expert pathologists. However, it is time-consuming for pathologists to identify the cellular and glandular patterns for Gleason grading in large TMA images. We used Gleason2019 Challenge dataset to build a convolutional neural network (CNN) model to segment TMA images to regions of different Gleason grades and predict the Gleason score according to the grading segmentation. We used a pre-trained model of prostate segmentation to increase the accuracy of the Gleason grade segmentation. The model achieved a mean Dice of 75.6% on the test cohort and ranked 4th in the Gleason2019 Challenge with a score of 0.778 combined of Cohen's kappa and the f1-score.

preprint2020arXiv

High-order minibands and interband Landau level reconstruction in graphene moire superlattice

The propagation of Dirac fermions in graphene through a long-period periodic potential would result in a band folding together with the emergence of a series of cloned Dirac points (DPs). In highly aligned graphene/hexagonal boron nitride (G/hBN) heterostructures, the lattice mismatch between the two atomic crystals generates a unique kind of periodic structure known as a moiré superlattice. Of particular interests is the emergent phenomena related to the reconstructed band-structure of graphene, such as the Hofstadter butterfly, topological currents, gate dependent pseudospin mixing, and ballistic miniband conduction. However, most studies so far have been limited to the lower-order minibands, e.g. the 1st and 2nd minibands counted from charge neutrality, and consequently the fundamental nature of the reconstructed higher-order miniband spectra still remains largely unknown. Here we report on probing the higher-order minibands of precisely aligned graphene moiré superlattices by transport spectroscopy. Using dual electrostatic gating, the edges of these high-order minibands, i.e. the 3rd and 4th minibands, can be reached. Interestingly, we have observed interband Landau level (LL) crossinginducing gap closures in a multiband magneto-transport regime, which originates from band overlap between the 2nd and 3rd minibands. As observed high-order minibands and LL reconstruction qualitatively match our simulated results. Our findings highlight the synergistic effect of minibands in transport, thus presenting a new opportunity for graphene electronic devices.

preprint2020arXiv

JarKA: Modeling Attribute Interactions for Cross-lingual Knowledge Alignment

Abstract. Cross-lingual knowledge alignment is the cornerstone in building a comprehensive knowledge graph (KG), which can benefit various knowledge-driven applications. As the structures of KGs are usually sparse, attributes of entities may play an important role in aligning the entities. However, the heterogeneity of the attributes across KGs prevents from accurately embedding and comparing entities. To deal with the issue, we propose to model the interactions between attributes, instead of globally embedding an entity with all the attributes. We further propose a joint framework to merge the alignments inferred from the attributes and the structures. Experimental results show that the proposed model outperforms the state-of-art baselines by up to 38.48% HitRatio@1. The results also demonstrate that our model can infer the alignments between attributes, relationships and values, in addition to entities.

preprint2020arXiv

Learning Noise-Aware Encoder-Decoder from Noisy Labels by Alternating Back-Propagation for Saliency Detection

In this paper, we propose a noise-aware encoder-decoder framework to disentangle a clean saliency predictor from noisy training examples, where the noisy labels are generated by unsupervised handcrafted feature-based methods. The proposed model consists of two sub-models parameterized by neural networks: (1) a saliency predictor that maps input images to clean saliency maps, and (2) a noise generator, which is a latent variable model that produces noises from Gaussian latent vectors. The whole model that represents noisy labels is a sum of the two sub-models. The goal of training the model is to estimate the parameters of both sub-models, and simultaneously infer the corresponding latent vector of each noisy label. We propose to train the model by using an alternating back-propagation (ABP) algorithm, which alternates the following two steps: (1) learning back-propagation for estimating the parameters of two sub-models by gradient ascent, and (2) inferential back-propagation for inferring the latent vectors of training noisy examples by Langevin Dynamics. To prevent the network from converging to trivial solutions, we utilize an edge-aware smoothness loss to regularize hidden saliency maps to have similar structures as their corresponding images. Experimental results on several benchmark datasets indicate the effectiveness of the proposed model.

preprint2020arXiv

Magnetic mixed valent semimetal EuZnSb$_2$ with Dirac states in the band structure

We report discovery of new antiferromagnetic semimetal EuZnSb$_2$, obtained and studied in the form of single crystals. Electric resistivity, magnetic susceptibility and heat capacity indicate antiferromagnetic order of Eu with $T_N$ = 20 K. The effective moment of Eu$^{2+}$ inferred from the magnetization and specific heat measurement is 3.5 $μ_B$, smaller than the theoretical value of Eu$^{2+}$ due to presence of both Eu$^{3+}$ and Eu$^{2+}$. Magnetic field-dependent resistivity measurements suggest dominant quasi two dimensional Fermi surfaces whereas the first-principle calculations point to the presence of Dirac fermions. Therefore, EuZnSb$_2$ could represent the first platform to study the interplay of dynamical charge fluctuations, localized magnetic 4$f$ moments and Dirac states with Sb orbital character.

preprint2020arXiv

Measurement of the neutron beam profile of the Back-n white neutron facility at CSNS with a Micromegas detector

The Back-n white neutron beam line, which uses back-streaming white neutrons from the spallation target of the China Spallation Neutron Source, is used for nuclear data measurements. A Micromegas-based neutron detector with two variants was specially developed to measure the beam spot distribution for this beam line. In this article, the design, fabrication, and characterization of the detector are described. The results of the detector performance tests are presented, which include the relative electron transparency, the gain and the gain uniformity, and the neutron beam profile reconstruction capability. The result of the first measurement of the Back-n neutron beam spot distribution is also presented.

preprint2020arXiv

Memory-Gated Recurrent Networks

The essence of multivariate sequential learning is all about how to extract dependencies in data. These data sets, such as hourly medical records in intensive care units and multi-frequency phonetic time series, often time exhibit not only strong serial dependencies in the individual components (the "marginal" memory) but also non-negligible memories in the cross-sectional dependencies (the "joint" memory). Because of the multivariate complexity in the evolution of the joint distribution that underlies the data generating process, we take a data-driven approach and construct a novel recurrent network architecture, termed Memory-Gated Recurrent Networks (mGRN), with gates explicitly regulating two distinct types of memories: the marginal memory and the joint memory. Through a combination of comprehensive simulation studies and empirical experiments on a range of public datasets, we show that our proposed mGRN architecture consistently outperforms state-of-the-art architectures targeting multivariate time series.

preprint2020arXiv

MLBF-Net: A Multi-Lead-Branch Fusion Network for Multi-Class Arrhythmia Classification Using 12-Lead ECG

Automatic arrhythmia detection using 12-lead electrocardiogram (ECG) signal plays a critical role in early prevention and diagnosis of cardiovascular diseases. In the previous studies on automatic arrhythmia detection, most methods concatenated 12 leads of ECG into a matrix, and then input the matrix to a variety of feature extractors or deep neural networks for extracting useful information. Under such frameworks, these methods had the ability to extract comprehensive features (known as integrity) of 12-lead ECG since the information of each lead interacts with each other during training. However, the diverse lead-specific features (known as diversity) among 12 leads were neglected, causing inadequate information learning for 12-lead ECG. To maximize the information learning of multi-lead ECG, the information fusion of comprehensive features with integrity and lead-specific features with diversity should be taken into account. In this paper, we propose a novel Multi-Lead-Branch Fusion Network (MLBF-Net) architecture for arrhythmia classification by integrating multi-loss optimization to jointly learning diversity and integrity of multi-lead ECG. MLBF-Net is composed of three components: 1) multiple lead-specific branches for learning the diversity of multi-lead ECG; 2) cross-lead features fusion by concatenating the output feature maps of all branches for learning the integrity of multi-lead ECG; 3) multi-loss co-optimization for all the individual branches and the concatenated network. We demonstrate our MLBF-Net on China Physiological Signal Challenge 2018 which is an open 12-lead ECG dataset. The experimental results show that MLBF-Net obtains an average $F_1$ score of 0.855, reaching the highest arrhythmia classification performance. The proposed method provides a promising solution for multi-lead ECG analysis from an information fusion perspective.

preprint2020arXiv

Model-Driven DNN Decoder for Turbo Codes: Design, Simulation and Experimental Results

This paper presents a novel model-driven deep learning (DL) architecture, called TurboNet, for turbo decoding that integrates DL into the traditional max-log-maximum a posteriori (MAP) algorithm. The TurboNet inherits the superiority of the max-log-MAP algorithm and DL tools and thus presents excellent error-correction capability with low training cost. To design the TurboNet, the original iterative structure is unfolded as deep neural network (DNN) decoding units, where trainable weights are introduced to the max-log-MAP algorithm and optimized through supervised learning. To efficiently train the TurboNet, a loss function is carefully designed to prevent tricky gradient vanishing issue. To further reduce the computational complexity and training cost of the TurboNet, we can prune it into TurboNet+. Compared with the existing black-box DL approaches, the TurboNet+ has considerable advantage in computational complexity and is conducive to significantly reducing the decoding overhead. Furthermore, we also present a simple training strategy to address the overfitting issue, which enable efficient training of the proposed TurboNet+. Simulation results demonstrate TurboNet+'s superiority in error-correction ability, signal-to-noise ratio generalization, and computational overhead. In addition, an experimental system is established for an over-the-air (OTA) test with the help of a 5G rapid prototyping system and demonstrates TurboNet's strong learning ability and great robustness to various scenarios.

preprint2020arXiv

Multi-Target Deep Learning for Algal Detection and Classification

Water quality has a direct impact on industry, agriculture, and public health. Algae species are common indicators of water quality. It is because algal communities are sensitive to changes in their habitats, giving valuable knowledge on variations in water quality. However, water quality analysis requires professional inspection of algal detection and classification under microscopes, which is very time-consuming and tedious. In this paper, we propose a novel multi-target deep learning framework for algal detection and classification. Extensive experiments were carried out on a large-scale colored microscopic algal dataset. Experimental results demonstrate that the proposed method leads to the promising performance on algal detection, class identification and genus identification.

preprint2020arXiv

SK-Unet: an Improved U-net Model with Selective Kernel for the Segmentation of Multi-sequence Cardiac MR

In the clinical environment, myocardial infarction (MI) as one com-mon cardiovascular disease is mainly evaluated based on the late gadolinium enhancement (LGE) cardiac magnetic resonance images (CMRIs). The auto-matic segmentations of left ventricle (LV), right ventricle (RV), and left ven-tricular myocardium (LVM) in the LGE CMRIs are desired for the aided diag-nosis in clinic. To accomplish this segmentation task, this paper proposes a modified U-net architecture by combining multi-sequence CMRIs, including the cine, LGE, and T2-weighted CMRIs. The cine and T2-weighted CMRIs are used to assist the segmentation in the LGE CMRIs. In this segmentation net-work, the squeeze-and-excitation residual (SE-Res) and selective kernel (SK) modules are inserted in the down-sampling and up-sampling stages, respective-ly. The SK module makes the obtained feature maps more informative in both spatial and channel-wise space, and attains more precise segmentation result. The utilized dataset is from the MICCAI challenge (MS-CMRSeg 2019), which is acquired from 45 patients including three CMR sequences. The cine and T2-weighted CMRIs acquired from 35 patients and the LGE CMRIs acquired from 5 patients are labeled. Our method achieves the mean dice score of 0.922 (LV), 0.827 (LVM), and 0.874 (RV) in the LGE CMRIs.

preprint2020arXiv

Stationary and Closed Rainbow subsets

We study the structured rainbow Ramsey theory at uncountable cardinals. When compared to the usual rainbow Ramsey theory, the variation focuses on finding a rainbow subset that not only is of a certain cardinality but also satisfies certain structural constraints, such as being stationary or closed in its supremum. In the process of dealing with cardinals greater than $ω_1$, we uncover some connections between versions of Chang's Conjectures and instances of rainbow Ramsey partition relations, addressing a question raised in \cite{zhang}.

preprint2020arXiv

Structured Massive Access for Scalable Cell-Free Massive MIMO Systems

How to meet the demand for increasing number of users, higher data rates, and stringent quality-of-service (QoS) in the beyond fifth-generation (B5G) networks? Cell-free massive multiple-input multiple-output (MIMO) is considered as a promising solution, in which many wireless access points cooperate to jointly serve the users by exploiting coherent signal processing. However, there are still many unsolved practical issues in cell-free massive MIMO systems, whereof scalable massive access implementation is one of the most vital. In this paper, we propose a new framework for structured massive access in cell-free massive MIMO systems, which comprises one initial access algorithm, a partial large-scale fading decoding (P-LSFD) strategy, two pilot assignment schemes, and one fractional power control policy. New closed-form spectral efficiency (SE) expressions with maximum ratio (MR) combining are derived. The simulation results show that our proposed framework provides high SE when using local partial minimum mean-square error (LP-MMSE) and MR combining. Specifically, the proposed initial access algorithm and pilot assignment schemes outperform their corresponding benchmarks, P-LSFD achieves scalability with a negligible performance loss compared to the conventional optimal large-scale fading decoding (LSFD), and scalable fractional power control provides a controllable trade-off between user fairness and the average SE.

preprint2020arXiv

Super-resolved optical mapping of reactive sulfur-vacancy in 2D transition metal dichalcogenides

Transition metal dichalcogenides (TMDs) represent an entire new class of semiconducting 2D materials with exciting properties. Defects in 2D TMDs can crucially affect their physical and chemical properties. However, characterization of the presence and spatial distribution of defects is limited either in throughput or in resolution. Here, we demonstrate large area mapping of reactive sulfur-deficient defects in 2D-TMDs coupling single-molecule localization microscopy with fluorescence labeling using thiol chemistry. Our method, reminiscent of PAINT strategies, relies on the specific binding by reversible physisorption of fluorescent probes to sulfur-vacancies via a thiol group and their intermittent emission to apply localization of the labeled defects with a precision down to 15 nm. Tuning the distance between the fluorophore and the docking thiol site allows us to control Föster Resonance Energy Transfer (FRET) process and reveal large structural defects such as grain boundaries and line defects, due to the local irregular lattice structure. Our methodology provides a simple and fast alternative for large-scale mapping of non-radiative defects in 2D materials and paves the way for in-situ and spatially resolved monitoring of the interaction between chemical agent and the defects in 2D materials that has general implications for defect engineering in aqueous condition.

preprint2020arXiv

Synchronization in PT-symmetric optomechanical resonators

Synchronization has great impacts in various fields such as self-clocking, communication, neural networks, etc. Here we present a mechanism of synchronization for two mechanical modes in two coupled optomechanical resonators by introducing the so-called PT-symmetric structure. It is shown that the degree of synchronization between the two far-off-resonant mechanical modes can be increased by decreasing the coupling strength between the two optomechanical resonators. Additionally, when we consider the stochastic noises in the optomechanical resonators, we find that more noises can enhance the degree of synchronization of the system under particular parameter regime. Our results open up the new dimension of research for PT-symmetric systems and synchronization.

preprint2020arXiv

UC-Net: Uncertainty Inspired RGB-D Saliency Detection via Conditional Variational Autoencoders

In this paper, we propose the first framework (UCNet) to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection methods treat the saliency detection task as a point estimation problem, and produce a single saliency map following a deterministic learning pipeline. Inspired by the saliency data labeling process, we propose probabilistic RGB-D saliency detection network via conditional variational autoencoders to model human annotation uncertainty and generate multiple saliency maps for each input image by sampling in the latent space. With the proposed saliency consensus process, we are able to generate an accurate saliency map based on these multiple predictions. Quantitative and qualitative evaluations on six challenging benchmark datasets against 18 competing algorithms demonstrate the effectiveness of our approach in learning the distribution of saliency maps, leading to a new state-of-the-art in RGB-D saliency detection.

preprint2020arXiv

Uncertainty Inspired RGB-D Saliency Detection

We propose the first stochastic framework to employ uncertainty for RGB-D saliency detection by learning from the data labeling process. Existing RGB-D saliency detection models treat this task as a point estimation problem by predicting a single saliency map following a deterministic learning pipeline. We argue that, however, the deterministic solution is relatively ill-posed. Inspired by the saliency data labeling process, we propose a generative architecture to achieve probabilistic RGB-D saliency detection which utilizes a latent variable to model the labeling variations. Our framework includes two main models: 1) a generator model, which maps the input image and latent variable to stochastic saliency prediction, and 2) an inference model, which gradually updates the latent variable by sampling it from the true or approximate posterior distribution. The generator model is an encoder-decoder saliency network. To infer the latent variable, we introduce two different solutions: i) a Conditional Variational Auto-encoder with an extra encoder to approximate the posterior distribution of the latent variable; and ii) an Alternating Back-Propagation technique, which directly samples the latent variable from the true posterior distribution. Qualitative and quantitative results on six challenging RGB-D benchmark datasets show our approach's superior performance in learning the distribution of saliency maps. The source code is publicly available via our project page: https://github.com/JingZhang617/UCNet.

preprint2020arXiv

Unsupervised Domain Expansion from Multiple Sources

Given an existing system learned from previous source domains, it is desirable to adapt the system to new domains without accessing and forgetting all the previous domains in some applications. This problem is known as domain expansion. Unlike traditional domain adaptation in which the target domain is the domain defined by new data, in domain expansion the target domain is formed jointly by the source domains and the new domain (hence, domain expansion) and the label function to be learned must work for the expanded domain. Specifically, this paper presents a method for unsupervised multi-source domain expansion (UMSDE) where only the pre-learned models of the source domains and unlabelled new domain data are available. We propose to use the predicted class probability of the unlabelled data in the new domain produced by different source models to jointly mitigate the biases among domains, exploit the discriminative information in the new domain, and preserve the performance in the source domains. Experimental results on the VLCS, ImageCLEF_DA and PACS datasets have verified the effectiveness of the proposed method.

preprint2020arXiv

Vectorial ball Prolate spheroidal wave functions with the divergence free constraint

In this paper, we introduce one family of vectorial prolate spheroidal wave functions of real order $α>-1$ on the unit ball in $R^3$, which satisfy the divergence free constraint, thus are termed as divergence free vectorial ball PSWFs. They are vectorial eigenfunctions of an integral operator related to the finite Fourier transform, and solve the divergence free constrained maximum concentration problem in three dimensions, i.e., to what extent can the total energy of a band-limited divergence free vectorial function be concentrated on the unit ball? Interestingly, any optimally concentrated divergence free vectorial functions, when represented in series in vector spherical harmonics, shall be also concentrated in one of the three vectorial spherical harmonics modes. Moreover, divergence free ball PSWFs are exactly the vectorial eigenfunctions of the second order Sturm-Liouville differential operator which defines the scalar ball PSWFs. Indeed, the divergence free vectorial ball PSWFs possess a simple and close relation with the scalar ball PSWFs such that they share the same merits. Simultaneously, it turns out that the divergence free ball PSWFs solve another second order Sturm-Liouville eigen equation defined through the curl operator $\nabla\times $ instead of the gradient operator $\nabla$.

preprint2020arXiv

Weakly-Supervised Salient Object Detection via Scribble Annotations

Compared with laborious pixel-wise dense labeling, it is much easier to label data by scribbles, which only costs 1$\sim$2 seconds to label one image. However, using scribble labels to learn salient object detection has not been explored. In this paper, we propose a weakly-supervised salient object detection model to learn saliency from such annotations. In doing so, we first relabel an existing large-scale salient object detection dataset with scribbles, namely S-DUTS dataset. Since object structure and detail information is not identified by scribbles, directly training with scribble labels will lead to saliency maps of poor boundary localization. To mitigate this problem, we propose an auxiliary edge detection task to localize object edges explicitly, and a gated structure-aware loss to place constraints on the scope of structure to be recovered. Moreover, we design a scribble boosting scheme to iteratively consolidate our scribble annotations, which are then employed as supervision to learn high-quality saliency maps. As existing saliency evaluation metrics neglect to measure structure alignment of the predictions, the saliency map ranking metric may not comply with human perception. We present a new metric, termed saliency structure measure, to measure the structure alignment of the predicted saliency maps, which is more consistent with human perception. Extensive experiments on six benchmark datasets demonstrate that our method not only outperforms existing weakly-supervised/unsupervised methods, but also is on par with several fully-supervised state-of-the-art models. Our code and data is publicly available at https://github.com/JingZhang617/Scribble_Saliency.

preprint2019arXiv

Joint Estimation of OD Demands and Cost Functions in Transportation Networks from Data

Existing work has tackled the problem of estimating Origin-Destination (OD) demands and recovering travel latency functions in transportation networks under the Wardropian assumption. The ultimate objective is to derive an accurate predictive model of the network to enable optimization and control. However, these two problems are typically treated separately and estimation is based on parametric models. In this paper, we propose a method to jointly recover nonparametric travel latency cost functions and estimate OD demands using traffic flow data. We formulate the problem as a bilevel optimization problem and develop an iterative first-order optimization algorithm to solve it. A numerical example using the Braess Network is presented to demonstrate the effectiveness of our method.

preprint2019arXiv

Spatiotemporal graph states from a single optical parametric oscillator

An experimental scheme is proposed for building massively multipartite entangled states using both the spatial and the frequency modes of an optical parametric oscillator. We provide analytical forms of the entangled states using the squeezed eigenmodes of Heisenberg equations, a.k.a. the nullifiers of the corresponding graph state. This scheme can generate, in parallel, several cluster states described by sparsely connected, bicolorable graph states, usable for one-way quantum computing. We indicate the experimentally accessible quantum graphs, depending on the squeezing parameter.

preprint2017arXiv

Interference Minimization in 5G Heterogeneous Networks

In this paper, we focus on one of the representative 5G network scenarios, namely multi-tier heterogeneous cellular networks. User association is investigated in order to reduce the down-link co-channel interference. Firstly, in order to analyze the multi-tier heterogeneous cellular networks where the base stations in different tiers usually adopt different transmission powers, we propose a Transmission Power Normalization Model (TPNM), which is able to convert a multi-tier cellular network into a single-tier network, such that all base stations have the same normalized transmission power. Then using TPNM, the signal and interference received at any point in the complex multi-tier environment can be analyzed by considering the same point in the equivalent single-tier cellular network model, thus significantly simplifying the analysis. On this basis, we propose a new user association scheme in heterogeneous cellular networks, where the base station that leads to the smallest interference to other co-channel mobile stations is chosen from a set of candidate base stations that satisfy the quality-of-service (QoS) constraint for an intended mobile station. Numerical results show that the proposed user association scheme is able to significantly reduce the down-link interference compared with existing schemes while maintaining a reasonably good QoS.

preprint2016arXiv

An Improved Composite Hypothesis Test for Markov Models with Applications in Network Anomaly Detection

Recent work has proposed the use of a composite hypothesis Hoeffding test for statistical anomaly detection. Setting an appropriate threshold for the test given a desired false alarm probability involves approximating the false alarm probability. To that end, a large deviations asymptotic is typically used which, however, often results in an inaccurate setting of the threshold, especially for relatively small sample sizes. This, in turn, results in an anomaly detection test that does not control well for false alarms. In this paper, we develop a tighter approximation using the Central Limit Theorem (CLT) under Markovian assumptions. We apply our result to a network anomaly detection application and demonstrate its advantages over earlier work.

preprint2016arXiv

Countable tightness and $\mathfrak G$-bases on Free topological groups

Given a Tychonoff space $X$, let $F(X)$ and $A(X)$ be respectively the free topological group and the free Abelian topological group over $X$ in the sense of Markov. In this paper, we consider two topological properties of $F(X)$ or $A(X)$, namely the countable tightness and $\mathfrak G$-base. We provide some characterizations of the countable tightness and $\mathfrak G$-base of $F(X)$ and $A(X)$ for various special classes of spaces $X$. Furthermore, we also study the countable tightness and $\mathfrak G$-base of some $F_{n}(X)$ of $F(X)$.

preprint2016arXiv

Data-driven Estimation of Origin-Destination Demand and User Cost Functions for the Optimization of Transportation Networks

In earlier work (Zhang et al., 2016) we used actual traffic data from the Eastern Massachusetts transportation network in the form of spatial average speeds and road segment flow capacities in order to estimate Origin-Destination (OD) flow demand matrices for the network. Based on a Traffic Assignment Problem (TAP) formulation (termed "forward problem"), in this paper we use a scheme similar to our earlier work to estimate initial OD demand matrices and then propose a new inverse problem formulation in order to estimate user cost functions. This new formulation allows us to efficiently overcome numerical difficulties that limited our prior work to relatively small subnetworks and, assuming the travel latency cost functions are available, to adjust the values of the OD demands accordingly so that the flow observations are as close as possible to the solutions of the forward problem. We also derive sensitivity analysis results for the total user latency cost with respect to important parameters such as road capacities and minimum travel times. Finally, using the same actual traffic data from the Eastern Massachusetts transportation network, we quantify the Price of Anarchy (POA) for a much larger network than that in Zhang et al. (2016).

preprint2016arXiv

Discriminating the effects of collapse models from environmental diffusion with levitated nanospheres

Collapse models postulate the existence of intrinsic noise which modifies quantum mechanics and is responsible for the emergence of macroscopic classicality. Assessing the validity of these models is extremely challenging because it is nontrivial to discriminate unambiguously their presence in experiments where other hardly controllable sources of noise compete to the overall decoherence. Here we provide a simple procedure able to probe the hypothetical presence of the collapse noise with a levitated nanosphere in a Fabry-Perot cavity. We show that the stationary state of the system is particularly sensitive, under specific experimental conditions, to the interplay between the trapping frequency, the cavity size, and the momentum diffusion induced by the collapse models, allowing to detect them even in the presence of standard environmental noises.

preprint2016arXiv

Experimental observation of a topological band gap opening in ultracold Fermi gases with two-dimensional spin-orbit coupling

The recent experimental realization of synthetic spin-orbit coupling (SOC) opens a new avenue for exploring novel quantum states with ultracold atoms. However, in experiments for generating two-dimensional SOC (e.g., Rashba type), a perpendicular Zeeman field, which opens a band gap at the Dirac point and induces many topological phenomena, is still lacking. Here we theoretically propose and experimentally realize a simple scheme for generating two-dimension SOC and a perpendicular Zeeman field simultaneously in ultracold Fermi gases by tuning the polarization of three Raman lasers that couple three hyperfine ground states of atoms. The resulting band gap opening at the Dirac point is probed using spin injection radio-frequency spectroscopy. Our observation may pave the way for exploring topological transport and topological superfluids with exotic Majorana and Weyl fermion excitations in ultracold atoms.

preprint2016arXiv

Generation and Detection of Surface Plasmon Polaritons by Transition Metal Dichalcogenides for Chip-level Electronic-Photonic Integrated Circuits

The monolithic integration of electronics and photonics has attracted enormous attention due to its potential applications. However, the realization of such hybrid circuits has remained a challenge because it requires optical communication at nanometer scales. A major challenge to this integration is the identification of a suitable material. After discussing the material aspect of the challenge, we identified atomically thin transition metal dichalcogenides (TMDs) as a perfect material platform to implement the circuit. The selection of TMDs is based on their very distinct property: monolayer TMDs are able to emit and absorb light at the same wavelength determined by direct exciton transitions. To prove the concept, we fabricated simple devices consisting of silver nanowires as plasmonic waveguides and monolayer TMDs as active optoelectronic media. Using photoexcitation, direct optical imaging and spectral analysis, we demonstrated generation and detection of surface plasmon polaritons by monolayer TMDs. Regarded as novel materials for electronics and photonics, transition metal dichalcogenides are expected to find new applications in next generation integrated circuits.

preprint2016arXiv

Multi-user Massive MIMO Communication Systems Based on Irregular Antenna Arrays

In practical mobile communication engineering applications, surfaces of antenna array deployment regions are usually uneven. Therefore, massive multi-input-multi-output (MIMO) communication systems usually transmit wireless signals by irregular antenna arrays. To evaluate the performance of irregular antenna arrays, the matrix correlation coefficient and ergodic received gain are defined for massive MIMO communication systems with mutual coupling effects. Furthermore, the lower bound of the ergodic achievable rate, symbol error rate (SER) and average outage probability are firstly derived for multi-user massive MIMO communication systems using irregular antenna arrays. Asymptotic results are also derived when the number of antennas approaches infinity. Numerical results indicate that there exists a maximum achievable rate when the number of antennas keeps increasing in massive MIMO communication systems using irregular antenna arrays. Moreover, the irregular antenna array outperforms the regular antenna array in the achievable rate of massive MIMO communication systems when the number of antennas is larger than or equal to a given threshold.

preprint2016arXiv

Nighttime Haze Removal with Illumination Correction

Haze removal is important for computational photography and computer vision applications. However, most of the existing methods for dehazing are designed for daytime images, and cannot always work well in the nighttime. Different from the imaging conditions in the daytime, images captured in nighttime haze condition may suffer from non-uniform illumination due to artificial light sources, which exhibit low brightness/contrast and color distortion. In this paper, we present a new nighttime hazy imaging model that takes into account both the non-uniform illumination from artificial light sources and the scattering and attenuation effects of haze. Accordingly, we propose an efficient dehazing algorithm for nighttime hazy images. The proposed algorithm includes three sequential steps. i) It enhances the overall brightness by performing a gamma correction step after estimating the illumination from the original image. ii) Then it achieves a color-balance result by performing a color correction step after estimating the color characteristics of the incident light. iii) Finally, it remove the haze effect by applying the dark channel prior and estimating the point-wise environmental light based on the previous illumination-balance result. Experimental results show that the proposed algorithm can achieve illumination-balance and haze-free results with good color rendition ability.

preprint2016arXiv

On involutions in Weyl groups

Let $(W,S)$ be a Coxeter system and $\ast$ be an automorphism of $W$ with order $\leq 2$ such that $s^{\ast}\in S$ for any $s\in S$. Let $I_{\ast}$ be the set of twisted involutions relative to $\ast$ in $W$. In this paper we consider the case when $\ast=\text{id}$ and study the braid $I_\ast$-transformations between the reduced $I_\ast$-expressions of involutions. If $W$ is the Weyl group of type $B_n$ or $D_n$, we explicitly describe a finite set of basic braid $I_\ast$-transformations for all $n$ simultaneously, and show that any two reduced $I_\ast$-expressions for a given involution can be transformed into each other through a series of basic braid $I_\ast$-transformations. In both cases, these basic braid $I_\ast$-transformations consist of the usual basic braid transformations plus some natural "right end transformations" and plus exactly one extra transformation. The main result generalizes our previous work for the Weyl group of type $A_{n}$.

preprint2016arXiv

Oscillation and variation for Riesz transform associated with Bessel operators

Let $λ>0$ and $\triangle_λ:=-\frac{d^2}{dx^2}-\frac{2λ}{x} \frac d{dx}$ be the Bessel operator on $\mathbb R_+:=(0,\infty)$. We show that the oscillation operator $\mathcal{O}(R_{Δ_λ,\ast})$ and variation operator $\mathcal{V}_ρ(R_{Δ_λ,\ast})$ of the Riesz transform $R_{Δ_λ}$ associated with $Δ_λ$ are both bounded on $L^p(\mathbb R_+, dm_λ)$ for $p\in(1,\,\infty)$, from $L^1(\mathbb{R}_{+},dm_λ)$ to $L^{1,\,\infty}(\mathbb{R}_{+},dm_λ)$, and from $L^{\infty}(\mathbb{R}_{+},dm_λ)$ to $BMO(\mathbb{R}_{+},dm_λ)$, where $ρ\in (2,\infty)$ and $dm_λ(x):=x^{2λ}dx$. As an application, we give the corresponding $L^p$-estimates for $β$-jump operators and the number of up-crossing.

preprint2016arXiv

Oscillation and variation for semigroups associated with Bessel operators

Let $λ>0$ and $\triangle_λ:=-\frac{d^2}{dx^2}-\frac{2λ}{x} \frac d{dx}$ be the Bessel operator on $\mathbb R_+:=(0,\infty)$. We show that the oscillation operator ${\mathcal O(P^{[λ]}_\ast)}$ and variation operator ${\mathcal V}_ρ(P^{[λ]}_\ast)$ of the Poisson semigroup $\{P^{[λ]}_t\}_{t>0}$ associated with $Δ_λ$ are both bounded on $L^p(\mathbb R_+, dm_λ)$ for $p\in(1, \infty)$, $BMO({{\mathbb R}_+},dm_λ)$, from $L^1({{\mathbb R}_+},dm_λ)$ to $L^{1,\,\infty}({{\mathbb R}_+},dm_λ)$, and from $H^1({{\mathbb R}_+},dm_λ)$ to $L^1({{\mathbb R}_+},dm_λ)$, where $ρ\in(2, \infty)$ and $dm_λ(x):=x^{2λ}\,dx$. As an application, an equivalent characterization of $H^1({{\mathbb R}_+},dm_λ)$ in terms of ${\mathcal V}_ρ(P^{[λ]}_\ast)$ is also established. All these results hold if $\{P^{[λ]}_t\}_{t>0}$ is replaced by the heat semigroup $\{W^{[λ]}_t\}_{t>0}$. }

preprint2016arXiv

RGB-D-based Action Recognition Datasets: A Survey

Human action recognition from RGB-D (Red, Green, Blue and Depth) data has attracted increasing attention since the first work reported in 2010. Over this period, many benchmark datasets have been created to facilitate the development and evaluation of new algorithms. This raises the question of which dataset to select and how to use it in providing a fair and objective comparative evaluation against state-of-the-art methods. To address this issue, this paper provides a comprehensive review of the most commonly used action recognition related RGB-D video datasets, including 27 single-view datasets, 10 multi-view datasets, and 7 multi-person datasets. The detailed information and analysis of these datasets is a useful resource in guiding insightful selection of datasets for future research. In addition, the issues with current algorithm evaluation vis-á-vis limitations of the available datasets and evaluation protocols are also highlighted; resulting in a number of recommendations for collection of new datasets and use of evaluation protocols.

preprint2016arXiv

Study on the magnetic measurement results of the injection system for CSNS/RCS

A combination of the H- stripping and phase space painting method is used to accumulate a high intensity beam in the Rapid Cycling Synchrotron (RCS) of the China Spallation Neutron Source (CSNS). The injection system for CSNS/RCS consists of three kinds of magnets: four direct current magnets (BC1-BC4), eight alternating current magnets (BH1-BH4 and BV1-BV4), two septum magnets (ISEP1 and ISEP2). In this paper, the magnetic measurements of the injection system were introduced and the data analysis was processed. The field uniformity and magnetizing curves of these magnets were given, and then the magnetizing fitting equations were obtained.

preprint2016arXiv

Using almost-everywhere theorems from analysis to study randomness

We study algorithmic randomness notions via effective versions of almost-everywhere theorems from analysis and ergodic theory. The effectivization is in terms of objects described by a computably enumerable set, such as lower semicomputable functions. The corresponding randomness notions are slightly stronger than \ML\ (ML) randomness. We establish several equivalences. Given a ML-random real $z$, the additional randomness strengths needed for the following are equivalent. \n (1) all effectively closed classes containing $z$ have density $1$ at $z$. \n (2) all nondecreasing functions with uniformly left-c.e.\ increments are differentiable at $z$. \n (3) $z$ is a Lebesgue point of each lower semicomputable integrable function. We also consider convergence of left-c.e.\ martingales, and convergence in the sense of Birkhoff's pointwise ergodic theorem. Lastly we study randomness notions for density of $Π^0_n$ and $Σ^1_1$ classes.

preprint2015arXiv

5G green cellular networks considering power allocation schemes

It is important to assess the effect of transmit power allocation schemes on the energy consumption on random cellular networks. The energy efficiency of 5G green cellular networks with average and water-filling power allocation schemes is studied in this paper. Based on the proposed interference and achievable rate model, an energy efficiency model is proposed for MIMO random cellular networks. Furthermore, the energy efficiency with average and water-filling power allocation schemes are presented, respectively. Numerical results indicate that the maximum limits of energy efficiency are always there for MIMO random cellular networks with different intensity ratios of mobile stations (MSs) to base stations (BSs) and channel conditions. Compared with the average power allocation scheme, the water-filling scheme is shown to improve the energy efficiency of MIMO random cellular networks when channel state information (CSI) is attainable for both transmitters and receivers.

preprint2015arXiv

Deep Convolutional Neural Networks for Action Recognition Using Depth Map Sequences

Recently, deep learning approach has achieved promising results in various fields of computer vision. In this paper, a new framework called Hierarchical Depth Motion Maps (HDMM) + 3 Channel Deep Convolutional Neural Networks (3ConvNets) is proposed for human action recognition using depth map sequences. Firstly, we rotate the original depth data in 3D pointclouds to mimic the rotation of cameras, so that our algorithms can handle view variant cases. Secondly, in order to effectively extract the body shape and motion information, we generate weighted depth motion maps (DMM) at several temporal scales, referred to as Hierarchical Depth Motion Maps (HDMM). Then, three channels of ConvNets are trained on the HDMMs from three projected orthogonal planes separately. The proposed algorithms are evaluated on MSRAction3D, MSRAction3DExt, UTKinect-Action and MSRDailyActivity3D datasets respectively. We also combine the last three datasets into a larger one (called Combined Dataset) and test the proposed method on it. The results show that our approach can achieve state-of-the-art results on the individual datasets and without dramatical performance degradation on the Combined Dataset.

preprint2015arXiv

Dissociation of Feshbach molecules via spin-orbit coupling in ultracold Fermi gases

We study the dissociation of Feshbach molecules in ultracold Fermi gases with spin-orbit (SO) coupling. Since SO coupling can induce quantum transition between the Feshbach molecules and the fully polarized Fermi gas, the Feshbach molecules can be dissociated by the SO coupling. We experimentally realized this new type of dissociation in ultracold gases of 40K atoms with SO coupling created by Raman beams, and observed that the dissociation rate is highly non-monotonic on both the positive and negative Raman-detuning sides. Our results show that the dissociation of Feshbach molecules can be controlled by new degrees of freedoms, i.e., the SO-coupling intensity or the momenta of the Raman beams, as well as the detuning of the Raman beams.

preprint2015arXiv

Entanglement distribution over quantum code-division-multiple-access networks

We present a method for quantum entanglement distribution over a so-called code-division-multiple-access network, in which two pairs of users share the same quantum channel to transmit information. The main idea of this method is to use different broad-band chaotic phase shifts, generated by electro-optic modulators (EOMs) and chaotic Colpitts circuits, to encode the information-bearing quantum signals coming from different users, and then recover the masked quantum signals at the receiver side by imposing opposite chaotic phase shifts. The chaotic phase shifts given to different pairs of users are almost uncorrelated due to the randomness of chaos and thus the quantum signals from different pair of users can be distinguished even when they are sent via the same quantum channel. It is shown that two maximally-entangled states can be generated between two pairs of users by our method mediated by bright coherent lights, which can be more easily implemented in experiments compared with single-photon lights. Our method is robust under the channel noises if only the decay rates of the information-bearing fields induced by the channel noises are not quite high. Our study opens up new perspectives for addressing and transmitting quantum information in future quantum networks.

preprint2015arXiv

Experimental realization of a two-dimensional synthetic spin-orbit coupling in ultracold Fermi gases

Spin-orbit coupling (SOC) is central to many physical phenomena, including fine structures of atomic spectra and quantum topological matters. Whereas SOC is in general fixed in a physical system, atom-laser interaction provides physicists a unique means to create and control synthetic SOC for ultracold atoms \cite{Dalibard}. Though significant experimental progresses have been made, a bottleneck in current studies is the lack of a two-dimensional (2D) synthetic SOC, which is crucial for realizing high-dimensional topological matters. Here, we report the experimental realization of 2D SOC in ultracold $^{40}$K Fermi gases using three lasers, each of which dresses one atomic hyperfine spin state. Through spin injection radio-frequency (rf) spectroscopy, we probe the spin-resolved energy dispersions of dressed atoms, and observe a highly controllable Dirac point created by the 2D SOC. Our work paves the way for exploring high-dimensional topological matters in ultracold atoms using Raman schemes.

preprint2015arXiv

Experimental study of balanced optical homodyne and heterodyne detection by controlling sideband modulation

We experimentally study optical homodyne and heterodyne detections with a same setup, which is flexible to manipulate the signal sideband modulation. When the modulation only generate a single signal sideband, the light field measurement by mixing the single sideband at $ω_{0}+Ω$ with a strong local oscillator at the carrier frequency $ω_{0}$ on a beam splitter become balanced heterodyne detection. When two signal sidebands at $ω_{0}\pmΩ$ are generated and the relative phase of the two sidebands is locked, this measurement corresponds to optical balanced homodyne detection. With this setup, we may confirm directly that the signal-to-noise ratio with heterodyne detection is two-fold worse than that with homodyne detection. This work will have important applications in quantum state measurement and quantum information.

preprint2015arXiv

Giant nonlinearity via breaking parity-time symmetry: a route to low-threshold phonon diodes

Nonreciprocal devices that permit wave transmission in only one direction are indispensible in many fields of science including, e.g., electronics, optics, acoustics, and thermodynamics. Manipulating phonons using such nonreciprocal devices may have a range of applications such as phonon diodes, transistors, switches, etc. One way of achieving nonreciprocal phononic devices is to use materials with strong nonlinear response to phonons. However, it is not easy to obtain the required strong mechanical nonlinearity, especially for few-phonon situations. Here, we present a general mechanism to amplify nonlinearity using $\mathcal{PT}$-symmetric structures, and show that an on-chip micro-scale phonon diode can be fabricated using a $\mathcal{PT}$-symmetric mechanical system, in which a lossy mechanical-resonator with very weak mechanical nonlinearity is coupled to a mechanical resonator with mechanical gain but no mechanical nonlinearity. When this coupled system transits from the $\mathcal{PT}$-symmetric regime to the broken-$\mathcal{PT}$-symmetric regime, the mechanical nonlinearity is transferred from the lossy resonator to the one with gain, and the effective nonlinearity of the system is significantly enhanced. This enhanced mechanical nonlinearity is almost lossless because of the gain-loss balance induced by the $\mathcal{PT}$-symmetric structure. Such an enhanced lossless mechanical nonlinearity is then used to control the direction of phonon propagation, and can greatly decrease (by over three orders of magnitude) the threshold of the input-field intensity necessary to observe the unidirectional phonon transport. We propose an experimentally realizable lossless low-threshold phonon diode of this type. Our study opens up new perspectives for constructing on-chip few-phonon devices and hybrid phonon-photon components.

preprint2015arXiv

Locally $σ$-compact rectifiable spaces

A topological space $G$ is said to be a {\it rectifiable space} provided that there are a surjective homeomorphism $φ:G\times G\rightarrow G\times G$ and an element $e\in G$ such that $π_{1}\circ φ=π_{1}$ and for every $x\in G$, $φ(x, x)=(x, e)$, where $π_{1}: G\times G\rightarrow G$ is the projection to the first coordinate. In this paper, we first prove that each locally compact rectifiable space is paracompact, which gives an affirmative answer to Arhangel'skii and Choban's question (Arhangel'skii and Choban [3]). Then we prove that every locally $σ$-compact rectifiable space with a $bc$-base is locally compact or zero-dimensional, which improves Arhangel'skii and van Mill's result (Arhangel'skii and van Mill [4]). Finally, we prove that each $k_ω$-rectifiable space is rectifiable complete.

preprint2015arXiv

Long-lived nanosecond spin relaxation and spin coherence of electrons in monolayer MoS_2 and WS_2

The recently-discovered monolayer transition metal dichalcogenides (TMDCs) provide a fertile playground to explore new coupled spin-valley physics. Although robust spin and valley degrees of freedom are inferred from polarized photoluminescence (PL) experiments, PL timescales are necessarily constrained by short-lived (3-100ps) electron-hole recombination. Direct probes of spin/valley polarization dynamics of resident carriers in electron (or hole) doped TMDCs, which may persist long after recombination ceases, are at an early stage. Here we directly measure the coupled spin-valley dynamics in electron-doped MoS_2 and WS_2 monolayers using optical Kerr spectroscopy, and unambiguously reveal very long electron spin lifetimes exceeding 3ns at 5K (2-3 orders of magnitude longer than typical exciton recombination times). In contrast with conventional III-V or II-VI semiconductors, spin relaxation accelerates rapidly in small transverse magnetic fields. Supported by a model of coupled spin-valley dynamics, these results indicate a novel mechanism of itinerant electron spin dephasing in the rapidly-fluctuating internal spin-orbit field in TMDCs, driven by fast intervalley scattering. Additionally, a long-lived spin coherence is observed at lower energies, commensurate with localized states. These studies provide crucial insight into the physics underpinning spin and valley dynamics of resident electrons in atomically-thin TMDCs.

preprint2015arXiv

Metrology with $\mathcal{PT}$-symmetric cavities: Enhanced sensitivity near the $\mathcal{PT}$-phase transition

We propose and analyze a new approach based on parity-time ($\mathcal{PT}$) symmetric microcavities with balanced gain and loss to enhance the performance of cavity-assisted metrology. We identify the conditions under which $\mathcal{PT}$-symmetric microcavities allow to improve sensitivity beyond what is achievable in loss-only systems. We discuss its application to the detection of mechanical motion, and show that the sensitivity is significantly enhanced in the vicinity of the transition point from unbroken- to broken-$\mathcal{PT}$ regimes. We believe that our results open a new direction for $\mathcal{PT}$-symmetric physical systems and it may find use in ultra-high precision metrology and sensing.

preprint2015arXiv

Noise suppression of on-chip mechanical resonators by chaotic coherent feedback

We propose a method to decouple the nanomechanical resonator in optomechanical systems from the environmental noise by introducing a chaotic coherent feedback loop. We find that the chaotic controller in the feedback loop can modulate the dynamics of the controlled optomechanical system and induce a broadband response of the mechanical mode. This broadband response of the mechanical mode will cut off the coupling between the mechanical mode and the environment and thus suppress the environmental noise of the mechanical modes. As an application, we use the protected optomechanical system to act as a quantum memory. It's shown that the noise-decoupled optomechanical quantum memory is efficient for storing information transferred from coherent or squeezed light.

preprint2015arXiv

On involutions in symmetric groups and a conjecture of Lusztig

Let $(W, S)$ be a Coxeter system equipped with a fixed automorphism $\ast$ of order $\leq 2$ which preserves $S$. Lusztig (and with Vogan in some special cases) have shown that the space spanned by set of "twisted" involutions was naturally endowed with a module structure of the Hecke algebra of $(W, S)$. Lusztig has conjectured that this module is isomorphic to the right ideal of the Hecke algebra (with Hecke parameter $u^2$) associated to $(W,S)$ generated by the element $X_{\emptyset}:=\sum_{w^\ast=w}u^{-\ell(w)}T_w$. In this paper we prove this conjecture in the case when $\ast=\text{id}$ and $W$ is the symmetric group on $n$ letters.

preprint2015arXiv

Panther: Fast Top-k Similarity Search in Large Networks

Estimating similarity between vertices is a fundamental issue in network analysis across various domains, such as social networks and biological networks. Methods based on common neighbors and structural contexts have received much attention. However, both categories of methods are difficult to scale up to handle large networks (with billions of nodes). In this paper, we propose a sampling method that provably and accurately estimates the similarity between vertices. The algorithm is based on a novel idea of random path, and an extended method is also presented, to enhance the structural similarity when two vertices are completely disconnected. We provide theoretical proofs for the error-bound and confidence of the proposed algorithm. We perform extensive empirical study and show that our algorithm can obtain top-k similar vertices for any vertex in a network approximately 300x faster than state-of-the-art methods. We also use identity resolution and structural hole spanner finding, two important applications in social networks, to evaluate the accuracy of the estimated similarities. Our experimental results demonstrate that the proposed algorithm achieves clearly better performance than several alternative methods.

preprint2015arXiv

Probabilistic representations of solutions of elliptic boundary value problem and non-symmetric semigroups

In this paper, we use a probabilistic approach to show that there exists a unique, bounded continuous solution to the Dirichlet boundary value problem for a general class of second order non-symmetric elliptic operators $L$ with singular coefficients, which does not necessarily have the maximum principle. The theory of Dirichlet forms and heat kernel estimates play a crucial role in our approach. A probabilistic representation of the non-symmetric semigroup $\{T_t\}_{t\ge 0}$ generated by $L$ is also given.

preprint2015arXiv

Squeezed Optomechanics with Phase-matched Amplification and Dissipation

We investigate the nonlinear interaction between a squeezed cavity mode and a mechanical mode in an optomechanical system (OMS) that allows us to selectively obtain either a radiation-pressure coupling or a parametric-amplification process. The squeezing of the cavity mode can enhance the interaction strength into the single-photon strong-coupling regime, even when the OMS is originally in the weak-coupling regime. Moreover, the noise of the squeezed mode can be suppressed completely by introducing a broadband-squeezed vacuum that is phase-matched with the parametric amplification that squeezes the cavity mode. This proposal offers an alternative approach to control OMS using a squeezed cavity mode, which should allow single-photon quantum processes to be implemented with currently available optomechanical technology. Potential applications range from engineering single-photon sources to nonclassical phonon states.

preprint2014arXiv

Coherent-feedback-induced photon blockade and optical bistability by an optomechanical controller

It is well-known that some nonlinear phenomena such as strong photon blockade are hard to be observed in optomechanical system with current experimental technology. Here, we present a coherent feedback control strategy in which a linear cavity is coherently controlled by an optomechanical controller in a feedback manner. The coherent feedback loop transfers and enhances quantum nonlinearity from the controller to the controlled cavity, which makes it possible to observe strong nonlinear effects in either linear cavity or optomechanical cavity. More interestingly, we find that the strong photon blockade under single-photon optomechanical weak coupling condition could be observed in the quantum regime. Additionally, the coherent feedback loop leads to two-photon and multiphoton tunnelings for the controlled linear cavity, which are also typical quantum nonlinear phenomenon. We hope that our work can give new perspectives in engineering nonlinear quantum phenomena.

preprint2014arXiv

Engineering of nonclassical motional states in optomechanical systems

We propose to synthesize arbitrary nonclassical motional states in optomechanical systems by using sideband excitations and photon blockade. We first demonstrate that the Hamiltonian of the optomechanical systems can be reduced, in the strong single-photon optomechanical coupling regime when the photon blockade occurs, to one describing the interaction between a driven two-level trapped ion and the vibrating modes, and then show a method to generate target states by using a series of classical pulses with desired frequencies, phases, and durations. We further analyze the effect of the photon leakage, due to small anharmonicity, on the fidelity of the expected motional state, and study environment induced decoherence. Moreover, we also discuss the experimental feasibility and provide operational parameters using the possible experimental data.

preprint2014arXiv

More nonlocality with less entanglement in a tripartite atom-optomechanical system

We study quantum effects in hybrid atomic optomechanics in a system comprising a cloud of atoms and a mobile mirror mediated by a single-mode cavity. Tripartite nonlocality is observed in the atom-light-mirror system, as demonstrated by the violation of the Mermin-Klyshko (MK) inequality. It has been shown [C. Genes, et al., PRA 77, 050307 (R) (2008)] that tripartite entanglement is optimized when the cavity is resonant with the anti-Stokes sideband of the driving laser and the atomic frequency matches the Stokes one. However, we show that this is not the case for the nonlocality. The MK function achieves {\it minima} when the atoms are resonant with both the Stokes and anti-Stokes sidebands, and unexpectedly, we find violation of the MK inequality only in a parameter region where entanglement is far from being maximum. A negative relation exists between nonlocality and entanglement with consideration of the possibility of bipartite nonlocality in the violation of the MK inequality. We also study the non-classicality of the mirror by post-selected measurements, e.g. Geiger-like detection, on the cavity and/or the atoms. We show that with feasible parameters Geiger-like detection on the atoms can effectively induce mechanical non-classicality.

preprint2014arXiv

New results on Hunt's hypothesis (H) for Lévy processes

In this paper, we present new results on Hunt's hypothesis (H) for Lévy processes. We start with a comparison result on Lévy processes which implies that big jumps have no effect on the validity of (H). Based on this result and the Kanda-Forst-Rao theorem, we give examples of subordinators satisfying (H). Afterwards we give a new necessary and sufficient condition for (H) and obtain an extended Kanda-Forst-Rao theorem. By virtue of this theorem, we give a new class of Lévy processes satisfying (H). Finally, we construct a type of subordinators that does not satisfy Rao's condition.

preprint2014arXiv

Nonlinear quantum input-output analysis using Volterra series

Quantum input-output theory plays a very important role for analyzing the dynamics of quantum systems, especially large-scale quantum networks. As an extension of the input-output formalism of Gardiner and Collet, we develop a new approach based on the quantum version of the Volterra series which can be used to analyze nonlinear quantum input-output dynamics. By this approach, we can ignore the internal dynamics of the quantum input-output system and represent the system dynamics by a series of kernel functions. This approach has the great advantage of modelling weak-nonlinear quantum networks. In our approach, the number of parameters, represented by the kernel functions, used to describe the input-output response of a weak-nonlinear quantum network, increases linearly with the scale of the quantum network, not exponentially as usual. Additionally, our approach can be used to formulate the quantum network with both nonlinear and nonconservative components, e.g., quantum amplifiers, which cannot be modelled by the existing methods, such as the Hudson-Parthasarathy model and the quantum transfer function model. We apply our general method to several examples, including Kerr cavities, optomechanical transducers, and a particular coherent feedback system with a nonlinear component and a quantum amplifier in the feedback loop. This approach provides a powerful way to the modelling and control of nonlinear quantum networks.

preprint2014arXiv

Phonon amplification in two coupled cavities containing one mechanical resonator

We study a general theory of phonon lasing [I. S. Grudinin et al., Phys. Rev. Lett. 104, 083901 (2010)] in coupled optomechancial systems. We derive the dynamical equation of the phonon lasing using supermodes formed by two cavity modes. A general threshold condition for phonon lasing is obtained. We also show the differences between phonon lasing and photon lasing, generated by photonic supermodes and two-level atomic systems, respectively. We find that the phonon lasing can be realized in certain parameter regime near the threshold. The phase diagram and second-order correlation function of the phonon lasing are also studied to show some interesting phenomena that cannot be observed in the common photon lasing with the two-level systems.

preprint2014arXiv

PT-Symmetric Phonon Laser

By exploiting recent developments associated with coupled microcavities, we introduce the concept of PT-symmetric phonon laser with balanced gain and loss. This is accomplished by introducing gain to one of the microcavities such that it balances the passive loss of the other. In the vicinity of the gain-loss balance, a strong nonlinear relation emerges between the intracavity photon intensity and the input power. This then leads to a giant enhancement of both optical pressure and mechanical gain, resulting in a highly efficient phonon-lasing action. These results provide a promising approach for manipulating optomechanical systems through PT-symmetric concepts. Potential applications range from enhancing mechanical cooling to designing phonon-laser amplifiers.

preprint2014arXiv

Radio-Frequency Manipulation of Fano-Feshbach Resonances in an Ultracold Fermi Gas of $^{40}$K

Experimental control of magnetic Fano-Feshbach resonances in ultracold $^{40}$K Fermi gases, using radio-frequency (RF) fields, is demonstrated. Spectroscopic measurements are made of three molecular levels within 50 MHz of the atomic continuum, along with their variation with magnetic field. Modifying the scattering properties by an RF field is shown by measuring the loss profile versus magnetic field. This work provides the high accuracy locations of ground molecular states near the s-wave Fano-Feshbach resonance, which can be used to study the crossover regime from a Bose-Einstein condensate to a Bardeen-Cooper-Schrieffer superfluid in presence of an RF field.

preprint2014arXiv

Signal Flows in Non-Markovian Linear Quantum Feedback Networks

Enabled by rapidly developing quantum technologies, it is possible to network quantum systems at a much larger scale in the near future. To deal with non-Markovian dynamics that is prevalent in solid-state devices, we propose a general transfer function based framework for modeling linear quantum networks, in which signal flow graphs are applied to characterize the network topology by flow of quantum signals. We define a noncommutative ring $\mathbb{D}$ and use its elements to construct Hamiltonians, transformations and transfer functions for both active and passive systems. The signal flow graph obtained for direct and indirect coherent quantum feedback systems clearly show the feedback loop via bidirectional signal flows. Importantly, the transfer function from input to output field is derived for non-Markovian quantum systems with colored inputs, from which the Markovian input-output relation can be easily obtained as a limiting case. Moreover, the transfer function possesses a symmetry structure that is analogous to the well-know scattering transformation in \sd picture. Finally, we show that these transfer functions can be integrated to build complex feedback networks via interconnections, serial products and feedback, which may include either direct or indirect coherent feedback loops, and transfer functions between quantum signal nodes can be calculated by the Riegle's matrix gain rule. The theory paves the way for modeling, analyzing and synthesizing non-Markovian linear quantum feedback networks in the frequency-domain.

preprint2014arXiv

Spectrum and Energy Efficiency Evaluation of Two-Tier Femtocell networks With Partially Open Channels

Two-tier femtocell networks is an efficient communication architecture that significantly improves throughput in indoor environments with low power consumption. Traditionally, a femtocell network is usually configured to be either completely open or completely closed in that its channels are either made available to all users or used by its own users only. This may limit network flexibility and performance. It is desirable for owners of femtocell base stations if a femtocell can partially open its channels for external users access. In such scenarios, spectrum and energy efficiency becomes a critical issue in the design of femtocell network protocols and structure. In this paper, we conduct performance analysis for two-tier femtocell networks with partially open channels. In particular, we build a Markov chain to model the channel access in the femtocell network and then derive the performance metrics in terms of the blocking probabilities. Based on stationary state probabilities derived by Markov chain models, spectrum and energy efficiency are modeled and analyzed under different scenarios characterized by critical parameters, including number of femtocells in a macrocell, average number of users, and number of open channels in a femtocell. Numerical and Monte-Carlo (MC) simulation results indicate that the number of open channels in a femtocell has an adverse impact on the spectrum and energy efficiency of two-tier femtocell networks. Results in this paper provide guidelines for trading off spectrum and energy efficiency of two-tier femtocell networks by configuring different numbers of open channels in a femtocell.

preprint2014arXiv

The obstacle problem for quasilinear stochastic PDEs: Analytical approach

We prove an existence and uniqueness result for quasilinear Stochastic PDEs with obstacle (OSPDE in short). Our method is based on analytical technics coming from the parabolic potential theory. The solution is expressed as a pair $(u,ν)$ where $u$ is a predictable continuous process which takes values in a proper Sobolev space and $ν$ is a random regular measure satisfying the minimal Skohorod condition.

preprint2014arXiv

Transparency and amplification in a hybrid system of mechanical resonator and circuit QED

We theoretically study the transparency and amplification of a weak probe field applied to the cavity in hy- brid systems formed by a driven superconducting circuit QED system and a mechanical resonator, or a driven optomechanical system and a superconducting qubit. We find that both the mechanical resonator and the su- perconducting qubit can result in the transparency to a weak probe field in such hybrid systems when a strong driving field is applied to the cavity. We also find that the weak probe field can be amplified in some parameter regimes. We further study the statistical properties of the output field via the degrees of second-order coherence. We find that the nonclassicality of the output field strongly depends on the system parameters. Our studies show that one can control single-photon transmission in the optomechanical system via a tunable artificial atom or in the circuit QED system via a mechanical resonator.

preprint2014arXiv

Two-dimensional balanced sampling plans avoiding adjacent units

Hedayat et al. first introduced balanced sampling plans for the exclusion of contiguous units. Wright detailed the results of a preliminary investigation of two-dimensional balanced sampling plans avoiding adjacent units (2-BSAs), and pointed out explicitly three types of 2-BSAs, which have different adjacency scheme, namely "Row and Column", "Sharing a Border" and "Island". This paper will provide more details for the three types of 2-BSAs from the point of view of design theory.

preprint2013arXiv

Feedback-induced nonlinearity and superconducting on-chip quantum optics

Quantum coherent feedback has been proven to be an efficient way to tune the dynamics of quantum optical systems and, recently, those of solid-state quantum circuits. Here, inspired by the recent progress of quantum feedback experiments, especially those in mesoscopic circuits, we prove that superconducting circuit QED systems, shunted with a coherent feedback loop, can change the dynamics of a superconducting transmission line resonator, i.e., a linear quantum cavity, and lead to strong on-chip nonlinear optical phenomena. We find that bistability can occur under the semiclassical approximation, and photon anti-bunching can be shown in the quantum regime. Our study presents new perspectives for engineering nonlinear quantum dynamics on a chip.

preprint2013arXiv

In-situ EXAFS study on the thermal decomposition of TiH2

Thermal decomposition behaviors of TiH2 powder under a flowing helium atmosphere and in a low vacuum condition have been studied by using in-situ EXAFS technique. By an EXAFS analysis containing the multiple scattering paths including H atoms, the changes of hydrogen stoichiometric ratio and the phase transformation sequence are obtained. The results demonstrate that the initial decomposition temperature is dependent on experimental conditions, which occurs, respectively, at about 300 and 400 degree in a low vacuum condition and under a flowing helium atmosphere. During the decomposition process of TiH2 in a low vacuum condition, the sample experiences a phase change process: δ(TiH2) - δ(TiHx) - δ(TiHx)+β(TiHx) - δ(TiHx)+β(TiHx)+α(Ti) - β(TiHx)+α(Ti) - α(Ti)+β(Ti). This study offers a way to detect the structural information of hydrogen. A detailed discussion about the decomposition process of TiH2 is given in this paper.

preprint2013arXiv

Levy-Khintchine type representation of Dirichlet generators and Semi-Dirichlet forms

Let $U$ be an open set of $\mathbb{R}^n$, $m$ a positive Radon measure on $U$ such that ${\rm supp}[m]=U$, and $(P_t)_{t>0}$ a strongly continuous contraction sub-Markovian semigroup on $L^2(U;m)$. We investigate the structure of $(P_t)_{t>0}$. (i) Denote respectively by $(A,D(A))$ and $(\hat A,D(\hat A))$ the generator and the co-generator of $(P_t)_{t>0}$. Under the assumption that $C^{\infty}_0(U)\subset D(A)\cap D(\hat A)$, we give an explicit Lévy-Khintchine type representation of $A$ on $C^{\infty}_0(U)$. (ii) If $(P_t)_{t>0}$ is an analytic semigroup and hence is associated with a semi-Dirichlet form $({\cal E}, D({\cal E}))$, we give an explicit characterization of ${\cal E}$ on $C^{\infty}_0(U)$ under the assumption that $C^{\infty}_0(U)\subset D({\cal E})$. We also present a LeJan type transformation rule for the diffusion part of regular semi-Dirichlet forms on general state spaces.

preprint2013arXiv

Many-body \textit{T}-matrix theory of a strongly interacting spin-orbit coupled Fermi gas: Momentum-resolved radio-frequency spectroscopy and fermionic pairing

Interacting Fermi gases with spin-orbit coupling are responsible for many intriguing phenomena such as topological superfluids and Majorana fermions. Here we characterize theoretically fermionic pairing in a strongly interacting spin-orbit coupled Fermi gas, by using momentum-resolved radio-frequency spectroscopy. We develop a strong-coupling $T$-matrix theory and present a phase diagram near the unitary resonance limit. A smooth transition from atomic to molecular responses in the momentum-resolved spectroscopy is predicted, with a clear signature of anisotropic pairing at and below resonance. Our prediction with many-body pairing can be directly tested in a spin-orbit coupled Fermi gas of $^{40}$K or $^{6}$Li atoms near broad Feshbach resonances.

preprint2013arXiv

On $hp$-Convergence of PSWFs and A New Well-Conditioned Prolate-Collocation Scheme

The first purpose of this paper is to provide a rigorous proof for the nonconvergence of $h$-refinement in $hp$-approximation by the PSWFs, a surprising convergence property that was first observed by Boyd et al [J. Sci. Comput., 2013]. The second purpose is to offer a new basis that leads to spectral-collocation systems with condition numbers independent of $(c,N),$ the intrinsic bandwidth parameter and the number of collocation points. In addition, this work gives insights into the development of effective spectral algorithms using this non-polynomial basis. We in particular highlight that the collocation scheme together with a very practical rule for pairing up $(c,N)$ significantly outperforms the Legendre polynomial-based method (and likewise other Jacobi polynomial-based method) in approximating highly oscillatory bandlimited functions.

preprint2013arXiv

Optical control of a magnetic Feshbach resonance in ultracold Fermi gases

We use laser light near-resonant with a molecular bound-to-bound transition to control a magnetic Feshbach resonance in ultracold Fermi gases of $^{40}$K atoms. The spectrum of excited molecular states is measured by applying a laser field that couples the ground Feshbach molecular state to electronically excited molecular states. Nine strong bound-to-bound resonances are observed below the $^{2}P_{1/2}+^{2}S_{1/2}$ threshold. We use radio-frequency spectroscopy to characterize the laser-dressed bound state near a specific bound-to-bound resonance and show clearly the shift of the magnetic Feshbach resonance using light with negligible atomic loss. The demonstrated technology could be used to modify interatomic interactions with high spatial and temporal resolutions in the crossover regime from a Bose-Einstein condensate (BEC) to a Bardeen-Cooper-Schrieffer (BCS) superfluid.

preprint2013arXiv

Quantum internet using code division multiple access

A crucial open problem in large-scale quantum networks is how to efficiently transmit quantum data among many pairs of users via a common data-transmission medium. We propose a solution by developing a quantum code division multiple access (q-CDMA) approach in which quantum information is chaotically encoded to spread its spectral content, and then decoded via chaos synchronization to separate different sender-receiver pairs. In comparison to other existing approaches, such as frequency division multiple access (FDMA), the proposed q-CDMA can greatly increase the information rates per channel used, especially for very noisy quantum channels.

preprint2013arXiv

Radio-frequency spectroscopy of a strongly interacting spin-orbit coupled Fermi gas

We investigate experimentally and theoretically radio-frequency spectroscopy and pairing of a spin-orbit-coupled Fermi gas of $^{40}$K atoms near a Feshbach resonance at $B_{0}=202.2$ G. Experimentally, the integrated spectroscopy is measured, showing characteristic blue and red shifts in the atomic and molecular responses, respectively, with increasing spin-orbit coupling. Theoretically, a smooth transition from atomic to molecular responses in the momentum-resolved spectroscopy is predicted, with a clear signature of anisotropic pairing at and below resonance. Our many-body prediction agrees qualitatively well with the observed spectroscopy near the Feshbach resonance.

preprint2013arXiv

Spin-Orbit Coupling Induced Coherent Production of Feshbach Molecules in a Degenerate Fermi Gas

In this work we demonstrate a dynamic process in which SO coupling can coherently produce s-wave Feshbach molecules from a fully polarized Fermi gas, and can induce a coherent oscillation between Feshbach molecules and spin polarized gas. For comparison, we also show that such phenomena are absent if the inter-component coupling is momentum-independent. This demonstrates experimentally that SO coupling does provide finite matrix element between a singlet state and a triplet state, and therefore, implies the bound pairs of a system with SO coupling have triplet p-wave component, which can become topological superfluid by further cooling these pairs to condensation and confining them to lower dimension.

preprint2013arXiv

Topological insulators, spin, and the tight-binding method

As one of the first proposed topologically protected states, the quantum spin Hall effect in graphene relies critically on the existence of a spin-dependent gap at the K/K' points of the Brillouin zone. Using a tight-binding formulation based on the method of invariants, we identify the origin of such an intrinsic gap as the three-center interaction between the pi-orbitals caused by spin-orbit interactions. This methodology incorporates all symmetry compliant interactions previously neglected and has wider applications for comparisons between first-principle calculations and the tight-binding method. It also identifies a correction to the Haldane model and its generalization, which incorporates the spin degrees of freedom and reproduces all the salient features required for the quantum spin Hall effect in graphene.

preprint2012arXiv

A Literature Survey of Cooperative Caching in Content Distribution Networks

Content distribution networks (CDNs) which serve to deliver web objects (e.g., documents, applications, music and video, etc.) have seen tremendous growth since its emergence. To minimize the retrieving delay experienced by a user with a request for a web object, caching strategies are often applied - contents are replicated at edges of the network which is closer to the user such that the network distance between the user and the object is reduced. In this literature survey, evolution of caching is studied. A recent research paper [15] in the field of large-scale caching for CDN was chosen to be the anchor paper which serves as a guide to the topic. Research studies after and relevant to the anchor paper are also analyzed to better evaluate the statements and results of the anchor paper and more importantly, to obtain an unbiased view of the large scale collaborate caching systems as a whole.

preprint2012arXiv

Momentum-resolved Raman spectroscopy of bound molecules in ultracold Fermi gas

The binding energy of Feshbach molecules from a two component Fermi gas of $^{40}$K atoms has been experimentally measured with the momentum-resolved Raman spectroscopy. Comparing with the radio-frequency spectroscopy, in the present experiment the signal of unpaired (free atoms) and the bound molecules can be directly observed and the binding energy can be simultaneously determined in a single running experiment. The energy-momentum dispersion spectra of the strongly interacting ultracold Fermi gas in BEC side are also measured and reconstructed. The present experimental technology of the momentum-resolved Raman spectroscopy can be easily extended to perform spatially momentum-resolved Raman spectroscopy and to obtain the response spectra of a homogeneous system in the local density approximation.

preprint2012arXiv

Momentum-resolved Raman spectroscopy of non-interacting ultracold Fermi gas

We report the experiment on probing the one-body spectral function in a trapped non-interacting $^{40}$K Fermi gas by means of the momentum-resolved Raman spectroscopy The experimental result is in good agreement with the expected quadratic dispersion in the non-interacting regime. Through the comparison with the radio-frequency spectrum, we found that the Raman spectrum shows some new characteristics.

preprint2012arXiv

Non-Markovian quantum input-output networks

Quantum input-output response analysis is a useful method for modeling the dynamics of complex quantum networks, such as those for communication or quantum control via cascade connections. Non-Markovian effects have not yet been studied in such networks. Here we extend the Markovian input-output network formalism developed in optical systems to non-Markovian cascaded networks which can be used, e.g., to analyze the input-output response of mesoscopic quantum networks. We use this formalism to explore the behavior of superconducting qubit networks, where we examine the effect of finite cavity bandwidths. We also discuss its application to open- and closed-loop control networks, and show how these networks create effective Hamiltonians for the controlled system.

preprint2012arXiv

Quantum Coherent Nonlinear Feedbacks with Applications to Quantum Optics on Chip

In the control of classical mechanical systems, the feedback has been successfully applied to the production of the desired nonlinear dynamics. However, how much this can be done is still an open problem in quantum mechanical systems. This paper proposes a scheme of generating strong nonlinear quantum effects via the recently developed coherent feedback techniques, which can be shown to outperform the measurement-based quantum feedback scheme that can only generate pseudo-nonlinear quantum effects. Such advancement is demonstrated by two application examples in quantum optics on chip. In the first example, we show that the nonlinear Kerr effect can be generated and amplified to be comparable with the linear effect in a transmission line resonator (TLR). In the second example, we show that by tuning the gains of the quantum amplifiers in a TLR coherent feedback network, non-Gaussian "light" (microwave field) can be generated and manipulated via the nonlinear effects which exhibits fully quantum sub-Poisson photoncount statistics and photon antibunching phenomenon. The scheme opens promising applications in demonstrating strong nonlinear quantum optics on chip, which is extremely weak and inflexible in traditional quantum optical devices.

preprint2012arXiv

Radio-frequency spectroscopy of weakly bound molecules in spin-orbit coupled atomic Fermi gases

We investigate theoretically radio-frequency spectroscopy of weakly bound molecules in an ultracold spin-orbit-coupled atomic Fermi gas. We consider two cases with either equal Rashba and Dresselhaus coupling or pure Rashba coupling. The former system has been realized very recently at Shanxi University [Wang et al., arXiv:1204.1887] and MIT [Cheuk et al., arXiv:1205.3483]. We predict realistic radio-frequency signals for revealing the unique properties of anisotropic molecules formed by spin-orbit coupling.

preprint2012arXiv

Species Diversity in Rock-Paper-Scissors Game Coupling with Levy Flight

Rock-paper-scissors (RPS) game is a nice model to study the biodiversity in ecosystem. However, the previous studies only consider the nearest- neighbor- interaction among the species. In this paper, taking the long range migration into account, the effects of the interplay between nearest-neighbor-interaction and long-range-interaction of Levy flight obey the power law distance distribution with the exponent h (-0.3<h<-0.1) in spatial RPS game is investigated. Taking the probability of long range Levy flight and the power exponent as parameters, the coexistence conditions of three species are found. The critical curves for stable coexistence of three species in the parameters space are presented. It is also found that long-range-interaction with Levy flight has interesting effects on the final spatiotemporal pattern of the system. The results reveal that the long-range-interaction of Levy flight exhibit pronounced effects on biodiversity of ecosystem.

preprint2012arXiv

Spectral Analysis and Identification of Noises in Quantum Systems

In quantum information processing, knowledge of the noise in the system is crucial for high-precision manipulation and tomography of coherent quantum operations. Existing strategies for identifying this noise require the use of additional quantum devices or control pulses. We present a noise-identification method directly based on the system's non-Markovian response of an ensemble measurement to the noise. The noise spectrum is identified by reversing the response relationship in the frequency domain. For illustration, the method is applied to superconducting charge qubits, but it is equally applicable to any type of qubits. We find that the identification strategy recovers the well-known Fermi's golden rule under the lowest-order perturbation approximation, which corresponds to the Markovian limit when the measurement time is much longer than the noise correlation time. Beyond such approximation, it is possible to further improve the precision at the so-called optimal point by incorporating the transient response data in the non-Markovian regime. This method is verified with experimental data from coherent oscillations in a superconducting charge qubit.

preprint2012arXiv

Spin-Orbit Coupled Degenerate Fermi Gases

Spin-orbit coupling plays an increasingly important role in the modern condensed matter physics. For instance, it gives birth to topological insulators and topological superconductors. Quantum simulation of spin-orbit coupling using ultracold Fermi gases will offer opportunities to study these new phenomena in a more controllable setting. Here we report the first experimental study of a spin-orbit coupled Fermi gas. We observe spin dephasing in spin dynamics and momentum distribution asymmetry in the equilibrium state as hallmarks of spin-orbit coupling. We also observe evidences of Lifshitz transition where the topology of Fermi surfaces change. This serves as an important first step toward finding Majorana fermions in this system.

preprint2012arXiv

Spin-orbit interaction in the k \cdot p theory for cubic crystals

Recent work by Elder, Ward and Zhang [Phys. Rev. B83. 165210 (2011)] has shown need for correction and modification of current implementation of the k.p method and operator ordering scheme using the interaction parameters defined under double group consideration. This manuscript examines the difference in treatment of spin-orbit interaction under the single and double group formulations. We show that the restriction to the adapted double group bases, brought about by the imposition of single group selection rule in calculating the k.p interaction, is not appropriate. In addition, the unitary transformation employed in the literature to diagonalise the intra-band spin orbit interaction in semiconductors with diamond lattice can not remove the inter-band terms. It leads to a bases set for valence band ordered differently from D_3/2^+ in the O(3) group thus invalidating any correlation of magnetic quantum number to the z component of angular momentum. Under the double group consideration, spin-orbit interaction affects all the zone centre states and offers a mechanism for changing the relative positions of various zone centre states in the conduction band. In addition to the mixing caused by k independent spin orbit terms, formation of hybridised orbitals under double group rules also produce mixing. This can lead to inversion in materials such as α-tin. The re-arrangement of conduction band zone centre states leads to a negative γ_2 in most materials with respect to the valence band bases having the same order as D_3/2^+ in the O(3) group. The double group formulation is also required in the description of Zeeman interaction, including the S.B term, for the mixed zone centre states. It provides a correspondence between second order interaction parameters and the Luttinger invariants. ...

preprint2012arXiv

Suppressing nano-scale stick-slip motion by feedback

When a micro cantilever with a nano-scale tip is manipulated on a substrate with atomic-scale roughness, the periodic lateral frictional force and stochastic fluctuations may induce stick-slip motion of the cantilever tip, which greatly decreases the precision of the nano manipulation. This unwanted motion cannot be reduced by open-loop control especially when there exist parameter uncertainties in the system model, and thus needs to introduce feedback control. However, real-time feedback cannot be realized by the existing virtual reality virtual feedback techniques based on the position sensing capacity of the atomic force microscopy (AFM). To solve this problem, we propose a new method to design real-time feedback control based on the force sensing approach to compensate for the disturbances and thus reduce the stick-slip motion of the cantilever tip. Theoretical analysis and numerical simulations show that the controlled motion of the cantilever tip tracks the desired trajectory with much higher precision. Further investigation shows that our proposal is robust under various parameter uncertainties. Our study opens up new perspectives of real-time nano manipulation.

preprint2011arXiv

An improved distributed routing algorithm for Benes based optical NoC

Integrated optical interconnect is believed to be one of the main technologies to replace electrical wires. Optical Network-on-Chip (ONoC) has attracted more attentions nowadays. Benes topology is a good choice for ONoC for its rearrangeable non-blocking character, multistage feature and easy scalability. Routing algorithm plays an important role in determining the performance of ONoC. But traditional routing algorithms for Benes network are not suitable for ONoC communication, we developed a new distributed routing algorithm for Benes ONoC in this paper. Our algorithm selected the routing path dynamically according to network condition and enables more path choices for the message traveling in the network. We used OPNET to evaluate the performance of our routing algorithm and also compared it with a well-known bit-controlled routing algorithm. ETE delay and throughput were showed under different packet length and network sizes. Simulation results show that our routing algorithm can provide better performance for ONoC.

preprint2011arXiv

Block-based Bayesian epistasis association mapping with application to WTCCC type 1 diabetes data

Interactions among multiple genes across the genome may contribute to the risks of many complex human diseases. Whole-genome single nucleotide polymorphisms (SNPs) data collected for many thousands of SNP markers from thousands of individuals under the case--control design promise to shed light on our understanding of such interactions. However, nearby SNPs are highly correlated due to linkage disequilibrium (LD) and the number of possible interactions is too large for exhaustive evaluation. We propose a novel Bayesian method for simultaneously partitioning SNPs into LD-blocks and selecting SNPs within blocks that are associated with the disease, either individually or interactively with other SNPs. When applied to homogeneous population data, the method gives posterior probabilities for LD-block boundaries, which not only result in accurate block partitions of SNPs, but also provide measures of partition uncertainty. When applied to case--control data for association mapping, the method implicitly filters out SNP associations created merely by LD with disease loci within the same blocks. Simulation study showed that this approach is more powerful in detecting multi-locus associations than other methods we tested, including one of ours. When applied to the WTCCC type 1 diabetes data, the method identified many previously known T1D associated genes, including PTPN22, CTLA4, MHC, and IL2RA.

preprint2011arXiv

Bose-Einstein Condensate in a light-induced vector potential using the 1064 $nm$ optical dipole trap lasers

We present a simple experiment of creating an effective vector gauge potential for Bose-Einstein condensed $^{87}$Rb in the F=2 hyperfine ground state using two crossed 1064 $nm$ optical dipole trap lasers as the Raman beams. Due to the far-detuning from the single-photon resonance with the electronically excited state, the spontaneous emission is strongly reduced, at the same time, the moderate strength of the Raman coupling still can be achieved. The atoms at the far detuning of the Raman coupling are loaded adiabatically into the dressed states by ramping the homogeneous bias magnetic field to resonance and the different energy dressed states are studied. This experiment is easily extended to produce synthetic magnetic or electric field from a spatial or time dependence of the effective vector potential.

preprint2011arXiv

Chaos can act as a decoherence suppressor

We propose a strategy to suppress decoherence of a solid-state qubit coupled to non-Markovian noises by attaching the qubit to a chaotic setup with the broad power distribution in particular in the high-frequency domain. Different from the existing decoherence control methods such as the usual dynamics decoupling control, high-frequency components of our control are generated by the chaotic setup driven by a low-frequency field, and the generation of complex optimized control pulses is not necessary. We apply the scheme to superconducting quantum circuits and find that various noises in a wide frequency domain, including low-frequency $1/f$, high-frequency Ohmic, sub-Ohmic, and super-Ohmic noises, can be efficiently suppressed by coupling the qubits to a Duffing oscillator as the chaotic setup. Significantly, the decoherence time of the qubit is prolonged approximately $100$ times in magnitude.

preprint2011arXiv

Continuous-variable multipartite unlockable bound entangled Gaussian states

Continuous-variable (CV) multipartite unlockable bound-entangled states is investigated in this paper. Comparing with the qubit multipartite unlockable bound-entangled states, CV multipartite unlockable bound-entangled states present the new and different properties. CV multipartite unlockable bound-entangled states may serve as a useful quantum resource for new multiparty communication schemes. The experimental protocol for generating CV unlockable bound-entangled states is proposed with a setup that is at present accessible.

preprint2011arXiv

Coupled-resonator-induced transparency with a squeezed vacuum

We present the first experimental observation of quantum fluctuation spectra in two coupled optical cavities with an injected squeezed vacuum light. The quadrature components of the reflected squeezed vacuum spectra are measured by phase sensitive homodyne detector. The experimental results demonstrate coupled-resonator-induced transparency in the quantum regime, in which electromagnetically-induced-transparency-like characteristic of the absorption and dispersion properties of the coupled optical cavities determines the line-shape of the reflected quantum noise spectra.

preprint2011arXiv

Suppressing non-Markovian noises by coupling the qubit to a chaotic device

To suppress decoherence of solid-state qubits which are coupled to the non-Markovian noises, we propose a strategy to couple the qubit with a chaotic device, of which the broad power distribution in the high-frequency domain can be used to freeze the noises just like the dynamical decoupling control (DDC) method. Compared with the DDC, high-frequency components can be generated by the chaotic device even driven by a low-frequency field and we do not need to optimize the control fields to generate complex control pulses. As an application to superconducting circuits, we find that various noises in a wide frequency domain, including low-frequency $1/f$, high-frequency Ohmic, sub-Ohmic, and super-Ohmic noises, can be efficiently suppressed by coupling the qubit to a Duffing oscillator, and the decoherence rate of the qubit is efficiently decreased for about 100 times in magnitude.

preprint2010arXiv

Graphical rule of transforming continuous-variable graph states by local homodyne detection

Graphical rule, describing that any single-mode homodyne detection turns a given continuous-variable (CV) graph state into a new one, is presented. Employing two simple graphical rules: local complement operation and vertex deletion (single quadrature-amplitude $\hat{x}$ measurement), the graphical rule for any single-mode quadrature component measurement can be obtained. The shape of CV weighted graph state may be designed and constructed easily from a given larger graph state by applying this graphical rule.

preprint2010arXiv

Observation of collective atomic recoil motion in a momentum-squeezed, ultra-cold, degenerate fermion gas

We demonstrate clear collective atomic recoil motion in a dilute, momentum-squeezed, ultra-cold degenerate fermion gas by circumventing the effects of Pauli blocking. Although gain from bosonic stimulation is necessarily absent because the quantum gas obeys Fermi-Dirac statistics, collective atomic recoil motion from the underlying wave-mixing process is clearly visible. With a single pump pulse of the proper polarization, we observe two mutually-perpendicular wave-mixing processes occurring simultaneously. Our experiments also indicate that the red-blue pump detuning asymmetry observed with Bose-Einstein condensates does not occur with fermions.

preprint2010arXiv

Transition from weak to strong measurements by nonlinear quantum feedback control

We find that feedback control may induce "pseudo" nonlinear dynamics in a damped harmonic oscillator, whose centroid trajectory in the phase space behaves like a classical nonlinear system. Thus, similar to nonlinear amplifiers (e.g., rf-driven Josephson junctions), feedback control on the harmonic oscillator can induce nonlinear bifurcation, which can be used to amplify small signals and further to measure quantum states of qubits. Using the circuit QED systems as an example, we show how to apply our method to measure superconducting charge qubits.

preprint2010arXiv

Triply-resonant Optical Parametric Oscillator by Four-wave Mixing with Rubidium Vapor inside an Optical Cavity

We present an experimental demonstration of simultaneous above-threshold oscillations of the Stokes and anti-Stokes fields together with the single pumping beam with rubidium atoms inside an optical standing-wave cavity. The triple resonant conditions can be achieved easily by making use of the large dispersions due to two-photon transitions in the three-level atomic system. This work provides a way to achieve high efficient nonlinear frequency conversion and the generated bright Stokes and anti-Stokes cavity output beams are potential resource for applications in quantum information science.

preprint2009arXiv

Multi-normal-mode splitting of a cavity in the presence of atoms -- towards the superstrong coupling regime

Multi-normal-mode splitting peaks are experimentally observed in a system with Doppler-broadened two-level atoms inside a relatively long optical cavity. In this system, the atoms-cavity interaction can reach the ``superstrong coupling" condition with atoms-cavity coupling strength $g\sqrt{N}$ to be near or larger than the cavity free-spectral range $Δ_{FSR}$. In such case, normal-mode splitting can occur in many cavity longitudinal modes to generate the multi-normal-mode splitting peaks, which can be well explained by the linear dispersion enhancement due to the largely increased atomic density in the cavity. Many new interesting phenomena might come out of this superstrong atoms-cavity coupling regime.

preprint2007arXiv

Expansion of an ultra-cold lithium gas in the BEC-BCS crossover

We present an experimental study of the time of flight properties of a gas of ultra-cold fermions in the BEC-BCS crossover. Since interactions can be tuned by changing the value of the magnetic field, we are able to probe both non interacting and strongly interacting behaviors. These measurements allow us to characterize the momentum distribution of the system as well as its equation of state. We also demonstrate the breakdown of superfluid hydrodynamics in the weakly attractive region of the phase diagram, probably caused by pair breaking taking place during the expansion.

preprint2006arXiv

Experimental Preparation of Quadripartite Cluster and GHZ Entangled States for Continuous Variables

The cluster states and Greenberger-Horne-Zeilinger (GHZ) states are two different types of multipartite quantum entangled states. We present the first experimental results generating continuous variable quadripartite cluster and GHZ entangled states of electromagnetic fields. Utilizing four two-mode squeezed states of light and linearly optical transformations, the two types of entangled states for amplitude and phase quadratures of light are experimentally produced. The combinations of the measured quadrature variances prove the full inseparability of the generated four subsystems. The presented experimental schemes show that the multipartite entanglement of continuous variables can be deterministically generated with the relatively simple implementation.

preprint2005arXiv

Cluster States for Continuous-Variable Multipartite Entanglement

We introduce a new class of continuous-variable (CV) multipartite entangled states, the CV cluster states, which might be generated from squeezing and kerr-like interaction. The entanglement properties of these states are studied in terms of classical communication and local operations. The quantum teleportation network with cluster states is investigated. The graph states as the general forms of cluster states are presented, which may be used to generate CV Greenberger-Horne-Zeilinger states by simply local measurements and classical communication. A chain for one-dimensional example of cluster states can be readily experimentally produced only with squeezed light and beamsplitters.

preprint2001arXiv

Quantum Dense Coding Exploiting Bright EPR Beam

Highly efficient quantum dense coding for continuous variables has been experimentally accomplished by means of exploiting bright EPR beam with anticorrelation of amplitude quadratures and correlation of phase quadratures, which is generated from a nondegenerate optical parametric amplifier operating in the state of deamplification. Two bits of classical information are encoded on two quadratures of a half of bright EPR beam at the sender Alice and transmitted to the receiver Bob via one qubit of the shared quantum state after encoding. The amplitude and phase signals are simultaneously decoded with the other half of EPR beam by the direct measurement of the Bell-state at Bob. The signal to noise ratios of the simultaneously measured amplitude and phase signals are improved 5.4dB and 4.8dB with respect to that of the shot noise limit respectively. A high degree of immunity to unauthorized eavesdropping of the presented quantum communication scheme is experimentally demonstrated.