Catalog footprint

What is connected

49works
25topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

49 published item(s)

preprint2026arXiv

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

Vision-language models typically reason over post-ISP RGB images, although RGB rendering can clip, suppress, or quantize sensor evidence before inference. We study whether grounding improves when the visual interface is moved closer to the underlying camera measurement. We formulate measurement-grounded vision-language learning and instantiate it as PRISM-VL, which combines RAW-derived Meas.-XYZ inputs, camera-conditioned grounding, and Exposure-Bracketed Supervision Aggregation for transferring supervision from RGB proxies to measurement-domain observations. Using a quality-controlled 150K instruction-tuning set and a held-out benchmark targeting low-light, HDR, visibility-sensitive, and hallucination-sensitive cases, PRISM-VL-8B reaches 0.6120 BLEU, 0.4571 ROUGE-L, and 82.66\% LLM-Judge accuracy, improving over the RGB Qwen3-VL-8B baseline by +0.1074 BLEU, +0.1071 ROUGE-L, and +4.46 percentage points. These results suggest that part of VLM grounding error arises from information lost during RGB rendering, and that preserving measurement-domain evidence can improve multimodal reasoning.

preprint2023arXiv

A Penalized Functional Linear Cox Regression Model for Spatially-defined Environmental Exposure with an Estimated Buffer Distance

In environmental health research, it is of interest to understand the effect of the neighborhood environment on health. Researchers have shown a protective association between green space around a person's residential address and depression outcomes. In measuring exposure to green space, distance buffers are often used. However, buffer distances differ across studies. Typically, the buffer distance is determined by researchers a priori. It is unclear how to identify an appropriate buffer distance for exposure assessment. To address geographic uncertainty problem for exposure assessment, we present a domain selection algorithm based on the penalized functional linear Cox regression model. The theoretical properties of our proposed method are studied and simulation studies are conducted to evaluate finite sample performances of our method. The proposed method is illustrated in a study of associations of green space exposure with depression and/or antidepressant use in the Nurses' Health Study.

preprint2022arXiv

Design Strategies and Approximation Methods for High-Performance Computing Variability Management

Performance variability management is an active research area in high-performance computing (HPC). We focus on input/output (I/O) variability. To study the performance variability, computer scientists often use grid-based designs (GBDs) to collect I/O variability data, and use mathematical approximation methods to build a prediction model. Mathematical approximation models could be biased particularly if extrapolations are needed. Space-filling designs (SFDs) and surrogate models such as Gaussian process (GP) are popular for data collection and building predictive models. The applicability of SFDs and surrogates in the HPC variability needs investigation. We investigate their applicability in the HPC setting in terms of design efficiency, prediction accuracy, and scalability. We first customize the existing SFDs so that they can be applied in the HPC setting. We conduct a comprehensive investigation of design strategies and the prediction ability of approximation methods. We use both synthetic data simulated from three test functions and the real data from the HPC setting. We then compare different methods in terms of design efficiency, prediction accuracy, and scalability. In synthetic and real data analysis, GP with SFDs outperforms in most scenarios. With respect to approximation models, GP is recommended if the data are collected by SFDs. If data are collected using GBDs, both GP and Delaunay can be considered. With the best choice of approximation method, the performance of SFDs and GBD depends on the property of the underlying surface. For the cases in which SFDs perform better, the number of design points needed for SFDs is about half of or less than that of the GBD to achieve the same prediction accuracy. SFDs that can be tailored to high dimension and non-smooth surface are recommended especially when large numbers of input factors need to be considered in the model.

preprint2022arXiv

Global strong solutions of 3D Compressible Navier-Stokes equations with short pulse type initial data

Short pulse initial datum is referred to the one supported in the ball of radius $δ$ and with amplitude $δ^{\frac12}$ which looks like a pulse. It was first introduced by Christodoulou to prove the formation of black holes for Einstein equations and also to catch the shock formation for compressible Euler equations. The aim of this article is to consider the same type initial data, which allow the density of the fluid to have large amplitude $δ^{-\fracαγ}$ with $δ\in(0,1],$ for the compressible Navier-Stokes equations. We prove the global well-posedness and show that the initial bump region of the density with large amplitude will disappear within a very short time. As a consequence, we obtain the global dynamic behavior of the solutions and the boundedness of $\|\nabla u\|_{L^1([0,\infty);L^\infty)}$. The key ingredients of the proof lie in the new observations for the effective viscous flux and new decay estimates for the density via the Lagrangian coordinate.

preprint2022arXiv

Meta Spatio-Temporal Debiasing for Video Scene Graph Generation

Video scene graph generation (VidSGG) aims to parse the video content into scene graphs, which involves modeling the spatio-temporal contextual information in the video. However, due to the long-tailed training data in datasets, the generalization performance of existing VidSGG models can be affected by the spatio-temporal conditional bias problem. In this work, from the perspective of meta-learning, we propose a novel Meta Video Scene Graph Generation (MVSGG) framework to address such a bias problem. Specifically, to handle various types of spatio-temporal conditional biases, our framework first constructs a support set and a group of query sets from the training data, where the data distribution of each query set is different from that of the support set w.r.t. a type of conditional bias. Then, by performing a novel meta training and testing process to optimize the model to obtain good testing performance on these query sets after training on the support set, our framework can effectively guide the model to learn to well generalize against biases. Extensive experiments demonstrate the efficacy of our proposed framework.

preprint2022arXiv

Multiplicity and orbital stability of normalized solutions to non-autonomous Schrödinger equation with mixed nonlinearities

This paper studies the multiplicity of normalized solutions to the Schrödinger equation with mixed nonlinearities \begin{equation*} \begin{cases} -Δu=λu+h(εx)|u|^{q-2}u+η|u|^{p-2}u,\quad x\in \mathbb{R}^N, \\ \int_{\mathbb{R}^N}|u|^2dx=a^2, \end{cases} \end{equation*} where $a, ε, η>0$, $q$ is $L^2$-subcritical, $p$ is $L^2$-supercritical, $λ\in \mathbb{R}$ is an unknown parameter that appears as a Lagrange multiplier, $h$ is a positive and continuous function. It is proved that the numbers of normalized solutions are at least the numbers of global maximum points of $h$ when $ε$ is small enough. Moreover, the orbital stability of the solutions obtained is analyzed as well. In particular, our results cover the Sobolev critical case $p=2N/(N-2)$.

preprint2022arXiv

Near-MDS Codes from Maximal Arcs in PG$(2,q)$

The singleton defect of an $[n,k,d]$ linear code ${\cal C}$ is defined as $s({\cal C})=n-k+1-d$. Codes with $S({\cal C})=0$ are called maximum distance separable (MDS) codes, and codes with $S(\cal C)=S(\cal C ^{\bot})=1$ are called near maximum distance separable (NMDS) codes. Both MDS codes and NMDS codes have good representations in finite projective geometry. MDS codes over $F_q$ with length $n$ and $n$-arcs in PG$(k-1,q)$ are equivalent objects. When $k=3$, NMDS codes of length $n$ are equivalent to $(n,3)$-arcs in PG$(2,q)$. In this paper, we deal with the NMDS codes with dimension 3. By adding some suitable projective points in maximal arcs of PG$(2,q)$, we can obtain two classes of $(q+5,3)$-arcs (or equivalently $[q+5,3,q+2]$ NMDS codes) for any prime power $q$. We also determine the exact weight distribution and the locality of such NMDS codes and their duals. It turns out that the resultant NMDS codes and their duals are both distance-optimal and dimension-optimal locally recoverable codes.

preprint2022arXiv

NTIRE 2022 Challenge on High Dynamic Range Imaging: Methods and Results

This paper reviews the challenge on constrained high dynamic range (HDR) imaging that was part of the New Trends in Image Restoration and Enhancement (NTIRE) workshop, held in conjunction with CVPR 2022. This manuscript focuses on the competition set-up, datasets, the proposed methods and their results. The challenge aims at estimating an HDR image from multiple respective low dynamic range (LDR) observations, which might suffer from under- or over-exposed regions and different sources of noise. The challenge is composed of two tracks with an emphasis on fidelity and complexity constraints: In Track 1, participants are asked to optimize objective fidelity scores while imposing a low-complexity constraint (i.e. solutions can not exceed a given number of operations). In Track 2, participants are asked to minimize the complexity of their solutions while imposing a constraint on fidelity scores (i.e. solutions are required to obtain a higher fidelity score than the prescribed baseline). Both tracks use the same data and metrics: Fidelity is measured by means of PSNR with respect to a ground-truth HDR image (computed both directly and with a canonical tonemapping operation), while complexity metrics include the number of Multiply-Accumulate (MAC) operations and runtime (in seconds).

preprint2022arXiv

Prediction for Distributional Outcomes in High-Performance Computing I/O Variability

Although high-performance computing (HPC) systems have been scaled to meet the exponentially-growing demand for scientific computing, HPC performance variability remains a major challenge and has become a critical research topic in computer science. Statistically, performance variability can be characterized by a distribution. Predicting performance variability is a critical step in HPC performance variability management and is nontrivial because one needs to predict a distribution function based on system factors. In this paper, we propose a new framework to predict performance distributions. The proposed model is a modified Gaussian process that can predict the distribution function of the input/output (I/O) throughput under a specific HPC system configuration. We also impose a monotonic constraint so that the predicted function is nondecreasing, which is a property of the cumulative distribution function. Additionally, the proposed model can incorporate both quantitative and qualitative input variables. We evaluate the performance of the proposed method by using the IOzone variability data based on various prediction tasks. Results show that the proposed method can generate accurate predictions, and outperform existing methods. We also show how the predicted functional output can be used to generate predictions for a scalar summary of the performance distribution, such as the mean, standard deviation, and quantiles. Our methods can be further used as a surrogate model for HPC system variability monitoring and optimization.

preprint2022arXiv

Renyi Entropy Rate of Stationary Ergodic Processes

In this paper, we examine the Renyi entropy rate of stationary ergodic processes. For a special class of stationary ergodic processes, we prove that the Renyi entropy rate always exists and can be polynomially approximated by its defining sequence; moreover, using the Markov approximation method, we show that the Renyi entropy rate can be exponentially approximated by that of the Markov approximating sequence, as the Markov order goes to infinity. For the general case, by constructing a counterexample, we disprove the conjecture that the Renyi entropy rate of a general stationary ergodic process always converges to its Shannon entropy rate as α goes to 1.

preprint2022arXiv

SDRTV-to-HDRTV via Hierarchical Dynamic Context Feature Mapping

In this work, we address the task of SDR videos to HDR videos(SDRTV-to-HDRTV). Previous approaches use global feature modulation for SDRTV-to-HDRTV. Feature modulation scales and shifts the features in the original feature space, which has limited mapping capability. In addition, the global image mapping cannot restore detail in HDR frames due to the luminance differences in different regions of SDR frames. To resolve the appeal, we propose a two-stage solution. The first stage is a hierarchical Dynamic Context feature mapping (HDCFM) model. HDCFM learns the SDR frame to HDR frame mapping function via hierarchical feature modulation (HME and HM ) module and a dynamic context feature transformation (DCT) module. The HME estimates the feature modulation vector, HM is capable of hierarchical feature modulation, consisting of global feature modulation in series with local feature modulation, and is capable of adaptive mapping of local image features. The DCT module constructs a feature transformation module in conjunction with the context, which is capable of adaptively generating a feature transformation matrix for feature mapping. Compared with simple feature scaling and shifting, the DCT module can map features into a new feature space and thus has a more excellent feature mapping capability. In the second stage, we introduce a patch discriminator-based context generation model PDCG to obtain subjective quality enhancement of over-exposed regions. PDCG can solve the problem that the model is challenging to train due to the proportion of overexposed regions of the image. The proposed method can achieve state-of-the-art objective and subjective quality results. Specifically, HDCFM achieves a PSNR gain of 0.81 dB at a parameter of about 100K. The number of parameters is 1/14th of the previous state-of-the-art methods. The test code will be released soon.

preprint2021arXiv

LCD Codes from tridiagonal Toeplitz matrice

Double Toeplitz (DT) codes are codes with a generator matrix of the form $(I,T)$ with $T$ a Toeplitz matrix, that is to say constant on the diagonals parallel to the main. When $T$ is tridiagonal and symmetric we determine its spectrum explicitly by using Dickson polynomials, and deduce from there conditions for the code to be LCD. Using a special concatenation process, we construct optimal or quasi-optimal examples of binary and ternary LCD codes from DT codes over extension fields.

preprint2021arXiv

Modelling Universal Order Book Dynamics in Bitcoin Market

Understanding the emergence of universal features such as the stylized facts in markets is a long-standing challenge that has drawn much attention from economists and physicists. Most existing models, such as stochastic volatility models, focus mainly on price changes, neglecting the complex trading dynamics. Recently, there are increasing studies on order books, thanks to the availability of large-scale trading datasets, aiming to understand the underlying mechanisms governing the market dynamics. In this paper, we collect order-book datasets of Bitcoin platforms across three countries over millions of users and billions of daily turnovers. We find a 1+1D field theory, govern by a set of KPZ-like stochastic equations, predicts precisely the order book dynamics observed in empirical data. Despite the microscopic difference of markets, we argue the proposed effective field theory captures the correct universality class of market dynamics. We also show that the model agrees with the existing stochastic volatility models at the long-wavelength limit.

preprint2021arXiv

On isodual double Toeplitz codes

Double Toeplitz (shortly DT) codes are introduced here as a generalization of double circulant codes. We show that such a code is isodual, hence formally self-dual. Self-dual DT codes are characterized as double circulant or double negacirculant. Likewise, even DT binary codes are characterized as double circulants. Numerical examples obtained by exhaustive search show that the codes constructed have best-known minimum distance, up to one unit, amongst formally self-dual codes, and sometimes improve on the known values. Over $\F_4$ an explicit construction of DT codes, based on quadratic residues in a prime field, performs equally well. We show that DT codes are asymptotically good over $\F_q$. Specifically, we construct DT codes arbitrarily close to the asymptotic varshamov-Gilbert bound for codes of rate one half.

preprint2021arXiv

Sequential Design of Computer Experiments with Quantitative and Qualitative Factors in Applications to HPC Performance Optimization

Computer experiments with both qualitative and quantitative factors are widely used in many applications. Motivated by the emerging need of optimal configuration in the high-performance computing (HPC) system, this work proposes a sequential design, denoted as adaptive composite exploitation and exploration (CEE), for optimization of computer experiments with qualitative and quantitative factors. The proposed adaptive CEE method combines the predictive mean and standard deviation based on the additive Gaussian process to achieve a meaningful balance between exploitation and exploration for optimization. Moreover, the adaptiveness of the proposed sequential procedure allows the selection of next design point from the adaptive design region. Theoretical justification of the adaptive design region is provided. The performance of the proposed method is evaluated by several numerical examples in simulations. The case study of HPC performance optimization further elaborates the merits of the proposed method.

preprint2021arXiv

SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering

Medical visual question answering (Med-VQA) has tremendous potential in healthcare. However, the development of this technology is hindered by the lacking of publicly-available and high-quality labeled datasets for training and evaluation. In this paper, we present a large bilingual dataset, SLAKE, with comprehensive semantic labels annotated by experienced physicians and a new structural medical knowledge base for Med-VQA. Besides, SLAKE includes richer modalities and covers more human body parts than the currently available dataset. We show that SLAKE can be used to facilitate the development and evaluation of Med-VQA systems. The dataset can be downloaded from http://www.med-vqa.com/slake.

preprint2021arXiv

The number of the non-full-rank Steiner triple systems

The $p$-rank of a Steiner triple system $B$ is the dimension of the linear span of the set of characteristic vectors of blocks of $B$, over GF$(p)$. We derive a formula for the number of different Steiner triple systems of order $v$ and given $2$-rank $r_2$, $r_2<v$, and a formula for the number of Steiner triple systems of order $v$ and given $3$-rank $r_3$, $r_3<v-1$. Also, we prove that there are no Steiner triple systems of $2$-rank smaller than $v$ and, at the same time, $3$-rank smaller than $v-1$. Our results extend previous work on enumerating Steiner triple systems according to the rank of their codes, mainly by Tonchev, V.A.Zinoviev and D.V.Zinoviev for the binary case and by Jungnickel and Tonchev for the ternary case.

preprint2021arXiv

Tight upper bound on the quantum value of Svetlichny operators under local filtering and hidden genuine nonlocality

Nonlocal quantum correlations among the quantum subsystems play essential roles in quantum science. The violation of the Svetlichny inequality provides sufficient conditions of genuine tripartite nonlocality. We provide tight upper bounds on the maximal quantum value of the Svetlichny operators under local filtering operations, and present a qualitative analytical analysis on the hidden genuine nonlocality for three-qubit systems. We investigate in detail two classes of three-qubit states whose hidden genuine nonlocalities can be revealed by local filtering.

preprint2020arXiv

Construction of isodual codes from polycirculant matrices

Double polycirculant codes are introduced here as a generalization of double circulant codes. When the matrix of the polyshift is a companion matrix of a trinomial, we show that such a code is isodual, hence formally self-dual. Numerical examples show that the codes constructed have optimal or quasi-optimal parameters amongst formally self-dual codes. Self-duality, the trivial case of isoduality, can only occur over $ \F_2$ in the double circulant case. Building on an explicit infinite sequence of irreducible trinomials over $\F_2,$ we show that binary double polycirculant codes are asymptotically good.

preprint2020arXiv

Long time existence for a two-dimensional strongly dispersive Boussinesq system

We prove a long time existence result for the solutions of a two-dimensional Boussinesq system modeling the propagation of long, weakly nonlinear water waves. This system is exceptional in the sense that it is the only linearly well-posed system in the (abcd) family of Boussinesq systems whose eigenvalues of the linearized system have nontrivial zeroes. This new difficulty is solved by the use of "good unknowns " and of normal form techniques.

preprint2019arXiv

On the number of resolvable Steiner triple systems of small 3-rank

In a recent work, Jungnickel, Magliveras, Tonchev, and Wassermann derived an overexponential lower bound on the number of nonisomorphic resolvable Steiner triple systems (STS) of order $v$, where $v=3^k$, and $3$-rank $v-k$. We develop an approach to generalize this bound and estimate the number of isomorphism classes of STS$(v)$ of rank $v-k-1$ for an arbitrary $v$ of form $3^kT$.

preprint2016arXiv

Adaptive Epidemic Dynamics in Networks: Thresholds and Control

Theoretical modeling of computer virus/worm epidemic dynamics is an important problem that has attracted many studies. However, most existing models are adapted from biological epidemic ones. Although biological epidemic models can certainly be adapted to capture some computer virus spreading scenarios (especially when the so-called homogeneity assumption holds), the problem of computer virus spreading is not well understood because it has many important perspectives that are not necessarily accommodated in the biological epidemic models. In this paper we initiate the study of such a perspective, namely that of adaptive defense against epidemic spreading in arbitrary networks. More specifically, we investigate a non-homogeneous Susceptible-Infectious-Susceptible (SIS) model where the model parameters may vary with respect to time. In particular, we focus on two scenarios we call semi-adaptive defense and fully-adaptive} defense, which accommodate implicit and explicit dependency relationships between the model parameters, respectively. In the semi-adaptive defense scenario, the model's input parameters are given; the defense is semi-adaptive because the adjustment is implicitly dependent upon the outcome of virus spreading. For this scenario, we present a set of sufficient conditions (some are more general or succinct than others) under which the virus spreading will die out; such sufficient conditions are also known as epidemic thresholds in the literature. In the fully-adaptive defense scenario, some input parameters are not known (i.e., the aforementioned sufficient conditions are not applicable) but the defender can observe the outcome of virus spreading. For this scenario, we present adaptive control strategies under which the virus spreading will die out or will be contained to a desired level.

preprint2016arXiv

Local- and Holistic- Structure Preserving Image Super Resolution via Deep Joint Component Learning

Recently, machine learning based single image super resolution (SR) approaches focus on jointly learning representations for high-resolution (HR) and low-resolution (LR) image patch pairs to improve the quality of the super-resolved images. However, due to treat all image pixels equally without considering the salient structures, these approaches usually fail to produce visual pleasant images with sharp edges and fine details. To address this issue, in this work we present a new novel SR approach, which replaces the main building blocks of the classical interpolation pipeline by a flexible, content-adaptive deep neural networks. In particular, two well-designed structure-aware components, respectively capturing local- and holistic- image contents, are naturally incorporated into the fully-convolutional representation learning to enhance the image sharpness and naturalness. Extensively evaluations on several standard benchmarks (e.g., Set5, Set14 and BSD200) demonstrate that our approach can achieve superior results, especially on the image with salient structures, over many existing state-of-the-art SR methods under both quantitative and qualitative measures.

preprint2016arXiv

Look, Listen and Learn - A Multimodal LSTM for Speaker Identification

Speaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory signals, the robustness to handle severe quality degradations and unconstrained content variations are also indispensable. In this paper, we describe a novel multimodal Long Short-Term Memory (LSTM) architecture which seamlessly unifies both visual and auditory modalities from the beginning of each sequence input. The key idea is to extend the conventional LSTM by not only sharing weights across time steps, but also sharing weights across modalities. We show that modeling the temporal dependency across face and voice can significantly improve the robustness to content quality degradations and variations. We also found that our multimodal LSTM is robustness to distractors, namely the non-speaking identities. We applied our multimodal LSTM to The Big Bang Theory dataset and showed that our system outperforms the state-of-the-art systems in speaker identification with lower false alarm rate and higher recognition accuracy.

preprint2016arXiv

On global dynamics of three dimensional magnetohydrodynamics: nonlinear stability of Alfvén waves

We construct and study global solutions for the 3-dimensional incompressible MHD systems with arbitrary small viscosity. In particular, we provide a rigorous justification for the following dynamical phenomenon observed in many contexts: the solution initially behaves like non-dispersive waves and the shape of the solution persists for a very long time (proportional to the Reynolds number), thereafter, the solution will be damped due to the long-time accumulation of the diffusive effects, eventually, the total energy of the system becomes extremely small compared to the viscosity so that the diffusion takes over and the solution afterwards decays fast in time. We do not assume any condition on the symmetry or on the vorticity. The size of data and the a priori estimates do not depend on viscosity. The proof is builded upon a novel use of the basic energy identity and a geometric study of the characteristic hypersurfaces. The approach is partly inspired by Christodoulou-Klainerman's proof of the nonlinear stability of Minkowski space in general relativity.

preprint2016arXiv

Push- and Pull-based Epidemic Spreading in Networks: Thresholds and Deeper Insights

Understanding the dynamics of computer virus (malware, worm) in cyberspace is an important problem that has attracted a fair amount of attention. Early investigations for this purpose adapted biological epidemic models, and thus inherited the so-called homogeneity assumption that each node is equally connected to others. Later studies relaxed this often-unrealistic homogeneity assumption, but still focused on certain power-law networks. Recently, researchers investigated epidemic models in {\em arbitrary} networks (i.e., no restrictions on network topology). However, all these models only capture {\em push-based} infection, namely that an infectious node always actively attempts to infect its neighboring nodes. Very recently, the concept of {\em pull-based} infection was introduced but was not treated rigorously. Along this line of research, the present paper investigates push- and pull-based epidemic spreading dynamics in arbitrary networks, using a Non-linear Dynamical Systems approach. The paper advances the state of the art as follows: (1) It presents a more general and powerful sufficient condition (also known as epidemic threshold in the literature) under which the spreading will become stable. (2) It gives both upper and lower bounds on the global mean infection rate, regardless of the stability of the spreading. (3) It offers insights into, among other things, the estimation of the global mean infection rate through localized monitoring of a small {\em constant} number of nodes, {\em without} knowing the values of the parameters.

preprint2015arXiv

Compressive sensing based Bayesian sparse channel estimation for OFDM communication systems: high performance and low complexity

In orthogonal frequency division modulation (OFDM) communication systems, channel state information (CSI) is required at receiver due to the fact that frequency-selective fading channel leads to disgusting inter-symbol interference (ISI) over data transmission. Broadband channel model is often described by very few dominant channel taps and they can be probed by compressive sensing based sparse channel estimation (SCE) methods, e.g., orthogonal matching pursuit algorithm, which can take the advantage of sparse structure effectively in the channel as for prior information. However, these developed methods are vulnerable to both noise interference and column coherence of training signal matrix. In other words, the primary objective of these conventional methods is to catch the dominant channel taps without a report of posterior channel uncertainty. To improve the estimation performance, we proposed a compressive sensing based Bayesian sparse channel estimation (BSCE) method which can not only exploit the channel sparsity but also mitigate the unexpected channel uncertainty without scarifying any computational complexity. The propose method can reveal potential ambiguity among multiple channel estimators that are ambiguous due to observation noise or correlation interference among columns in the training matrix. Computer simulations show that propose method can improve the estimation performance when comparing with conventional SCE methods.

preprint2015arXiv

Deep Multimodal Speaker Naming

Automatic speaker naming is the problem of localizing as well as identifying each speaking character in a TV/movie/live show video. This is a challenging problem mainly attributes to its multimodal nature, namely face cue alone is insufficient to achieve good performance. Previous multimodal approaches to this problem usually process the data of different modalities individually and merge them using handcrafted heuristics. Such approaches work well for simple scenes, but fail to achieve high performance for speakers with large appearance variations. In this paper, we propose a novel convolutional neural networks (CNN) based learning framework to automatically learn the fusion function of both face and audio cues. We show that without using face tracking, facial landmark localization or subtitle/transcript, our system with robust multimodal feature extraction is able to achieve state-of-the-art speaker naming performance evaluated on two diverse TV series. The dataset and implementation of our algorithm are publicly available online.

preprint2015arXiv

Hierarchical Saliency Detection on Extended CSSD

Complex structures commonly exist in natural images. When an image contains small-scale high-contrast patterns either in the background or foreground, saliency detection could be adversely affected, resulting erroneous and non-uniform saliency assignment. The issue forms a fundamental challenge for prior methods. We tackle it from a scale point of view and propose a multi-layer approach to analyze saliency cues. Different from varying patch sizes or downsizing images, we measure region-based scales. The final saliency values are inferred optimally combining all the saliency cues in different scales using hierarchical inference. Through our inference model, single-scale information is selected to obtain a saliency map. Our method improves detection quality on many images that cannot be handled well traditionally. We also construct an extended Complex Scene Saliency Dataset (ECSSD) to include complex but general natural images.

preprint2015arXiv

Improved adaptive sparse channel estimation using mixed square/fourth error criterion

Sparse channel estimation problem is one of challenge technical issues in stable broadband wireless communications. Based on square error criterion (SEC), adaptive sparse channel estimation (ASCE) methods, e.g., zero-attracting least mean square error (ZA-LMS) algorithm and reweighted ZA-LMS (RZA-LMS) algorithm, have been proposed to mitigate noise interferences as well as to exploit the inherent channel sparsity. However, the conventional SEC-ASCE methods are vulnerable to 1) random scaling of input training signal; and 2) imbalance between convergence speed and steady state mean square error (MSE) performance due to fixed step-size of gradient descend method. In this paper, a mixed square/fourth error criterion (SFEC) based improved ASCE methods are proposed to avoid aforementioned shortcomings. Specifically, the improved SFEC-ASCE methods are realized with zero-attracting least mean square/fourth error (ZA-LMS/F) algorithm and reweighted ZA-LMS/F (RZA-LMS/F) algorithm, respectively. Firstly, regularization parameters of the SFEC-ASCE methods are selected by means of Monte-Carlo simulations. Secondly, lower bounds of the SFEC-ASCE methods are derived and analyzed. Finally, simulation results are given to show that the proposed SFEC-ASCE methods achieve better estimation performance than the conventional SEC-ASCE methods. 1

preprint2015arXiv

Improved Adaptive Sparse Channel Estimation Using Re-Weighted L1-norm Normalized Least Mean Fourth Algorithm

In next-generation wireless communications systems, accurate sparse channel estimation (SCE) is required for coherent detection. This paper studies SCE in terms of adaptive filtering theory, which is often termed as adaptive channel estimation (ACE). Theoretically, estimation accuracy could be improved by either exploiting sparsity or adopting suitable error criterion. It motivates us to develop effective adaptive sparse channel estimation (ASCE) methods to improve estimation performance. In our previous research, two ASCE methods have been proposed by combining forth-order error criterion based normalized least mean fourth (NLMF) and L1-norm penalized functions, i.e., zero-attracting NLMF (ZA-NLMF) algorithm and reweighted ZA-NLMF (RZA-NLMF) algorithm. Motivated by compressive sensing theory, an improved ASCE method is proposed by using reweighted L1-norm NLMF (RL1-NLMF) algorithm where RL1 can exploit more sparsity information than ZA and RZA. Specifically, we construct the cost function of RL1-NLMF and hereafter derive its update equation. In addition, intuitive figure is also given to verify that RL1 is more efficient than conventional two sparsity constraints. Finally, simulation results are provided to confirm this study.

preprint2015arXiv

Iterative-Promoting Variable Step Size Least Mean Square Algorithm for Accelerating Adaptive Channel Estimation

Invariable step size based least-mean-square error (ISS-LMS) was considered as a very simple adaptive filtering algorithm and hence it has been widely utilized in many applications, such as adaptive channel estimation. It is well known that the convergence speed of ISS-LMS is fixed by the initial step-size. In the channel estimation scenarios, it is very hard to make tradeoff between convergence speed and estimation performance. In this paper, we propose an iterative-promoting variable step size based least-mean-square error (VSS-LMS) algorithm to control the convergence speed as well as to improve the estimation performance. Simulation results show that the proposed algorithm can achieve better estimation performance than previous ISS-LMS while without sacrificing convergence speed.

preprint2015arXiv

Iterative-Promoting Variable Step-size Least Mean Square Algorithm For Adaptive Sparse Channel Estimation

Least mean square (LMS) type adaptive algorithms have attracted much attention due to their low computational complexity. In the scenarios of sparse channel estimation, zero-attracting LMS (ZA-LMS), reweighted ZA-LMS (RZA-LMS) and reweighted -norm LMS (RL1-LMS) have been proposed to exploit channel sparsity. However, these proposed algorithms may hard to make tradeoff between convergence speed and estimation performance with only one step-size. To solve this problem, we propose three sparse iterative-promoting variable step-size LMS (IP-VSS-LMS) algorithms with sparse constraints, i.e. ZA, RZA and RL1. These proposed algorithms are termed as ZA-IPVSS-LMS, RZA-IPVSS-LMS and RL1-IPVSS-LMS respectively. Simulation results are provided to confirm effectiveness of the proposed sparse channel estimation algorithms.

preprint2015arXiv

Maximum correntropy criterion based sparse adaptive filtering algorithms for robust channel estimation under non-Gaussian environments

Sparse adaptive channel estimation problem is one of the most important topics in broadband wireless communications systems due to its simplicity and robustness. So far many sparsity-aware channel estimation algorithms have been developed based on the well-known minimum mean square error (MMSE) criterion, such as the zero-attracting least mean square (ZALMS), which are robust under Gaussian assumption. In non-Gaussian environments, however, these methods are often no longer robust especially when systems are disturbed by random impulsive noises. To address this problem, we propose in this work a robust sparse adaptive filtering algorithm using correntropy induced metric (CIM) penalized maximum correntropy criterion (MCC) rather than conventional MMSE criterion for robust channel estimation. Specifically, MCC is utilized to mitigate the impulsive noise while CIM is adopted to exploit the channel sparsity efficiently. Both theoretical analysis and computer simulations are provided to corroborate the proposed methods.

preprint2015arXiv

On Vectorization of Deep Convolutional Neural Networks for Vision Tasks

We recently have witnessed many ground-breaking results in machine learning and computer vision, generated by using deep convolutional neural networks (CNN). While the success mainly stems from the large volume of training data and the deep network architectures, the vector processing hardware (e.g. GPU) undisputedly plays a vital role in modern CNN implementations to support massive computation. Though much attention was paid in the extent literature to understand the algorithmic side of deep CNN, little research was dedicated to the vectorization for scaling up CNNs. In this paper, we studied the vectorization process of key building blocks in deep CNNs, in order to better understand and facilitate parallel implementation. Key steps in training and testing deep CNNs are abstracted as matrix and vector operators, upon which parallelism can be easily achieved. We developed and compared six implementations with various degrees of vectorization with which we illustrated the impact of vectorization on the speed of model training and testing. Besides, a unified CNN framework for both high-level and low-level vision tasks is provided, along with a vectorized Matlab implementation with state-of-the-art speed performance.

preprint2015arXiv

Regularization Parameter Selection Method for Sign LMS with Reweighted L1-Norm Constriant Algorithm

Broadband frequency-selective fading channels usually have the inherent sparse nature. By exploiting the sparsity, adaptive sparse channel estimation (ASCE) algorithms, e.g., least mean square with reweighted L1-norm constraint (LMS-RL1) algorithm, could bring a considerable performance gain under assumption of additive white Gaussian noise (AWGN). In practical scenario of wireless systems, however, channel estimation performance is often deteriorated by unexpected non-Gaussian mixture noises which include AWGN and impulsive noises. To design stable communication systems, sign LMS-RL1 (SLMS-RL1) algorithm is proposed to remove the impulsive noise and to exploit channel sparsity simultaneously. It is well known that regularization parameter (REPA) selection of SLMS-RL1 is a very challenging issue. In the worst case, inappropriate REPA may even result in unexpected instable convergence of SLMS-RL1 algorithm. In this paper, Monte Carlo based selection method is proposed to select suitable REPA so that SLMS-RL1 can achieve two goals: stable convergence as well as usage sparsity information. Simulation results are provided to corroborate our studies.

preprint2015arXiv

ROSA: Robust sparse adaptive channel estimation in the presence of impulsive noises

Based on the assumption of Gaussian noise model, conventional adaptive filtering algorithms for reconstruction sparse channels were proposed to take advantage of channel sparsity due to the fact that broadband wireless channels usually have the sparse nature. However, state-of-the-art algorithms are vulnerable to deteriorate under the assumption of non-Gaussian noise models (e.g., impulsive noise) which often exist in many advanced communications systems. In this paper, we study the problem of RObust Sparse Adaptive channel estimation (ROSA) in the environment of impulsive noises using variable step-size affine projection sign algorithm (VSS-APSA). Specifically, standard VSS-APSA algorithm is briefly reviewed and three sparse VSS-APSA algorithms are proposed to take advantage of channel sparsity with different sparse constraints. To fairly evaluate the performance of these proposed algorithms, alpha-stable noise is considered to approximately model the realistic impulsive noise environments. Simulation results show that the proposed algorithms can achieve better performance than standard VSS-APSA algorithm in different impulsive environments.

preprint2014arXiv

Adaptive MIMO Channel Estimation using Sparse Variable Step-Size NLMS Algorithms

To estimate multiple-input multiple-output (MIMO) channels, invariable step-size normalized least mean square (ISSNLMS) algorithm was applied to adaptive channel estimation (ACE). Since the MIMO channel is often described by sparse channel model due to broadband signal transmission, such sparsity can be exploited by adaptive sparse channel estimation (ASCE) methods using sparse ISS-NLMS algorithms. It is well known that step-size is a critical parameter which controls three aspects: algorithm stability, estimation performance and computational cost. The previous approaches can exploit channel sparsity but their step-sizes are keeping invariant which unable balances well the three aspects and easily cause either estimation performance loss or instability. In this paper, we propose two stable sparse variable step-size NLMS (VSS-NLMS) algorithms to improve the accuracy of MIMO channel estimators. First, ASCE for estimating MIMO channels is formulated in MIMO systems. Second, different sparse penalties are introduced to VSS-NLMS algorithm for ASCE. In addition, difference between sparse ISSNLMS algorithms and sparse VSS-NLMS ones are explained. At last, to verify the effectiveness of the proposed algorithms for ASCE, several selected simulation results are shown to prove that the proposed sparse VSS-NLMS algorithms can achieve better estimation performance than the conventional methods via mean square error (MSE) and bit error rate (BER) metrics.

preprint2014arXiv

Affine Combination of Two Adaptive Sparse Filters for Estimating Large Scale MIMO Channels

Large scale multiple-input multiple-output (MIMO) system is considered one of promising technologies for realizing next-generation wireless communication system (5G) to increasing the degrees of freedom in space and enhancing the link reliability while considerably reducing the transmit power. However, large scale MIMO system design also poses a big challenge to traditional one-dimensional channel estimation techniques due to high complexity and curse of dimensionality problems which are caused by long delay spread as well as large number antenna. Since large scale MIMO channels often exhibit sparse or/and cluster-sparse structure, in this paper, we propose a simple affine combination of adaptive sparse channel estimation method for reducing complexity and exploiting channel sparsity in the large scale MIMO system. First, problem formulation and standard affine combination of adaptive least mean square (LMS) algorithm are introduced. Then we proposed an effective affine combination method with two sparse LMS filters and designed an approximate optimum affine combiner according to stochastic gradient search method as well. Later, to validate the proposed algorithm for estimating large scale MIMO channel, computer simulations are provided to confirm effectiveness of the proposed algorithm which can achieve better estimation performance than the conventional one as well as traditional method.

preprint2014arXiv

An Evasion and Counter-Evasion Study in Malicious Websites Detection

Malicious websites are a major cyber attack vector, and effective detection of them is an important cyber defense task. The main defense paradigm in this regard is that the defender uses some kind of machine learning algorithms to train a detection model, which is then used to classify websites in question. Unlike other settings, the following issue is inherent to the problem of malicious websites detection: the attacker essentially has access to the same data that the defender uses to train its detection models. This 'symmetry' can be exploited by the attacker, at least in principle, to evade the defender's detection models. In this paper, we present a framework for characterizing the evasion and counter-evasion interactions between the attacker and the defender, where the attacker attempts to evade the defender's detection models by taking advantage of this symmetry. Within this framework, we show that an adaptive attacker can make malicious websites evade powerful detection models, but proactive training can be an effective counter-evasion defense mechanism. The framework is geared toward the popular detection model of decision tree, but can be adapted to accommodate other classifiers.

preprint2014arXiv

Block Bayesian Sparse Learning Algorithms With Application to Estimating Channels in OFDM Systems

Cluster-sparse channels often exist in frequencyselective fading broadband communication systems. The main reason is received scattered waveform exhibits cluster structure which is caused by a few reflectors near the receiver. Conventional sparse channel estimation methods have been proposed for general sparse channel model which without considering the potential cluster-sparse structure information. In this paper, we investigate the cluster-sparse channel estimation (CS-CE) problems in the state of the art orthogonal frequencydivision multiplexing (OFDM) systems. Novel Bayesian clustersparse channel estimation (BCS-CE) methods are proposed to exploit the cluster-sparse structure by using block sparse Bayesian learning (BSBL) algorithm. The proposed methods take advantage of the cluster correlation in training matrix so that they can improve estimation performance. In addition, different from our previous method using uniform block partition information, the proposed methods can work well when the prior block partition information of channels is unknown. Computer simulations show that the proposed method has a superior performance when compared with the previous methods.

preprint2014arXiv

Extra Gain:Improved Sparse Channel Estimation Using Reweighted l_1-norm Penalized LMS/F Algorithm

The channel estimation is one of important techniques to ensure reliable broadband signal transmission. Broadband channels are often modeled as a sparse channel. Comparing with traditional dense-assumption based linear channel estimation methods, e.g., least mean square/fourth (LMS/F) algorithm, exploiting sparse structure information can get extra performance gain. By introducing l_1-norm penalty, two sparse LMS/F algorithms, (zero-attracting LMSF, ZA-LMS/F and reweighted ZA-LMSF, RZA-LMSF), have been proposed [1]. Motivated by existing reweighted l_1-norm (RL1) sparse algorithm in compressive sensing [2], we propose an improved channel estimation method using RL1 sparse penalized LMS/F (RL1-LMS/F) algorithm to exploit more efficient sparse structure information. First, updating equation of RL1-LMS/F is derived. Second, we compare their sparse penalize strength via figure example. Finally, computer simulation results are given to validate the superiority of proposed method over than conventional two methods.

preprint2014arXiv

Novel Realization of Adaptive Sparse Sensing with Sparse Least Mean Fourth Algorithm

Nonlinear sparse sensing (NSS) techniques have been adopted for realizing compressive sensing (CS) in many applications such as Radar imaging and sparse channel estimation. Unlike the NSS, in this paper, we propose an adaptive sparse sensing (ASS) approach using reweighted zero-attracting normalized least mean fourth (RZA-NLMF) algorithm which depends on several given parameters, i.e., reweighted factor, regularization parameter and initial step-size. First, based on the independent assumption, Cramer Rao lower bound (CRLB) is derived as for the performance comparisons. In addition, reweighted factor selection method is proposed for achieving robust estimation performance. Finally, to verify the algorithm, Monte Carlo based computer simulations are given to show that the ASS achieves much better mean square error (MSE) performance than the NSS.

preprint2014arXiv

RZA-NLMF algorithm based adaptive sparse sensing for realizing compressive sensing problems

Nonlinear sparse sensing (NSS) techniques have been adopted for realizing compressive sensing in many applications such as Radar imaging. Unlike the NSS, in this paper, we propose an adaptive sparse sensing (ASS) approach using reweighted zero-attracting normalized least mean fourth (RZA-NLMF) algorithm which depends on several given parameters, i.e., reweighted factor, regularization parameter and initial step-size. First, based on the independent assumption, Cramer Rao lower bound (CRLB) is derived as for the trademark of performance comparisons. In addition, reweighted factor selection method is proposed for achieving robust estimation performance. Finally, to verify the algorithm, Monte Carlo based computer simulations are given to show that the ASS achieves much better mean square error (MSE) performance than the NSS.

preprint2013arXiv

A Graph Theoretical Approach to Network Encoding Complexity

Consider an acyclic directed network $G$ with sources $S_1, S_2,..., S_l$ and distinct sinks $R_1, R_2,..., R_l$. For $i=1, 2,..., l$, let $c_i$ denote the min-cut between $S_i$ and $R_i$. Then, by Menger's theorem, there exists a group of $c_i$ edge-disjoint paths from $S_i$ to $R_i$, which will be referred to as a group of Menger's paths from $S_i$ to $R_i$ in this paper. Although within the same group they are edge-disjoint, the Menger's paths from different groups may have to merge with each other. It is known that by choosing Menger's paths appropriately, the number of mergings among different groups of Menger's paths is always bounded by a constant, which is independent of the size and the topology of $G$. The tightest such constant for the all the above-mentioned networks is denoted by $\mathcal{M}(c_1, c_2,..., c_2)$ when all $S_i$'s are distinct, and by $\mathcal{M}^*(c_1, c_2,..., c_2)$ when all $S_i$'s are in fact identical. It turns out that $\mathcal{M}$ and $\mathcal{M}^*$ are closely related to the network encoding complexity for a variety of networks, such as multicast networks, two-way networks and networks with multiple sessions of unicast. Using this connection, we compute in this paper some exact values and bounds in network encoding complexity using a graph theoretical approach.

preprint2013arXiv

Dense Scattering Layer Removal

We propose a new model, together with advanced optimization, to separate a thick scattering media layer from a single natural image. It is able to handle challenging underwater scenes and images taken in fog and sandstorm, both of which are with significantly reduced visibility. Our method addresses the critical issue -- this is, originally unnoticeable impurities will be greatly magnified after removing the scattering media layer -- with transmission-aware optimization. We introduce non-local structure-aware regularization to properly constrain transmission estimation without introducing the halo artifacts. A selective-neighbor criterion is presented to convert the unconventional constrained optimization problem to an unconstrained one where the latter can be efficiently solved.

preprint2013arXiv

Global small solutions to 2-D incompressible MHD system

In this paper, we consider the global wellposedness of 2-D incompressible magneto-hydrodynamical system with small and smooth initial data. It is a coupled system between the Navier-Stokes equations and a free transport equation with an universal nonlinear coupling structure. The main difficulty of the proof lies in exploring the dissipative mechanism of the system due to the fact that there is a free transport equation in the system. To achieve this and to avoid the difficulty of propagating anisotropic regularity for the free transport equation, we first reformulate our system \eqref{1.1} in the Lagrangian coordinates \eqref{a14}. Then we employ anisotropic Littlewood-Paley analysis to establish the key {\it a priori} $L^1(\R^+; Lip(\R^2))$ estimate to the Lagrangian velocity field $Y_t$. With this estimate, we prove the global wellposedness of \eqref{a14} with smooth and small initial data by using the energy method. We emphasize that the algebraic structure of \eqref{a14} is crucial for the proofs to work. The global wellposedness of the original system \eqref{1.1} then follows by a suitable change of variables.

preprint2013arXiv

Global small solutions to three-dimensional incompressible MHD system

In this paper, we consider the global wellposedness of 3-D incompressible magneto-hydrodynamical system with small and smooth initial data. The main difficulty of the proof lies in establishing the global in time $L^1$ estimate for the velocity field due to the strong degeneracy and anisotropic spectral properties of the linearized system. To achieve this and to avoid the difficulty of propagating anisotropic regularity for the transport equation, we first write our system \eqref{B1} in the Lagrangian formulation \eqref{B11}. Then we employ anisotropic Littlewood-Paley analysis to establish the key $L^1$ in time estimates to the velocity and the gradient of the pressure in the Lagrangian coordinate. With those estimates, we prove the global wellposedness of \eqref{B11} with smooth and small initial data by using the energy method. Toward this, we will have to use the algebraic structure of \eqref{B11} in a rather crucial way. The global wellposedness of the original system \eqref{B1} then follows by a suitable change of variables together with a continuous argument. We should point out that compared with the linearized systems of 2-D MHD equations in \cite{XLZMHD1} and that of the 3-D modified MHD equations in \cite{LZ}, our linearized system \eqref{B19} here is much more degenerate, moreover, the formulation of the initial data for \eqref{B11} is more subtle than that in \cite{XLZMHD1}.