Topic overview

Applications

3567 works11054 researchers

Map preview

Start with the graph, then narrow the list

3567works
11054researchers

Next steps

Use the topic as a working map

Open the full map for clusters, then return here to scan ranked papers and people.

Topic graph

See the topic as a live network

Open full explorer

Inspect nearby papers, researchers, institutions and communities without opening a separate graph page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Papers in this area

24 paper(s) to start with

preprint2017arXiv

Assignment of endogenous retrovirus integration sites using a mixture model

Structural variation occurs in the genomes of individuals because of the different positions occupied by repetitive genome elements like endogenous retroviruses, or ERVs. The presence or absence of ERVs can be determined by identifying the junction with the host genome using high-throughput sequence technology and a clustering algorithm. The resulting data give the number of sequence reads assigned to each ERV-host junction sequence for each sampled individual. Variability in the number of reads from an individual integration site makes it difficult to determine whether a site is present for low read counts. We present a novel two-component mixture of negative binomial distributions to model these counts and assign a probability that a given ERV is present in a given individual. We explain how our approach is superior to existing alternatives, including another form of two-component mixture model and the much more common approach of selecting a threshold count for declaring the presence of an ERV. We apply our method to a data set of ERV integrations in mule deer [Odocoileus hemionus], a species for which no genomic resources are available, and demonstrate that the discovered patte

preprint2016arXiv

Information theoretical noninvasive damage detection in bridge structures

Damage detection of mechanical structures such as bridges is an important research problem in civil engineering. Using spatially distributed sensor time series data collected from a recent experiment on a local bridge in upper state New York, we study noninvasive damage detection using information-theoretical methods. Several findings are in order. First, the time series data, which represent accelerations measured at the sensors, more closely follow Laplace distribution than normal distribution, allowing us to develop parameter estimators for various information-theoretic measures such as entropy and mutual information. Secondly, as damage is introduced by the removal of bolts of the first diaphragm connection, the interaction between spatially nearby sensors as measured by mutual information become weaker, suggesting that the bridge is "loosened". Finally, using a proposed oMII procedure to prune away indirect interactions, we found that the primary direction of interaction or influence aligns with the traffic direction on the bridge even after damaging the bridge.

preprint2016arXiv

Bayesian Transformed GARMA Models

Transformed Generalized Autoregressive Moving Average (TGARMA) models were recently proposed to deal with non-additivity, non-normality and heteroscedasticity in real time series data. In this paper, a Bayesian approach is proposed for TGARMA models, thus extending the original model. We conducted a simulation study to investigate the performance of Bayesian estimation and Bayesian model selection criteria. In addition, a real dataset was analysed using the proposed approach.

preprint2016arXiv

Fast Bayesian whole-brain fMRI analysis with spatial 3D priors

Spatial whole-brain Bayesian modeling of task-related functional magnetic resonance imaging (fMRI) is a great computational challenge. Most of the currently proposed methods therefore do inference in subregions of the brain separately or do approximate inference without comparison to the true posterior distribution. A popular such method, which is now the standard method for Bayesian single subject analysis in the SPM software, is introduced in Penny et al. (2005b). The method processes the data slice-by-slice and uses an approximate variational Bayes (VB) estimation algorithm that enforces posterior independence between activity coefficients in different voxels. We introduce a fast and practical Markov chain Monte Carlo (MCMC) scheme for exact inference in the same model, both slice-wise and for the whole brain using a 3D prior on activity coefficients. The algorithm exploits sparsity and uses modern techniques for efficient sampling from high-dimensional Gaussian distributions, leading to speed-ups without which MCMC would not be a practical option. Using MCMC, we are for the first time able to evaluate the approximate VB posterior against the exact MCMC posterior, and show that

preprint2016arXiv

Predicting Pediatric Surgical Durations

Effective management of operating room resources relies on accurate predictions of surgical case durations. This prediction problem is known to be particularly difficult in pediatric hospitals due to the extreme variation in pediatric patient populations. We propose a novel metric for measuring accuracy of predictions which captures key issues relevant to hospital operations. With this metric in mind we propose several tree-based prediction models. Some are automated (they do not require input from surgeons) while others are semi-automated (they do require input from surgeons). We see that many of our automated methods generally outperform currently used algorithms and even achieve the same performance as surgeons. Our semi-automated methods can outperform surgeons by a significant margin. We gain insights into the predictive value of different features and suggest avenues of future work.

preprint2016arXiv

Counterfactual Prediction with Deep Instrumental Variables Networks

We are in the middle of a remarkable rise in the use and capability of artificial intelligence. Much of this growth has been fueled by the success of deep learning architectures: models that map from observables to outputs via multiple layers of latent representations. These deep learning algorithms are effective tools for unstructured prediction, and they can be combined in AI systems to solve complex automated reasoning problems. This paper provides a recipe for combining ML algorithms to solve for causal effects in the presence of instrumental variables -- sources of treatment randomization that are conditionally independent from the response. We show that a flexible IV specification resolves into two prediction tasks that can be solved with deep neural nets: a first-stage network for treatment prediction and a second-stage network whose loss function involves integration over the conditional treatment distribution. This Deep IV framework imposes some specific structure on the stochastic gradient descent routine used for training, but it is general enough that we can take advantage of off-the-shelf ML capabilities and avoid extensive algorithm customization. We outline how to ob

preprint2016arXiv

Data-Driven Forecast of Dengue Outbreaks in Brazil: A Critical Assessment of Climate Conditions for Different Capitals

Local climate conditions play a major role in the development of the mosquito population responsible for transmitting Dengue Fever. Since the {\em Aedes Aegypti} mosquito is also a primary vector for the recent Zika and Chikungunya epidemics across the Americas, a detailed monitoring of periods with favorable climate conditions for mosquito profusion may improve the timing of vector-control efforts and other urgent public health strategies. We apply dimensionality reduction techniques and machine-learning algorithms to climate time series data and analyze their connection to the occurrence of Dengue outbreaks for seven major cities in Brazil. Specifically, we have identified two key variables and a period during the annual cycle that are highly predictive of epidemic outbreaks. The key variables are the frequency of precipitation and temperature during an approximately two month window of the winter season preceding the outbreak. Thus simple climate signatures may be influencing Dengue outbreaks even months before their occurrence. Some of the more challenging datasets required usage of compressive-sensing procedures to estimate missing entries for temperature and precipitation rec

preprint2016arXiv

Two-Qubit Separability Probabilities as Joint Functions of the Bloch Radii of the Qubit Subsystems

We detect a certain pattern of behavior of separability probabilities $p(r_A,r_B)$ for two-qubit systems endowed with Hilbert-Schmidt, and more generally, random induced measures, where $r_A$ and $r_B$ are the Bloch radii ($0 \leq r_A,r_B \leq 1$) of the qubit reduced states ($A,B$). We observe a relative repulsion of radii effect, that is $p(r_A,r_A) < p(r_A,1-r_A)$, except for rather narrow "crossover" intervals $[\tilde{r}_A,\frac{1}{2}]$. Among the seven specific cases we study are, firstly, the "toy" seven-dimensional $X$-states model and, then, the fifteen-dimensional two-qubit states obtained by tracing over the pure states in $4 \times K$-dimensions, for $K=3, 4, 5$, with $K=4$ corresponding to Hilbert-Schmidt (flat/Euclidean) measure. We also examine the real (two-rebit) $K=4$, the $X$-states $K=5$, and Bures (minimal monotone)--for which no nontrivial crossover behavior is observed--instances. In the two $X$-states cases, we derive analytical results, for $K=3, 4$, we propose formulas that well-fit our numerical results, and for the other scenarios, rely presently upon large numerical analyses. The separability probability crossover regions found expand in

preprint2016arXiv

lfda: An R Package for Local Fisher Discriminant Analysis and Visualization

Local Fisher discriminant analysis is a localized variant of Fisher discriminant analysis and it is popular for supervised dimensionality reduction method. lfda is an R package for performing local Fisher discriminant analysis, including its variants such as kernel local Fisher discriminant analysis and semi-supervised local Fisher discriminant analysis. It also provides visualization functions to easily visualize the dimension reduction results by using either rgl for 3D visualization or ggfortify for 2D visualization in ggplot2 style.

preprint2016arXiv

A Latent-class Model for Estimating Product-choice Probabilities from Clickstream Data

This paper analyzes customer product-choice behavior based on the recency and frequency of each customer's page views on e-commerce sites. Recently, we devised an optimization model for estimating product-choice probabilities that satisfy monotonicity, convexity, and concavity constraints with respect to recency and frequency. This shape-restricted model delivered high predictive performance even when there were few training samples. However, typical e-commerce sites deal in many different varieties of products, so the predictive performance of the model can be further improved by integration of such product heterogeneity. For this purpose, we develop a novel latent-class shape-restricted model for estimating product-choice probabilities for each latent class of products. We also give a tailored expectation-maximization algorithm for parameter estimation. Computational results demonstrate that higher predictive performance is achieved with our latent-class model than with the previous shape-restricted model and common latent-class logistic regression.

preprint2016arXiv

Fused Mean-variance Filter for Feature Screening

This paper proposes a novel model-free screening procedure for ultrahigh dimensional data analysis. By utilizing slicing technique which has been successfully ap- plied to continuous variables, we construct a new index called the fused mean-variance for feature screening. This method has the following merits: (i) it is model-free, i.e., without specifying regression form of predictors and response variable; (ii) it can be used to analyze various types of variables including discrete, categorical and continuous vari- ables; (iii) it still works well even when the covariates/random errors are heavy-tailed or the predictors are strongly dependent. Under some regularity conditions, we establish the sure screening and rank consistency. Simulation studies are conducted to assess the performance of the proposed approach. A real data is used to illustrate the proposed method.

preprint2016arXiv

EEG in the classroom: Synchronised neural recordings during video presentation

We performed simultaneous recordings of electroencephalography (EEG) from multiple students in a classroom, and measured the inter-subject correlation (ISC) of activity evoked by a common video stimulus. The neural reliability, as quantified by ISC, has been linked to engagement and attentional modulation in earlier studies that used high-grade equipment in laboratory settings. Here we reproduce many of the results from these studies using portable low-cost equipment, focusing on the robustness of using ISC for subjects experiencing naturalistic stimuli. The present data shows that stimulus-evoked neural responses, known to be modulated by attention, can be tracked in for groups of students with synchronized EEG acquisition. This is a step towards real-time inference of engagement in the classroom.

preprint2016arXiv

Should we opt for the Black Friday discounted price or wait until the Boxing Day?

We derive an optimal strategy for minimizing the expected loss in the two-period economy when a pivotal decision needs to be made during the first time period and cannot be subsequently reversed. Our interest in the problem has been motivated by the classical shopper's dilemma during the Black Friday promotion period, and our solution crucially relies on the pioneering work of McDonnell and Abbott on the two-envelope paradox.

preprint2016arXiv

Modeling Tangential Vector Fields on a Sphere

Physical processes that manifest as tangential vector fields on a sphere are common in geophysical and environmental sciences. These naturally occurring vector fields are often subject to physical constraints, such as being curl-free or divergence-free. We construct a new class of parametric models for cross-covariance functions of curl-free and divergence-free vector fields that are tangential to the unit sphere. These models are constructed by applying the surface gradient or the surface curl operator to scalar random potential fields defined on the unit sphere. We propose a likelihood-based estimation procedure for the model parameters and show that fast computation is possible even for large data sets when the observations are on a regular latitude-longitude grid. Characteristics and utility of the proposed methodology are illustrated through simulation studies and by applying it to an ocean surface wind velocity data set collected through satellite-based scatterometry remote sensing. We also compare the performance of the proposed model with a class of bivariate Matérn models in terms of estimation and prediction, and demonstrate that the proposed model is superior in capturin

preprint2016arXiv

Benchmark Dose Estimation using a Family of Link Functions

This article proposes a method of estimating benchmark dose (BMD) using a family of link functions in binomial response models dealing with model uncertainty problems. Researchers usually estimate the BMD using binomial response models with a single link function. Several forms of link function have been proposed to fit dose response models to estimate the BMD and the corresponding benchmark dose lower bound (BMDL). However, if the assumed link is not correct, then the estimated BMD and BMDL from the fitted model may not be accurate. To account for model uncertainty, model averaging (MA) methods are proposed to estimate BMD averaging over a model space containing a finite number of standard models. Usual model averaging focuses on a pre-specified list of parametric models leading to pitfalls when none of the models in the list is the correct model. Here, an alternative which augments an initial list of parametric models with an infinite number of additional models having varying links has been proposed. In addition, different methods for estimating BMDL based on the family of link functions are derived. The proposed approach is compared with MA in a simulation study and applied to

preprint2016arXiv

Is the familywise error rate in genomics controlled by methods based on the effective number of independent tests?

In genome-wide association (GWA) studies the goal is to detect association between one or more genetic markers and a given phenotype. The number of genetic markers in a GWA study can be in the order hundreds of thousands and therefore multiple testing methods are needed. This paper presents a set of popular methods to be used to correct for multiple testing in GWA studies. All are based on the concept of estimating an effective number of independent tests. We compare these methods using simulated data and data from the TOP study, and show that the effective number of independent tests is not additive over blocks of independent genetic markers unless we assume a common value for the local significance level. We also show that the reviewed methods based on estimating the effective number of independent tests in general do not control the familywise error rate.

preprint2016arXiv

Bayesian Non-Central Chi Regression For Neuroimaging

We propose a regression model for non-central $χ$ (NC-$χ$) distributed functional magnetic resonance imaging (fMRI) and diffusion weighted imaging (DWI) data, with the heteroscedastic Rician regression model as a prominent special case. The model allows both parameters in the NC-$χ$ distribution to be linked to explanatory variables, with the relevant covariates automatically chosen by Bayesian variable selection. A highly efficient Markov chain Monte Carlo (MCMC) algorithm is proposed for simulating from the joint Bayesian posterior distribution of all model parameters and the binary covariate selection indicators. Simulated fMRI data is used to demonstrate that the Rician model is able to localize brain activity much more accurately than the traditionally used Gaussian model at low signal-to-noise ratios. Using a diffusion dataset from the Human Connectome Project, it is also shown that the commonly used approximate Gaussian noise model underestimates the mean diffusivity (MD) and the fractional anisotropy (FA) in the single-diffusion tensor model compared to the theoretically correct Rician model.

preprint2016arXiv

Impact of model choice on LR assessment in case of rare haplotype match (frequentist approach)

The likelihood ratio (LR) measures the relative weight of forensic data regarding two hypotheses. Several levels of uncertainty arise if frequentist methods are chosen for its assessment: the assumed population model only approximates the true one and its parameters are estimated through a database. Moreover, it may be wise to discard part of data, especially that only indirectly related to the hypotheses. Different reductions define different LRs. Therefore, it is more sensible to talk about "a" LR instead of "the" LR, and the error involved in the estimation should be quantified. Two frequentist methods are proposed in the light of these points for the `rare type match problem', that is when a match between the perpetrator's and the suspect's DNA profile, never observed before in the database of reference, is to be evaluated.

preprint2016arXiv

A Point-process Response Model for Spike Trains from Single Neurons in Neural Circuits under Optogenetic Stimulation

Optogenetics is a new tool to study neuronal circuits that have been genetically modified to allow stimulation by flashes of light. We study recordings from single neurons within neural circuits under optogenetic stimulation. The data from these experiments present a statistical challenge of modeling a high frequency point process (neuronal spikes) while the input is another high frequency point process (light flashes). We further develop a generalized linear model approach to model the relationships between two point processes, employing additive point-process response functions. The resulting model, Point-process Responses for Optogenetics (PRO), provides explicit nonlinear transformations to link the input point process with the output one. Such response functions may provide important and interpretable scientific insights into the properties of the biophysical process that governs neural spiking in response to optogenetic stimulation. We validate and compare the PRO model using a real dataset and simulations, and our model yields a superior area-under-the- curve value as high as 93% for predicting every future spike. For our experiment on the recurrent layer V circuit in the pr

preprint2016arXiv

Low-Cost Energy Meter Calibration Method for Measurement and Verification

Energy meters need to be calibrated for use in Measurement and Verification (M&V) projects. However, calibration can be prohibitively expensive and affect project feasibility negatively. This study presents a novel low-cost in-situ meter data calibration technique using a relatively low accuracy commercial energy meter as a calibrator. Calibration is achieved by combining two machine learning tools: the SIMulation EXtrapolation (SIMEX) Measurement Error Model, and Bayesian regression. The model is trained or calibrated on half-hourly building energy data for 24 hours. Measurements are then compared to the true values over the following months to verify the method. Results show that the hybrid method significantly improves parameter estimates and goodness of fit when compared to Ordinary Least Squares regression or standard SIMEX. This study also addresses the effect of mismeasurement in energy monitoring, and implements a powerful technique for mitigating the bias that arises because of it. Meters calibrated by the technique presented have satisfactory accuracy for most M&V applications, at a significantly lower cost.

preprint2016arXiv

A Note on a Sum of Lognormals

This note considers the applicability of Gauss-Hermite quadrature and direct numerical quadrature for computation of moment generating function (mgf) and the derivatives. A preprocessing using the asymptotic technique is employed while computing the characteristic function (chf) using Gauss Hermite quadrature while this is optional for mgf. The mgf of the low and high amplitude regions of a single lognormal variable and the derivatives is examined and attention is drawn to the effect of variance. The problem of inversion of the mgf/chf of a sum of lognormals to obtain the CDF/pdf is considered with special reference to methods related to Post Widder technique, Gaussian quadrature and the Fourier series method. The method based on the complex exponential integral which makes use of the derivative of the cumulant is an alternative. Segmentation of the mgf/chf on the basis of the derivative structure which indicates activity rate is shown to be useful.

preprint2016arXiv

Statistically validated network of portfolio overlaps and systemic risk

Common asset holding by financial institutions, namely portfolio overlap, is nowadays regarded as an important channel for financial contagion with the potential to trigger fire sales and thus severe losses at the systemic level. In this paper we propose a method to assess the statistical significance of the overlap between pairs of heterogeneously diversified portfolios, which then allows us to build a validated network of financial institutions where links indicate potential contagion channels due to realized portfolio overlaps. The method is implemented on a historical database of institutional holdings ranging from 1999 to the end of 2013, but can be in general applied to any bipartite network where the presence of similar sets of neighbors is of interest. We find that the proportion of validated network links (i.e., of statistically significant overlaps) increased steadily before the 2007-2008 global financial crisis and reached a maximum when the crisis occurred. We argue that the nature of this measure implies that systemic risk from fire sales liquidation was maximal at that time. After a sharp drop in 2008, systemic risk resumed its growth in 2009, with a notable accelerat

preprint2016arXiv

Fast and Adaptive Sparse Precision Matrix Estimation in High Dimensions

This paper proposes a new method for estimating sparse precision matrices in the high dimensional setting. It has been popular to study fast computation and adaptive procedures for this problem. We propose a novel approach, called Sparse Column-wise Inverse Operator, to address these two issues. We analyze an adaptive procedure based on cross validation, and establish its convergence rate under the Frobenius norm. The convergence rates under other matrix norms are also established. This method also enjoys the advantage of fast computation for large-scale problems, via a coordinate descent algorithm. Numerical merits are illustrated using both simulated and real datasets. In particular, it performs favorably on an HIV brain tissue dataset and an ADHD resting-state fMRI dataset.

preprint2016arXiv

Probabilistic graphical model based approach for water mapping using GaoFen-2 (GF-2) high resolution imagery and Landsat 8 time series

The objective of this paper is to evaluate the potential of Gaofen-2 (GF-2) high resolution multispectral sensor (MS) and panchromatic (PAN) imagery on water mapping. Difficulties of water mapping on high resolution data includes: 1) misclassification between water and shadows or other low-reflectance ground objects, which is mostly caused by the spectral similarity within the given band range; 2) small water bodies with size smaller than the spatial resolution of MS image. To solve the confusion between water and low-reflectance objects, the Landsat 8 time series with two shortwave infrared (SWIR) bands is added because water has extremely strong absorption in SWIR. In order to integrate the three multi-sensor, multi-resolution data sets, the probabilistic graphical model (PGM) is utilized here with conditional probability distribution defined mainly based on the size of each object. For comparison, results from the SVM classifier on the PCA fused and MS data, thresholding method on the PAN image, and water index method on the Landsat data are computed. The confusion matrices are calculated for all the methods. The results demonstrate that the PGM method can achieve the best perfo

People in this topic

12 visible researcher(s)