Topic overview

physics.data-an

1229 works4631 researchers

Map preview

Start with the graph, then narrow the list

1229works
4631researchers

Next steps

Use the topic as a working map

Open the full map for clusters, then return here to scan ranked papers and people.

Topic graph

See the topic as a live network

Open full explorer

Inspect nearby papers, researchers, institutions and communities without opening a separate graph page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Papers in this area

24 paper(s) to start with

preprint2017arXiv

Degree-degree distribution in a power law random intersection graph with clustering

The bivariate distribution of degrees of adjacent vertices (degree-degree distribution) is an important network characteristic defining the statistical dependencies between degrees of adjacent vertices. We show the asymptotic degree-degree distribution of a sparse inhomogeneous random intersection graph and discuss its relation to the clustering and power law properties of the graph.

preprint2017arXiv

Don't choose theories: Normative inductive reasoning and the status of physical theories

Evaluating theories in physics used to be easy. Our theories provided very distinct predictions. Experimental accuracy was so small that worrying about epistemological problems was not necessary. That is no longer the case. The underdeterminacy problem between string theory and the standard model for current possible experimental energies is one example. We need modern inductive methods for this problem, Bayesian methods or the equivalent Solomonoff induction. To illustrate the proper way to work with induction problems I will use the concepts of Solomoff induction to study the status of string theory. Previous attempts have focused on the Bayesian solution. And they run into the question of why string theory is widely accepted with no data backing it. Logically unsupported additions to the Bayesian method were proposed. I will show here that, by studying the problem from the point of view of the Solomonoff induction those additions can be understood much better. They are not ways to update probabilities. Instead, they are considerations about the priors as well as heuristics to attempt to deal with our finite resources. For the general problem, Solomonoff induction also makes it c

preprint2016arXiv

Fractional Brownian motion time-changed by gamma and inverse gamma process

Many real time-series exhibit behavior adequate to long range dependent data. Additionally very often these time-series have constant time periods and also have characteristics similar to Gaussian processes although they are not Gaussian. Therefore there is need to consider new classes of systems to model these kind of empirical behavior. Motivated by this fact in this paper we analyze two processes which exhibit long range dependence property and have additional interesting characteristics which may be observed in real phenomena. Both of them are constructed as the superposition of fractional Brownian motion (FBM) and other process. In the first case the internal process, which plays role of the time, is the gamma process while in the second case the internal process is its inverse. We present in detail their main properties paying main attention to the long range dependence property. Moreover, we show how to simulate these processes and estimate their parameters. We propose to use a novel method based on rescaled modified cumulative distribution function for estimation of parameters of the second considered process. This method is very useful in description of rounded data, like

preprint2016arXiv

Modified cumulative distribution function in application to waiting time analysis in CTRW scenario

The continuous time random walk model plays an important role in modeling of so called anomalous diffusion behaviour. One of the specific property of such model are constant time periods visible in trajectory. In the continuous time random walk approach they are realizations of the sequence called waiting times. The main attention of the paper is paid on the analysis of waiting times distribution. We introduce here novel methods of estimation and statistical investigation of such distribution. The methods are based on the modified cumulative distribution function. In this paper we consider three special cases of waiting time distributions, namely $α$-stable, tempered stable and gamma. However the proposed methodology can be applied to broad set of distributions - in general it may serve as a method of fitting any distribution function if the observations are rounded. The new statistical techniques we apply to the simulated data as well as to the real data describing $CO_2$ concentration in indoor air.

preprint2016arXiv

Analysis of multichannel measurements of rare processes with uncertain expected background and acceptance

A typical experiment in high energy physics is considered. The result of the experiment is assumed to be a histogram consisting of bins or channels with numbers of corresponding registered events. The expected background and expected signal shape or acceptance are measured in separate auxiliary experiments, or calculated by the Monte Carlo method with finite sample size, and hence with finite precision. An especially complex situation occurs when the expected background in some of the channels happens to be zero due to either a fluctuation of the auxiliary measurement (or simulation) or because it is truly zero. Different statistical methods give different confidence intervals for the full signal rate and different significances of the signal+background hypothesis versus the pure background hypothesis. Detailed analysis and numerical tests are presented.

preprint2016arXiv

Indigenization of Urban Mobility

The identification of urban mobility patterns is very important for predicting and controlling spatial events. In this study, we analyzed millions of geographical check-ins crawled from a leading Chinese location-based social networking service (Jiepang.com), which contains demographic information that facilitates group-specific studies. We determined the distinct mobility patterns of natives and non-natives in all five large cities that we considered. We used a mixed method to assign different algorithms to natives and non-natives, which greatly improved the accuracy of location prediction compared with the basic algorithms. We also propose so-called indigenization coefficients to quantify the extent to which an individual behaves like a native, which depends only on their check-in behavior, rather than requiring demographic information. Surprisingly, the hybrid algorithm weighted using the indigenization coefficients outperformed a mixed algorithm that used additional demographic information, suggesting the advantage of behavioral data in characterizing individual mobility compared with the demographic information. The present location prediction algorithms can find applications

preprint2016arXiv

Sparsity enabled cluster reduced-order models for control

Characterizing and controlling nonlinear, multi-scale phenomena play important roles in science and engineering. Cluster-based reduced-order modeling (CROM) was introduced to exploit the underlying low-dimensional dynamics of complex systems. CROM builds a data-driven discretization of the Perron-Frobenius operator, resulting in a probabilistic model for ensembles of trajectories. A key advantage of CROM is that it embeds nonlinear dynamics in a linear framework, and uncertainty can be managed with data assimilation. CROM is typically computed on high-dimensional data, however, access to and computations on this full-state data limit the online implementation of CROM for prediction and control. Here, we address this key challenge by identifying a small subset of critical measurements to learn an efficient CROM, referred to as sparsity-enabled CROM. In particular, we leverage compressive measurements to faithfully embed the cluster geometry and preserve the probabilistic dynamics. Further, we show how to identify fewer optimized sensor locations tailored to a specific problem that outperform random measurements. Both of these sparsity-enabled sensing strategies significantly reduce

preprint2016arXiv

Simultaneous Estimation of Noise Variance and Number of Peaks in Bayesian Spectral Deconvolution

The heuristic identification of peaks from noisy complex spectra often leads to misunderstanding of the physical and chemical properties of matter. In this paper, we propose a framework based on Bayesian inference, which enables us to separate multipeak spectra into single peaks statistically and consists of two steps. The first step is estimating both the noise variance and the number of peaks as hyperparameters based on Bayes free energy, which generally is not analytically tractable. The second step is fitting the parameters of each peak function to the given spectrum by calculating the posterior density, which has a problem of local minima and saddles since multipeak models are nonlinear and hierarchical. Our framework enables the escape from local minima or saddles by using the exchange Monte Carlo method and calculates Bayes free energy via the multiple histogram method. We discuss a simulation demonstrating how efficient our framework is and show that estimating both the noise variance and the number of peaks prevents overfitting, overpenalizing, and misunderstanding the precision of parameter estimation.

preprint2016arXiv

Rising Above Chaotic Likelihoods

Berliner (Likelihood and Bayesian prediction for chaotic systems, J. Am. Stat. Assoc. 1991) identified a number of difficulties in using the likelihood function within the Bayesian paradigm which arise both for state estimation and for parameter estimation of chaotic systems. Even when the equations of the system are given, he demonstrated "chaotic likelihood functions" both of initial conditions and of parameter values in the Logistic Map. Chaotic likelihood functions, while ultimately smooth, have such complicated small scale structure as to cast doubt on the possibility of identifying high likelihood states in practice. In this paper, the challenge of chaotic likelihoods is overcome by embedding the observations in a higher dimensional sequence-space; this allows good state estimation with finite computational power. An importance sampling approach is introduced, where Pseudo-orbit Data Assimilation is employed in the sequence-space, first to identify relevant pseudo-orbits and then relevant trajectories. Estimates are identified with likelihoods orders of magnitude higher than those previously identified in the examples given by Berliner. Pseudo-orbit Data Assimilation

preprint2016arXiv

Explaining the Prevalence, Scaling and Variance of Urban Phenomena

The prevalence of many urban phenomena changes systematically with population size. We propose a theory that unifies models of economic complexity and cultural evolution to derive urban scaling. The theory accounts for the difference in scaling exponents and average prevalence across phenomena, as well as the difference in the variance within phenomena across cities of similar size. The central ideas are that a number of necessary complementary factors must be simultaneously present for a phenomenon to occur, and that the diversity of factors is logarithmically related to population size. The model reveals that phenomena that require more factors will be less prevalent, scale more superlinearly and show larger variance across cities of similar size. The theory applies to data on education, employment, innovation, disease and crime, and it entails the ability to predict the prevalence of a phenomenon across cities, given information about the prevalence in a single city.

preprint2016arXiv

Machine learning and multivariate goodness of fit

Multivariate goodness-of-fit and two-sample tests are important components of many nuclear and particle physics analyses. While a variety of powerful methods are available if the dimensionality of the feature space is small, such tests rapidly lose power as the dimensionality increases and the data inevitably become sparse. Machine learning classifiers are powerful tools capable of reducing highly multivariate problems into univariate ones, on which commonly used tests such as $χ^2$ or Kolmogorov-Smirnov may be applied. We explore applying both traditional and machine-learning-based tests to several example problems, and study how the power of each approach depends on the dimensionality. A pedagogical discussion is provided on which types of problems are best suited to using traditional versus machine-learning-based tests, and on the how to properly employ the machine-learning-based approach.

preprint2016arXiv

Covariance and correlation estimators in bipartite complex systems with a double heterogeneity

We present a weighted estimator of the covariance and correlation in bipartite complex systems with a double layer of heterogeneity. The advantage provided by the weighted estimators lies in the fact that the unweighted sample covariance and correlation can be shown to possess a bias. Indeed, such a bias affects real bipartite systems, and, for example, we report its effects on two empirical systems, one social and the other biological. On the contrary, our newly proposed weighted estimators remove the bias and are better suited to describe such systems.

preprint2016arXiv

A Neural Network Approach for the Peak Profile Characterization

The neural network-based approach, presented in this paper, was developed for the analysis of peak profiles and for the prediction of base profile characteristics, such as width, asymmetry, asymptotic ("peak tales"), etc. of the observed distributions. The obtained parameters can be used as the initial parameters in the peak decomposition applications. The neural network architecture, presented here, was designed for the analysis of one particular type of peak profiles, the Voigt type distributions (symmetrical and asymmetrical), and is suitable for a variety of applications, such as x-ray and neutron powder diffraction, x-ray spectroscopy, etc. The approach itself, however, is not limited to the demonstrated case, but is applicable to other types of peak profile distributions. The approach was successfully tested on experimentally collected x-ray powder diffraction data.

preprint2016arXiv

Nonlinear FM Waveform Design to Reduction of sidelobe level in Autocorrelation Function

This paper will design non-linear frequency modulation (NLFM) signal for Chebyshev, Kaiser, Taylor, and raised-cosine power spectral densities (PSDs). Then, the variation of peak sidelobe level with regard to mainlobe width for these four different window functions are analyzed. It has been demonstrated that reduction of sidelobe level in NLFM signal can lead to increase in mainlobe width of autocorrelation function. Furthermore, the results of power spectral density obtained from the simulation and the desired PSD are compared. Finally, error percentage between simulated PSD and desired PSD for different peak sidelobe level are illustrated. The stationary phase concept is the possible source for this error.

preprint2016arXiv

On the detection of superdiffusive behaviour in time series

We present a new method for detecting superdiffusive behaviour and for determining rates of superdiffusion in time series data. Our method applies equally to stochastic and deterministic time series data (with no prior knowledge required of the nature of the data) and relies on one realisation (ie one sample path) of the process. Linear drift effects are automatically removed without any preprocessing. We show numerical results for time series constructed from i.i.d. $α$-stable random variables and from deterministic weakly chaotic maps. We compare our method with the standard method of estimating the growth rate of the mean-square displacement as well as the $p$-variation method, maximum likelihood, quantile matching and linear regression of the empirical characteristic function.

preprint2016arXiv

Revisiting Causality Inference in Memory-less Transition Networks

Several methods exist to infer causal networks from massive volumes of observational data. However, almost all existing methods require a considerable length of time series data to capture cause and effect relationships. In contrast, memory-less transition networks or Markov Chain data, which refers to one-step transitions to and from an event, have not been explored for causality inference even though such data is widely available. We find that causal network can be inferred from characteristics of four unique distribution zones around each event. We call this Composition of Transitions and show that cause, effect, and random events exhibit different behavior in their compositions. We applied machine learning models to learn these different behaviors and to infer causality. We name this new method Causality Inference using Composition of Transitions (CICT). To evaluate CICT, we used an administrative inpatient healthcare dataset to set up a network of patients transitions between different diagnoses. We show that CICT is highly accurate in inferring whether the transition between a pair of events is causal or random and performs well in identifying the direction of causality in a

preprint2016arXiv

Sequential motif profile of natural visibility graphs

The concept of sequential visibility graph motifs -subgraphs appearing with characteristic frequencies in the visibility graphs associated to time series- has been advanced recently along with a theoretical framework to compute analytically the motif profiles associated to Horizontal Visibility Graphs (HVGs). Here we develop a theory to compute the profile of sequential visibility graph motifs in the context of Natural Visibility Graphs (VGs). This theory gives exact results for deterministic aperiodic processes with a smooth invariant density or stochastic processes that fulfil the Markov property and have a continuous marginal distribution. The framework also allows for a linear time numerical estimation in the case of empirical time series. A comparison between the HVG and the VG case (including evaluation of their robustness for short series polluted with measurement noise) is also presented.

preprint2016arXiv

Reweighting with Boosted Decision Trees

Machine learning tools are commonly used in modern high energy physics (HEP) experiments. Different models, such as boosted decision trees (BDT) and artificial neural networks (ANN), are widely used in analyses and even in the software triggers. In most cases, these are classification models used to select the "signal" events from data. Monte Carlo simulated events typically take part in training of these models. While the results of the simulation are expected to be close to real data, in practical cases there is notable disagreement between simulated and observed data. In order to use available simulation in training, corrections must be introduced to generated data. One common approach is reweighting - assigning weights to the simulated events. We present a novel method of event reweighting based on boosted decision trees. The problem of checking the quality of reweighting step in analyses is also discussed.

preprint2016arXiv

Predicting dataset popularity for the CMS experiment

The CMS experiment at the LHC accelerator at CERN relies on its computing infrastructure to stay at the frontier of High Energy Physics, searching for new phenomena and making discoveries. Even though computing plays a significant role in physics analysis we rarely use its data to predict the system behavior itself. A basic information about computing resources, user activities and site utilization can be really useful for improving the throughput of the system and its management. In this paper, we discuss a first CMS analysis of dataset popularity based on CMS meta-data which can be used as a model for dynamic data placement and provide the foundation of data-driven approach for the CMS computing infrastructure.

preprint2016arXiv

CoinCalc -- A new R package for quantifying simultaneities of event series

We present the new R package CoinCalc for performing event coincidence analysis (ECA), a novel statistical method to quantify the simultaneity of events contained in two series of observations, either as simultaneous or lagged coincidences within a user-specific temporal tolerance window. The package also provides different analytical as well as surrogate-based significance tests (valid under different assumptions about the nature of the observed event series) as well as an intuitive visualization of the identified coincidences. We demonstrate the usage of CoinCalc based on two typical geoscientific example problems addressing the relationship between meteorological extremes and plant phenology as well as that between soil properties and land cover.

preprint2016arXiv

Manifold Learning with Contracting Observers for Data-driven Time-series Analysis

Analyzing signals arising from dynamical systems typically requires many modeling assumptions and parameter estimation. In high dimensions, this modeling is particularly difficult due to the "curse of dimensionality". In this paper, we propose a method for building an intrinsic representation of such signals in a purely data-driven manner. First, we apply a manifold learning technique, diffusion maps, to learn the intrinsic model of the latent variables of the dynamical system, solely from the measurements. Second, we use concepts and tools from control theory and build a linear contracting observer to estimate the latent variables in a sequential manner from new incoming measurements. The effectiveness of the presented framework is demonstrated by applying it to a toy problem and to a music analysis application. In these examples we show that our method reveals the intrinsic variables of the analyzed dynamical systems.

preprint2011arXiv

Detecting series periodicity with horizontal visibility graphs

The horizontal visibility algorithm has been recently introduced as a mapping between time series and networks. The challenge lies in characterizing the structure of time series (and the processes that generated those series) using the powerful tools of graph theory. Recent works have shown that the visibility graphs inherit several degrees of correlations from their associated series, and therefore such graph theoretical characterization is in principle possible. However, both the mathematical grounding of this promising theory and its applications are on its infancy. Following this line, here we address the question of detecting hidden periodicity in series polluted with a certain amount of noise. We first put forward some generic properties of horizontal visibility graphs which allow us to define a (graph theoretical) noise reduction filter. Accordingly, we evaluate its performance for the task of calculating the period of noisy periodic signals, and compare our results with standard time domain (autocorrelation) methods. Finally, potentials, limitations and applications are discussed.

preprint2008arXiv

Surrogates with random Fourier Phases

The method of surrogates is widely used in the field of nonlinear data analysis for testing for weak nonlinearities. The two most commonly used algorithms for generating surrogates are the amplitude adjusted Fourier transform (AAFT) and the iterated amplitude adjusted Fourier transfom (IAAFT) algorithm. Both the AAFT and IAAFT algorithm conserve the amplitude distribution in real space and reproduce the power spectrum (PS) of the original data set very accurately. The basic assumption in both algorithms is that higher-order correlations can be wiped out using a Fourier phase randomization procedure. In both cases, however, the randomness of the Fourier phases is only imposed before the (first) Fourier back tranformation. Until now, it has not been studied how the subsequent remapping and iteration steps may affect the randomness of the phases. Using the Lorenz system as an example, we show that both algorithms may create surrogate realizations containing Fourier phase correlations. We present two new iterative surrogate data generating methods being able to control the randomization of Fourier phases at every iteration step. The resulting surrogate realizations which are truly line

preprint2016arXiv

Support Vector Machines and Generalisation in HEP

We review the concept of support vector machines (SVMs) and discuss examples of their use. One of the benefits of SVM algorithms, compared with neural networks and decision trees is that they can be less susceptible to over fitting than those other algorithms are to over training. This issue is related to the generalisation of a multivariate algorithm (MVA); a problem that has often been overlooked in particle physics. We discuss cross validation and how this can be used to improve the generalisation of a MVA in the context of High Energy Physics analyses. The examples presented use the Toolkit for Multivariate Analysis (TMVA) based on ROOT and describe our improvements to the SVM functionality and new tools introduced for cross validation within this framework.

People in this topic

12 visible researcher(s)