Source author record

Eduardo G. Altmann

Eduardo G. Altmann appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

29works
21topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

29 published item(s)

preprint2020arXiv

Scaling laws and dynamics of hashtags on Twitter

In this paper we quantify the statistical properties and dynamics of the frequency of hashtag use on Twitter. Hashtags are special words used in social media to attract attention and to organize content. Looking at the collection of all hashtags used in a period of time, we identify the scaling laws underpinning the hashtag frequency distribution (Zipf's law), the number of unique hashtags as a function of sample size (Heaps' law), and the fluctuations around expected values (Taylor's law). While these scaling laws appear to be universal, in the sense that similar exponents are observed irrespective of when the sample is gathered, the volume and nature of the hashtags depends strongly on time, with the appearance of bursts at the minute scale, fat-tailed noise, and long-range correlations. We quantify this dynamics by computing the Jensen-Shannon divergence between hashtag distributions obtained $τ$ times apart and we find that the speed of change decays roughly as $1/τ$. Our findings are based on the analysis of 3.5 billion hashtags used between 2015 and 2016.

preprint2020arXiv

Spatial interactions in urban scaling laws

Analyses of urban scaling laws assume that observations in different cities are independent of the existence of nearby cities. Here we introduce generative models and data-analysis methods that overcome this limitation by modelling explicitly the effect of interactions between individuals at different locations. Parameters that describe the scaling law and the spatial interactions are inferred from data simultaneously, allowing for rigorous (Bayesian) model comparison and overcoming the problem of defining the boundaries of urban regions. Results in five different datasets show that including spatial interactions typically leads to better models and a change in the exponent of the scaling law. Data and codes are provided in Ref. [1].

preprint2020arXiv

Structure of resonance eigenfunctions for chaotic systems with partial escape

Physical systems are often neither completely closed nor completely open, but instead they are best described by dynamical systems with partial escape or absorption. In this paper we introduce classical measures that explain the main properties of resonance eigenfunctions of chaotic quantum systems with partial escape. We construct a family of conditionally-invariant measures with varying decay rates by interpolating between the natural measures of the forward and backward dynamics. Numerical simulations in a representative system show that our classical measures correctly describe the main features of the quantum eigenfunctions: their multi-fractal phase space distribution, their product structure along stable/unstable directions, and their dependence on the decay rate. The (Jensen-Shannon) distance between classical and quantum measures goes to zero in the semiclassical limit for long- and short-lived eigenfunctions, while it remains finite for intermediate cases.

preprint2016arXiv

Impact of lexical and sentiment factors on the popularity of scientific papers

We investigate how textual properties of scientific papers relate to the number of citations they receive. Our main finding is that correlations are non-linear and affect differently most-cited and typical papers. For instance, we find that in most journals short titles correlate positively with citations only for the most cited papers, for typical papers the correlation is in most cases negative. Our analysis of 6 different factors, calculated both at the title and abstract level of 4.3 million papers in over 1500 journals, reveals the number of authors, and the length and complexity of the abstract, as having the strongest (positive) influence on the number of citations.

preprint2016arXiv

Similarity of symbol frequency distributions with heavy tails

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to texts, the distribution of symbol frequencies is characterized by heavy-tailed distributions (e.g., Zipf's law). The large number of low-frequency symbols in these distributions poses major difficulties to the estimation of the similarity between sequences, e.g., they hinder an accurate finite-size estimation of entropies. Here we show analytically how the systematic (bias) and statistical (fluctuations) errors in these estimations depend on the sample size~$N$ and on the exponent~$γ$ of the heavy-tailed distribution. Our results are valid for the Shannon entropy $(α=1)$, its corresponding similarity measures (e.g., the Jensen-Shanon divergence), and also for measures based on the generalized entropy of order $α$. For small $α$'s, including $α=1$, the errors decay slower than the $1/N$-decay observed in short-tailed distributions. For $α$ larger than a critical value $α^* = 1+1/γ\leq 2$, the $1/N$-decay is recovered. We show the practical significance of our results by quantifying the evolution of the English language over the last two centuries using a complete $α$-spectrum of measures. We find that frequent words change more slowly than less frequent words and that $α=2$ provides the most robust measure to quantify language change.

preprint2015arXiv

Chaotic Explosions

We investigate chaotic dynamical systems for which the intensity of trajectories might grow unlimited in time. We show that (i) the intensity grows exponentially in time and is distributed spatially according to a fractal measure with an information dimension smaller than that of the phase space,(ii) such exploding cases can be described by an operator formalism similar to the one applied to chaotic systems with absorption (decaying intensities), but (iii) the invariant quantities characterizing explosion and absorption are typically not directly related to each other, e.g., the decay rate and fractal dimensions of absorbing maps typically differ from the ones computed in the corresponding inverse (exploding) maps. We illustrate our general results through numerical simulation in the cardioid billiard mimicking a lasing optical cavity, and through analytical calculations in the baker map.

preprint2015arXiv

Quantum signatures of classical multifractal measures

A clear signature of classical chaoticity in the quantum regime is the fractal Weyl law, which connects the density of eigenstates to the dimension $D_0$ of the classical invariant set of open systems. Quantum systems of interest are often {\it partially} open (e.g., cavities in which trajectories are partially reflected/absorbed). In the corresponding classical systems $D_0$ is trivial (equal to the phase-space dimension), and the fractality is manifested in the (multifractal) spectrum of Rényi dimensions $D_q$. In this paper we investigate the effect of such multifractality on the Weyl law. Our numerical simulations in area-preserving maps show for a wide range of configurations and system sizes $M$ that (i) the Weyl law is governed by a dimension different from $D_0=2$ and (ii) the observed dimension oscillates as a function of $M$ and other relevant parameters. We propose a classical model which considers an undersampled measure of the chaotic invariant set, explains our two observations, and predicts that the Weyl law is governed by a non-trivial dimension $D_\mathrm{asymptotic} < D_0$ in the semi-classical limit $M\rightarrow\infty$.

preprint2015arXiv

Sampling motif-constrained ensembles of networks

The statistical significance of network properties is conditioned on null models which satisfy spec- ified properties but that are otherwise random. Exponential random graph models are a principled theoretical framework to generate such constrained ensembles, but which often fail in practice, either due to model inconsistency, or due to the impossibility to sample networks from them. These problems affect the important case of networks with prescribed clustering coefficient or number of small connected subgraphs (motifs). In this paper we use the Wang-Landau method to obtain a multicanonical sampling that overcomes both these problems. We sample, in polynomial time, net- works with arbitrary degree sequences from ensembles with imposed motifs counts. Applying this method to social networks, we investigate the relation between transitivity and homophily, and we quantify the correlation between different types of motifs, finding that single motifs can explain up to 60% of the variation of motif profiles.

preprint2015arXiv

Statistical laws in linguistics

Zipf's law is just one out of many universal laws proposed to describe statistical regularities in language. Here we review and critically discuss how these laws can be statistically interpreted, fitted, and tested (falsified). The modern availability of large databases of written text allows for tests with an unprecedent statistical accuracy and also a characterization of the fluctuations around the typical behavior. We find that fluctuations are usually much larger than expected based on simplifying statistical assumptions (e.g., independence and lack of correlations between observations).These simplifications appear also in usual statistical tests so that the large fluctuations can be erroneously interpreted as a falsification of the law. Instead, here we argue that linguistic laws are only meaningful (falsifiable) if accompanied by a model for which the fluctuations can be computed (e.g., a generative model of the text). The large fluctuations we report show that the constraints imposed by linguistic laws on the creativity process of text generation are not as tight as one could expect.

preprint2015arXiv

Stochastic perturbations in open chaotic systems: random versus noisy maps

We investigate the effects of random perturbations on fully chaotic open systems. Perturbations can be applied to each trajectory independently (white noise) or simultaneously to all trajectories (random map). We compare these two scenarios by generalizing the theory of open chaotic systems and introducing a time-dependent conditionally-map-invariant measure. For the same perturbation strength we show that the escape rate of the random map is always larger than that of the noisy map. In random maps we show that the escape rate $κ$ and dimensions $D$ of the relevant fractal sets often depend nonmonotonically on the intensity of the random perturbation. We discuss the accuracy (bias) and precision (variance) of finite-size estimators of $κ$ and $D$, and show that the improvement of the precision of the estimations with the number of trajectories $N$ is extremely slow ($\propto 1/\ln N$). We also argue that the finite-size $D$ estimators are typically biased. General theoretical results are combined with analytical calculations and numerical simulations in area-preserving baker maps.

preprint2015arXiv

Temporal-varying failures of nodes in networks

We consider networks in which random walkers are removed because of the failure of specific nodes. We interpret the rate of loss as a measure of the importance of nodes, a notion we denote as failure-centrality. We show that the degree of the node is not sufficient to determine this measure and that, in a first approximation, the shortest loops through the node have to be taken into account. We propose approximations of the failure-centrality which are valid for temporal-varying failures and we dwell on the possibility of externally changing the relative importance of nodes in a given network, by exploiting the interference between the loops of a node and the cycles of the temporal pattern of failures. In the limit of long failure cycles we show analytically that the escape in a node is larger than the one estimated from a stochastic failure with the same failure probability. We test our general formalism in two real-world networks (air-transportation and e-mail users) and show how communities lead to deviations from predictions for failures in hubs.

preprint2014arXiv

Efficiency of Monte Carlo Sampling in Chaotic Systems

In this paper we investigate how the complexity of chaotic phase spaces affect the efficiency of importance sampling Monte Carlo simulations. We focus on a flat-histogram simulation of the distribution of finite-time Lyapunov exponent in a simple chaotic system and obtain analytically that the computational effort of the simulation: (i) scales polynomially with the finite-time, a tremendous improvement over the exponential scaling obtained in usual uniform sampling simulations; and (ii) the polynomial scaling is sub-optimal, a phenomenon known as critical slowing down. We show that critical slowing down appears because of the limited possibilities to issue a local proposal on the Monte Carlo procedure in chaotic systems. These results remain valid in other methods and show how generic properties of chaotic systems limit the efficiency of Monte Carlo simulations.

preprint2014arXiv

Extracting information from S-curves of language change

It is well accepted that adoption of innovations are described by S-curves (slow start, accelerating period, and slow end). In this paper, we analyze how much information on the dynamics of innovation spreading can be obtained from a quantitative description of S-curves. We focus on the adoption of linguistic innovations for which detailed databases of written texts from the last 200 years allow for an unprecedented statistical precision. Combining data analysis with simulations of simple models (e.g., the Bass dynamics on complex networks) we identify signatures of endogenous and exogenous factors in the S-curves of adoption. We propose a measure to quantify the strength of these factors and three different methods to estimate it from S-curves. We obtain cases in which the exogenous factors are dominant (in the adoption of German orthographic reforms and of one irregular verb) and cases in which endogenous factors are dominant (in the adoption of conventions for romanization of Russian names and in the regularization of most studied verbs). These results show that the shape of S-curve is not universal and contains information on the adoption mechanism. (published at "J. R. Soc. Interface, vol. 11, no. 101, (2014) 1044"; DOI: http://dx.doi.org/10.1098/rsif.2014.1044)

preprint2014arXiv

Predictability of extreme events in social media

It is part of our daily social-media experience that seemingly ordinary items (videos, news, publications, etc.) unexpectedly gain an enormous amount of attention. Here we investigate how unexpected these events are. We propose a method that, given some information on the items, quantifies the predictability of events, i.e., the potential of identifying in advance the most successful items defined as the upper bound for the quality of any prediction based on the same information. Applying this method to different data, ranging from views in YouTube videos to posts in Usenet discussion groups, we invariantly find that the predictability increases for the most extreme events. This indicates that, despite the inherently stochastic collective dynamics of users, efficient prediction is possible for the most extreme events.

preprint2014arXiv

Scaling laws and fluctuations in the statistics of word frequencies

In this paper we combine statistical analysis of large text databases and simple stochastic models to explain the appearance of scaling laws in the statistics of word frequencies. Besides the sublinear scaling of the vocabulary size with database size (Heaps' law), here we report a new scaling of the fluctuations around this average (fluctuation scaling analysis). We explain both scaling laws by modeling the usage of words by simple stochastic processes in which the overall distribution of word-frequencies is fat tailed (Zipf's law) and the frequency of a single word is subject to fluctuations across documents (as in topic models). In this framework, the mean and the variance of the vocabulary size can be expressed as quenched averages, implying that: i) the inhomogeneous dissemination of words cause a reduction of the average vocabulary size in comparison to the homogeneous case, and ii) correlations in the co-occurrence of words lead to an increase in the variance and the vocabulary size becomes a non-self-averaging quantity. We address the implications of these observations to the measurement of lexical richness. We test our results in three large text databases (Google-ngram, Enlgish Wikipedia, and a collection of scientific articles).

preprint2013arXiv

Chaotic Systems with Absorption

Motivated by applications in optics and acoustics we develop a dynamical-system approach to describe absorption in chaotic systems. We introduce an operator formalism from which we obtain (i) a general formula for the escape rate $κ$ in terms of the natural conditionally-invariant measure of the system; (ii) an increased multifractality when compared to the spectrum of dimensions $D_q$ obtained without taking absorption and return times into account; and (iii) a generalization of the Kantz-Grassberger formula that expresses $D_1$ in terms of $κ$, the positive Lyapunov exponent, the average return time, and a new quantity, the reflection rate. Simulations in the cardioid billiard confirm these results.

preprint2013arXiv

Identifying trends in word frequency dynamics

The word-stock of a language is a complex dynamical system in which words can be created, evolve, and become extinct. Even more dynamic are the short-term fluctuations in word usage by individuals in a population. Building on the recent demonstration that word niche is a strong determinant of future rise or fall in word frequency, here we introduce a model that allows us to distinguish persistent from temporary increases in frequency. Our model is illustrated using a 10^8-word database from an online discussion group and a 10^11-word collection of digitized books. The model reveals a strong relation between changes in word dissemination and changes in frequency. Aside from their implications for short-term word frequency dynamics, these observations are potentially important for language evolution as new words must survive in the short term in order to survive in the long term.

preprint2013arXiv

Leaking Chaotic Systems

There are numerous physical situations in which a hole or leak is introduced in an otherwise closed chaotic system. The leak can have a natural origin, it can mimic measurement devices, and it can also be used to reveal dynamical properties of the closed system. In this paper we provide an unified treatment of leaking systems and we review applications to different physical problems, both in the classical and quantum pictures. Our treatment is based on the transient chaos theory of open systems, which is essential because real leaks have finite size and therefore estimations based on the closed system differ essentially from observations. The field of applications reviewed is very broad, ranging from planetary astronomy and hydrodynamical flows, to plasma physics and quantum fidelity. The theory is expanded and adapted to the case of partial leaks (partial absorption/transmission) with applications to room acoustics and optical microcavities in mind. Simulations in the lima .con family of billiards illustrate the main text. Regarding billiard dynamics, we emphasize that a correct discrete time representation can only be given in terms of the so- called true-time maps, while traditional Poincar é maps lead to erroneous results. We generalize Perron-Frobenius-type operators so that they describe true-time maps with partial leaks.

preprint2013arXiv

Monte Carlo Sampling in Fractal Landscapes

We propose a flat-histogram Monte Carlo method to efficiently sample fractal landscapes such as escape time functions of open chaotic systems. This is achieved by using a random-walk step which depends on the height of the landscape via the largest Lyapunov exponent of the associated chaotic system. By generalizing the Wang-Landau algorithm, we obtain a method which simultaneously constructs the density of states (escape time distribution) and the correct step-length distribution. As a result, averages are obtained in polynomial computational time, a dramatic improvement over the exponential scaling of traditional uniform sampling. Our results are not limited by the dimensionality of the phase space and are confirmed numerically for dimensions as large as 30.

preprint2013arXiv

Optimal noise maximizes collective motion in heterogeneous media

We study the effect of spatial heterogeneity on the collective motion of self-propelled particles (SPPs). The heterogeneity is modeled as a random distribution of either static or diffusive obstacles, which the SPPs avoid while trying to align their movements. We find that such obstacles have a dramatic effect on the collective dynamics of usual SPP models. In particular, we report about the existence of an optimal (angular) noise amplitude that maximizes collective motion. We also show that while at low obstacle densities the system exhibits long-range order, in strongly heterogeneous media collective motion is quasi-long-range and exists only for noise values in between two critical noise values, with the system being disordered at both, large and low noise amplitudes. Since most real system have spatial heterogeneities, the finding of an optimal noise intensity has immediate practical and fundamental implications for the design and evolution of collective motion strategies.

preprint2013arXiv

Probing the statistical properties of unknown texts: application to the Voynich Manuscript

While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed investigating the properties of statistical measurements across different languages and texts. In this study we propose a framework that aims at determining if a text is compatible with a natural language and which languages are closest to it, without any knowledge of the meaning of the words. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing text, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for key-words of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.

preprint2013arXiv

Stochastic model for the vocabulary growth in natural languages

We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of words: (i) a finite number of core-words which have higher frequency and do not affect the probability of a new word to be used; and (ii) the remaining virtually infinite number of noncore-words which have lower frequency and once used reduce the probability of a new word to be used in the future. Our model relies on a careful analysis of the google-ngram database of books published in the last centuries and its main consequence is the generalization of Zipf's and Heaps' law to two scaling regimes. We confirm that these generalizations yield the best simple description of the data among generic descriptive models and that the two free parameters depend only on the language but not on the database. From the point of view of our model the main change on historical time scales is the composition of the specific words included in the finite list of core-words, which we observe to decay exponentially in time with a rate of approximately 30 words per year for English.

preprint2012arXiv

Analysis of an information-theoretic model for communication

We study the cost-minimization problem posed by Ferrer i Cancho and Solé in their model of communication that aimed at explaining the origin of Zipf's law [PNAS 100, 788 (2003)]. Direct analysis shows that the minimum cost is $\min {λ, 1-λ}$, where $λ$ determines the relative weights of speaker's and hearer's costs in the total, as shown in several previous works using different approaches. The nature and multiplicity of the minimizing solution changes discontinuously at $λ=1/2$, being qualitatively different for $λ< 1/2$, $λ> 1/2$, and $λ=1/2$. Zipf's law is found only in a vanishing fraction of the minimum-cost solutions at $λ= 1/2$ and therefore is not explained by this model. Imposing the further condition of equal costs yields distributions substantially closer to Zipf's law, but significant differences persist. We also investigate the solutions reached by the previously used minimization algorithm and find that they correctly recover global minimum states at the transition.

preprint2012arXiv

Effect of noise in open chaotic billiards

We investigate the effect of white-noise perturbations on chaotic trajectories in open billiards. We focus on the temporal decay of the survival probability for generic mixed-phase-space billiards. The survival probability has a total of five different decay regimes that prevail for different intermediate times. We combine new calculations and recent results on noise perturbed Hamiltonian systems to characterize the origin of these regimes, and to compute how the parameters scale with noise intensity and billiard openness. Numerical simulations in the annular billiard support and illustrate our results.

preprint2012arXiv

Faster than expected escape for a class of fully chaotic maps

We investigate the dependence of the escape rate on the position of a hole placed in uniformly hyperbolic systems admitting a finite Markov partition. We derive an exact periodic orbit formula for finite size Markov holes which differs from other periodic expansions in the literature and can account for additional distortion to maps with piecewise constant expansion rate. Using asymptotic expansions in powers of hole size we show that for systems conjugate to the binary shift, the average escape rate is always larger than the expectation based on the hole size. Moreover, we show that in the small hole limit the difference between the two decays like a known constant times the square of the hole size. Finally, we relate this problem to the random choice of hole positions and we discuss possible extensions of our results to non-Markov holes as well as applications to leaky dynamical networks.

preprint2012arXiv

On the origin of long-range correlations in texts

The complexity of human interactions with social and natural phenomena is mirrored in the way we describe our experiences through natural language. In order to retain and convey such a high dimensional information, the statistical properties of our linguistic output has to be highly correlated in time. An example are the robust observations, still largely not understood, of correlations on arbitrary long scales in literary texts. In this paper we explain how long-range correlations flow from highly structured linguistic levels down to the building blocks of a text (words, letters, etc..). By combining calculations and data analysis we show that correlations take form of a bursty sequence of events once we approach the semantically relevant topics of the text. The mechanisms we identify are fairly general and can be equally applied to other hierarchical settings.

preprint2011arXiv

Comparing intermittency and network measurements of words and their dependency on authorship

Many features from texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper we quantify how topological properties of word co-occurrence networks and intermittency (or burstiness) in word distribution depend on the style of authors. Our database contains 40 books from 8 authors who lived in the 19th and 20th centuries, for which the following network measurements were obtained: clustering coefficient, average shortest path lengths, and betweenness. We found that the two factors with stronger dependency on the authors were the skewness in the distribution of word intermittency and the average shortest paths. Other factors such as the betweeness and the Zipf's law exponent show only weak dependency on authorship. Also assessed was the contribution from each measurement to authorship recognition using three machine learning methods. The best performance was a ca. 65 % accuracy upon combining complex network and intermittency features with the nearest neighbor algorithm. From a detailed analysis of the interdependence of the various metrics it is concluded that the methods used here are complementary for providing short- and long-scale perspectives of texts, which are useful for applications such as identification of topical words and information retrieval.

preprint2011arXiv

Niche as a determinant of word fate in online groups

Patterns of word use both reflect and influence a myriad of human activities and interactions. Like other entities that are reproduced and evolve, words rise or decline depending upon a complex interplay between {their intrinsic properties and the environments in which they function}. Using Internet discussion communities as model systems, we define the concept of a word niche as the relationship between the word and the characteristic features of the environments in which it is used. We develop a method to quantify two important aspects of the size of the word niche: the range of individuals using the word and the range of topics it is used to discuss. Controlling for word frequency, we show that these aspects of the word niche are strong determinants of changes in word frequency. Previous studies have already indicated that word frequency itself is a correlate of word success at historical time scales. Our analysis of changes in word frequencies over time reveals that the relative sizes of word niches are far more important than word frequencies in the dynamics of the entire vocabulary at shorter time scales, as the language adapts to new concepts and social groupings. We also distinguish endogenous versus exogenous factors as additional contributors to the fates of words, and demonstrate the force of this distinction in the rise of novel words. Our results indicate that short-term nonstationarity in word statistics is strongly driven by individual proclivities, including inclinations to provide novel information and to project a distinctive social identity.

preprint2010arXiv

Noise-enhanced trapping in chaotic scattering

We show that noise enhances the trapping of trajectories in scattering systems. In fully chaotic systems, the decay rate can decrease with increasing noise due to a generic mismatch between the noiseless escape rate and the value predicted by the Liouville measure of the exit set. In Hamiltonian systems with mixed phase space we show that noise leads to a slower algebraic decay due to trajectories performing a random walk inside Kolmogorov-Arnold-Moser islands. We argue that these noise-enhanced trapping mechanisms exist in most scattering systems and are likely to be dominant for small noise intensities, which is confirmed through a detailed investigation in the Henon map. Our results can be tested in fluid experiments, affect the fractal Weyl's law of quantum systems, and modify the estimations of chemical reaction rates based on phase-space transition state theory.