Source author record

Filippo Petroni

Filippo Petroni appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

21works
14topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

21 published item(s)

preprint2026arXiv

Intraday Limit Order Price Change Transition Dynamics Across Market Capitalizations Through Markov Analysis

Quantitative understanding of stochastic dynamics in limit order price changes is essential for execution strategy design. We analyze intraday transition dynamics of ask and bid orders across market capitalization tiers using high-frequency NASDAQ100 tick data. Employing a discrete-time Markov chain framework, we categorize consecutive price changes into nine states and estimate transition probability matrices (TPMs) for six intraday intervals across High ($\mathtt{HMC}$), Medium ($\mathtt{MMC}$), and Low ($\mathtt{LMC}$) market cap stocks. Element-wise TPM comparison reveals systematic patterns: price inertia peaks during opening and closing hours, stabilizing midday. A capitalization gradient is observed: $\mathtt{HMC}$ stocks exhibit the strongest inertia, while $\mathtt{LMC}$ stocks show lower stability and wider spreads. Markov metrics, including spectral gap, entropy rate, and mean recurrence times, quantify these dynamics. Clustering analysis identifies three distinct temporal phases on the bid side -- Opening, Midday, and Closing, and four phases on the ask side by distinguishing Opening, Midday, Pre-Close, and Close. This indicates that sellers initiate end-of-day positioning earlier than buyers. Stationary distributions show limit order dynamics are dominated by neutral and mild price changes. Jensen-Shannon divergence confirms the closing hour as the most distinct phase, with capitalization modulating temporal contrasts and bid-ask asymmetry. These findings support capitalization-aware and time-adaptive execution algorithms.

preprint2026arXiv

Regime Discovery and Intra-Regime Return Dynamics in Global Equity Markets

Financial markets alternate between tranquil periods and episodes of stress, and return dynamics can change substantially across these regimes. We study regime-dependent dynamics in developed and developing equity indices using a data-driven Hilbert--Huang-based regime identification and profiling pipeline, followed by variable-length Markov modeling of categorized returns. Market regimes are identified using an Empirical Mode Decomposition-based Hilbert--Huang Transform, where instantaneous energy from the Hilbert spectrum separates Normal, High, and Extreme regimes. We then profile each regime using Holo--Hilbert Spectral Analysis, which jointly resolves carrier frequencies, amplitude-modulation frequencies, and amplitude-modulation energy (AME). AME, interpreted as volatility intensity, declines monotonically from Extreme to High to Normal regimes. This decline is markedly sharper in developed markets, while developing markets retain higher baseline volatility intensity even in Normal regimes. Building on these regime-specific volatility signatures, we discretize daily returns into five quintile states $\mathtt{R}_1$ to $\mathtt{R}_5$ and estimate Variable-Length Markov Chains via context trees within each regime. Unconditional state probabilities show tail states dominate in Extreme regimes and recede as regimes stabilize, alongside persistent downside asymmetry. Entropy peaks in High regimes, indicating maximum unpredictability during moderate-volatility periods. Conditional transition dynamics, evaluated over contexts of length up to three days from the context-tree estimates, indicate that developed markets normalize more effectively as stress subsides, whereas developing markets retain residual tail dependence and downside persistence even in Normal regimes, consistent with a coexistence of continuation and burst-like shifts.

preprint2020arXiv

A micro-to-macro approach to returns, volumes and waiting times

Fundamental variables in financial market are not only price and return but a very important role is also played by trading volumes. Here we propose a new multivariate model that takes into account price returns, logarithmic variation of trading volumes and also waiting times, the latter to be intended as the time interval between changes in trades, price, and volume of stocks. Our approach is based on a generalization of semi-Markov chains where an endogenous index process is introduced. We also take into account the dependence structure between the above mentioned variables by means of copulae. The proposed model is motivated by empirical evidences which are known in financial literature and that are also confirmed in this work by analysing real data from Italian stock market in the period August 2015 - August 2017. By using Monte Carlo simulations, we show that the model reproduces all these empirical evidences.

preprint2015arXiv

Observability of Market Daily Volatility

We study the price dynamics of 65 stocks from the Dow Jones Composite Average from 1973 until 2014. We show that it is possible to define a Daily Market Volatility $σ(t)$ which is directly observable from data. This quantity is usually indirectly defined by $r(t)=σ(t) ω(t)$ where the $r(t)$ are the daily returns of the market index and the $ω(t)$ are i.i.d. random variables with vanishing average and unitary variance. The relation $r(t)=σ(t) ω(t)$ alone is unable to give an operative definition of the index volatility, which remains unobservable. On the contrary, we show that using the whole information available in the market, the index volatility can be operatively defined and detected.

preprint2015arXiv

Tornadoes and related damage costs: statistical modeling with a semi-Markov approach

We propose a statistical approach to tornadoes modeling for predicting and simulating occurrences of tornadoes and accumulated cost distributions over a time interval. This is achieved by modeling the tornadoes intensity, measured with the Fujita scale, as a stochastic process. Since the Fujita scale divides tornadoes intensity into six states, it is possible to model the tornadoes intensity by using Markov and semi-Markov models. We demonstrate that the semi-Markov approach is able to reproduce the duration effect that is detected in tornadoes occurrence. The superiority of the semi-Markov model as compared to the Markov chain model is also affirmed by means of a statistical test of hypothesis. As an application we compute the expected value and the variance of the costs generated by the tornadoes over a given time interval in a given area. he paper contributes to the literature by demonstrating that semi-Markov models represent an effective tool for physical analysis of tornadoes as well as for the estimation of the economic damages to human things.

preprint2014arXiv

Operational risk of a wind farm energy production by Extreme Value Theory and Copulas

In this paper we use risk management techniques to evaluate the potential effects of those operational risks that affect the energy production of a wind farm. We concentrate our attention on three major risk factors: wind speed uncertainty, wind turbine reliability and interactions of wind turbines due mainly to their placement. As a first contribution, we show that the Weibull distribution, commonly used to fit recorded wind speed data, underestimates rare events. Therefore, in order to achieve a better estimation of the tail of the wind speed distribution, we advance a Generalized Pareto distribution. The wind turbines reliability is considered by modeling the failures events as a compound Poisson process. Finally, the use of Copula able us to consider the correlation between wind turbines that compose the wind farm. Once this procedure is set up, we show a sensitivity analysis and we also compare the results from the proposed procedure with those obtained by ignoring the aforementioned risk factors.

preprint2014arXiv

Threshold Model for Triggered Avalanches on Networks

Based on a theoretical model for opinion spreading on a network, through avalanches, the effect of external field is now considered, by using methods from non-equilibrium statistical mechanics. The original part contains the implementation that the avalanche is only triggered when a local variable (a so called awareness) reaches and goes above a threshold. The dynamical rules are constrained to be as simple as possible, in order to sort out the basic features, though more elaborated variants are proposed. Several results are obtained for a Erdös-Rényi network and interpreted through simple analytical laws, scale free or logistic map-like, i.e., (i) the sizes, durations, and number of avalanches, including the respective distributions, (ii) the number of times the external field is applied to one possible node before all nodes are found to be above the threshold, (iii) the number of nodes still below the threshold and the number of hot nodes (close to threshold) at each time step.

preprint2013arXiv

Forecasting wind speed financial return

The prediction of wind speed is very important when dealing with the production of energy through wind turbines. In this paper, we show a new nonparametric model, based on semi-Markov chains, to predict wind speed. Particularly we use an indexed semi-Markov model that has been shown to be able to reproduce accurately the statistical behavior of wind speed. The model is used to forecast, one step ahead, wind speed. In order to check the validity of the model we show, as indicator of goodness, the root mean square error and mean absolute error between real data and predicted ones. We also compare our forecasting results with those of a persistence model. At last, we show an application of the model to predict financial indicators like the Internal Rate of Return, Duration and Convexity.

preprint2013arXiv

Multivariate high-frequency financial data via semi-Markov processes

In this paper we propose a bivariate generalization of a weighted indexed semi-Markov chains to study the high frequency price dynamics of traded stocks. We assume that financial returns are described by a weighted indexed semi-Markov chain model. We show, through Monte Carlo simulations, that the model is able to reproduce important stylized facts of financial time series like the persistence of volatility and at the same time it can reproduce the correlation between stocks. The model is applied to data from Italian stock market from 1 January 2007 until the end of December 2010.

preprint2013arXiv

Performability analysis of the second order semi-Markov chains: an application to wind energy production

In this paper a general second order semi-Markov reward model is presented. Equations for the higher order moments of the reward process are presented for the first time and applied to wind energy production. The application is executed by considering a database, freely available from the web, that includes wind speed data taken from L.S.I. - Lastem station (Italy) and sampled every 10 minutes. We compute the expectation and the variance of the total energy produced by using the commercial blade Aircon HAWT - 10 kW.

preprint2013arXiv

Reliability measures for indexed semi-Markov chains applied to wind energy production

The computation of the dependability measures is a crucial point in the planning and development of a wind farm. In this paper we address the issue of energy production by wind turbine by using an indexed semi-Markov chain as a model of wind speed. We present the mathematical model, we describe the data and technical characteristics of a commercial wind turbine (Aircon HAWT-10kW). We show how to compute some of the main dependability measures such as reliability, availability and maintainability functions. We compare the results of the model with real energy production obtained from data available in the Lastem station (Italy) and sampled every 10 minutes.

preprint2013arXiv

Reliability measures of second order semi-Markov chain applied to wind energy production

In this paper we consider the problem of wind energy production by using a second order semi-Markov chain in state and duration as a model of wind speed. The model used in this paper is based on our previous work where we have showed the ability of second order semi-Markov process in reproducing statistical features of wind speed. Here we briefly present the mathematical model and describe the data and technical characteristics of a commercial wind turbine (Aircon HAWT-10kW). We show how, by using our model, it is possible to compute some of the main dependability measures such as reliability, availability and maintainability functions. We compare, by means of Monte Carlo simulations, the results of the model with real energy production obtained from data available in the Lastem station (Italy) and sampled every 10 minutes. The computation of the dependability measures is a crucial point in the planning and development of a wind farm. Through our model, we show how the values of this quantity can be obtained both analytically and computationally.

preprint2013arXiv

Wind speed forecasting at different time scales: a non parametric approach

The prediction of wind speed is one of the most important aspects when dealing with renewable energy. In this paper we show a new nonparametric model, based on semi-Markov chains, to predict wind speed. Particularly we use an indexed semi-Markov model, that reproduces accurately the statistical behavior of wind speed, to forecast wind speed one step ahead for different time scales and for very long time horizon maintaining the goodness of prediction. In order to check the main features of the model we show, as indicator of goodness, the root mean square error between real data and predicted ones and we compare our forecasting results with those of a persistence model.

preprint2012arXiv

First and second order semi-Markov chains for wind speed modeling

The increasing interest in renewable energy, particularly in wind, has given rise to the necessity of accurate models for the generation of good synthetic wind speed data. Markov chains are often used with this purpose but better models are needed to reproduce the statistical properties of wind speed data. We downloaded a database, freely available from the web, in which are included wind speed data taken from L.S.I. -Lastem station (Italy) and sampled every 10 minutes. With the aim of reproducing the statistical properties of this data we propose the use of three semi-Markov models. We generate synthetic time series for wind speed by means of Monte Carlo simulations. The time lagged autocorrelation is then used to compare statistical properties of the proposed models with those of real data and also with a synthetic time series generated though a simple Markov chain.

preprint2012arXiv

Weighted-indexed semi-Markov models for modeling financial returns

In this paper we propose a new stochastic model based on a generalization of semi-Markov chains to study the high frequency price dynamics of traded stocks. We assume that the financial returns are described by a weighted indexed semi-Markov chain model. We show, through Monte Carlo simulations, that the model is able to reproduce important stylized facts of financial time series as the first passage time distributions and the persistence of volatility. The model is applied to data from Italian and German stock market from first of January 2007 until end of December 2010.

preprint2012arXiv

Wind speed modeled as an indexed semi-Markov process

The increasing interest in renewable energy, particularly in wind, has given rise to the necessity of accurate models for the generation of good synthetic wind speed data. Markov chains are often used with this purpose but better models are needed to reproduce the statistical properties of wind speed data. In a previous paper we showed that semi-Markov processes are more appropriate for this purpose but to reach an accurate reproduction of real data features high order model should be used. In this work we introduce an indexed semi-Markov process that is able to fit real data. We downloaded a database, freely available from the web, in which are included wind speed data taken from L.S.I. -Lastem station (Italy) and sampled every 10 minutes. We then generate synthetic time series for wind speed by means of Monte Carlo simulations. The time lagged autocorrelation is then used to compare statistical properties of the proposed model with those of real data and also with a synthetic time series generated though a simple semi-Markov process.

preprint2011arXiv

A semi-Markov model for price returns

We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday return are described by a discrete time homogeneous semi-Markov process and the overnight returns are modeled by a Markov chain. Based on this assumptions we derived the equations for the first passage time distribution and the volatility autocorreletion function. Theoretical results have been compared with empirical findings from real data. In particular we analyzed high frequency data from the Italian stock market from first of January 2007 until end of December 2010. The semi-Markov hypothesis is also tested through a nonparametric test of hypothesis.

preprint2011arXiv

A semi-Markov model with memory for price changes

We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday returns are described by a discrete time homogeneous semi-Markov which depends also on a memory index. The index is introduced to take into account periods of high and low volatility in the market. First of all we derive the equations governing the process and then theoretical results have been compared with empirical findings from real data. In particular we analyzed high frequency data from the Italian stock market from first of January 2007 until end of December 2010.

preprint2009arXiv

Lexical evolution rates by automated stability measure

Phylogenetic trees can be reconstructed from the matrix which contains the distances between all pairs of languages in a family. Recently, we proposed a new method which uses normalized Levenshtein distances among words with same meaning and averages on all the items of a given list. Decisions about the number of items in the input lists for language comparison have been debated since the beginning of glottochronology. The point is that words associated to some of the meanings have a rapid lexical evolution. Therefore, a large vocabulary comparison is only apparently more accurate then a smaller one since many of the words do not carry any useful information. In principle, one should find the optimal length of the input lists studying the stability of the different items. In this paper we tackle the problem with an automated methodology only based on our normalized Levenshtein distance. With this approach, the program of an automated reconstruction of languages relationships is completed.

preprint2009arXiv

Measures of lexical distance between languages

The idea of measuring distance between languages seems to have its roots in the work of the French explorer Dumont D'Urville \cite{Urv}. He collected comparative words lists of various languages during his voyages aboard the Astrolabe from 1826 to 1829 and, in his work about the geographical division of the Pacific, he proposed a method to measure the degree of relation among languages. The method used by modern glottochronology, developed by Morris Swadesh in the 1950s, measures distances from the percentage of shared cognates, which are words with a common historical origin. Recently, we proposed a new automated method which uses normalized Levenshtein distance among words with the same meaning and averages on the words contained in a list. Recently another group of scholars \cite{Bak, Hol} proposed a refined of our definition including a second normalization. In this paper we compare the information content of our definition with the refined version in order to decide which of the two can be applied with greater success to resolve relationships among languages.

preprint2008arXiv

Statistical dynamics of religion evolutions

A religion affiliation can be considered as a "degree of freedom" of an agent on the human genre network. A brief review is given on the state of the art in data analysis and modelization of religious "questions" in order to suggest and if possible initiate further research, ... after using a "statistical physics filter". We present a discussion of the evolution of 18 so called religions, as measured through their number of adherents between 1900 and 2000. Some emphasis is made on a few cases presenting a minimum or a maximum in the investigated time range, - thereby suggesting a competitive ingredient to be considered, beside the well accepted "at birth" attachement effect. The importance of the "external field" is still stressed through an Avrami late stage crystal growth-like parameter. The observed features and some intuitive interpretations point to opinion based models with vector, rather than scalar, like agents.