Source author record

Vladimir Vovk

Vladimir Vovk appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

44works
11topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

44 published item(s)

preprint2026arXiv

Aggregation in conformal e-classification

Aggregating conformal predictors is a standard way of balancing their predictive and computational efficiency while retaining their validity, at least approximately. An important advantage of conformal e-predictors is that they are easier to aggregate without sacrificing their validity. This paper studies experimentally cross-conformal e-prediction, which is an existing method of aggregating conformal e-predictors, and its modifications that are conceptually simpler and more flexible.

preprint2026arXiv

Inductive Venn-Abers and related regressors

Venn-Abers predictors are probabilistic predictors that enjoy appealing properties of validity, but their major limitation is that they are applicable only to the case of binary classification, with a recent extension to bounded regression. We generalize them to the case of unbounded regression, which requires adding an element of conformal prediction. In our simulation and empirical studies we investigate the predictive efficiency of point regressors derived from Venn-Abers regressors and argue that they somewhat improve the predictive efficiency of standard regressors for larger training sets.

preprint2021arXiv

Retrain or not retrain: Conformal test martingales for change-point detection

We argue for supplementing the process of training a prediction algorithm by setting up a scheme for detecting the moment when the distribution of the data changes and the algorithm needs to be retrained. Our proposed schemes are based on exchangeability martingales, i.e., processes that are martingales under any exchangeable distribution for the data. Our method, based on conformal prediction, is general and can be applied on top of any modern prediction algorithm. Its validity is guaranteed, and in this paper we make first steps in exploring its efficiency.

preprint2020arXiv

A note on data splitting with e-values: online appendix to my comment on Glenn Shafer's "Testing by betting"

This note reanalyzes Cox's idealized example of testing with data splitting using e-values (Shafer's betting scores). Cox's exciting finding was that the method of data splitting, while allowing flexible data analysis, achieves quite high efficiencies, of about 80%. The most serious objection to the method was that it involves splitting data at random, and so different people analyzing the same data may get very different answers. Using e-values instead of p-values remedies this disadvantage.

preprint2020arXiv

Testing randomness

The hypothesis of randomness is fundamental in statistical machine learning and in many areas of nonparametric statistics; it says that the observations are assumed to be independent and coming from the same unknown probability distribution. This hypothesis is close, in certain respects, to the hypothesis of exchangeability, which postulates that the distribution of the observations is invariant with respect to their permutations. This paper reviews known methods of testing the two hypotheses concentrating on the online mode of testing, when the observations arrive sequentially. All known online methods for testing these hypotheses are based on conformal martingales, which are defined and studied in detail. The paper emphasizes conceptual and practical aspects and states two kinds of results. Validity results limit the probability of a false alarm or the frequency of false alarms for various procedures based on conformal martingales, including conformal versions of the CUSUM and Shiryaev-Roberts procedures. Efficiency results establish connections between randomness, exchangeability, and conformal martingales.

preprint2020arXiv

Training conformal predictors

Efficiency criteria for conformal prediction, such as \emph{observed fuzziness} (i.e., the sum of p-values associated with false labels), are commonly used to \emph{evaluate} the performance of given conformal predictors. Here, we investigate whether it is possible to exploit efficiency criteria to \emph{learn} classifiers, both conformal predictors and point classifiers, by using such criteria as training objective functions. The proposed idea is implemented for the problem of binary classification of hand-written digits. By choosing a 1-dimensional model class (with one real-valued free parameter), we can solve the optimization problems through an (approximate) exhaustive search over (a discrete version of) the parameter space. Our empirical results suggest that conformal predictors trained by minimizing their observed fuzziness perform better than conformal predictors trained in the traditional way by minimizing the \emph{prediction error} of the corresponding point classifier. They also have a reasonable performance in terms of their prediction error on the test set.

preprint2019arXiv

Non-algorithmic theory of randomness

This paper proposes an alternative language for expressing results of the algorithmic theory of randomness. The language is more precise in that it does not involve unspecified additive or multiplicative constants, making mathematical results, in principle, applicable in practice. Our main testing ground for the proposed language is the problem of defining Bernoulli sequences, which was of great interest to Andrei Kolmogorov and his students.

preprint2016arXiv

Another example of duality between game-theoretic and measure-theoretic probability

This paper makes a small step towards a non-stochastic version of superhedging duality relations in the case of one traded security with a continuous price path. Namely, we prove the coincidence of game-theoretic and measure-theoretic expectation for lower semicontinuous positive functionals. We consider a new broad definition of game-theoretic probability, leaving the older narrower definitions for future work.

preprint2016arXiv

Criteria of efficiency for conformal prediction

We study optimal conformity measures for various criteria of efficiency of classification in an idealised setting. This leads to an important class of criteria of efficiency that we call probabilistic; it turns out that the most standard criteria of efficiency used in literature on conformal prediction are not probabilistic unless the problem of classification is binary. We consider both unconditional and label-conditional conformal prediction.

preprint2016arXiv

Purely pathwise probability-free Ito integral

This paper gives several simple constructions of the pathwise Ito integral $\int_0^tϕdω$ for an integrand $ϕ$ and a price path $ω$ as integrator, with $ϕ$ and $ω$ satisfying various topological and analytical conditions. The definitions are purely pathwise in that neither $ϕ$ nor $ω$ are assumed to be paths of stochastic processes, and the Ito integral exists almost surely in a non-probabilistic financial sense. For example, one of the results shows the existence of $\int_0^tϕdω$ for a cadlag integrand $ϕ$ and a cadlag integrator $ω$ with jumps bounded in a predictable manner.

preprint2016arXiv

Rough paths in idealized financial markets

This paper considers possible price paths of a financial security in an idealized market. Its main result is that the variation index of typical price paths is at most 2, in this sense, typical price paths are not rougher than typical paths of Brownian motion. We do not make any stochastic assumptions and only assume that the price path is positive and right-continuous. The qualification "typical" means that there is a trading strategy (constructed explicitly in the proof) that risks only one monetary unit but brings infinite capital when the variation index of the realized price path exceeds 2. The paper also reviews some known results for continuous price paths and lists several open problems.

preprint2015arXiv

Continuous-time trading and the emergence of probability

This paper establishes a non-stochastic analogue of the celebrated result by Dubins and Schwarz about reduction of continuous martingales to Brownian motion via time change. We consider an idealized financial security with continuous price path, without making any stochastic assumptions. It is shown that typical price paths possess quadratic variation, where "typical" is understood in the following game-theoretic sense: there exists a trading strategy that earns infinite capital without risking more than one monetary unit if the process of quadratic variation does not exist. Replacing time by the quadratic variation process, we show that the price path becomes Brownian motion. This is essentially the same conclusion as in the Dubins-Schwarz result, except that the probabilities (constituting the Wiener measure) emerge instead of being postulated. We also give an elegant statement, inspired by Peter McCullagh's unpublished work, of this result in terms of game-theoretic probability theory.

preprint2015arXiv

Large-scale probabilistic predictors with and without guarantees of validity

This paper studies theoretically and empirically a method of turning machine-learning algorithms into probabilistic predictors that automatically enjoys a property of validity (perfect calibration) and is computationally efficient. The price to pay for perfect calibration is that these probabilistic predictors produce imprecise (in practice, almost precise for large data sets) probabilities. When these imprecise probabilities are merged into precise probabilities, the resulting predictors, while losing the theoretical property of perfect calibration, are consistently more accurate than the existing methods in empirical studies.

preprint2015arXiv

The fundamental nature of the log loss function

The standard loss functions used in the literature on probabilistic prediction are the log loss function, the Brier loss function, and the spherical loss function; however, any computable proper loss function can be used for comparison of prediction algorithms. This note shows that the log loss function is most selective in that any prediction algorithm that is optimal for a given data sequence (in the sense of the algorithmic theory of randomness) under the log loss function will be optimal under any computable proper mixable loss function; on the other hand, there is a data sequence and a prediction algorithm that is optimal for that sequence under either of the two other standard loss functions but not under the log loss function.

preprint2014arXiv

Efficiency of conformalized ridge regression

Conformal prediction is a method of producing prediction sets that can be applied on top of a wide range of prediction algorithms. The method has a guaranteed coverage probability under the standard IID assumption regardless of whether the assumptions (often considerably more restrictive) of the underlying algorithm are satisfied. However, for the method to be really useful it is desirable that in the case where the assumptions of the underlying algorithm are satisfied, the conformal predictor loses little in efficiency as compared with the underlying algorithm (whereas being a conformal predictor, it has the stronger guarantee of validity). In this paper we explore the degree to which this additional requirement of efficiency is satisfied in the case of Bayesian ridge regression; we find that asymptotically conformal prediction sets differ little from ridge regression prediction intervals when the standard Bayesian assumptions are satisfied.

preprint2014arXiv

Ito calculus without probability in idealized financial markets

We consider idealized financial markets in which price paths of the traded securities are cadlag functions, imposing mild restrictions on the allowed size of jumps. We prove the existence of quadratic variation for typical price paths, where the qualification "typical" means that there is a trading strategy that risks only one monetary unit and brings infinite capital if quadratic variation does not exist. This result allows one to apply numerous known results in pathwise Ito calculus to typical price paths; we give a brief overview of such results.

preprint2014arXiv

Prediction with Advice of Unknown Number of Experts

In the framework of prediction with expert advice, we consider a recently introduced kind of regret bounds: the bounds that depend on the effective instead of nominal number of experts. In contrast to the Normal- Hedge bound, which mainly depends on the effective number of experts but also weakly depends on the nominal one, we obtain a bound that does not contain the nominal number of experts at all. We use the defensive forecasting method and introduce an application of defensive forecasting to multivalued supermartingales.

preprint2014arXiv

Regression Conformal Prediction with Nearest Neighbours

In this paper we apply Conformal Prediction (CP) to the k-Nearest Neighbours Regression (k-NNR) algorithm and propose ways of extending the typical nonconformity measure used for regression so far. Unlike traditional regression methods which produce point predictions, Conformal Predictors output predictive regions that satisfy a given confidence level. The regions produced by any Conformal Predictor are automatically valid, however their tightness and therefore usefulness depends on the nonconformity measure used by each CP. In effect a nonconformity measure evaluates how strange a given example is compared to a set of other examples based on some traditional machine learning algorithm. We define six novel nonconformity measures based on the k-Nearest Neighbours Regression algorithm and develop the corresponding CPs following both the original (transductive) and the inductive CP approaches. A comparison of the predictive regions produced by our measures with those of the typical regression measure suggests that a major improvement in terms of predictive region tightness is achieved by the new measures.

preprint2014arXiv

Venn-Abers predictors

This paper continues study, both theoretical and empirical, of the method of Venn prediction, concentrating on binary prediction problems. Venn predictors produce probability-type predictions for the labels of test objects which are guaranteed to be well calibrated under the standard assumption that the observations are generated independently from the same distribution. We give a simple formalization and proof of this property. We also introduce Venn-Abers predictors, a new class of Venn predictors based on the idea of isotonic regression, and report promising empirical results both for Venn-Abers predictors and for their more computationally efficient simplified version.

preprint2012arXiv

Conditional validity of inductive conformal predictors

Conformal predictors are set predictors that are automatically valid in the sense of having coverage probability equal to or exceeding a given confidence level. Inductive conformal predictors are a computationally efficient version of conformal predictors satisfying the same property of validity. However, inductive conformal predictors have been only known to control unconditional coverage probability. This paper explores various versions of conditional validity and various ways to achieve them using inductive conformal predictors and their modifications.

preprint2012arXiv

On-line Prediction with Kernels and the Complexity Approximation Principle

The paper describes an application of Aggregating Algorithm to the problem of regression. It generalizes earlier results concerned with plain linear regression to kernel techniques and presents an on-line algorithm which performs nearly as well as any oblivious kernel predictor. The paper contains the derivation of an estimate on the performance of this algorithm. The estimate is then used to derive an application of the Complexity Approximation Principle to kernel methods.

preprint2012arXiv

Plug-in martingales for testing exchangeability on-line

A standard assumption in machine learning is the exchangeability of data, which is equivalent to assuming that the examples are generated from the same probability distribution independently. This paper is devoted to testing the assumption of exchangeability on-line: the examples arrive one by one, and after receiving each example we would like to have a valid measure of the degree to which the assumption of exchangeability has been falsified. Such measures are provided by exchangeability martingales. We extend known techniques for constructing exchangeability martingales and show that our new method is competitive with the martingales introduced before. Finally we investigate the performance of our testing method on two benchmark datasets, USPS and Statlog Satellite data; for the former, the known techniques give satisfactory results, but for the latter our new more flexible method becomes necessary.

preprint2011arXiv

A simplified Capital Asset Pricing Model

We consider a Black-Scholes market in which a number of stocks and an index are traded. The simplified Capital Asset Pricing Model is the conjunction of the usual Capital Asset Pricing Model, or CAPM, and the statement that the appreciation rate of the index is equal to its squared volatility plus the interest rate. (The mathematical statement of the conjunction is simpler than that of the usual CAPM.) Our main result is that either we can outperform the index or the simplified CAPM holds.

preprint2011arXiv

Losing money with a high Sharpe ratio

A simple example shows that losing all money is compatible with a very high Sharpe ratio (as computed after losing all money). However, the only way that the Sharpe ratio can be high while losing money is that there is a period in which all or almost all money is lost. This note explores the best achievable Sharpe and Sortino ratios for investors who lose money but whose one-period returns are bounded below (or both below and above) by a known constant.

preprint2011arXiv

On-line predictive linear regression

We consider the on-line predictive version of the standard problem of linear regression; the goal is to predict each consecutive response given the corresponding explanatory variables and all the previous observations. We are mainly interested in prediction intervals rather than point predictions. The standard treatment of prediction intervals in linear regression analysis has two drawbacks: (1) the classical prediction intervals guarantee that the probability of error is equal to the nominal significance level epsilon, but this property per se does not imply that the long-run frequency of error is close to epsilon; (2) it is not suitable for prediction of complex systems as it assumes that the number of observations exceeds the number of parameters. We state a general result showing that in the on-line protocol the frequency of error for the classical prediction intervals does equal the nominal significance level, up to statistical fluctuations. We also describe alternative regression models in which informative prediction intervals can be found before the number of observations exceeds the number of parameters. One of these models, which only assumes that the observations are independent and identically distributed, is popular in machine learning but greatly underused in the statistical theory of regression.

preprint2011arXiv

Probability-free pricing of adjusted American lookbacks

Consider an American option that pays G(X^*_t) when exercised at time t, where G is a positive increasing function, X^*_t := \sup_{s\le t}X_s, and X_s is the price of the underlying security at time s. Assuming zero interest rates, we show that the seller of this option can hedge his position by trading in the underlying security if he begins with initial capital X_0\int_{X_0}^{\infty}G(x)x^{-2}dx (and this is the smallest initial capital that allows him to hedge his position). This leads to strategies for trading that are always competitive both with a given strategy's current performance and, to a somewhat lesser degree, with its best performance so far. It also leads to methods of statistical testing that avoid sacrificing too much of the maximum statistical significance that they achieve in the course of accumulating data.

preprint2011arXiv

Test Martingales, Bayes Factors and $p$-Values

A nonnegative martingale with initial value equal to one measures evidence against a probabilistic hypothesis. The inverse of its value at some stopping time can be interpreted as a Bayes factor. If we exaggerate the evidence by considering the largest value attained so far by such a martingale, the exaggeration will be limited, and there are systematic ways to eliminate it. The inverse of the exaggerated value at some stopping time can be interpreted as a $p$-value. We give a simple characterization of all increasing functions that eliminate the exaggeration.

preprint2011arXiv

The Capital Asset Pricing Model as a corollary of the Black-Scholes model

We consider a financial market in which two securities are traded: a stock and an index. Their prices are assumed to satisfy the Black-Scholes model. Besides assuming that the index is a tradable security, we also assume that it is efficient, in the following sense: we do not expect a prespecified self-financing trading strategy whose wealth is almost surely nonnegative at all times to outperform the index greatly. We show that, for a long investment horizon, the appreciation rate of the stock has to be close to the interest rate (assumed constant) plus the covariance between the volatility vectors of the stock and the index. This contains both a version of the Capital Asset Pricing Model and our earlier result that the equity premium is close to the squared volatility of the index.

preprint2011arXiv

The efficient index hypothesis and its implications in the BSM model

This note studies the behavior of an index I_t which is assumed to be a tradable security, to satisfy the BSM model dI_t/I_t = μdt + σdW_t, and to be efficient in the following sense: we do not expect a prespecified trading strategy whose value is almost surely always nonnegative to outperform the index greatly. The efficiency of the index imposes severe restrictions on its growth rate; in particular, for a long investment horizon we should have μ\approx r+σ^2, where r is the interest rate. This provides another partial solution to the equity premium puzzle. All our mathematical results are extremely simple.

preprint2010arXiv

Insuring against loss of evidence in game-theoretic probability

We consider the game-theoretic scenario of testing the performance of Forecaster by Sceptic who gambles against the forecasts. Sceptic's current capital is interpreted as the amount of evidence he has found against Forecaster. Reporting the maximum of Sceptic's capital so far exaggerates the evidence. We characterize the set of all increasing functions that remove the exaggeration. This result can be used for insuring against loss of evidence.

preprint2010arXiv

Prediction with Advice of Unknown Number of Experts

In the framework of prediction with expert advice, we consider a recently introduced kind of regret bounds: the bounds that depend on the effective instead of nominal number of experts. In contrast to the NormalHedge bound, which mainly depends on the effective number of experts and also weakly depends on the nominal one, we obtain a bound that does not contain the nominal number of experts at all. We use the defensive forecasting method and introduce an application of defensive forecasting to multivalued supermartingales.

preprint2010arXiv

Supermartingales in Prediction with Expert Advice

We apply the method of defensive forecasting, based on the use of game-theoretic supermartingales, to prediction with expert advice. In the traditional setting of a countable number of experts and a finite number of outcomes, the Defensive Forecasting Algorithm is very close to the well-known Aggregating Algorithm. Not only the performance guarantees but also the predictions are the same for these two methods of fundamentally different nature. We discuss also a new setting where the experts can give advice conditional on the learner's future decision. Both the algorithms can be adapted to the new setting and give the same performance guarantees as in the traditional setting. Finally, we outline an application of defensive forecasting to a setting with several loss functions.

preprint2007arXiv

Continuous-time trading and emergence of randomness

A new definition of events of game-theoretic probability zero in continuous time is proposed and used to prove results suggesting that trading in financial markets results in the emergence of properties usually associated with randomness. This paper concentrates on "qualitative" results, stated in terms of order (or order topology) rather than in terms of the precise values taken by the price processes (assumed continuous).

preprint2006arXiv

Hedging predictions in machine learning

Recent advances in machine learning make it possible to design efficient prediction algorithms for data sets with huge numbers of parameters. This paper describes a new technique for "hedging" the predictions output by many such algorithms, including support vector machines, kernel ridge regression, kernel nearest neighbours, and by many other state-of-the-art methods. The hedged predictions for the labels of new objects include quantitative measures of their own accuracy and reliability. These measures are provably valid under the assumption of randomness, traditional in machine learning: the objects and their labels are assumed to be generated independently from the same probability distribution. In particular, it becomes possible to control (up to statistical fluctuations) the number of erroneous predictions by selecting a suitable confidence level. Validity being achieved automatically, the remaining goal of hedged prediction is efficiency: taking full account of the new objects' features and other available information to produce as accurate predictions as possible. This can be done successfully using the powerful machinery of modern machine learning.