Source author record

James Allen Fill

James Allen Fill appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

15works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

15 published item(s)

preprint2025arXiv

Convergence of the QuickVal Residual

QuickSelect (aka Find), introduced by Hoare (1961), is a randomized algorithm for selecting a specified order statistic from an input sequence of $n$ objects, or rather their identifying labels usually known as keys. The keys can be numeric or symbol strings, or indeed any labels drawn from a given linearly ordered set. We discuss various ways in which the cost of comparing two keys can be measured, and we can measure the efficiency of the algorithm by the total cost of such comparisons. We define and discuss a closely related algorithm known as QuickVal and a natural probabilistic model for the input to this algorithm; QuickVal searches (almost surely unsuccessfully) for a specified population quantile $α\in [0, 1]$ in an input sample of size $n$. Call the total cost of comparisons for this algorithm $S_n$. We discuss a natural way to define the random variables $S_1, S_2, \ldots$ on a common probability space. For a general class of cost functions, Fill and Nakama (2013) proved under mild assumptions that the scaled cost $S_n / n$ of QuickVal converges in $L^p$ and almost surely to a limit random variable $S$. For a general cost function, we consider what we term the QuickVal residual: \[ρ_n := \frac{S_n}n - S.\] The residual is of natural interest, especially in light of the previous analogous work on the sorting algorithm QuickSort. In the case $α= 0$ of QuickMin with unit cost per key-comparison, we are able to calculate -- à la Bindjeme and Fill (2012) for QuickSort -- the exact (and asymptotic) $L^2$-norm of the residual. We take the result as motivation for the scaling factor $\sqrt{n}$ for the QuickVal residual for general population quantiles and for general cost. We then prove in general (under mild conditions on the cost function) that $\sqrt{n}\,ρ_n$ converges in law to a scale-mixture of centered Gaussians, and we also prove convergence of moments.

preprint2015arXiv

Strong Stationary Duality for Diffusion Processes

We develop the theory of strong stationary duality for diffusion processes on compact intervals. We analytically derive the generator and boundary behavior of the dual process and recover a central tenet of the classical Markov chain theory in the diffusion setting by linking the separation distance in the primal diffusion to the absorption time in the dual diffusion. We also exhibit our strong stationary dual as the natural limiting process of the strong stationary dual sequence of a well chosen sequence of approximating birth-and-death Markov chains, allowing for simultaneous numerical simulations of our primal and dual diffusion processes. Lastly, we show how our new definition of diffusion duality allows the spectral theory of cutoff phenomena to extend naturally from birth-and-death Markov chains to the present diffusion context.

preprint2013arXiv

Comparison inequalities and fastest-mixing Markov chains

We introduce a new partial order on the class of stochastically monotone Markov kernels having a given stationary distribution $π$ on a given finite partially ordered state space $\mathcal{X}$. When $K\preceq L$ in this partial order we say that $K$ and $L$ satisfy a comparison inequality. We establish that if $K_1,\ldots,K_t$ and $L_1,\ldots,L_t$ are reversible and $K_s\preceq L_s$ for $s=1,\ldots,t$, then $K_1\cdots K_t\preceq L_1\cdots L_t$. In particular, in the time-homogeneous case we have $K^t\preceq L^t$ for every $t$ if $K$ and $L$ are reversible and $K\preceq L$, and using this we show that (for suitable common initial distributions) the Markov chain $Y$ with kernel $K$ mixes faster than the chain $Z$ with kernel $L$, in the strong sense that at every time $t$ the discrepancy - measured by total variation distance or separation or $L^2$-distance - between the law of $Y_t$ and $π$ is smaller than that between the law of $Z_t$ and $π$. Using comparison inequalities together with specialized arguments to remove the stochastic monotonicity restriction, we answer a question of Persi Diaconis by showing that, among all symmetric birth-and-death kernels on the path $\mathcal{X}=\{0,\ldots,n\}$, the one (we call it the uniform chain) that produces fastest convergence from initial state 0 to the uniform distribution has transition probability 1/2 in each direction along each edge of the path, with holding probability 1/2 at each endpoint.

preprint2013arXiv

Distributional convergence for the number of symbol comparisons used by QuickSort

Most previous studies of the sorting algorithm QuickSort have used the number of key comparisons as a measure of the cost of executing the algorithm. Here we suppose that the n independent and identically distributed (i.i.d.) keys are each represented as a sequence of symbols from a probabilistic source and that QuickSort operates on individual symbols, and we measure the execution cost as the number of symbol comparisons. Assuming only a mild "tameness" condition on the source, we show that there is a limiting distribution for the number of symbol comparisons after normalization: first centering by the mean and then dividing by n. Additionally, under a condition that grows more restrictive as p increases, we have convergence of moments of orders p and smaller. In particular, we have convergence in distribution and convergence of moments of every order whenever the source is memoryless, that is, whenever each key is generated as an infinite string of i.i.d. symbols. This is somewhat surprising; even for the classical model that each key is an i.i.d. string of unbiased ("fair") bits, the mean exhibits periodic fluctuations of order n.

preprint2012arXiv

Distributional convergence for the number of symbol comparisons used by QuickSelect

When the search algorithm QuickSelect compares keys during its execution in order to find a key of target rank, it must operate on the keys' representations or internal structures, which were ignored by the previous studies that quantified the execution cost for the algorithm in terms of the number of required key comparisons. In this paper, we analyze running costs for the algorithm that take into account not only the number of key comparisons but also the cost of each key comparison. We suppose that keys are represented as sequences of symbols generated by various probabilistic sources and that QuickSelect operates on individual symbols in order to find the target key. We identify limiting distributions for the costs and derive integral and series expressions for the expectations of the limiting distributions. These expressions are used to recapture previously obtained results on the number of key comparisons required by the algorithm.

preprint2012arXiv

Exact L^2-distance from the limit for QuickSort key comparisons (extended abstract)

Using a recursive approach, we obtain a simple exact expression for the L^2-distance from the limit in Régnier's (1989) classical limit theorem for the number of key comparisons required by QuickSort. A previous study by Fill and Janson (2002) using a similar approach found that the d_2-distance is of order between n^{-1} log n and n^{-1/2}, and another by Neininger and Ruschendorf (2002) found that the Zolotarev zeta_3-distance is of exact order n^{-1} log n. Our expression reveals that the L^2-distance is asymptotically equivalent to (2 n^{-1} ln n)^{1/2}.

preprint2012arXiv

Hitting times and interlacing eigenvalues: a stochastic approach using intertwinings

We develop a systematic matrix-analytic approach, based on intertwinings of Markov semigroups, for proving theorems about hitting-time distributions for finite-state Markov chains -- an approach that (sometimes) deepens understanding of the theorems by providing corresponding sample-path-by-sample-path stochastic constructions. We employ our approach to give new proofs and constructions for two theorems due to Mark Brown, theorems giving two quite different representations of hitting-time distributions for finite-state Markov chains started in stationarity. The proof, and corresponding construction, for one of the two theorems elucidates an intriguing connection between hitting-time distributions and the interlacing eigenvalues theorem for bordered symmetric matrices.

preprint2012arXiv

The limiting distribution for the number of symbol comparisons used by QuickSort is nondegenerate (extended abstract)

In a continuous-time setting, Fill (2010) proved, for a large class of probabilistic sources, that the number of symbol comparisons used by QuickSort, when centered by subtracting the mean and scaled by dividing by time, has a limiting distribution, but proved little about that limiting random variable Y -- not even that it is nondegenerate. We establish the nondegeneracy of Y. The proof is perhaps surprisingly difficult.

preprint2012arXiv

The number of bit comparisons used by Quicksort: an average-case analysis

The analyses of many algorithms and data structures (such as digital search trees) for searching and sorting are based on the representation of the keys involved as bit strings and so count the number of bit comparisons. On the other hand, the standard analyses of many other algorithms (such as Quicksort) are performed in terms of the number of key comparisons. We introduce the prospect of a fair comparison between algorithms of the two types by providing an average-case analysis of the number of bit comparisons required by Quicksort. Counting bit comparisons rather than key comparisons introduces an extra logarithmic factor to the asymptotic average total. We also provide a new algorithm, "BitsQuick", that reduces this factor to constant order by eliminating needless bit comparisons.

preprint2010arXiv

On vertex, edge, and vertex-edge random graphs

We consider three classes of random graphs: edge random graphs, vertex random graphs, and vertex-edge random graphs. Edge random graphs are Erdos-Renyi random graphs, vertex random graphs are generalizations of geometric random graphs, and vertex-edge random graphs generalize both. The names of these three types of random graphs describe where the randomness in the models lies: in the edges, in the vertices, or in both. We show that vertex-edge random graphs, ostensibly the most general of the three models, can be approximated arbitrarily closely by vertex random graphs, but that the two categories are distinct.

preprint2008arXiv

Two-player Knock 'em Down

We analyze the two-player game of Knock 'em Down, asymptotically as the number of tokens to be knocked down becomes large. Optimal play requires mixed strategies with deviations of order sqrt(n) from the naive law-of-large numbers allocation. Upon rescaling by sqrt(n) and sending n to infinity, we show that optimal play's random deviations always have bounded support and have marginal distributions that are absolutely continuous with respect to Lebesgue measure.

preprint2000arXiv

Mixing times for Markov chains on wreath products and related homogeneous spaces

We develop a method for analyzing the mixing times for a quite general class of Markov chains on the complete monomial group G \wr S_n (the wreath product of a group G with the permutation group S_n) and a quite general class of Markov chains on the homogeneous space (G \wr S_n) / (S_r \times S_{n - r}). We derive an exact formula for the L^2 distance in terms of the L^2 distances to uniformity for closely related random walks on the symmetric groups S_j for 1 \leq j \leq n or for closely related Markov chains on the homogeneous spaces S_{i + j} / (S_i \times S_j) for various values of i and j, respectively. Our results are consistent with those previously known, but our method is considerably simpler and more general.