Source author record

Steven J. Miller

Steven J. Miller appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

114works
18topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

114 published item(s)

preprint2022arXiv

$k$-Diophantine $m$-tuples in Finite Fields

In this paper, we define a $k$-Diophantine $m$-tuple to be a set of $m$ positive integers such that the product of any $k$ distinct positive integers is one less than a perfect square. We study these sets in finite fields $\mathbb{F}_p$ for odd prime $p$ and guarantee the existence of a $k$-Diophantine m-tuple provided $p$ is larger than some explicit lower bound. We also give a formula for the number of 3-Diophantine triples in $\mathbb{F}_p$ as well as an asymptotic formula for the number of $k$-Diophantine $k$-tuples.

preprint2022arXiv

Bounding Vanishing at the Central Point of Cuspidal Newforms

The Katz-Sarnak Density Conjecture states that zeros of families of $L$-functions are well-modeled by eigenvalues of random matrix ensembles. For suitably restricted test functions, this correspondence yields upper bounds for the families' order of vanishing at the central point. We generalize previous results on the $n$\textsuperscript{th} centered moment of the distribution of zeros to allow arbitrary test functions. On the computational side, we use our improved formulas to obtain significantly better bounds on the order of vanishing for cuspidal newforms, setting world records for the quality of the bounds. We also discover better test functions that further optimize our bounds. We see improvement as early as the $5$\textsuperscript{th} order, and our bounds improve rapidly as the rank grows (more than one order of magnitude better for rank 10 and more than four orders of magnitude for rank 50).

preprint2022arXiv

Class Numbers and Pell's Equation $x^2 + 105y^2 = z^2$

Two well-studied Diophantine equations are those of Pythagorean triples and elliptic curves, for the first we have a parametrization through rational points on the unit circle, and for the second we have a structure theorem for the group of rational solutions. Recently, Yekutieli discussed a connection between these two problems, and described the group structure of Pythagorean triples and the number of triples for a given hypotenuse. In arXiv:2112.03663 we generalized these methods and results to Pell's equation. We find a similar group structure and count on the number of solutions for a given $z$ to $x^2 + Dy^2 = z^2$ when $D$ is 1 or 2 modulo 4 and the class group of $\mathbb{Q}[\sqrt{-D}]$ is a free $\mathbb{Z}_2$ module, which always happens if the class number is at most 2. In this paper, we discuss the main results of arXiv:2112.03663 using some concrete examples in the case of $D=105$.

preprint2022arXiv

Modeling Random Walks to Infinity on Primes in $\mathbb{Z}[\sqrt{2}]$

An interesting question, known as the Gaussian moat problem, asks whether it is possible to walk to infinity on Gaussian primes with steps of bounded length. Our work examines a similar situation in the real quadratic integer ring $\mathbb{Z}[\sqrt{2}]$ whose primes cluster near the asymptotes $y = \pm x/\sqrt{2}$ as compared to Gaussian primes, which cluster near the origin. We construct a probabilistic model of primes in $\mathbb{Z}[\sqrt{2}]$ by applying the prime number theorem and a combinatorial theorem for counting the number of lattice points whose absolute values of their norms are at most $r^2$. We then prove that it is impossible to walk to infinity if the walk remains within some bounded distance from the asymptotes. Lastly, we perform a few moat calculations to show that the longest walk is likely to stay close to the asymptotes; hence, we conjecture that there is no walk to infinity on $\mathbb{Z}[\sqrt{2}]$ primes with steps of bounded length.

preprint2022arXiv

Walking to Infinity on the Fibonacci Sequence

An interesting open problem in number theory asks whether it is possible to walk to infinity on primes, where each term in the sequence has one more digit than the previous. In this paper, we study its variation where we walk on the Fibonacci sequence. We prove that all walks starting with a Fibonacci number and the following terms are Fibonacci numbers obtained by appending exactly one digit at a time to the right have a length of at most two. In the more general case where we append at most a bounded number of digits each time, we give a formula for the length of the longest walk.

preprint2021arXiv

A Refined Conjecture for the Variance of Gaussian Primes Across Sectors

We derive a refined conjecture for the variance of Gaussian primes across sectors, with a power saving error term, by applying the L-functions Ratios Conjecture. We observe a bifurcation point in the main term, consistent with the Random Matrix Theory (RMT) heuristic previously proposed by Rudnick and Waxman. Our model also identifies a second bifurcation point, undetected by the RMT model, that emerges upon taking into account lower order terms. For sufficiently small sectors, we moreover prove an unconditional result that is consistent with our conjecture down to lower order terms.

preprint2021arXiv

Distribution of Eigenvalues of Matrix Ensembles arising from Wigner and Palindromic Toeplitz Blocks

Random Matrix Theory (RMT) has successfully modeled diverse systems, from energy levels of heavy nuclei to zeros of $L$-functions; this correspondence has allowed RMT to successfully predict many number theoretic behaviors. However there are some operations which to date have no RMT analogue. Our motivation is to find an RMT analogue of Rankin-Selberg convolution, which constructs a new $L$-functions from an input pair. We report one such attempt; while it does not appear to model convolution, it does create new ensembles with properties hybridizing those of its constituents. For definiteness we concentrate on the ensemble of palindromic real symmetric Toeplitz (PST) matrices and the ensemble of real symmetric matrices, whose limiting spectral measures are the Gaussian and semi-circular distributions, respectively; these were chosen as they are the two extreme cases in terms of moment calculations. For a PST matrix $A$ and a real symmetric matrix $B$, we construct an ensemble of random real symmetric block matrices whose first row is $\lbrace A, B \rbrace$ and whose second row is $\lbrace B, A \rbrace$. By Markov's Method of Moments and the use of free probability, we show this ensemble converges weakly and almost surely to a new, universal distribution with a hybrid of Gaussian and semi-circular behaviors. We extend this construction by considering an iterated concatenation of matrices from an arbitrary pair of random real symmetric sub-ensembles with different limiting spectral measures. We prove that finite iterations converge to new, universal distributions with hybrid behavior, and that infinite iterations converge to the limiting spectral measure of the dominant component matrix.

preprint2021arXiv

Limiting Spectral Distributions of Families of Block Matrix Ensembles

We introduce a new matrix operation on a pair of matrices, $\text{swirl}(A,X),$ and discuss its implications on the limiting spectral distribution. In a special case, the resultant ensemble converges almost surely to the Rayleigh distribution. In proving this, we provide a novel combinatorial proof that the random matrix ensemble of circulant Hankel matrices converges almost surely to the Rayleigh distribution, using the method of moments.

preprint2021arXiv

Tinkering with Lattices: A New Take on the Erdős Distance Problem

The Erdős distance problem concerns the least number of distinct distances that can be determined by $N$ points in the plane. The integer lattice with $N$ points is known as \textit{near-optimal}, as it spans $Θ(N/\sqrt{\log(N)})$ distinct distances, the lower bound for a set of $N$ points (Erdős, 1946). The only previous non-asymptotic work related to the Erdős distance problem that has been done was for $N \leq 13$. We take a new non-asymptotic approach to this problem in a model case, studying the distance distribution, or in other words, the plot of frequencies of each distance of the $N\times N$ integer lattice. In order to fully characterize this distribution, we adapt previous number-theoretic results from Fermat and Erdős in order to relate the frequency of a given distance on the lattice to the sum-of-squares formula. We study the distance distributions of all the lattice's possible subsets; although this is a restricted case, the structure of the integer lattice allows for the existence of subsets which can be chosen so that their distance distributions have certain properties, such as emulating the distribution of randomly distributed sets of points for certain small subsets, or emulating that of the larger lattice itself. We define an error which compares the distance distribution of a subset with that of the full lattice. The structure of the integer lattice allows us to take subsets with certain geometric properties in order to maximize error; we show these geometric constructions explicitly. Further, we calculate explicit upper bounds for the error when the number of points in the subset is $4$, $5$, $9$ or $\left \lceil N^2/2\right\rceil$ and prove a lower bound in cases with a small number of points.

preprint2020arXiv

Bounds on Zeckendorf Games

Zeckendorf proved that every positive integer $n$ can be written uniquely as the sum of non-adjacent Fibonacci numbers. We use this decomposition to construct a two-player game. Given a fixed integer $n$ and an initial decomposition of $n=n F_1$, the two players alternate by using moves related to the recurrence relation $F_{n+1}=F_n+F_{n-1}$, and whoever moves last wins. The game always terminates in the Zeckendorf decomposition; depending on the choice of moves the length of the game and the winner can vary, though for $n\ge 2$ there is a non-constructive proof that Player 2 has a winning strategy. Initially the lower bound of the length of a game was order $n$ (and known to be sharp) while the upper bound was of size $n \log n$. Recent work decreased the upper bound to of size $n$, but with a larger constant than was conjectured. We improve the upper bound and obtain the sharp bound of $\frac{\sqrt{5}+3}{2}\ n - IZ(n) - \frac{1+\sqrt{5}}{2}Z(n)$, which is of order $n$ as $Z(n)$ is the number of terms in the Zeckendorf decomposition of $n$ and $IZ(n)$ is the sum of indices in the Zeckendorf decomposition of $n$ (which are at most of sizes $\log n$ and $\log^2 n$ respectively). We also introduce a greedy algorithm that realizes the upper bound, and show that the longest game on any $n$ is achieved by applying splitting moves whenever possible.

preprint2020arXiv

Central Limit Theorems for Compound Paths on the 2-Dimensional Lattice

Zeckendorf proved that every integer can be written uniquely as a sum of non-consecutive Fibonacci numbers $\{F_n\}$, and later researchers showed that the distribution of the number of summands needed for such decompositions of integers in $[F_n, F_{n+1})$ converges to a Gaussian as $n\to\infty$. Decomposition problems have been studied extensively for a variety of different sequences and notions of a legal decompositions; for the Fibonacci numbers, a legal decomposition is one for which each summand is used at most once and no two consecutive summands may be chosen. Recently, Chen et al. [CCGJMSY] generalized earlier work to $d$-dimensional lattices of positive integers; there, a legal decomposition is a path such that every point chosen had each component strictly less than the component of the previous chosen point in the path. They were able to prove Gaussianity results despite the lack of uniqueness of the decompositions; however, their results should hold in the more general case where some components are identical. The strictly decreasing assumption was needed in that work to obtain simple, closed form combinatorial expressions, which could then be well approximated and led to the limiting behavior. In this work we remove that assumption through inclusion-exclusion arguments. These lead to more involved combinatorial sums; using generating functions and recurrence relations we obtain tractable forms in $2$ dimensions and prove Gaussianity again; a more involved analysis should work in higher dimensions.

preprint2020arXiv

Constructions of Generalized MSTD Sets in Higher Dimensions

Let $A$ be a set of finite integers, define $$A+A \ = \ \{a_1+a_2: a_1,a_2 \in A\}, \ \ \ A-A \ = \ \{a_1-a_2: a_1,a_2 \in A\},$$ and for non-negative integers $s$ and $d$ define $$sA-dA\ =\ \underbrace{A+\cdots+A}_{s} -\underbrace{A-\cdots-A}_{d}.$$ A More Sums than Differences (MSTD) set is an $A$ where $|A+A| > |A-A|$. It was initially thought that the percentage of subsets of $[0,n]$ that are MSTD would go to zero as $n$ approaches infinity as addition is commutative and subtraction is not. However, in a surprising 2006 result, Martin and O'Bryant proved that a positive percentage of sets are MSTD, although this percentage is extremely small, about $10^{-4}$ percent. This result was extended by Iyer, Lazarev, Miller, ans Zhang [ILMZ] who showed that a positive percentage of sets are generalized MSTD sets, sets for $\{s_1,d_1\} \neq \{s_2, d_2\}$ and $s_1+d_1=s_2+d_2$ with $|s_1A-d_1A| > |s_2A-d_2A|$, and that in $d$-dimensions, a positive percentage of sets are MSTD. For many such results, establishing explicit MSTD sets in $1$-dimensions relies on the specific choice of the elements on the left and right fringes of the set to force certain differences to be missed while desired sums are attained. In higher dimensions, the geometry forces a more careful assessment of what elements have the same behavior as $1$-dimensional fringe elements. We study fringes in $d$-dimensions and use these to create new explicit constructions. We prove the existence of generalized MSTD sets in $d$-dimensions and the existence of $k$-generational sets, which are sets where $|cA+cA|>|cA-cA|$ for all $1\leq c \leq k$. We then prove that under certain conditions, there are no sets with $|kA+kA|>|kA-kA|$ for all $k \in \mathbb{N}.$

preprint2020arXiv

Deterministic Zeckendorf Games

Zeckendorf proved that every positive integer can be written uniquely as the sum of non-adjacent Fibonacci numbers. We further explore a two-player Zeckendorf game introduced in Baird-Smith, Epstein, Flint, and Miller: Given a fixed integer $n$ and an initial decomposition of $n = nF_1$, players alternate using moves related to the recurrence relation $F_{n+1} = F_n + F_{n_1}$, and the last player to move wins. We improve the upper bound on the number of moves possible and show that it is of the same order in $n$ as the lower bound; this is an improvement by a logarithm over previous work. The new upper bound is $3n - 3Z(n) - IZ(n) + 1$, and the existing lower bound is sharp at $n - Z(n)$ moves, where $Z(n)$ is the number of terms in the Zeckendorf decomposition of $n$ and $IZ(n)$ is the sum of indices in the same Zeckendorf decomposition of $n$. We also studied four deterministic variants of the game, where there was a fixed order on which available move one takes: Combine Largest, Split Largest, Combine Smallest and Split Smallest. We prove that Combine Largest and Split Largest realize the lower bound. Split Smallest has the largest number of moves over all possible games, and is close to the new upper bound. For Combine Split games, the number of moves grows linearly with $n$.

preprint2020arXiv

Distribution of missing differences in diffsets

Lazarev, Miller and O'Bryant investigated the distribution of $|S+S|$ for $S$ chosen uniformly at random from $\{0, 1, \dots, n-1\}$, and proved the existence of a divot at missing 7 sums (the probability of missing exactly 7 sums is less than missing 6 or missing 8 sums). We study related questions for $|S-S|$, and shows some divots from one end of the probability distribution, $P(|S-S|=k)$, as well as a peak at $k=4$ from the other end, $P(2n-1-|S-S|=k)$. A corollary of our results is an asymptotic bound for the number of complete rulers of length $n$.

preprint2020arXiv

Extensions of Autocorrelation Inequalities with Applications to Additive Combinatorics

In a 2019 paper, Barnard and Steinerberger show that for $f\in L^1(\mathbf{R})$, the following autocorrelation inequality holds: \begin{equation*} \min_{0 \leq t \leq 1} \int_\mathbf{R} f(x) f(x+t)\ \mathrm{d}x \ \leq\ 0.411 ||f||_{L^1}^2, \end{equation*} where the constant $0.411$ cannot be replaced by $0.37$. In addition to being interesting and important in their own right, inequalities such as these have applications in additive combinatorics where some problems, such as those of minimal difference basis, can be encapsulated by a convolution inequality similar to the above integral. Barnard and Steinerberger suggest that future research may focus on the existence of functions extremizing the above inequality (which is itself related to Brascamp-Lieb type inequalities). We show that for $f$ to be extremal under the above, we must have \begin{equation*} \max_{x_1 \in \mathbf{R} }\min_{0 \leq t \leq 1} \left[ f(x_1-t)+f(x_1+t) \right] \ \leq\ \min_{x_2 \in \mathbf{R} } \max_{0 \leq t \leq 1} \left[ f(x_2-t)+f(x_2+t) \right] . \end{equation*} Our central technique for deriving this result is local perturbation of $f$ to increase the value of the autocorrelation, while leaving $||f||_{L^1}$ unchanged. These perturbation methods can be extended to examine a more general notion of autocorrelation. Let $d,n \in \mathbb{Z}^+$, $f \in L^1$, $A$ be a $d \times n$ matrix with real entries and columns $a_i$ for $1 \leq i \leq n$, and $C$ be a constant. For a broad class of matrices $A$, we prove necessary conditions for $f$ to extremize autocorrelation inequalities of the form \begin{equation*} \min_{ \mathbf{t} \in [0,1]^d } \int_{\mathbf{R}} \prod_{i=1}^n\ f(x+ \mathbf{t} \cdot a_i)\ \mathrm{d}x\ \leq\ C ||f||_{L^1}^n. \end{equation*}

preprint2020arXiv

On the sum of $k$-th powers in terms of earlier sums

For $k$ a positive integer let $S_k(n) = 1^k + 2^k + \cdots + n^k$, i.e., $S_k(n)$ is the sum of the first $k$-th powers. Faulhaber conjectured (later proved by Jacobi) that for $k$ odd, $S_k(n)$ could be written as a polynomial of $S_1(n)$; for example $S_3(n) = S_1(n)^2$. We extend this result and prove that for any $k$ there is a polynomial $g_k(x,y)$ such that $S_k(n) = g(S_1(n), S_2(n))$. The proof yields a recursive formula to evaluate $S_k(n)$ as a polynomial of $n$ that has roughly half the number of terms as the classical one.

preprint2016arXiv

A Probabilistic Approach to Generalized Zeckendorf Decompositions

Generalized Zeckendorf decompositions are expansions of integers as sums of elements of solutions to recurrence relations. The simplest cases are base-$b$ expansions, and the standard Zeckendorf decomposition uses the Fibonacci sequence. The expansions are finite sequences of nonnegative integer coefficients (satisfying certain technical conditions to guarantee uniqueness of the decomposition) and which can be viewed as analogs of sequences of variable-length words made from some fixed alphabet. In this paper we present a new approach and construction for uniform measures on expansions, identifying them as the distribution of a Markov chain conditioned not to hit a set. This gives a unified approach that allows us to easily recover results on the expansions from analogous results for Markov chains, and in this paper we focus on laws of large numbers, central limit theorems for sums of digits, and statements on gaps (zeros) in expansions. We expect the approach to prove useful in other similar contexts.

preprint2016arXiv

Central Limit Theorems for Gaps of Generalized Zeckendorf Decompositions

Zeckendorf proved that every integer can be written uniquely as a sum of non-adjacent Fibonacci numbers $\{1,2,3,5,\dots\}$. This has been extended to many other recurrence relations $\{G_n\}$ (with their own notion of a legal decomposition) and to proving that the distribution of the number of summands of an $M \in [G_n, G_{n+1})$ converges to a Gaussian as $n\to\infty$. We prove that for any non-negative integer $g$ the average number of gaps of size $g$ in many generalized Zeckendorf decompositions is $C_μn+d_μ+o(1)$ for constants $C_μ> 0$ and $d_μ$ depending on $g$ and the recurrence, the variance of the number of gaps of size $g$ is similarly $C_σn + d_σ+ o(1)$ with $C_σ> 0$, and the number of gaps of size $g$ of an $M\in[G_n,G_{n+1})$ converges to a Gaussian as $n\to\infty$. The proof is by analysis of an associated two-dimensional recurrence; we prove a general result on when such behavior converges to a Gaussian, and additionally re-derive other results in the literature.

preprint2016arXiv

Legal Decompositions Arising from Non-positive Linear Recurrences

Zeckendorf's theorem states that any positive integer can be written uniquely as a sum of non-adjacent Fibonacci numbers; this result has been generalized to many recurrence relations, especially those arising from linear recurrences with leading term positive. We investigate legal decompositions arising from two new sequences: the $(s,b)$-Generacci sequence and the Fibonacci Quilt sequence. Both satisfy recurrence relations with leading term zero, and thus previous results and techniques do not apply. These sequences exhibit drastically different behavior. We show that the $(s,b)$-Generacci sequence leads to unique legal decompositions, whereas not only do we have non-unique legal decompositions with the Fibonacci Quilt sequence, we also have that in this case the average number of legal decompositions grows exponentially. Another interesting difference is that while in the $(s,b)$-Generacci case the greedy algorithm always leads to a legal decomposition, in the Fibonacci Quilt setting the greedy algorithm leads to a legal decomposition (approximately) 93\% of the time. In the $(s,b)$-Generacci case, we again have Gaussian behavior in the number of summands as well as for the Fibonacci Quilt sequence when we restrict to decompositions resulting from a modified greedy algorithm.

preprint2016arXiv

New Behavior in Legal Decompositions Arising from Non-positive Linear Recurrences

Zeckendorf's theorem states every positive integer has a unique decomposition as a sum of non-adjacent Fibonacci numbers. This result has been generalized to many sequences $\{a_n\}$ arising from an integer positive linear recurrence, each of which has a corresponding notion of a legal decomposition. Previous work proved the number of summands in decompositions of $m \in [a_n, a_{n+1})$ becomes normally distributed as $n\to\infty$, and the individual gap measures associated to each $m$ converge to geometric random variables, when the leading coefficient in the recurrence is positive. We explore what happens when this assumption is removed in two special sequences. In one we regain all previous results, including unique decomposition; in the other the number of legal decompositions exponentially grows and the natural choice for the legal decomposition (the greedy algorithm) only works approximately 92.6\% of the time (though a slight modification always works). We find a connection between the two sequences, which explains why the distribution of the number of summands and gaps between summands behave the same in the two examples. In the course of our investigations we found a new perspective on dealing with roots of polynomials associated to the characteristic polynomials. This allows us to remove the need for the detailed technical analysis of their properties which greatly complicated the proofs of many earlier results in the subject, as well as handle new cases beyond the reach of existing techniques.

preprint2016arXiv

On the Asymptotic Behavior of Variance of PLRS Decompositions

A positive linear recurrence sequence is of the form $H_{n+1} = c_1 H_n + \cdots + c_L H_{n+1-L}$ with each $c_i \ge 0$ and $c_1 c_L > 0$, with appropriately chosen initial conditions. There is a notion of a legal decomposition (roughly, given a sum of terms in the sequence we cannot use the recurrence relation to reduce it) such that every positive integer has a unique legal decomposition using terms in the sequence; this generalizes the Zeckendorf decomposition, which states any positive integer can be written uniquely as a sum of non-adjacent Fibonacci numbers. Previous work proved not only that a decomposition exists, but that the number of summands $K_n(m)$ in legal decompositions of $m \in [H_n, H_{n+1})$ converges to a Gaussian. Using partial fractions and generating functions it is easy to show the mean and variance grow linearly in $n$: $a n + b + o(1)$ and $C n + d + o(1)$, respectively; the difficulty is proving $a$ and $C$ are positive. Previous approaches relied on delicate analysis of polynomials related to the generating functions and characteristic polynomials, and is algebraically cumbersome. We introduce new, elementary techniques that bypass these issues. The key insight is to use induction and bootstrap bounds through conditional probability expansions to show the variance is unbounded, and hence $C > 0$ (the mean is handled easily through a simple counting argument).

preprint2016arXiv

Some Results in the Theory of Low-lying Zeros: Determining the 1-level density, identifying the group symmetry and the arithmetic of moments of Satake parameters

While Random Matrix Theory has successfully modeled many quantities of families of L-functions, it frequently cannot see the family's arithmetic. In some situations this requires an extended theory that inserts arithmetic factors depending on the family, while in other cases these factors result in contributions which vanish in the limit, and are thus not detected. We review the general theory associated to one of the most important statistics, the n-level density of zeros near the central point. According to the Katz-Sarnak density conjecture, to each family of L-functions there is a corresponding symmetry group such that the behavior of zeros near the central point as the conductors tend to infinity agrees with the behavior of eigenvalues near 1 as the matrix size tends to infinity. We show how these calculations are done, emphasizing the techniques, methods and obstructions to improving the results, by considering in full detail a family of Dirichlet characters. We then describe how we may associate a symmetry constant to each family, and how to determine the symmetry group of a compound family in terms of the symmetries of the constituents. These calculations explain the remarkable universality of behavior, where the main terms are independent of the arithmetic (only the first two moments of the Satake parameters contribute in the limit; similar to the Central Limit Theorem, the higher moments are only felt in the rate of convergence). We end by exploring lower order terms in families of elliptic curves. We present evidence supporting a conjecture that the average second moment in one-parameter families without complex multiplication has, when appropriately viewed, a negative bias, and end with a discussion of the consequences of this bias on the distribution of low-lying zeros, in particular relations between such a bias and the observed excess rank in families.

preprint2016arXiv

Subsets of $\mathbb{F}_q[x]$ free of 3-term geometric progressions

Several recent papers have considered the Ramsey-theoretic problem of how large a subset of integers can be without containing any 3-term geometric progressions. This problem has also recently been generalized to number fields, determining bounds on the greatest possible density of ideals avoiding geometric progressions. We study the analogous problem over $\mathbb{F}_q[x]$, first constructing a set greedily which avoids these progressions and calculating its density, and then considering bounds on the upper density of subsets of $\mathbb{F}_q[x]$ which avoid 3-term geometric progressions. This new setting gives us a parameter $q$ to vary and study how our bounds converge to 1 as it changes, and finite characteristic introduces some extra combinatorial structure that increases the tractibility of common questions in this area.

preprint2015arXiv

A Generalization of Zeckendorf's Theorem via Circumscribed $m$-gons

Zeckendorf's theorem states that every positive integer can be uniquely decomposed as a sum of nonconsecutive Fibonacci numbers, where the Fibonacci numbers satisfy $F_n=F_{n-1}+F_{n-2}$ for $n\geq 3$, $F_1=1$ and $F_2=2$. The distribution of the number of summands in such decomposition converges to a Gaussian, the gaps between summands converges to geometric decay, and the distribution of the longest gap is similar to that of the longest run of heads in a biased coin; these results also hold more generally, though for technical reasons previous work needed to assume the coefficients in the recurrence relation are non-negative and the first term is positive. We extend these results by creating an infinite family of integer sequences called the $m$-gonal sequences arising from a geometric construction using circumscribed $m$-gons. They satisfy a recurrence where the first $m+1$ leading terms vanish, and thus cannot be handled by existing techniques. We provide a notion of a legal decomposition, and prove that the decompositions exist and are unique. We then examine the distribution of the number of summands used in the decompositions and prove that it displays Gaussian behavior. There is geometric decay in the distribution of gaps, both for gaps taken from all integers in an interval and almost surely in distribution for the individual gap measures associated to each integer in the interval. We end by proving that the distribution of the longest gap between summands is strongly concentrated about its mean, behaving similarly as in the longest run of heads in tosses of a coin.

preprint2015arXiv

Crescent configurations

In 1989, Erdős conjectured that for a sufficiently large $n$ it is impossible to place $n$ points in general position in a plane such that for every $1\le i \le n-1$ there is a distance that occurs exactly $i$ times. For small $n$ this is possible and in his paper he provided constructions for $n\leq 8$. The one for $n=5$ was due to Pomerance while Palásti came up with the constructions for $n=7,8$. Constructions for $n=9$ and above remain undiscovered, and little headway has been made toward a proof that for sufficiently large $n$ no configuration exists. In this paper we consider a natural generalization to higher dimensions and provide a construction which shows that for any given $n$ there exists a sufficiently large dimension $d$ such that there is a configuration in $d$-dimensional space meeting Erdős' criteria.

preprint2015arXiv

Determining Optimal Test Functions for Bounding the Average Rank in Families of $L$-Functions

Given an $L$-function, one of the most important questions concerns its vanishing at the central point; for example, the Birch and Swinnerton-Dyer conjecture states that the order of vanishing there of an elliptic curve $L$-function equals the rank of the Mordell-Weil group. The Katz and Sarnak Density Conjecture states that this and other behavior is well-modeled by random matrix ensembles. This correspondence is known for many families when the test functions are suitably restricted. For appropriate choices, we obtain bounds on the average order of vanishing at the central point in families. In this note we report on progress in determining the optimal test functions for the various classical compact groups for different support restrictions, and discuss how this relates to improved rank bounds.

preprint2015arXiv

Equipartitions and a Distribution for Numbers: A Statistical Model for Benford's Law

A statistical model for the fragmentation of a conserved quantity is analyzed, using the principle of maximum entropy and the theory of partitions. Upper and lower bounds for the restricted partitioning problem are derived and applied to the distribution of fragments. The resulting power law directly leads to Benford's law for the first digits of the parts.

preprint2015arXiv

Gaps between zeros of GL(2) $L$-functions

Let $L(s,f)$ be an $L$-function associated to a primitive (holomorphic or Maass) cusp form $f$ on GL(2) over $\mathbb{Q}$. Combining mean-value estimates of Montgomery and Vaughan with a method of Ramachandra, we prove a formula for the mixed second moments of derivatives of $L(1/2+it,f)$ and, via a method of Hall, use it to show that there are infinitely many gaps between consecutive zeros of $L(s,f)$ along the critical line that are at least $\sqrt 3 = 1.732...$ times the average spacing. Using general pair correlation results due to Murty and Perelli in conjunction with a technique of Montgomery, we also prove the existence of small gaps between zeros of any primitive $L$-function of the Selberg class. In particular, when $f$ is a primitive holomorphic cusp form on GL(2) over $\mathbb{Q}$, we prove that there are infinitely many gaps between consecutive zeros of $L(s,f)$ along the critical line that are at most $< 0.823$ times the average spacing.

preprint2015arXiv

Gaussian Distribution of the Number of Summands in Generalized Zeckendorf Decompositions in Small Intervals

Zeckendorf's theorem states that every positive integer can be written uniquely as a sum of non-consecutive Fibonacci numbers ${F_n}$, with initial terms $F_1 = 1, F_2 = 2$. Previous work proved that as $n \to \infty$ the distribution of the number of summands in the Zeckendorf decompositions of $m \in [F_n, F_{n+1})$, appropriately normalized, converges to the standard normal. The proofs crucially used the fact that all integers in $[F_n, F_{n+1})$ share the same potential summands and hold for more general positive linear recurrence sequences $\{G_n\}$. We generalize these results to subintervals of $[G_n, G_{n+1})$ as $n \to \infty$ for certain sequences. The analysis is significantly more involved here as different integers have different sets of potential summands. Explicitly, fix an integer sequence $α(n) \to \infty$. As $n \to \infty$, for almost all $m \in [G_n, G_{n+1})$ the distribution of the number of summands in the generalized Zeckendorf decompositions of integers in the subintervals $[m, m + G_{α(n)})$, appropriately normalized, converges to the standard normal. The proof follows by showing that, with probability tending to $1$, $m$ has at least one appropriately located large gap between indices in its decomposition. We then use a correspondence between this interval and $[0, G_{α(n)})$ to obtain the result, since the summands are known to have Gaussian behavior in the latter interval.

preprint2015arXiv

Geometric-progression-free sets over quadratic number fields

A problem of recent interest has been to study how large subsets of the natural numbers can be while avoiding 3-term geometric progressions. Building on recent progress on this problem, we consider the analogous problem over quadratic number fields. We first construct high-density subsets of the algebraic integers of an imaginary quadratic number field that avoid 3-term geometric progressions. When unique factorization fails or over a real quadratic number field, we instead look at subsets of ideals of the ring of integers. Our approach here is to construct sets "greedily," a generalization of the greedy set of rational integers considered by Rankin. We then describe the densities of these sets in terms of values of the Dedekind zeta function. Next, we consider geometric-progression-free sets with large upper density. We generalize an argument by Riddell to obtain upper bounds for the upper density of geometric-progression-free subsets, and construct sets avoiding geometric progressions with high upper density to obtain lower bounds for the supremum of the upper density of all such subsets. Both arguments depend critically on the elements with small norm in the ring of integers.

preprint2015arXiv

Leading Digit Laws on Linear Lie Groups

We determine the leading digit laws for the matrix components of a linear Lie group $G$. These laws generalize the observations that the normalized Haar measure of the Lie group $\mathbb{R}^+$ is $dx/x$ and that the scale invariance of $dx/x$ implies the distribution of the digits follow Benford's law, which is the probability of observing a significand base $B$ of at most $s$ is $\log_B(s)$; thus the first digit is $d$ with probability $\log_B(1 + 1/d)$). Viewing this scale invariance as left invariance of Haar measure, we determine the power laws in significands from one matrix component of various such $G$. We also determine the leading digit distribution of a fixed number of components of a unit sphere, and find periodic behavior when the dimension of the sphere tends to infinity in a certain progression.

preprint2015arXiv

Ramsey Theory Problems over the Integers: Avoiding Generalized Progressions

Two well studied Ramsey-theoretic problems consider subsets of the natural numbers which either contain no three elements in arithmetic progression, or in geometric progression. We study generalizations of this problem, by varying the kinds of progressions to be avoided and the metrics used to evaluate the density of the resulting subsets. One can view a 3-term arithmetic progression as a sequence $x, f_n(x), f_n(f_n(x))$, where $f_n(x) = x + n$, $n$ a nonzero integer. Thus avoiding three-term arithmetic progressions is equivalent to containing no three elements of the form $x, f_n(x), f_n(f_n(x))$ with $f_n \in\mathcal{F}_{\rm t}$, the set of integer translations. One can similarly construct related progressions using different families of functions. We investigate several such families, including geometric progressions ($f_n(x) = nx$ with $n > 1$ a natural number) and exponential progressions ($f_n(x) = x^n$). Progression-free sets are often constructed "greedily," including every number so long as it is not in progression with any of the previous elements. Rankin characterized the greedy geometric-progression-free set in terms of the greedy arithmetic set. We characterize the greedy exponential set and prove that it has asymptotic density 1, and then discuss how the optimality of the greedy set depends on the family of functions used to define progressions. Traditionally, the size of a progression-free set is measured using the (upper) asymptotic density, however we consider several different notions of density, including the uniform and exponential densities.

preprint2015arXiv

Spherical Matrix Ensembles

The spherical orthogonal, unitary, and symplectic ensembles (SOE/SUE/SSE) $S_β(N,r)$ consist of $N \times N$ real symmetric, complex hermitian, and quaternionic self-adjoint matrices of Frobenius norm $r$, made into a probability space with the uniform measure on the sphere. For each of these ensembles, we determine the joint eigenvalue distribution for each $N$, and we prove the empirical spectral measures rapidly converge to the semicircular distribution as $N \to \infty$. In the unitary case ($β=2$), we also find an explicit formula for the empirical spectral density for each $N$.

preprint2015arXiv

The emergence of 4-cycles in polynomial maps over the extended integers

Let $f(x) \in \mathbb{Z}[x]$; for each integer $α$ it is interesting to consider the number of iterates $n_α$, if possible, needed to satisfy $f^{n_α}(α) = α$. The sets $\{α, f(α), \ldots, f^{n_α - 1}(α), α\}$ generated by the iterates of $f$ are called cycles. For $\mathbb{Z}[x]$ it is known that cycles of length 1 and 2 occur, and no others. While much is known for extensions to number fields, we concentrate on extending $\mathbb{Z}$ by adjoining reciprocals of primes. Let $\mathbb{Z}[1/p_1, \ldots, 1/p_n]$ denote $\mathbb{Z}$ extended by adding in the reciprocals of the $n$ primes $p_1, \ldots, p_n$ and all their products and powers with each other and the elements of $\mathbb{Z}$. Interestingly, cycles of length 4, called 4-cycles, emerge for polynomials in $\mathbb{Z}\left[1/p_1, \ldots, 1/p_n\right][x]$ under the appropriate conditions. The problem of finding criteria under which 4-cycles emerge is equivalent to determining how often a sum of four terms is zero, where the terms are $\pm 1$ times a product of elements from the list of $n$ primes. We investigate conditions on sets of primes under which 4-cycles emerge. We characterize when 4-cycles emerge if the set has one or two primes, and (assuming a generalization of the ABC conjecture) find conditions on sets of primes guaranteed not to cause 4-cycles to emerge.

preprint2015arXiv

The M&M Game: From Morsels to Modern Mathematics

To an adult, it's obvious that the day of someone's death is not precisely determined by the day of birth, but it's a very different story for a child. When the third named author was four years old he asked his father, the fifth named author: If two people are born on the same day, do they die on the same day? While this could easily be demonstrated through murder, such a proof would greatly diminish the possibility of teaching additional lessons, and thus a different approach was taken. With the help of the fourth named author they invented what we'll call \emph{the M\&M Game}: Given $k$ people, each simultaneously flips a fair coin, with each eating an M\&M on a head and not eating on a tail. The process then continues until all \mandms\ are consumed, and two people are deemed to die at the same time if they run out of \mandms\ together\footnote{Is one really living without \mandms?}. This led to a great concrete demonstration of randomness appropriate for little kids; it also led to a host of math problems which have been used in probability classes and math competitions. There are many ways to determine the probability of a tie, which allow us in this article to use this problem as a springboard to a lot of great mathematics, including memoryless process, combinatorics, statistical inference, graph theory, and hypergeometric functions.

preprint2014arXiv

A Generalization of Fibonacci Far-Difference Representations and Gaussian Behavior

A natural generalization of base B expansions is Zeckendorf's Theorem: every integer can be uniquely written as a sum of non-consecutive Fibonacci numbers $\{F_n\}$, with $F_{n+1} = F_n + F_{n-1}$ and $F_1=1, F_2=2$. If instead we allow the coefficients of the Fibonacci numbers in the decomposition to be zero or $\pm 1$, the resulting expression is known as the far-difference representation. Alpert proved that a far-difference representation exists and is unique under certain restraints that generalize non-consecutiveness, specifically that two adjacent summands of the same sign must be at least 4 indices apart and those of opposite signs must be at least 3 indices apart. We prove that a far-difference representation can be created using sets of Skipponacci numbers, which are generated by recurrence relations of the form $S^{(k)}_{n+1} = S^{(k)}_{n} + S^{(k)}_{n-k}$ for $k \ge 0$. Every integer can be written uniquely as a sum of the $\pm S^{(k)}_n $'s such that every two terms of the same sign differ in index by at least 2k+2, and every two terms of opposite signs differ in index by at least k+2. Additionally, we prove that the number of positive and negative terms in given Skipponacci decompositions converges to a Gaussian, with a computable correlation coefficient that is a rational function of the smallest root of the characteristic polynomial of the recurrence. The proof uses recursion to obtain the generating function for having a fixed number of summands, which we prove converges to the generating function of a Gaussian. We next explore the distribution of gaps between summands, and show that for any k the probability of finding a gap of length $j \ge 2k+2$ decays geometrically, with decay ratio equal to the largest root of the given k-Skipponacci recurrence. We conclude by finding sequences that have an (s,d) far-difference representation for any positive integers s,d.

preprint2014arXiv

Benford Behavior of Generalized Zeckendorf Decompositions

We prove connections between Zeckendorf decompositions and Benford's law. Recall that if we define the Fibonacci numbers by $F_1 = 1, F_2 = 2$ and $F_{n+1} = F_n + F_{n-1}$, every positive integer can be written uniquely as a sum of non-adjacent elements of this sequence; this is called the Zeckendorf decomposition, and similar unique decompositions exist for sequences arising from recurrence relations of the form $G_{n+1}=c_1G_n+\cdots+c_LG_{n+1-L}$ with $c_i$ positive and some other restrictions. Additionally, a set $S \subset \mathbb{Z}$ is said to satisfy Benford's law base 10 if the density of the elements in $S$ with leading digit $d$ is $\log_{10}{(1+\frac{1}{d})}$; in other words, smaller leading digits are more likely to occur. We prove that as $n\to\infty$ for a randomly selected integer $m$ in $[0, G_{n+1})$ the distribution of the leading digits of the summands in its generalized Zeckendorf decomposition converges to Benford's law almost surely. Our results hold more generally: one obtains similar theorems to those regarding the distribution of leading digits when considering how often values in sets with density are attained in the summands in the decompositions.

preprint2014arXiv

Benford Behavior of Zeckendorf Decompositions

A beautiful theorem of Zeckendorf states that every integer can be written uniquely as the sum of non-consecutive Fibonacci numbers $\{ F_i \}_{i = 1}^{\infty}$. A set $S \subset \mathbb{Z}$ is said to satisfy Benford's law if the density of the elements in $S$ with leading digit $d$ is $\log_{10}{(1+\frac{1}{d})}$; in other words, smaller leading digits are more likely to occur. We prove that, as $n\to\infty$, for a randomly selected integer $m$ in $[0, F_{n+1})$ the distribution of the leading digits of the Fibonacci summands in its Zeckendorf decomposition converge to Benford's law almost surely. Our results hold more generally, and instead of looking at the distribution of leading digits one obtains similar theorems concerning how often values in sets with density are attained.

preprint2014arXiv

Continued fraction digit averages an Maclaurin's inequalities

A classical result of Khinchin says that for almost all real numbers $α$, the geometric mean of the first $n$ digits $a_i(α)$ in the continued fraction expansion of $α$ converges to a number $K = 2.6854520\ldots$ (Khinchin's constant) as $n \to \infty$. On the other hand, for almost all $α$, the arithmetic mean of the first $n$ continued fraction digits $a_i(α)$ approaches infinity as $n \to \infty$. There is a sequence of refinements of the AM-GM inequality, Maclaurin's inequalities, relating the $1/k$-th powers of the $k$-th elementary symmetric means of $n$ numbers for $1 \leq k \leq n$. On the left end (when $k=n$) we have the geometric mean, and on the right end ($k=1$) we have the arithmetic mean. We analyze what happens to the means of continued fraction digits of a typical real number in the limit as one moves $f(n)$ steps away from either extreme. We prove sufficient conditions on $f(n)$ to ensure to ensure divergence when one moves $f(n)$ steps away from the arithmetic mean and convergence when one moves $f(n)$ steps away from the geometric mean. For typical $α$ we conjecture the behavior for $f(n)=cn$, $0<c<1$. We also study the limiting behavior of such means for quadratic irrational $α$, providing rigorous results, as well as numerically supported conjectures.

preprint2014arXiv

Gaussian Behavior of the Number of Summands in Zeckendorf Decompositions in Small Intervals

Zeckendorf's theorem states that every positive integer can be written uniquely as a sum of non-consecutive Fibonacci numbers ${F_n}$, with initial terms $F_1 = 1, F_2 = 2$. We consider the distribution of the number of summands involved in such decompositions. Previous work proved that as $n \to \infty$ the distribution of the number of summands in the Zeckendorf decompositions of $m \in [F_n, F_{n+1})$, appropriately normalized, converges to the standard normal. The proofs crucially used the fact that all integers in $[F_n, F_{n+1})$ share the same potential summands. We generalize these results to subintervals of $[F_n, F_{n+1})$ as $n \to \infty$; the analysis is significantly more involved here as different integers have different sets of potential summands. Explicitly, fix an integer sequence $α(n) \to \infty$. As $n \to \infty$, for almost all $m \in [F_n, F_{n+1})$ the distribution of the number of summands in the Zeckendorf decompositions of integers in the subintervals $[m, m + F_{α(n)})$, appropriately normalized, converges to the standard normal. The proof follows by showing that, with probability tending to $1$, $m$ has at least one appropriately located large gap between indices in its decomposition. We then use a correspondence between this interval and $[0, F_{α(n)})$ to obtain the result, since the summands are known to have Gaussian behavior in the latter interval. % We also prove the same result for more general linear recurrences.

preprint2014arXiv

Generalizing Zeckendorf's Theorem: The Kentucky Sequence

By Zeckendorf's theorem, an equivalent definition of the Fibonacci sequence (appropriately normalized) is that it is the unique sequence of increasing integers such that every positive number can be written uniquely as a sum of non-adjacent elements; this is called a legal decomposition. Previous work examined the distribution of the number of summands and the spacings between them, in legal decompositions arising from the Fibonacci numbers and other linear recurrence relations with non-negative integral coefficients. Many of these results were restricted to the case where the first term in the defining recurrence was positive. We study a generalization of the Fibonacci numbers with a simple notion of legality which leads to a recurrence where the first term vanishes. We again have unique legal decompositions, Gaussian behavior in the number of summands, and geometric decay in the distribution of gaps.

preprint2014arXiv

Irrationality measure and lower bounds for pi(x)

In this note we show how the irrationality measure of $ζ(s) = π^2/6$ can be used to obtain explicit lower bounds for $π(x)$. We analyze the key ingredients of the proof of the finiteness of the irrationality measure, and show how to obtain good lower bounds for $π(x)$ from these arguments as well. While versions of some of the results here have been done by other authors, our arguments are more elementary and yield a lower bound of order $x/\log x$ as a natural boundary.

preprint2014arXiv

Limiting Spectral Measures for Random Matrix Ensembles with a Polynomial Link Function

Consider the ensembles of real symmetric Toeplitz matrices and real symmetric Hankel matrices whose entries are i.i.d. random variables chosen from a fixed probability distribution p of mean 0, variance 1, and finite higher moments. Previous work on real symmetric Toeplitz matrices shows that the spectral measures, or densities of normalized eigenvalues, converge almost surely to a universal near-Gaussian distribution, while previous work on real symmetric Hankel matrices shows that the spectral measures converge almost surely to a universal non-unimodal distribution. Real symmetric Toeplitz matrices are constant along the diagonals, while real symmetric Hankel matrices are constant along the skew diagonals. We generalize the Toeplitz and Hankel matrices to study matrices that are constant along some curve described by a real-valued bivariate polynomial. Using the Method of Moments and an analysis of the resulting Diophantine equations, we show that the spectral measures associated with linear bivariate polynomials converge in probability and almost surely to universal non-semicircular distributions. We prove that these limiting distributions approach the semicircle in the limit of large values of the polynomial coefficients. We then prove that the spectral measures associated with the sum or difference of any two real-valued polynomials with different degrees converge in probability and almost surely to a universal semicircular distribution.

preprint2014arXiv

Low-lying zeroes of Maass form $L$-functions

The Katz-Sarnak density conjecture states that the scaling limits of the distributions of zeros of families of automorphic $L$-functions agree with the scaling limits of eigenvalue distributions of classical subgroups of the unitary groups $U(N)$. This conjecture is often tested by way of computing particular statistics, such as the one-level density, which evaluates a test function with compactly supported Fourier transform at normalized zeros near the central point. Iwaniec, Luo, and Sarnak studied the one-level densities of cuspidal newforms of weight $k$ and level $N$. They showed in the limit as $kN \to\infty$ that these families have one-level densities agreeing with orthogonal type for test functions with Fourier transform supported in $(-2,2)$. Exceeding $(-1,1)$ is important as the three orthogonal groups are indistinguishable for support up to $(-1,1)$ but are distinguishable for any larger support. We study the other family of ${\rm GL}_2$ automorphic forms over $\mathbb{Q}$: Maass forms. To facilitate the analysis, we use smooth weight functions in the Kuznetsov formula which, among other restrictions, vanish to order $M$ at the origin. For test functions with Fourier transform supported inside $\left(-2 + \frac{2}{2M+1}, 2 - \frac{2}{2M+1}\right)$, we unconditionally prove the one-level density of the low-lying zeros of level 1 Maass forms, as the eigenvalues tend to infinity, agrees only with that of the scaling limit of orthogonal matrices.

preprint2014arXiv

Maass waveforms and low-lying zeros

The Katz-Sarnak Density Conjecture states that the behavior of zeros of a family of $L$-functions near the central point (as the conductors tend to zero) agrees with the behavior of eigenvalues near 1 of a classical compact group (as the matrix size tends to infinity). Using the Petersson formula, Iwaniec, Luo and Sarnak proved that the behavior of zeros near the central point of holomorphic cusp forms agrees with the behavior of eigenvalues of orthogonal matrices for suitably restricted test functions $ϕ$. We prove similar results for families of cuspidal Maass forms, the other natural family of ${\rm GL}_2/\mathbb{Q}$ $L$-functions. For suitable weight functions on the space of Maass forms, the limiting behavior agrees with the expected orthogonal group. We prove this for $\Supp(\widehatϕ)\subseteq (-3/2, 3/2)$ when the level $N$ tends to infinity through the square-free numbers; if the level is fixed the support decreases to being contained in $(-1,1)$, though we still uniquely specify the symmetry type by computing the 2-level density.

preprint2014arXiv

Most Subsets are Balanced in Finite Groups

The sumset is one of the most basic and central objects in additive number theory. Many of the most important problems (such as Goldbach's conjecture and Fermat's Last theorem) can be formulated in terms of the sumset $S + S = \{x+y : x,y\in S\}$ of a set of integers $S$. A finite set of integers $A$ is sum-dominated if $|A+A| > |A-A|$. Though it was believed that the percentage of subsets of $\{0,...,n\}$ that are sum-dominated tends to zero, in 2006 Martin and O'Bryant proved a very small positive percentage are sum-dominated if the sets are chosen uniformly at random (through work of Zhao we know this percentage is approximately $4.5 \cdot 10^{-4}$). While most sets are difference-dominated in the integer case, this is not the case when we take subsets of many finite groups. We show that if we take subsets of larger and larger finite groups uniformly at random, then not only does the probability of a set being sum-dominated tend to zero but the probability that $|A+A|=|A-A|$ tends to one, and hence a typical set is balanced in this case. The cause of this marked difference in behavior is that subsets of $\{0,..., n\}$ have a fringe, whereas finite groups do not. We end with a detailed analysis of dihedral groups, where the results are in striking contrast to what occurs for subsets of integers.

preprint2014arXiv

Pythagoras at the Bat

The Pythagorean formula is one of the most popular ways to measure the true ability of a team. It is very easy to use, estimating a team's winning percentage from the runs they score and allow. This data is readily available on standings pages; no computationally intensive simulations are needed. Normally accurate to within a few games per season, it allows teams to determine how much a run is worth in different situations. This determination helps solve some of the most important economic decisions a team faces: How much is a player worth, which players should be pursued, and how much should they be offered. We discuss the formula and these applications in detail, and provide a theoretical justification, both for the formula as well as simpler linear estimators of a team's winning percentage. The calculations and modeling are discussed in detail, and when possible multiple proofs are given. We analyze the 2012 season in detail, and see that the data for that and other recent years support our modeling conjectures. We conclude with a discussion of work in progress to generalize the formula and increase its predictive power \emph{without} needing expensive simulations, though at the cost of requiring play-by-play data.

preprint2014arXiv

Relieving and Readjusting Pythagoras

Bill James invented the Pythagorean expectation in the late 70's to predict a baseball team's winning percentage knowing just their runs scored and allowed. His original formula estimates a winning percentage of ${\rm RS}^2/({\rm RS}^2+{\rm RA}^2)$, where ${\rm RS}$ stands for runs scored and ${\rm RA}$ for runs allowed; later versions found better agreement with data by replacing the exponent 2 with numbers near 1.83. Miller and his colleagues provided a theoretical justification by modeling runs scored and allowed by independent Weibull distributions. They showed that a single Weibull distribution did a very good job of describing runs scored and allowed, and led to a predicted won-loss percentage of $({\rm RS_{\rm obs}}-1/2)^γ/ (({\rm RS_{\rm obs}}-1/2)^γ+ ({\rm RA_{\rm obs}}-1/2)^γ)$, where ${\rm RS_{\rm obs}}$ and ${\rm RA_{\rm obs}}$ are the observed runs scored and allowed and $γ$ is the shape parameter of the Weibull (typically close to 1.8). We show a linear combination of Weibulls more accurately determines a team's run production and increases the prediction accuracy of a team's winning percentage by an average of about 25% (thus while the currently used variants of the original predictor are accurate to about four games a season, the new combination is accurate to about three). The new formula is more involved computationally; however, it can be easily computed on a laptop in a matter of minutes from publicly available season data. It performs as well (or slightly better) than the related Pythagorean formulas in use, and has the additional advantage of having a theoretical justification for its parameter values (and not just an optimization of parameters to minimize prediction error).

preprint2014arXiv

Sets Characterized by Missing Sums and Differences in Dilating Polytopes

A sum-dominant set is a finite set $A$ of integers such that $|A+A| > |A-A|$. As a typical pair of elements contributes one sum and two differences, we expect sum-dominant sets to be rare in some sense. In 2006, however, Martin and O'Bryant showed that the proportion of sum-dominant subsets of $\{0,\dots,n\}$ is bounded below by a positive constant as $n\to\infty$. Hegarty then extended their work and showed that for any prescribed $s,d\in\mathbb{N}_0$, the proportion $ρ^{s,d}_n$ of subsets of $\{0,\dots,n\}$ that are missing exactly $s$ sums in $\{0,\dots,2n\}$ and exactly $2d$ differences in $\{-n,\dots,n\}$ also remains positive in the limit. We consider the following question: are such sets, characterized by their sums and differences, similarly ubiquitous in higher dimensional spaces? We generalize the integers in a growing interval to the lattice points in a dilating polytope. Specifically, let $P$ be a polytope in $\mathbb{R}^D$ with vertices in $\mathbb{Z}^D$, and let $ρ_n^{s,d}$ now denote the proportion of subsets of $L(nP)$ that are missing exactly $s$ sums in $L(nP)+L(nP)$ and exactly $2d$ differences in $L(nP)-L(nP)$. As it turns out, the geometry of $P$ has a significant effect on the limiting behavior of $ρ_n^{s,d}$. We define a geometric characteristic of polytopes called local point symmetry, and show that $ρ_n^{s,d}$ is bounded below by a positive constant as $n\to\infty$ if and only if $P$ is locally point symmetric. We further show that the proportion of subsets in $L(nP)$ that are missing exactly $s$ sums and at least $2d$ differences remains positive in the limit, independent of the geometry of $P$. A direct corollary of these results is that if $P$ is additionally point symmetric, the proportion of sum-dominant subsets of $L(nP)$ also remains positive in the limit.

preprint2014arXiv

Sums and differences of correlated random sets

Many fundamental questions in additive number theory (such as Goldbach's conjecture, Fermat's last theorem, and the Twin Primes conjecture) can be expressed in the language of sum and difference sets. As a typical pair of elements contributes one sum and two differences, we expect that $|A-A| > |A+A|$ for a finite set $A$. However, in 2006 Martin and O'Bryant showed that a positive proportion of subsets of $\{0, \dots, n\}$ are sum-dominant, and Zhao later showed that this proportion converges to a positive limit as $n \to \infty$. Related problems, such as constructing explicit families of sum-dominant sets, computing the value of the limiting proportion, and investigating the behavior as the probability of including a given element in $A$ to go to zero, have been analyzed extensively. We consider many of these problems in a more general setting. Instead of just one set $A$, we study sums and differences of pairs of \emph{correlated} sets $(A,B)$. Specifically, we place each element $a \in \{0,\dots, n\}$ in $A$ with probability $p$, while $a$ goes in $B$ with probability $ρ_1$ if $a \in A$ and probability $ρ_2$ if $a \not \in A$. If $|A+B| > |(A-B) \cup (B-A)|$, we call the pair $(A,B)$ a \emph{sum-dominant $(p,ρ_1, ρ_2)$-pair}. We prove that for any fixed $\vecρ=(p, ρ_1, ρ_2)$ in $(0,1)^3$, $(A,B)$ is a sum-dominant $(p,ρ_1, ρ_2)$-pair with positive probability, and show that this probability approaches a limit $P(\vecρ)$. Furthermore, we show that the limit function $P(\vecρ)$ is continuous. We also investigate what happens as $p$ decays with $n$, generalizing results of Hegarty-Miller on phase transitions. Finally, we find the smallest sizes of MSTD pairs.

preprint2014arXiv

Surpassing the Ratios Conjecture in the 1-level density of Dirichlet $L$-functions

We study the $1$-level density of low-lying zeros of Dirichlet $L$-functions in the family of all characters modulo $q$, with $Q/2 < q\leq Q$. For test functions whose Fourier transform is supported in $(-3/2, 3/2)$, we calculate this quantity beyond the square-root cancellation expansion arising from the $L$-function Ratios Conjecture of Conrey, Farmer and Zirnbauer. We discover the existence of a new lower-order term which is not predicted by this powerful conjecture. This is the first family where the 1-level density is determined well enough to see a term which is not predicted by the Ratios Conjecture, and proves that the exponent of the error term $Q^{-\frac 12 +ε}$ in the Ratios Conjecture is best possible. We also give more precise results when the support of the Fourier Transform of the test function is restricted to the interval $[-1,1]$. Finally we show how natural conjectures on the distribution of primes in arithmetic progressions allow one to extend the support. The most powerful conjecture is Montgomery's, which implies that the Ratios Conjecture's prediction holds for any finite support up to an error $Q^{-\frac 12 +ε}$.

preprint2014arXiv

The Distribution of Gaps between Summands in Generalized Zeckendorf Decompositions

Zeckendorf proved that any integer can be decomposed uniquely as a sum of non-adjacent Fibonacci numbers, $F_n$. Using continued fractions, Lekkerkerker proved the average number of summands of an $m \in [F_n, F_{n+1})$ is essentially $n/(φ^2 +1)$, with $φ$ the golden ratio. Miller-Wang generalized this by adopting a combinatorial perspective, proving that for any positive linear recurrence the number of summands in decompositions for integers in $[G_n, G_{n+1})$ converges to a Gaussian distribution. We prove the probability of a gap larger than the recurrence length converges to decaying geometrically, and that the distribution of the smaller gaps depends in a computable way on the coefficients of the recurrence. These results hold both for the average over all $m \in [G_n, G_{n+1})$, as well as holding almost surely for the gap measure associated to individual $m$. The techniques can also be used to determine the distribution of the longest gap between summands, which we prove is similar to the distribution of the longest gap between heads in tosses of a biased coin. It is a double exponential strongly concentrated about the mean, and is on the order of $\log n$ with computable constants depending on the recurrence.

preprint2014arXiv

The effect of convolving families of L-functions on the underlying group symmetries

L-functions for GL_n(A_Q) and GL_m(A_Q), respectively, such that, as N,M --> oo, the statistical behavior (1-level density) of the low-lying zeros of L-functions in F_N (resp., G_M) agrees with that of the eigenvalues near 1 of matrices in G_1 (resp., G_2) as the size of the matrices tend to infinity, where each G_i is one of the classical compact groups (unitary, symplectic or orthogonal). Assuming that the convolved families of L-functions F_N x G_M are automorphic, we study their 1-level density. (We also study convolved families of the form f x G_M for a fixed f.) Under natural assumptions on the families (which hold in many cases) we can associate to each family L of L-functions a symmetry constant c_L equal to 0 (resp., 1 or -1) if the corresponding low-lying zero statistics agree with those of the unitary (resp., symplectic or orthogonal) group. Our main result is that c_{F x G} = c_G * c_G: the symmetry type of the convolved family is the product of the symmetry types of the two families. A similar statement holds for the convolved families f x G_M. We provide examples built from Dirichlet L-functions and holomorphic modular forms and their symmetric powers. An interesting special case is to convolve two families of elliptic curves with rank. In this case the symmetry group of the convolution is independent of the ranks, in accordance with the general principle of multiplicativity of the symmetry constants (but the ranks persist, before taking the limit N,M --> oo, as lower-order terms).

preprint2014arXiv

The n-level densities of low-lying zeros of quadratic Dirichlet L-functions

Previous work by Rubinstein and Gao computed the n-level densities for families of quadratic Dirichlet L-functions for test functions where the sum of the supports of the Fourier transforms is at most 2, and showed agreement with random matrix theory predictions in this range for n < 4 but only in a restricted range for larger n. We extend these results and show agreement for n < 8, and reduce higher n to a Fourier transform identity. The proof involves adopting a new combinatorial perspective to convert all terms to a canonical form, which facilitates the comparison of the two sides.

preprint2014arXiv

The Weibull Distribution and Benford's Law

Benford's law states that many data sets have a bias towards lower leading digits (about $30\%$ are 1s). There are numerous applications, from designing efficient computers to detecting tax, voter and image fraud. It's important to know which common probability distributions are almost Benford. We show the Weibull distribution, for many values of its parameters, is close to Benford's law, quantifying the deviations. As the Weibull distribution arises in many problems, especially survival analysis, our results provide additional arguments for the prevalence of Benford behavior. The proof is by Poisson summation, a powerful technique to attack such problems.

preprint2014arXiv

Zeros of Dirichlet L-functions over Function Fields

Random matrix theory has successfully modeled many systems in physics and mathematics, and often the analysis and results in one area guide development in the other. Hughes and Rudnick computed $1$-level density statistics for low-lying zeros of the family of primitive Dirichlet $L$-functions of fixed prime conductor $Q$, as $Q \to \infty$, and verified the unitary symmetry predicted by random matrix theory. We compute $1$- and $2$-level statistics of the analogous family of Dirichlet $L$-functions over $\mathbb{F}_q(T)$. Whereas the Hughes-Rudnick results were restricted by the support of the Fourier transform of their test function, our test function is periodic and our results are only restricted by a decay condition on its Fourier coefficients. We show the main terms agree with unitary symmetry, and also isolate error terms. In concluding, we discuss an $\mathbb{F}_q(T)$-analogue of Montgomery's Hypothesis on the distribution of primes in arithmetic progressions, which Fiorilli and Miller show would remove the restriction on the Hughes-Rudnick results.

preprint2013arXiv

Benford's Law and Continuous Dependent Random Variables

Many systems exhibit a digit bias. For example, the first digit base 10 of the Fibonacci numbers, or of $2^n$, equals 1 not 10% or 11% of the time, as one would expect if all digits were equally likely, but about 30% of the time. This phenomenon, known as Benford's Law, has many applications, ranging from detecting tax fraud for the IRS to analyzing round-off errors in computer science. The central question is determining which data sets follow Benford's law. Inspired by natural processes such as particle decay, our work examines models for the decomposition of conserved quantities. We prove that in many instances the distribution of lengths of the resulting pieces converges to Benford behavior as the number of divisions grow. The main difficulty is that the resulting random variables are dependent, which we handle by a careful analysis of the dependencies and tools from Fourier analysis to obtain quantified convergence rates.

preprint2013arXiv

Explicit Constructions of Large Families of Generalized More Sums Than Differences Sets

A More Sums Than Differences (MSTD) set is a set of integers A contained in {0, ..., n-1} whose sumset A+A is larger than its difference set A-A. While it is known that as n tends to infinity a positive percentage of subsets of {0, ..., n-1} are MSTD sets, the methods to prove this are probabilistic and do not yield nice, explicit constructions. Recently Miller, Orosz and Scheinerman gave explicit constructions of a large family of MSTD sets; though their density is less than a positive percentage, their family's density among subsets of {0, ..., n-1} is at least C/n^4 for some C>0, significantly larger than the previous constructions, which were on the order of 1 / 2^{n/2}. We generalize their method and explicitly construct a large family of sets A with |A+A+A+A| > |(A+A)-(A+A)|. The additional sums and differences allow us greater freedom than in Miller, Orosz and Scheinerman, and we find that for any epsilon>0 the density of such sets is at least C / n^epsilon. In the course of constructing such sets we find that for any integer k there is an A such that |A+A+A+A| - |A+A-A-A| = k, and show that the minimum span of such a set is 30.

preprint2013arXiv

Generalizing Zeckendorf's Theorem to f-decompositions

A beautiful theorem of Zeckendorf states that every positive integer can be uniquely decomposed as a sum of non-consecutive Fibonacci numbers $\{F_n\}$, where $F_1 = 1$, $F_2 = 2$ and $F_{n+1} = F_n + F_{n-1}$. For general recurrences $\{G_n\}$ with non-negative coefficients, there is a notion of a legal decomposition which again leads to a unique representation, and the number of summands in the representations of uniformly randomly chosen $m \in [G_n, G_{n+1})$ converges to a normal distribution as $n \to \infty$. We consider the converse question: given a notion of legal decomposition, is it possible to construct a sequence $\{a_n\}$ such that every positive integer can be decomposed as a sum of terms from the sequence? We encode a notion of legal decomposition as a function $f:\N_0\to\N_0$ and say that if $a_n$ is in an "$f$-decomposition", then the decomposition cannot contain the $f(n)$ terms immediately before $a_n$ in the sequence; special choices of $f$ yield many well known decompositions (including base-$b$, Zeckendorf and factorial). We prove that for any $f:\N_0\to\N_0$, there exists a sequence $\{a_n\}_{n=0}^\infty$ such that every positive integer has a unique $f$-decomposition using $\{a_n\}$. Further, if $f$ is periodic, then the unique increasing sequence $\{a_n\}$ that corresponds to $f$ satisfies a linear recurrence relation. Previous research only handled recurrence relations with no negative coefficients. We find a function $f$ that yields a sequence that cannot be described by such a recurrence relation. Finally, for a class of functions $f$, we prove that the number of summands in the $f$-decomposition of integers between two consecutive terms of the sequence converges to a normal distribution.

preprint2013arXiv

Newman's conjecture in various settings

De Bruijn and Newman introduced a deformation of the Riemann zeta function $ζ(s)$, and found a real constant $Λ$ which encodes the movement of the zeros of $ζ(s)$ under the deformation. The Riemann hypothesis (RH) is equivalent to $Λ\le 0$. Newman made the conjecture that $Λ\ge 0$ along with the remark that "the new conjecture is a quantitative version of the dictum that the Riemann hypothesis, if true, is only barely so." Newman's conjecture is still unsolved, and previous work could only handle the Riemann zeta function and quadratic Dirichlet $L$-functions, obtaining lower bounds very close to zero (for example, for $ζ(s)$ the bound is at least $-1.14541 \cdot 10^{-11}$, and for quadratic Dirichlet $L$-functions it is at least $-1.17 \cdot 10^{-7}$). We generalize the techniques to apply to automorphic $L$-functions as well as function field $L$-functions. We further determine the limit of these techniques by studying linear combinations of $L$-functions, proving that these methods are insufficient. We explicitly determine the Newman constants in various function field settings, which has strong implications for Newman's quantitative version of RH. In particular, let $\mathcal D \in \bbZ[T]$ be a square-free polynomial of degree 3. Let $D_p$ be the polynomial in $\bbF_p[T]$ obtained by reducing $\mathcal D$ modulo $p$. Then the Newman constant $Λ_{D_p}$ equals $\log \frac{|a_p(\mathcal D)|}{2\sqrt{p}}$; by Sato--Tate (if the curve is non-CM) there exists a sequence of primes such that $\lim_{n \to\infty} Λ_{D_{p_n}} = 0$. We end by discussing connections with random matrix theory.

preprint2013arXiv

On the spectral distribution of large weighted random regular graphs

McKay proved that the limiting spectral measures of the ensembles of $d$-regular graphs with $N$ vertices converge to Kesten's measure as $N\to\infty$. In this paper we explore the case of weighted graphs. More precisely, given a large $d$-regular graph we assign random weights, drawn from some distribution $\mathcal{W}$, to its edges. We study the relationship between $\mathcal{W}$ and the associated limiting spectral distribution obtained by averaging over the weighted graphs. Among other results, we establish the existence of a unique `eigendistribution', i.e., a weight distribution $\mathcal{W}$ such that the associated limiting spectral distribution is a rescaling of $\mathcal{W}$. Initial investigations suggested that the eigendistribution was the semi-circle distribution, which by Wigner's Law is the limiting spectral measure for real symmetric matrices. We prove this is not the case, though the deviation between the eigendistribution and the semi-circular density is small (the first seven moments agree, and the difference in each higher moment is $O(1/d^2)$). Our analysis uses combinatorial results about closed acyclic walks in large trees, which may be of independent interest.

preprint2013arXiv

Special Sets of Primes in Function Fields

When investigating the distribution of the Euler totient function, one encounters sets of primes P where if p is in P then r is in P for all r|(p-1). While it is easy to construct finite sets of such primes, the only infinite set known is the set of all primes. We translate this problem into the function field setting and construct an infinite such set in F_p[x] whenever p is equivalent to 2 or 5 modulo 9.

preprint2013arXiv

The Pythagorean Won-Loss Formula and Hockey: A Statistical Justification for Using the Classic Baseball Formula as an Evaluative Tool in Hockey

Originally devised for baseball, the Pythagorean Won-Loss formula estimates the percentage of games a team should have won at a particular point in a season. For decades, this formula had no mathematical justification. In 2006, Steven Miller provided a statistical derivation by making some heuristic assumptions about the distributions of runs scored and allowed by baseball teams. We make a similar set of assumptions about hockey teams and show that the formula is just as applicable to hockey as it is to baseball. We hope that this work spurs research in the use of the Pythagorean Won-Loss formula as an evaluative tool for sports outside baseball.

preprint2013arXiv

When Generalized Sumsets are Difference Dominated

We study the relationship between the number of minus signs in a generalized sumset, $A+...+A-...-A$, and its cardinality; without loss of generality we may assume there are at least as many positive signs as negative signs. As addition is commutative and subtraction is not, we expect that for most $A$ a combination with more minus signs has more elements than one with fewer; however, recently Iyer, Lazarev, Miller and Zhang proved that a positive percentage of the time the combination with fewer minus signs can have more elements. Their analysis involves choosing sets $A$ uniformly at random from ${0,...,N}$; this is equivalent to independently choosing each element of ${0,...,N}$ to be in $A$ with probability 1/2. We investigate what happens when instead each element is chosen with probability $p(N)$, with $\lim_{N\to\infty} p(N) =0$. We prove that the set with more minus signs is larger with probability 1 as $N\to\infty$ if $p(N)=cN^{-δ}$ for $δ\ge\frac{h-1}{h}$, where $h$ is the number of total summands in $A+...+A-...-A$, and explicitly quantify their relative sizes. The results generalize earlier work of Hegarty and Miller, and we see a phase transition in the behavior of the cardinalities when $δ= \frac{h-1}{h}$.

preprint2012arXiv

Coordinate sum and difference sets of $d$-dimensional modular hyperbolas

Many problems in additive number theory, such as Fermat's last theorem and the twin prime conjecture, can be understood by examining sums or differences of a set with itself. A finite set $A \subset \mathbb{Z}$ is considered sum-dominant if $|A+A|>|A-A|$. If we consider all subsets of ${0, 1, ..., n-1}$, as $n\to\infty$ it is natural to expect that almost all subsets should be difference-dominant, as addition is commutative but subtraction is not; however, Martin and O'Bryant in 2007 proved that a positive percentage are sum-dominant as $n\to\infty$. This motivates the study of "coordinate sum dominance". Given $V \subset (\Z/n\Z)^2$, we call $S:={x+y: (x,y) \in V}$ a coordinate sumset and $D:=\{x-y: (x,y) \in V\}$ a coordinate difference set, and we say $V$ is coordinate sum dominant if $|S|>|D|$. An arithmetically interesting choice of $V$ is $\bar{H}_2(a;n)$, which is the reduction modulo $n$ of the modular hyperbola $H_2(a;n) := {(x,y): xy \equiv a \bmod n, 1 \le x,y < n}$. In 2009, Eichhorn, Khan, Stein, and Yankov determined the sizes of $S$ and $D$ for $V=\bar{H}_2(1;n)$ and investigated conditions for coordinate sum dominance. We extend their results to reduced $d$-dimensional modular hyperbolas $\bar{H}_d(a;n)$ with $a$ coprime to $n$.

preprint2012arXiv

Distribution of Eigenvalues of Weighted, Structured Matrix Ensembles

The limiting distribution of eigenvalues of N x N random matrices has many applications. One of the most studied ensembles are real symmetric matrices with independent entries iidrv; the limiting rescaled spectral measure (LRSM) $\widetildeμ$ is the semi-circle. Studies have determined the LRSMs for many structured ensembles, such as Toeplitz and circulant matrices. These have very different behavior; the LRSM for both have unbounded support. Given a structured ensemble such that (i) each random variable occurs o(N) times in each row and (ii) the LRSM exists, we introduce a parameter to continuously interpolate between these behaviors. We fix a p in [1/2, 1] and study the ensemble of signed structured matrices by multiplying the (i,j)-th and (j,i)-th entries of a matrix by a randomly chosen epsilon_ij in {1, -1}, with Prob(epsilon_ij = 1) = p (i.e., the Hadamard product). For p = 1/2 we prove that the limiting signed rescaled spectral measure is the semi-circle. For all other p, the limiting measure has bounded (resp., unbounded) support if $\widetildeμ$ has bounded (resp., unbounded) support, and converges to $\widetildeμ$ as p -> 1. Notably, these results hold for Toeplitz and circulant matrix ensembles. The proofs are by the Method of Moments. The analysis involves the pairings of 2k vertices on a circle. The contribution of each in the signed case is weighted by a factor depending on p and the number of vertices involved in at least one crossing. These numbers appear in combinatorics and knot theory. The number of configurations with no vertices involved in a crossing is well-studied, and are the Catalan numbers. We prove similar formulas for configurations with up to 10 vertices in at least one crossing. We derive a closed-form expression for the expected value and determine the asymptotics for the variance for the number of vertices in at least one crossing.

preprint2012arXiv

Distribution of Missing Sums in Sumsets

For any finite set of integers X, define its sumset X+X to be {x+y: x, y in X}. In a recent paper, Martin and O'Bryant investigated the distribution of |A+A| given the uniform distribution on subsets A of {0, 1, ..., n-1}. They also conjectured the existence of a limiting distribution for |A+A| and showed that the expectation of |A+A| is 2n - 11 + O((3/4)^{n/2}). Zhao proved that the limits m(k) := lim_{n --> oo} Prob(2n-1-|A+A|=k) exist, and that sum_{k >= 0} m(k)=1. We continue this program and give exponentially decaying upper and lower bounds on m(k), and sharp bounds on m(k) for small k. Surprisingly, the distribution is at least bimodal; sumsets have an unexpected bias against missing exactly 7 sums. The proof of the latter is by reduction to questions on the distribution of related random variables, with large scale numerical computations a key ingredient in the analysis. We also derive an explicit formula for the variance of |A+A| in terms of Fibonacci numbers, finding Var(|A+A|) is approximately 35.9658. New difficulties arise in the form of weak dependence between events of the form {x in A+A}, {y in A+A}. We surmount these obstructions by translating the problem to graph theory. This approach also yields good bounds on the probability for A+A missing a consecutive block of length k.

preprint2012arXiv

First Order Approximations of the Pythagorean Won-Loss Formula for Predicting MLB Teams' Winning Percentages

We mathematically prove that an existing linear predictor of baseball teams' winning percentages (Jones and Tappin 2005) is simply just a first-order approximation to Bill James' Pythagorean Won-Loss formula and can thus be written in terms of the formula's well-known exponent. We estimate the linear model on twenty seasons of Major League Baseball data and are able to verify that the resulting coefficient estimate, with 95% confidence, is virtually identical to the empirically accepted value of 1.82. Our work thus helps explain why this simple and elegant model is such a strong linear predictor.

preprint2012arXiv

The Average Gap Distribution for Generalized Zeckendorf Decompositions

An interesting characterization of the Fibonacci numbers is that, if we write them as $F_1 = 1$, $F_2 = 2$, $F_3 = 3$, $F_4 = 5, ...$, then every positive integer can be written uniquely as a sum of non-adjacent Fibonacci numbers. This is now known as Zeckendorf's theorem [21], and similar decompositions exist for many other sequences ${G_{n+1} = c_1 G_{n} + ... + c_L G_{n+1-L}}$ arising from recurrence relations. Much more is known. Using continued fraction approaches, Lekkerkerker [15] proved the average number of summands needed for integers in $[G_n, G_{n+1})$ is on the order of $C_{\rm Lek} n$ for a non-zero constant; this was improved by others to show the number of summands has Gaussian fluctuations about this mean. Kolo$\breve{\rm g}$lu, Kopp, Miller and Wang [17, 18] recently recast the problem combinatorially, reproving and generalizing these results. We use this new perspective to investigate the distribution of gaps between summands. We explore the average behavior over all $m \in [G_n, G_{n+1})$ for special choices of the $c_i$'s. Specifically, we study the case where each $c_i \in {0,1}$ and there is a $g$ such that there are always exactly $g-1$ zeros between two non-zero $c_i$'s; note this includes the Fibonacci, Tribonacci and many other important special cases. We prove there are no gaps of length less than $g$, and the probability of a gap of length $j > g$ decays geometrically, with the decay ratio equal to the largest root of the recurrence relation. These methods are combinatorial and apply to related problems; we end with a discussion of similar results for far-difference (i.e., signed) decompositions.

preprint2011arXiv

A Random Matrix Model for Elliptic Curve L-Functions of Finite Conductor

We propose a random matrix model for families of elliptic curve L-functions of finite conductor. A repulsion of the critical zeros of these L-functions away from the center of the critical strip was observed numerically by S. J. Miller in 2006; such behaviour deviates qualitatively from the conjectural limiting distribution of the zeros (for large conductors this distribution is expected to approach the one-level density of eigenvalues of orthogonal matrices after appropriate rescaling).Our purpose here is to provide a random matrix model for Miller's surprising discovery. We consider the family of even quadratic twists of a given elliptic curve. The main ingredient in our model is a calculation of the eigenvalue distribution of random orthogonal matrices whose characteristic polynomials are larger than some given value at the symmetry point in the spectra. We call this sub-ensemble of SO(2N) the excised orthogonal ensemble. The sieving-off of matrices with small values of the characteristic polynomial is akin to the discretization of the central values of L-functions implied by the formula of Waldspurger and Kohnen-Zagier.The cut-off scale appropriate to modeling elliptic curve L-functions is exponentially small relative to the matrix size N. The one-level density of the excised ensemble can be expressed in terms of that of the well-known Jacobi ensemble, enabling the former to be explicitly calculated. It exhibits an exponentially small (on the scale of the mean spacing) hard gap determined by the cut-off value, followed by soft repulsion on a much larger scale. Neither of these features is present in the one-level density of SO(2N). When N tends to infinity we recover the limiting orthogonal behaviour. Our results agree qualitatively with Miller's discrepancy. Choosing the cut-off appropriately gives a model in good quantitative agreement with the number-theoretical data.

preprint2011arXiv

An elliptic curve test of the L-Functions Ratios Conjecture

We compare the L-Function Ratios Conjecture's prediction with number theory for the family of quadratic twists of a fixed elliptic curve with prime conductor, and show agreement in the 1-level density up to an error term of size X^{-(1-sigma)/2} for test functions supported in (-sigma, sigma); this gives us a power-savings for σ<1. This test of the Ratios Conjecture introduces complications not seen in previous cases (due to the level of the elliptic curve). Further, the results here are one of the key ingredients in the companion paper [DHKMS2], where they are used to determine the effective matrix size for modeling zeros near the central point for this family. The resulting model beautifully describes the behavior of these low lying zeros for finite conductors, explaining the data observed by Miller in [Mil3]. A key ingredient in our analysis is a generalization of Jutila's bound for sums of quadratic characters with the additional restriction that the fundamental discriminant be congruent to a non-zero square modulo a square-free integer M. This bound is needed for two purposes. The first is to analyze the terms in the explicit formula corresponding to characters raised to an odd power. The second is to determine the main term in the 1-level density of quadratic twists of a fixed form on GL_n. Such an analysis was performed by Rubinstein [Rub], who implicitly assumed that Jutila's bound held with the additional restriction on the fundamental discriminants; in this paper we show that assumption is justified.

preprint2011arXiv

Finding and Counting MSTD sets

We review the basic theory of More Sums Than Differences (MSTD) sets, specifically their existence, simple constructions of infinite families, the proof that a positive percentage of sets under the uniform binomial model are MSTD but not if the probability that each element is chosen tends to zero, and 'explicit' constructions of large families of MSTD sets. We conclude with some new constructions and results of generalized MSTD sets, including among other items results on a positive percentage of sets having a given linear combination greater than another linear combination, and a proof that a positive percentage of sets are $k$-generational sum-dominant (meaning $A$, $A+A$, $...$, $kA = A + ...+A$ are each sum-dominant).

preprint2011arXiv

From Fibonacci Numbers to Central Limit Type Theorems

A beautiful theorem of Zeckendorf states that every integer can be written uniquely as a sum of non-consecutive Fibonacci numbers $\{F_n\}_{n=1}^{\infty}$. Lekkerkerker proved that the average number of summands for integers in $[F_n, F_{n+1})$ is $n/(ϕ^2 + 1)$, with $ϕ$ the golden mean. This has been generalized to the following: given nonnegative integers $c_1,c_2,...,c_L$ with $c_1,c_L>0$ and recursive sequence $\{H_n\}_{n=1}^{\infty}$ with $H_1=1$, $H_{n+1} =c_1H_n+c_2H_{n-1}+...+c_nH_1+1$ $(1\le n< L)$ and $H_{n+1}=c_1H_n+c_2H_{n-1}+...+c_LH_{n+1-L}$ $(n\geq L)$, every positive integer can be written uniquely as $\sum a_iH_i$ under natural constraints on the $a_i$'s, the mean and the variance of the numbers of summands for integers in $[H_{n}, H_{n+1})$ are of size $n$, and the distribution of the numbers of summands converges to a Gaussian as $n$ goes to the infinity. Previous approaches used number theory or ergodic theory. We convert the problem to a combinatorial one. In addition to re-deriving these results, our method generalizes to a multitude of other problems (in the sequel paper \cite{BM} we show how this perspective allows us to determine the distribution of gaps between summands in decompositions). For example, it is known that every integer can be written uniquely as a sum of the $\pm F_n$'s, such that every two terms of the same (opposite) sign differ in index by at least 4 (3). The presence of negative summands introduces complications and features not seen in previous problems. We prove that the distribution of the numbers of positive and negative summands converges to a bivariate normal with computable, negative correlation, namely $-(21-2ϕ)/(29+2ϕ) \approx -0.551058$.

preprint2011arXiv

Gaussian Behavior in Generalized Zeckendorf Decompositions

A beautiful theorem of Zeckendorf states that every integer can be written uniquely as a sum of non-consecutive Fibonacci numbers $\{F_n\}_{n=1}^{\infty}$; Lekkerkerker proved that the average number of summands for integers in $[F_n, F_{n+1})$ is $n/(ϕ^2 + 1)$, with $ϕ$ the golden mean. Interestingly, the higher moments seem to have been ignored. We discuss the proof that the distribution of the number of summands converges to a Gaussian as $n \to \infty$, and comment on generalizations to related decompositions. For example, every integer can be written uniquely as a sum of the $\pm F_n$'s, such that every two terms of the same (opposite) sign differ in index by at least 4 (3). The distribution of the numbers of positive and negative summands converges to a bivariate normal with computable, negative correlation, namely $-(21-2ϕ)/(29+2ϕ) \approx -0.551058$.

preprint2011arXiv

Generalized More Sums Than Differences Sets

A More Sums Than Differences (MSTD, or sum-dominant) set is a finite set $A\subset \mathbb{Z}$ such that $|A+A|<|A-A|$. Though it was believed that the percentage of subsets of $\{0,...,n\}$ that are sum-dominant tends to zero, in 2006 Martin and O'Bryant \cite{MO} proved a positive percentage are sum-dominant. We generalize their result to the many different ways of taking sums and differences of a set. We prove that $|ε_1A+...+ε_kA|>|δ_1A+...+δ_kA|$ a positive percent of the time for all nontrivial choices of $ε_j,δ_j\in \{-1,1\}$. Previous approaches proved the existence of infinitely many such sets given the existence of one; however, no method existed to construct such a set. We develop a new, explicit construction for one such set, and then extend to a positive percentage of sets. We extend these results further, finding sets that exhibit different behavior as more sums/differences are taken. For example, notation as above we prove that for any $m$, $|ε_1A + ... + ε_kA| - |δ_1A + ... + δ_kA| = m$ a positive percentage of the time. We find the limiting behavior of $kA=A+...+A$ for an arbitrary set $A$ as $k\to\infty$ and an upper bound of $k$ for such behavior to settle down. Finally, we say $A$ is $k$-generational sum-dominant if $A$, $A+A$, ...,$kA$ are all sum-dominant. Numerical searches were unable to find even a 2-generational set (heuristics indicate the probability is at most $10^{-9}$, and almost surely significantly less). We prove the surprising result that for any $k$ a positive percentage of sets are $k$-generational, and no set can be $k$-generational for all $k$.

preprint2011arXiv

Generalized Ramanujan Primes

In 1845, Bertrand conjectured that for all integers $x\ge2$, there exists at least one prime in $(x/2, x]$. This was proved by Chebyshev in 1860, and then generalized by Ramanujan in 1919. He showed that for any $n\ge1$, there is a (smallest) prime $R_n$ such that $π(x)- π(x/2) \ge n$ for all $x \ge R_n$. In 2009 Sondow called $R_n$ the $n$th Ramanujan prime and proved the asymptotic behavior $R_n \sim p_{2n}$ (where $p_m$ is the $m$th prime). In the present paper, we generalize the interval of interest by introducing a parameter $c \in (0,1)$ and defining the $n$th $c$-Ramanujan prime as the smallest integer $R_{c,n}$ such that for all $x\ge R_{c,n}$, there are at least $n$ primes in $(cx,x]$. Using consequences of strengthened versions of the Prime Number Theorem, we prove that $R_{c,n}$ exists for all $n$ and all $c$, that $R_{c,n} \sim p_{\frac{n}{1-c}}$ as $n\to\infty$, and that the fraction of primes which are $c$-Ramanujan converges to $1-c$. We then study finer questions related to their distribution among the primes, and see that the $c$-Ramanujan primes display striking behavior, deviating significantly from a probabilistic model based on biased coin flipping; this was first observed by Sondow, Nicholson, and Noe in the case $c = 1/2$. This model is related to the Cramer model, which correctly predicts many properties of primes on large scales, but has been shown to fail in some instances on smaller scales.

preprint2011arXiv

Low-lying Zeros of Cuspidal Maass Forms

The Katz-Sarnak Density Conjecture states that the behavior of zeros of a family of $L$-functions near the central point (as the conductors tend to zero) agree with the behavior of eigenvalues near 1 of a classical compact group (as the matrix size tends to infinity). Using the Petersson formula, Iwaniec, Luo and Sarnak \cite{ILS} proved that the behavior of zeros near the central point of holomorphic cusp forms agree with the behavior of eigenvalues of orthogonal matrices for suitably restricted test functions. We prove a similar result for level 1 cuspidal Maass forms, the other natural family of ${\rm GL}_2$ $L$-functions. We use the explicit formula to relate sums of our test function at scaled zeros to sums of the Fourier transform at the primes weighted by the $L$-function coefficients, and then use the Kuznetsov trace formula to average the Fourier coefficients over the family. There are numerous technical obstructions in handling the terms in the trace formula, which are surmounted through the use of smooth weight functions for the Maass eigenvalues and results on Kloosterman sums and Bessel and hyperbolic functions.

preprint2011arXiv

The Limiting Spectral Measure for Ensembles of Symmetric Block Circulant Matrices

Given an ensemble of NxN random matrices, a natural question to ask is whether or not the empirical spectral measures of typical matrices converge to a limiting spectral measure as N --> oo. While this has been proved for many thin patterned ensembles sitting inside all real symmetric matrices, frequently there is no nice closed form expression for the limiting measure. Further, current theorems provide few pictures of transitions between ensembles. We consider the ensemble of symmetric m-block circulant matrices with entries i.i.d.r.v. These matrices have toroidal diagonals periodic of period m. We view m as a "dial" we can "turn" from the thin ensemble of symmetric circulant matrices, whose limiting eigenvalue density is a Gaussian, to all real symmetric matrices, whose limiting eigenvalue density is a semi-circle. The limiting eigenvalue densities f_m show a visually stunning convergence to the semi-circle as m tends to infinity, which we prove. In contrast to most studies of patterned matrix ensembles, our paper gives explicit closed form expressions for the densities. We prove that f_m is the product of a Gaussian and a degree 2m-2 polynomial; the formula equals that of the m x m Gaussian Unitary Ensemble (GUE). The proof is by the moments. The new feature, which allows us to obtain closed form expressions, is converting the central combinatorial problem in the moment calculation into an equivalent counting problem in algebraic topology. We end with a generalization of the m-block circulant pattern, dropping the assumption that the m random variables be distinct. We prove that the limiting spectral distribution exists and is determined by the pattern of the independent elements within an m-period, depending on not only the frequency at which each element appears, but also the way the elements are arranged.

preprint2011arXiv

Virus Dynamics on Starlike Graphs

The field of epidemiology has presented fascinating and relevant questions for mathematicians, primarily concerning the spread of viruses in a community. The importance of this research has greatly increased over time as its applications have expanded to also include studies of electronic and social networks and the spread of information and ideas. We study virus propagation on a non-linear hub and spoke graph (which models well many airline networks). We determine the long-term behavior as a function of the cure and infection rates, as well as the number of spokes n. For each n we prove the existence of a critical threshold relating the two rates. Below this threshold, the virus always dies out; above this threshold, all non-trivial initial conditions iterate to a unique non-trivial steady state. We end with some generalizations to other networks.

preprint2010arXiv

A unitary test of the Ratios Conjecture

The Ratios Conjecture of Conrey, Farmer and Zirnbauer predicts the answers to numerous questions in number theory, ranging from n-level densities and correlations to mollifiers to moments and vanishing at the central point. The conjecture gives a recipe to generate these answers, which are believed to be correct up to square-root cancelation. These predictions have been verified, for suitably restricted test functions, for the 1-level density of orthogonal and symplectic families of L-functions. In this paper we verify the conjecture's predictions for the unitary family of all Dirichlet $L$-functions with prime conductor; we show square-root agreement between prediction and number theory if the support of the Fourier transform of the test function is in (-1,1), and for support up to (-2,2) we show agreement up to a power savings in the family's cardinality.

preprint2010arXiv

Distribution of Eigenvalues of Highly Palindromic Toeplitz Matrices

Consider the ensemble of real symmetric Toeplitz matrices whose entries are i.i.d random variables chosen from a fixed probability distribution p of mean 0, variance 1 and finite higher moments. Previous work [BDJ,HM] showed that the limiting spectral measures (the density of normalized eigenvalues) converge weakly and almost surely to a universal distribution almost that of the Gaussian, independent of p. The deficit from the Gaussian distribution is due to obstructions to solutions of Diophantine equations and can be removed (see [MMS]) by making the first row palindromic. In this paper, we study the case where there is more than one palindrome in the first row of a real symmetric Toeplitz matrix. Using the method of moments and an analysis of the resulting Diophantine equations, we show that the moments of this ensemble converge to an universal distribution with a fatter tail than any previously seen limiting spectral measure.

preprint2010arXiv

Effective equidistribution and the Sato-Tate law for families of elliptic curves

Extending recent work of others, we provide effective bounds on the family of all elliptic curves and one-parameter families of elliptic curves modulo p (for p prime tending to infinity) obeying the Sato-Tate Law. We present two methods of proof. Both use the framework of Murty-Sinha; the first involves only knowledge of the moments of the Fourier coefficients of the L-functions and combinatorics, and saves a logarithm, while the second requires a Sato-Tate law. Our purpose is to illustrate how the caliber of the result depends on the error terms of the inputs and what combinatorics must be done.

preprint2010arXiv

Low-lying Zeros of Number Field $L$-functions

One of the most important statistics in studying the zeros of L-functions is the 1-level density, which measures the concentration of zeros near the central point. Fouvry and Iwaniec [FI] proved that the 1-level density for L-functions attached to imaginary quadratic fields agrees with results predicted by random matrix theory. In this paper, we show a similar agreement with random matrix theory occurring in more general sequences of number fields. We first show that the main term agrees with random matrix theory, and similar to all other families studied to date, is independent of the arithmetic of the fields. We then derive the first lower order term of the 1-level density, and see the arithmetic enter.

preprint2010arXiv

Modeling Convolutions of $L$-Functions

A number of mathematical methods have been shown to model the zeroes of $L$-functions with remarkable success, including the Ratios Conjecture and Random Matrix Theory. In order to understand the structure of convolutions of families of $L$-functions, we investigate how well these methods model the zeros of such functions. Our primary focus is the convolution of the $L$-function associated to Ramanujan's tau function with the family of quadratic Dirichlet $L$-functions, for which J.B. Conrey and N.C. Snaith computed the Ratios Conjecture's prediction. Our main result is performing the number theory calculations and verifying these predictions for the one-level density for suitably restricted test functions up to square-root error term. Unlike Random Matrix Theory, which only predicts the main term, the Ratios Conjecture detects the arithmetic of the family and makes detailed predictions about their dependence in the lower order terms. Interestingly, while Random Matrix Theory is frequently used to model behavior of L-functions (or at least the main terms), there has been little if any work on the analogue of convolving families of L-functions by convolving random matrix ensembles. We explore one possibility by considering Kronecker products; unfortunately, it appears that this is not the correct random matrix analogue to convolving families.

preprint2010arXiv

On the number of summands in Zeckendorf decompositions

Zeckendorf proved that every positive integer has a unique representation as a sum of non-consecutive Fibonacci numbers. Once this has been shown, it's natural to ask how many summands are needed. Using a continued fraction approach, Lekkerkerker proved that the average number of such summands needed for integers in $[F_n, F_{n+1})$ is $n / (φ^2 + 1) + O(1)$, where $φ= \frac{1+\sqrt{5}}2$ is the golden mean. Surprisingly, no one appears to have investigated the distribution of the number of summands; our main result is that this converges to a Gaussian as $n\to\infty$. Moreover, such a result holds not just for the Fibonacci numbers but many other problems, such as linear recurrence relation with non-negative integer coefficients (which is a generalization of base $B$ expansions of numbers) and far-difference representations. In general the proofs involve adopting a combinatorial viewpoint and analyzing the resulting generating functions through partial fraction expansions and differentiating identities. The resulting arguments become quite technical; the purpose of this paper is to concentrate on the special and most interesting case of the Fibonacci numbers, where the obstructions vanish and the proofs follow from some combinatorics and Stirling's formula; see [MW] for proofs in the general case.

preprint2010arXiv

The lowest eigenvalue of Jacobi random matrix ensembles and Painlevé VI

We present two complementary methods, each applicable in a different range, to evaluate the distribution of the lowest eigenvalue of random matrices in a Jacobi ensemble. The first method solves an associated Painleve VI nonlinear differential equation numerically, with suitable initial conditions that we determine. The second method proceeds via constructing the power-series expansion of the Painleve VI function. Our results are applied in a forthcoming paper in which we model the distribution of the first zero above the central point of elliptic curve L-function families of finite conductor and of conjecturally orthogonal symmetry.

preprint2009arXiv

An Orthogonal Test of the $L$-functions Ratios Conjecture, II

Recently Conrey, Farmer, and Zirnbauer developed the L-functions Ratios conjecture, which gives a recipe that predicts a wealth of statistics, from moments to spacings between adjacent zeros and values of L-functions. The problem with this method is that several of its steps involve ignoring error terms of size comparable to the main term; amazingly, the errors seem to cancel and the resulting prediction is expected to be accurate up to square-root cancellation. We prove the accuracy of the Ratios Conjecture's prediction for the 1-level density of families of cuspidal newforms of constant sign (up to square-root agreement for support in (-1,1), and up to a power savings in (-2,2)), and discuss the arithmetic significance of the lower order terms. This is the most involved test of the Ratios Conjecture's predictions to date, as it is known that the error terms dropped in some of the steps do not cancel, but rather contribute a main term! Specifically, these are the non-diagonal terms in the Petersson formula, which lead to a Bessel-Kloosterman sum which contributes only when the support of the Fourier transform of the test function exceeds (-1, 1).

preprint2009arXiv

Lower order terms in the 1-level density for families of holomorphic cuspidal newforms

The Katz-Sarnak density conjecture states that, in the limit as the conductors tend to infinity, the behavior of normalized zeros near the central point of families of L-functions agree with the N -> oo scaling limits of eigenvalues near 1 of subgroups of U(N). Evidence for this has been found for many families by studying the n-level densities; for suitably restricted test functions the main terms agree with random matrix theory. In particular, all one-parameter families of elliptic curves with rank r over Q(T) and the same distribution of signs of functional equations have the same limiting behavior. We break this universality and find family dependent lower order correction terms in many cases; these lower order terms have applications ranging from excess rank to modeling the behavior of zeros near the central point, and depend on the arithmetic of the family. We derive an alternate form of the explicit formula for GL(2) L-functions which simplifies comparisons, replacing sums over powers of Satake parameters by sums of the moments of the Fourier coefficients lambda_f(p). Our formula highlights the differences that we expect to exist from families whose Fourier coefficients obey different laws (for example, we expect Sato-Tate to hold only for non-CM families of elliptic curves). Further, by the work of Rosen and Silverman we expect lower order biases to the Fourier coefficients in families of elliptic curves with rank over Q(T); these biases can be seen in our expansions. We analyze several families of elliptic curves and see different lower order corrections, depending on whether or not the family has complex multiplication, a forced torsion point, or non-zero rank over Q(T).

preprint2009arXiv

Nuclei, Primes and the Random Matrix Connection

In this article, we discuss the remarkable connection between two very different fields, number theory and nuclear physics. We describe the essential aspects of these fields, the quantities studied, and how insights in one have been fruitfully applied in the other. The exciting branch of modern mathematics, random matrix theory, provides the connection between the two fields. We assume no detailed knowledge of number theory, nuclear physics, or random matrix theory; all that is required is some familiarity with linear algebra and probability theory, as well as some results from complex analysis. Our goal is to provide the inquisitive reader with a sound overview of the subjects, placing them in their historical context in a way that is not traditionally given in the popular and technical surveys.

preprint2008arXiv

An Orthogonal Test of the L-Functions Ratios Conjecture

We test the predictions of the L-functions Ratios Conjecture for the family of cuspidal newforms of weight k and level N, with either k fixed and N --> oo through the primes or N=1 and k --> oo. We study the main and lower order terms in the 1-level density. We provide evidence for the Ratios Conjecture by computing and confirming its predictions up to a power savings in the family's cardinality, at least for test functions whose Fourier transforms are supported in (-2, 2). We do this both for the weighted and unweighted 1-level density (where in the weighted case we use the Petersson weights), thus showing that either formulation may be used. These two 1-level densities differ by a term of size 1 / log(k^2 N). Finally, we show that there is another way of extending the sums arising in the Ratios Conjecture, leading to a different answer (although the answer is such a lower order term that it is hopeless to observe which is correct).

preprint2008arXiv

Chains of distributions, hierarchical Bayesian models and Benford's Law

Kossovsky recently conjectured that the distribution of leading digits of a chain of probability distributions converges to Benford's law as the length of the chain grows. We prove his conjecture in many cases, and provide an interpretation in terms of products of independent random variables and a central limit theorem. An interesting consequence is that in hierarchical Bayesian models priors tend to satisfy Benford's Law as the number of levels of the hierarchy increases, which allows us to develop some simple tests (based on Benford's law) to test proposed models. We give explicit formulas for the error terms as sums of Mellin transforms, which converges extremely rapidly as the number of terms in the chain grows. We may interpret our results as showing that certain Markov chain Monte Carlo processes are rapidly mixing to Benford's law.

preprint2008arXiv

Explicit constructions of infinite families of MSTD sets

We explicitly construct infinite families of MSTD (more sums than differences) sets. There are enough of these sets to prove that there exists a constant C such that at least C / r^4 of the 2^r subsets of {1,...,r} are MSTD sets; thus our family is significantly denser than previous constructions (whose densities are at most f(r)/2^{r/2} for some polynomial f(r)). We conclude by generalizing our method to compare linear forms epsilon_1 A + ... + epsilon_n A with epsilon_i in {-1,1}.

preprint2008arXiv

Order Statistics and Benford's Law

Fix a base B and let zeta have the standard exponential distribution; the distribution of digits of zeta base B is known to be very close to Benford's Law. If there exists a C such that the distribution of digits of C times the elements of some set is the same as that of zeta, we say that set exhibits shifted exponential behavior base B (with a shift of log_B C \bmod 1). Let X_1, >..., X_N be independent identically distributed random variables. If the X_i's are drawn from the uniform distribution on [0,L], then as N\to\infty the distribution of the digits of the differences between adjacent order statistics converges to shifted exponential behavior (with a shift of \log_B L/N \bmod 1). By differentiating the cumulative distribution function of the logarithms modulo 1, applying Poisson Summation and then integrating the resulting expression, we derive rapidly converging explicit formulas measuring the deviations from Benford's Law. Fix a delta in (0,1) and choose N independent random variables from any compactly supported distribution with uniformly bounded first and second derivatives and a second order Taylor series expansion at each point. The distribution of digits of any N^δconsecutive differences \emph{and} all N-1 normalized differences of the order statistics exhibit shifted exponential behavior. We derive conditions on the probability density which determine whether or not the distribution of the digits of all the un-normalized differences converges to Benford's Law, shifted exponential behavior, or oscillates between the two, and show that the Pareto distribution leads to oscillating behavior.

preprint2008arXiv

The Distribution of the Largest Non-trivial Eigenvalues in Families of Random Regular Graphs

Recently Friedman proved Alon's conjecture for many families of d-regular graphs, namely that given any epsilon > 0 `most' graphs have their largest non-trivial eigenvalue at most 2 sqrt{d-1}+epsilon in absolute value; if the absolute value of the largest non-trivial eigenvalue is at most 2 sqrt{d-1} then the graph is said to be Ramanujan. These graphs have important applications in communication network theory, allowing the construction of superconcentrators and nonblocking networks, coding theory and cryptography. As many of these applications depend on the size of the largest non-trivial positive and negative eigenvalues, it is natural to investigate their distributions. We show these are well-modeled by the beta=1 Tracy-Widom distribution for several families. If the observed growth rates of the mean and standard deviation as a function of the number of vertices holds in the limit, then in the limit approximately 52% of d-regular graphs from bipartite families should be Ramanujan, and about 27% from non-bipartite families (assuming the largest positive and negative eigenvalues are independent).

preprint2008arXiv

When almost all sets are difference dominated

We investigate the relationship between the sizes of the sum and difference sets attached to a subset of {0,1,...,N}, chosen randomly according to a binomial model with parameter p(N), with N^{-1} = o(p(N)). We show that the random subset is almost surely difference dominated, as N --> oo, for any choice of p(N) tending to zero, thus confirming a conjecture of Martin and O'Bryant. The proofs use recent strong concentration results. Furthermore, we exhibit a threshold phenomenon regarding the ratio of the size of the difference- to the sumset. If p(N) = o(N^{-1/2}) then almost all sums and differences in the random subset are almost surely distinct, and in particular the difference set is almost surely about twice as large as the sumset. If N^{-1/2} = o(p(N)) then both the sum and difference sets almost surely have size (2N+1) - O(p(N)^{-2}), and so the ratio in question is almost surely very close to one. If p(N) = c N^{-1/2} then as c increases from zero to infinity (i.e., as the threshold is crossed), the same ratio almost surely decreases continuously from two to one according to an explicitly given function of c. We also extend our results to the comparison of the generalized difference sets attached to an arbitrary pair of binary linear forms. For certain pairs of forms f and g, we show that there in fact exists a sharp threshold at c_{f,g} N^{-1/2}, for some computable constant c_{f,g}, such that one form almost surely dominates below the threshold, and the other almost surely above it. The heart of our approach involves using different tools to obtain strong concentration of the sizes of the sum and difference sets about their mean values, for various ranges of the parameter p.

preprint2007arXiv

A Symplectic Test of the L-Functions Ratios Conjecture

Recently Conrey, Farmer and Zirnbauer conjectured formulas for the averages over a family of ratios of products of shifted L-functions. Their L-functions Ratios Conjecture predicts both the main and lower order terms for many problems, ranging from n-level correlations and densities to mollifiers and moments to vanishing at the central point. There are now many results showing agreement between the main terms of number theory and random matrix theory; however, there are very few families where the lower order terms are known. These terms often depend on subtle arithmetic properties of the family, and provide a way to break the universality of behavior. The L-functions Ratios Conjecture provides a powerful and tractable way to predict these terms. We test a specific case here, that of the 1-level density for the symplectic family of quadratic Dirichlet characters arising from even fundamental discriminants d \le X. For test functions supported in (-1/3, 1/3) we calculate all the lower order terms up to size O(X^{-1/2+epsilon}) and observe perfect agreement with the conjecture (for test functions supported in (-1, 1) we show agreement up to errors of size O(X^{-epsilon}) for any epsilon). Thus for this family and suitably restricted test functions, we completely verify the Ratios Conjecture's prediction for the 1-level density.

preprint2007arXiv

The Modulo 1 Central Limit Theorem and Benford's Law for Products

We derive a necessary and sufficient condition for the sum of M independent continuous random variables modulo 1 to converge to the uniform distribution in L^1([0,1]), and discuss generalizations to discrete random variables. A consequence is that if X_1, ..., X_M are independent continuous random variables with densities f_1, ..., f_M, for any base B as M \to \infty for many choices of the densities the distribution of the digits of X_1 * ... * X_M converges to Benford's law base B. The rate of convergence can be quantified in terms of the Fourier coefficients of the densities, and provides an explanation for the prevalence of Benford behavior in many diverse systems.

preprint2006arXiv

A Derivation of the Pythagorean Won-Loss Formula in Baseball

It has been noted that in many professional sports leagues a good predictor of a team's won-loss percentage is Bill James' Pythagorean Formula RSobs^c / (RSobs^c + RAobs^c), where RSobs (resp. RAobs) is the observed average number of runs scored (allowed) per game and c is a constant for the league; for baseball the best agreement is when c is about 1.82. We provide a theoretical justification for this formula and value of c by modelling the number of runs scored and allowed in baseball games as independent random variables drawn from Weibull distributions with the same b and c but different a; the probability density f(x;a,b,c) is 0 for x < b and is (c/a) ((x-b)/a)^{c-1} exp(-((x-b)/a)^c) otherwise. This model leads to a predicted won-loss percentage of (RS-b)^c / ((RS-b)^c + (RA-b)^c); here RS (resp. RA) is the mean of the random variable corresponding to runs scored (allowed), and RS - b (resp. RA - b) is an estimator of RSobs (resp. RAobs). An analysis of the 14 American League teams from the 2004 baseball season shows that (1) given that the runs scored and allowed in a game cannot be equal, the runs scored and allowed are statistically independent; (2) the best fit Weibull parameters attained from a least squares or a maximum likelihood analysis give good fits; least squares gives a mean value of c of 1.79 with a standard deviation of .09, and maximum likelihood gives a mean value of c of 1.74 with a standard deviation of .06, which agree beautifully with the observed best value of 1.82 attained by fitting RSobs^c / (RSobs^c + RAobs^c) to the observed winning percentages.

preprint2006arXiv

Closed-Form Bayesian Inferences for the Logit Model via Polynomial Expansions

Articles in Marketing and choice literatures have demonstrated the need for incorporating person-level heterogeneity into behavioral models (e.g., logit models for multiple binary outcomes as studied here). However, the logit likelihood extended with a population distribution of heterogeneity doesn't yield closed-form inferences, and therefore numerical integration techniques are relied upon (e.g., MCMC methods). We present here an alternative, closed-form Bayesian inferences for the logit model, which we obtain by approximating the logit likelihood via a polynomial expansion, and then positing a distribution of heterogeneity from a flexible family that is now conjugate and integrable. For problems where the response coefficients are independent, choosing the Gamma distribution leads to rapidly convergent closed-form expansions; if there are correlations among the coefficients one can still obtain rapidly convergent closed-form expansions by positing a distribution of heterogeneity from a Multivariate Gamma distribution. The solution then comes from the moment generating function of the Multivariate Gamma distribution or in general from the multivariate heterogeneity distribution assumed. Closed-form Bayesian inferences, derivatives (useful for elasticity calculations), population distribution parameter estimates (useful for summarization) and starting values (useful for complicated algorithms) are hence directly available. Two simulation studies demonstrate the efficacy of our approach.

preprint2006arXiv

Distribution of Eigenvalues of Real Symmetric Palindromic Toeplitz Matrices and Circulant Matrices

Consider the ensemble of real symmetric Toeplitz matrices, each independent entry an i.i.d. random variable chosen from a fixed probability distribution p of mean 0, variance 1, and finite higher moments. Previous investigations showed that the limiting spectral measure (the density of normalized eigenvalues) converges weakly and almost surely, independent of p, to a distribution which is almost the standard Gaussian. The deviations from Gaussian behavior can be interpreted as arising from obstructions to solutions of Diophantine equations. We show that these obstructions vanish if instead one considers real symmetric palindromic Toeplitz matrices, matrices where the first row is a palindrome. A similar result was previously proved for a related circulant ensemble through an analysis of the explicit formulas for eigenvalues. By Cauchy's interlacing property and the rank inequality, this ensemble has the same limiting spectral distribution as the palindromic Toeplitz matrices; a consequence of combining the two approaches is a version of the almost sure Central Limit Theorem. Thus our analysis of these Diophantine equations provides an alternate technique for proving limiting spectral measures for certain ensembles of circulant matrices.

preprint2006arXiv

Low lying zeros of L-functions with orthogonal symmetry

We investigate the moments of a smooth counting function of the zeros near the central point of L-functions of weight k cuspidal newforms of prime level N. We split by the sign of the functional equations and show that for test functions whose Fourier transform is supported in (-1/n, 1/n), as N --> oo the first n centered moments are Gaussian. By extending the support to (-1/n-1, 1/n-1), we see non-Gaussian behavior; in particular the odd centered moments are non-zero for such test functions. If we do not split by sign, we obtain Gaussian behavior for support in (-2/n, 2/n) if 2k >= n. The nth centered moments agree with Random Matrix Theory in this extended range, providing additional support for the Katz-Sarnak conjectures. The proof requires calculating multidimensional integrals of the non-diagonal terms in the Bessel-Kloosterman expansion of the Petersson formula. We convert these multidimensional integrals to one-dimensional integrals already considered in the work of Iwaniec-Luo-Sarnak, and derive a new and more tractable expression for the nth centered moments for such test functions. This new formula facilitates comparisons between number theory and random matrix theory for test functions supported in (-1/n-1, 1/n-1) by simplifying the combinatorial arguments. As an application we obtain bounds for the percentage of such cusp forms with a given order of vanishing at the central point.

preprint2006arXiv

The Low Lying Zeros of a GL(4) and a GL(6) family of L-functions

We investigate the large weight (k --> oo) limiting statistics for the low lying zeros of a GL(4) and a GL(6) family of L-functions, {L(s,phi x f): f in H_k(1)} and {L(s,phi times sym^2 f): f in H_k(1)}; here phi is a fixed even Hecke-Maass cusp form and H_k(1) is a Hecke eigenbasis for the space H_k(1) of holomorphic cusp forms of weight k for the full modular group. Katz and Sarnak conjecture that the behavior of zeros near the central point should be well modeled by the behavior of eigenvalues near 1 of a classical compact group. By studying the 1- and 2-level densities, we find evidence of underlying symplectic and SO(even) symmetry, respectively. This should be contrasted with previous results of Iwaniec-Luo-Sarnak for the families {L(s,f): f in H_k(1)} and {L(s,sym^2f): f in H_k(1)}, where they find evidence of orthogonal and symplectic symmetry, respectively. The present examples suggest a relation between the symmetry type of a family and that of its twistings, which will be further studied in a subsequent paper. Both the GL(4) and the GL(6) families above have all even functional equations, and neither is naturally split from an orthogonal family. A folklore conjecture states that such families must be symplectic, which is true for the first family but false for the second. Thus the theory of low lying zeros is more than just a theory of signs of functional equations. An analysis of these families suggest that it is the second moment of the Satake parameters that determines the symmetry group.

preprint2005arXiv

Incomplete Quadratic Exponential Sums in Several Variables

We consider incomplete exponential sums in several variables of the form S(f,n,m) = \frac{1}{2^n} \sum_{x_1 \in \{-1,1\}} ... \sum_{x_n \in \{-1,1\}} x_1 ... x_n e^{2πi f(x)/p}, where m>1 is odd and f is a polynomial of degree d with coefficients in Z/mZ. We investigate the conjecture, originating in a problem in computational complexity, that for each fixed d and m the maximum norm of S(f,n,m) converges exponentially fast to 0 as n grows to infinity. The conjecture is known to hold in the case when m=3 and d=2, but existing methods for studying incomplete exponential sums appear to be insufficient to resolve the question for an arbitrary odd modulus m, even when d=2. In the present paper we develop three separate techniques for studying the problem in the case of quadratic f, each of which establishes a different special case of the conjecture. We show that a bound of the required sort holds for almost all quadratic polynomials, a stronger form of the conjecture holds for all quadratic polynomials with no more than 10 variables, and for arbitrarily many variables the conjecture is true for a class of quadratic polynomials having a special form.

preprint2005arXiv

Investigations of Zeros Near the Central Point of Elliptic Curve L-Functions

We explore the effect of zeros at the central point on nearby zeros of elliptic curve L-functions, especially for one-parameter families of rank r over Q. By the Birch and Swinnerton Dyer Conjecture and Silverman's Specialization Theorem, for t sufficiently large the L-function of each curve E_t in the family has r zeros (called the family zeros) at the central point. We observe experimentally a repulsion of the zeros near the central point, and the repulsion increases with r. There is greater repulsion in the subset of curves of rank r+2 than in the subset of curves of rank r in a rank r family. For curves with comparable conductors, the behavior of rank 2 curves in a rank 0 one-parameter family over Q is statistically different from that of rank 2 curves from a rank 2 family. Unlike excess rank calculations, the repulsion decreases markedly as the conductors increase, and we conjecture that the r family zeros do not repel in the limit. Finally, the differences between adjacent normalized zeros near the central point are statistically independent of the repulsion, family rank and rank of the curves in the subset. Specifically, the differences between adjacent normalized zeros are statistically equal for all curves investigated with rank 0, 2 or 4 and comparable conductors from one-parameter families of rank 0 or 2 over Q.

preprint2005arXiv

Variation in the number of points on elliptic curves and applications to excess rank

Michel proved that for a one-parameter family of elliptic curves over Q(T) with non-constant j(T) that the second moment of the number of solutions modulo p is p^2 + O(p^{3/2}). We show this bound is sharp by studying y^2 = x^3 + Tx^2 + 1. Lower order terms for such moments in a family are related to lower order terms in the n-level densities of Katz and Sarnak, which describe the behavior of the zeros near the central point of the associated L-functions. We conclude by investigating similar families and show how the lower order terms in the second moment may affect the expected bounds for the average rank of families in numerical investigations.

preprint2004arXiv

Constructing Elliptic Curves over $\mathbb{Q}(T)$ with Moderate Rank

We give several new constructions for moderate rank elliptic curves over $\mathbb{Q}(T)$. In particular we construct infinitely many rational elliptic surfaces (not in Weierstrass form) of rank 6 over $\mathbb{Q}$ using polynomials of degree two in $T$. While our method generates linearly independent points, we are able to show the rank is exactly 6 \emph{without} having to verify the points are independent. The method generalizes; however, the higher rank surfaces are not rational, and we need to check that the constructed points are linearly independent.

preprint2003arXiv

1- and 2-Level Densities for Rational Families of Elliptic Curves: Evidence for the Underlying Group Symmetries

Following Katz-Sarnak, Iwaniec-Luo-Sarnak, and Rubinstein, we use the 1- and 2-level densities to study the distribution of low lying zeros for one-parameter rational families of elliptic curves over Q(t). Modulo standard conjectures, for small support the densities agree with Katz and Sarnak's predictions. Further, the densities confirm that the curves' L-functions behave in a manner consistent with having r zeros at the critical point, as predicted by the Birch and Swinnerton-Dyer conjecture. By studying the 2-level densities of some constant sign families, we find the first examples of families of elliptic curves where we can distinguish SO(even) from SO(odd) symmetry.

preprint2003arXiv

Eigenvalue Spacing Distribution for the Ensemble of Real Symmetric Toeplitz Matrices

Consider the ensemble of Real Symmetric Toeplitz Matrices, each entry iidrv from a fixed probability distribution p of mean 0, variance 1, and finite higher moments. The limiting spectral measure (the density of normalized eigenvalues) converges weakly to a new universal distribution with unbounded support, independent of p. This distribution's moments are almost those of the Gaussian's; the deficit may be interpreted in terms of Diophantine obstructions. With a little more work, we obtain almost sure convergence. An investigation of spacings between adjacent normalized eigenvalues looks Poissonian, and not GOE.