Source author record

Chris Peterson

Chris Peterson appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

13works
14topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

13 published item(s)

preprint2025arXiv

A Granular Grassmannian Clustering Framework via the Schubert Variety of Best Fit

In many classification and clustering tasks, it is useful to compute a geometric representative for a dataset or a cluster, such as a mean or median. When datasets are represented by subspaces, these representatives become points on the Grassmann or flag manifold, with distances induced by their geometry, often via principal angles. We introduce a subspace clustering algorithm that replaces subspace means with a trainable prototype defined as a Schubert Variety of Best Fit (SVBF) - a subspace that comes as close as possible to intersecting each cluster member in at least one fixed direction. Integrated in the Linde-Buzo-Grey (LBG) pipeline, this SVBF-LBG scheme yields improved cluster purity on synthetic, image, spectral, and video action data, while retaining the mathematical structure required for downstream analysis.

preprint2022arXiv

Quasi-polynomial growth of numerical and affine semigroups with constrained gaps

A common tool in the theory of numerical semigroups is to interpret a desired class of semigroups as the integer lattice points in a rational polyhedron in order to leverage computational and enumerative techniques from polyhedral geometry. Most arguments of this type make use of a parametrization of numerical semigroups with fixed multiplicity $m$ in terms of their $m$-Apéry sets, giving a representation called Kunz coordinates which obey a collection of inequalities defining the Kunz polyhedron. In this work, we introduce a new class of polyhedra describing numerical semigroups in terms of a truncated addition table of their sporadic elements. Applying a classical theorem of Ehrhart to slices of these polyhedra, we prove that the number of numerical semigroups with $n$ sporadic elements and Frobenius number $f$ is polynomial up to periodicity, or quasi-polynomial, as a function of $f$ for fixed $n$. We also generalize this approach to higher dimensions to demonstrate quasi-polynomial growth of the number of affine semigroups with a fixed number of elements, and all gaps, contained in an integer dilation of a fixed polytope.

preprint2022arXiv

The Flag Median and FlagIRLS

Finding prototypes (e.g., mean and median) for a dataset is central to a number of common machine learning algorithms. Subspaces have been shown to provide useful, robust representations for datasets of images, videos and more. Since subspaces correspond to points on a Grassmann manifold, one is led to consider the idea of a subspace prototype for a Grassmann-valued dataset. While a number of different subspace prototypes have been described, the calculation of some of these prototypes has proven to be computationally expensive while other prototypes are affected by outliers and produce highly imperfect clustering on noisy data. This work proposes a new subspace prototype, the flag median, and introduces the FlagIRLS algorithm for its calculation. We provide evidence that the flag median is robust to outliers and can be used effectively in algorithms like Linde-Buzo-Grey (LBG) to produce improved clusterings on Grassmannians. Numerical experiments include a synthetic dataset, the MNIST handwritten digits dataset, the Mind's Eye video dataset and the UCF YouTube action dataset. The flag median is compared the other leading algorithms for computing prototypes on the Grassmannian, namely, the $\ell_2$-median and to the flag mean. We find that using FlagIRLS to compute the flag median converges in $4$ iterations on a synthetic dataset. We also see that Grassmannian LBG with a codebook size of $20$ and using the flag median produces at least a $10\%$ improvement in cluster purity over Grassmannian LBG using the flag mean or $\ell_2$-median on the Mind's Eye dataset.

preprint2020arXiv

Metric thickenings and group actions

Let $G$ be a group acting properly and by isometries on a metric space $X$; it follows that the quotient or orbit space $X/G$ is also a metric space. We study the Vietoris-Rips and Čech complexes of $X/G$. Whereas (co)homology theories for metric spaces let the scale parameter of a Vietoris-Rips or Čech complex go to zero, and whereas geometric group theory requires the scale parameter to be sufficiently large, we instead consider intermediate scale parameters (neither tending to zero nor to infinity). As a particular case, we study the Vietoris-Rips and Čech thickenings of projective spaces at the first scale parameter where the homotopy type changes.

preprint2020arXiv

The apolar algebra of a product of linear forms

Apolarity is an important tool in commutative algebra and algebraic geometry which studies a form, $f$, by the action of polynomial differential operators on $f$. The quotient of all polynomial differential operators by those which annihilate $f$ is called the apolar algebra of $f$. In general, the apolar algebra of a form is useful for determining its Waring rank, which can be seen as the problem of decomposing the supersymmetric tensor, associated to the form, minimally as a sum of rank one supersymmetric tensors. In this article we study the apolar algebra of a product of linear forms, which generalizes the case of monomials and connects to the geometry of hyperplane arrangements. In the first part of the article we provide a bound on the Waring rank of a product of linear forms under certain genericity assumptions; for this we use the defining equations of so-called star configurations due to Geramita, Harbourne, and Migliore. In the second part of the article we use the computer algebra system Bertini, which operates by homotopy continuation methods, to solve certain rank equations for catalecticant matrices. Our computations suggest that, up to a change of variables, there are exactly six homogeneous polynomials of degree six in three variables which factor completely as a product of linear forms defining an irreducible multi-arrangement and whose apolar algebras have dimension six in degree three. As a consequence of these calculations, we find six cases of such forms with cactus rank six, five of which also have Waring rank six. Among these are products defining subarrangements of the braid and Hessian arrangements.

preprint2020arXiv

The flag manifold as a tool for analyzing and comparing data sets

The shape and orientation of data clouds reflect variability in observations that can confound pattern recognition systems. Subspace methods, utilizing Grassmann manifolds, have been a great aid in dealing with such variability. However, this usefulness begins to falter when the data cloud contains sufficiently many outliers corresponding to stray elements from another class or when the number of data points is larger than the number of features. We illustrate how nested subspace methods, utilizing flag manifolds, can help to deal with such additional confounding factors. Flag manifolds, which are parameter spaces for nested subspaces, are a natural geometric generalization of Grassmann manifolds. To make practical comparisons on a flag manifold, algorithms are proposed for determining the distances between points $[A], [B]$ on a flag manifold, where $A$ and $B$ are arbitrary orthogonal matrix representatives for $[A]$ and $[B]$, and for determining the initial direction of these minimal length geodesics. The approach is illustrated in the context of (hyper) spectral imagery showing the impact of ambient dimension, sample dimension, and flag structure.

preprint2019arXiv

A fractal dimension for measures via persistent homology

We use persistent homology in order to define a family of fractal dimensions, denoted $\mathrm{dim}_{\mathrm{PH}}^i(μ)$ for each homological dimension $i\ge 0$, assigned to a probability measure $μ$ on a metric space. The case of $0$-dimensional homology ($i=0$) relates to work by Michael J Steele (1988) studying the total length of a minimal spanning tree on a random sampling of points. Indeed, if $μ$ is supported on a compact subset of Euclidean space $\mathbb{R}^m$ for $m\ge2$, then Steele's work implies that $\mathrm{dim}_{\mathrm{PH}}^0(μ)=m$ if the absolutely continuous part of $μ$ has positive mass, and otherwise $\mathrm{dim}_{\mathrm{PH}}^0(μ)<m$. Experiments suggest that similar results may be true for higher-dimensional homology $0<i<m$, though this is an open question. Our fractal dimension is defined by considering a limit, as the number of points $n$ goes to infinity, of the total sum of the $i$-dimensional persistent homology interval lengths for $n$ random points selected from $μ$ in an i.i.d. fashion. To some measures $μ,$ we are able to assign a finer invariant, a curve measuring the limiting distribution of persistent homology interval lengths as the number of points goes to infinity. We prove this limiting curve exists in the case of $0$-dimensional homology when $μ$ is the uniform distribution over the unit interval, and conjecture that it exists when $μ$ is the rescaled probability measure for a compact set in Euclidean space with positive Lebesgue measure.

preprint2016arXiv

Eigenschemes and the Jordan canonical form

We study the eigenscheme of a matrix which encodes information about the eigenvectors and generalized eigenvectors of a square matrix. The two main results in this paper are this decomposition encodes the numeric data of the Jordan canonical form of the matrix. We also describe how the eigenscheme can be interpreted as the zero locus of a global section of the tangent bundle on projective space. This interpretation allows one to see eigenvectors and generalized eigenvectors of matrices from an alternative viewpoint.

preprint2016arXiv

Persistent Homology on Grassmann Manifolds for Analysis of Hyperspectral Movies

The existence of characteristic structure, or shape, in complex data sets has been recognized as increasingly important for mathematical data analysis. This realization has motivated the development of new tools such as persistent homology for exploring topological invariants, or features, in large data sets. In this paper we apply persistent homology to the characterization of gas plumes in time dependent sequences of hyperspectral cubes, i.e. the analysis of 4-way arrays. We investigate hyperspectral movies of Long-Wavelength Infrared data monitoring an experimental release of chemical simulant into the air. Our approach models regions of interest within the hyperspectral data cubes as points on the real Grassmann manifold $G(k, n)$ (whose points parameterize the $k$-dimensional subspaces of $\mathbb{R}^n$), contrasting our approach with the more standard framework in Euclidean space. An advantage of this approach is that it allows a sequence of time slices in a hyperspectral movie to be collapsed to a sequence of points in such a way that some of the key structure within and between the slices is encoded by the points on the Grassmann manifold. This motivates the search for topological features, associated with the evolution of the frames of a hyperspectral movie, within the corresponding points on the Grassmann manifold. The proposed mathematical model affords the processing of large data sets while retaining valuable discriminatory information. In this paper, we discuss how embedding our data in the Grassmann manifold, together with topological data analysis, captures dynamical events that occur as the chemical plume is released and evolves.

preprint2016arXiv

Random fields and the enumerative geometry of lines on real and complex hypersurfaces

We derive a formula expressing the average number $E_n$ of real lines on a random hypersurface of degree $2n-3$ in $\mathbb{R}\textrm{P}^n$ in terms of the expected modulus of the determinant of a special random matrix. In the case $n=3$ we prove that the average number of real lines on a random cubic surface in $\mathbb{R}\textrm{P}^3$ equals: $$E_3=6\sqrt{2}-3.$$ Our technique can also be used to express the number $C_n$ of complex lines on a generic hypersurface of degree $2n-3$ in $\mathbb{C}\textrm{P}^n$ in terms of the determinant of a random Hermitian matrix. As a special case we obtain a new proof of the classical statement $C_3=27.$ We determine, at the logarithmic scale, the asymptotic of the quantity $E_n$, by relating it to $C_n$ (whose asymptotic has been recently computed D. Zagier). Specifically we prove that: $$\lim_{n\to \infty}\frac{\log E_n}{\log C_n}=\frac{1}{2}.$$ Finally we show that this approach can be used to compute the number $R_n=(2n-3)!!$ of real lines, counted with their intrinsic signs, on a generic real hypersurface of degree $2n-3$ in $\mathbb{R}\textrm{P}^n$.

preprint2016arXiv

Stratifying High Dimensional Data Based on Proximity to the Convex Hull Boundary

The convex hull of a set of points, $C$, serves to expose extremal properties of $C$ and can help identify elements in $C$ of high interest. For many problems, particularly in the presence of noise, the true vertex set (and facets) may be difficult to determine. One solution is to expand the list of high interest candidates to points lying near the boundary of the convex hull. We propose a quadratic program for the purpose of stratifying points in a data cloud based on proximity to the boundary of the convex hull. For each data point, a quadratic program is solved to determine an associated weight vector. We show that the weight vector encodes geometric information concerning the point's relationship to the boundary of the convex hull. The computation of the weight vectors can be carried out in parallel, and for a fixed number of points and fixed neighborhood size, the overall computational complexity of the algorithm grows linearly with dimension. As a consequence, meaningful computations can be completed on reasonably large, high dimensional data sets.

preprint2012arXiv

Locally Linear Embedding Clustering Algorithm for Natural Imagery

The ability to characterize the color content of natural imagery is an important application of image processing. The pixel by pixel coloring of images may be viewed naturally as points in color space, and the inherent structure and distribution of these points affords a quantization, through clustering, of the color information in the image. In this paper, we present a novel topologically driven clustering algorithm that permits segmentation of the color features in a digital image. The algorithm blends Locally Linear Embedding (LLE) and vector quantization by mapping color information to a lower dimensional space, identifying distinct color regions, and classifying pixels together based on both a proximity measure and color content. It is observed that these techniques permit a significant reduction in color resolution while maintaining the visually important features of images.