Source author record

Max A. Alekseyev

Max A. Alekseyev appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

16works
18topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

16 published item(s)

preprint2026arXiv

Classification of integral modular data up to rank 13

This paper classifies all modular data of integral modular fusion categories up to rank 13. Furthermore, it also classifies all integral half-Frobenius fusion rings up to rank 12. We find that each perfect integral modular fusion category up to rank 13, as well as every perfect integral half-Frobenius fusion ring up to rank 12, is trivial. We have also refined the non-pointed odd-dimensional modular data at ranks below 25 to three items, all of rank 17, FPdim 225, and type [[1,3],[3,8],[5,6]], filling gaps in the literature. For rank 25, we have narrowed down the perfect case to 3 open types. Our initial key insight is that the Egyptian fractions, which are typically employed to list possible types, can be chosen with squared denominators. We then develop several type criteria as initial filters. To obtain the fusion rings, we solve the dimension and associativity equations using new features on Normaliz created specifically for this purpose. The S-matrices (if they exist) are obtained by self-transposing the character table, while the T-matrices are derived by solving the Anderson-Moore-Vafa equations. Finally, we verify the extended axioms of modular data. From rank 13 onward, the types were further restricted by additional properties unique to the modular case, which involved the universal grading, congruence representations of the modular group and Galois action, leading to critical arithmetic constraints. In particular, we get that, up to rank 21, a prime divisor of the global FPdim does not exceed the rank, and more strongly up to rank 15 in the non-pointed case, does not exceed half the rank. Ultimately, we narrowed down the classification at rank 14 to 35 possible types, 8 of which are non-perfect.

preprint2019arXiv

Orienting Ordered Scaffolds: Complexity and Algorithms

Despite the recent progress in genome sequencing and assembly, many of the currently available assembled genomes come in a draft form. Such draft genomes consist of a large number of genomic fragments (scaffolds), whose order and/or orientation (i.e., strand) in the genome are unknown. There exist various scaffold assembly methods, which attempt to determine the order and orientation of scaffolds along the genome chromosomes. Some of these methods (e.g., based on FISH physical mapping, chromatin conformation capture, etc.) can infer the order of scaffolds, but not necessarily their orientation. This leads to a special case of the scaffold orientation problem (i.e., deducing the orientation of each scaffold) with a known order of the scaffolds. We address the problem of orientating ordered scaffolds as an optimization problem based on given weighted orientations of scaffolds and their pairs (e.g., coming from pair-end sequencing reads, long reads, or homologous relations). We formalize this problem using notion of a scaffold graph (i.e., a graph, where vertices correspond to the assembled contigs or scaffolds and edges represent connections between them). We prove that this problem is NP-hard, and present a polynomial-time algorithm for solving its special case, where orientation of each scaffold is imposed relatively to at most two other scaffolds. We further develop an FPT algorithm for the general case of the OOS problem.

preprint2016arXiv

Combinatorial Scoring of Phylogenetic Networks

Construction of phylogenetic trees and networks for extant species from their characters represents one of the key problems in phylogenomics. While solution to this problem is not always uniquely defined and there exist multiple methods for tree/network construction, it becomes important to measure how well the constructed networks capture the given character relationship across the species. In the current study, we propose a novel method for measuring the specificity of a given phylogenetic network in terms of the total number of distributions of character states at the leaves that the network may impose. While for binary phylogenetic trees, this number has an exact formula and depends only on the number of leaves and character states but not on the tree topology, the situation is much more complicated for non-binary trees or networks. Nevertheless, we develop an algorithm for combinatorial enumeration of such distributions, which is applicable for arbitrary trees and networks under some reasonable assumptions.

preprint2016arXiv

Computing the inverses, their power sums, and extrema for Euler's totient and other multiplicative functions

We propose a generic algorithm for computing the inverses of a multiplicative function under the assumption that the set of inverses is finite. More generally, our algorithm can compute certain functions of the inverses, such as their power sums (e.g., cardinality) or extrema, without direct enumeration of the inverses. We illustrate our algorithm with Euler's totient function $φ(\cdot)$ and the $k$-th power sum of divisors $σ_k(\cdot)$. For example, we can establish that the number of solutions to $σ_1(x) = 10^{1000}$ is 15,512,215,160,488,452,125,793,724,066,873,737,608,071,476, while it is intractable to iterate over the actual solutions.

preprint2016arXiv

Weighted de Bruijn Graphs for the Menage Problem and Its Generalizations

We address the problem of enumeration of seating arrangements of married couples around a circular table such that no spouses sit next to each other and no k consecutive persons are of the same gender. While the case of k=2 corresponds to the classical problème des ménages with a well-studied solution, no closed-form expression for the number of seating arrangements is known when k>=3. We propose a novel approach for this type of problems based on enumeration of walks in certain algebraically weighted de Bruijn graphs. Our approach leads to new expressions for the menage numbers and their exponential generating function and allows one to efficiently compute the number of seating arrangements in general cases, which we illustrate in detail for the ternary case of k=3.

preprint2015arXiv

A Computational Method for the Rate Estimation of Evolutionary Transpositions

Genome rearrangements are evolutionary events that shuffle genomic architectures. Most frequent genome rearrangements are reversals, translocations, fusions, and fissions. While there are some more complex genome rearrangements such as transpositions, they are rarely observed and believed to constitute only a small fraction of genome rearrangements happening in the course of evolution. The analysis of transpositions is further obfuscated by intractability of the underlying computational problems. We propose a computational method for estimating the rate of transpositions in evolutionary scenarios between genomes. We applied our method to a set of mammalian genomes and estimated the transpositions rate in mammalian evolution to be around 0.26.

preprint2014arXiv

On integral points on biquadratic curves and near-multiples of squares in Lucas sequences

We describe an algorithmic reduction of the search for integral points on a curve y^2 = ax^4 + bx^2 + c with nonzero ac(b^2-4ac) to solving a finite number of Thue equations. While existence of such reduction is anticipated from arguments of algebraic number theory, our algorithm is elementary and to best of our knowledge is the first published algorithm of this kind. In combination with other methods and powered by existing software Thue equations solvers, it allows one to efficiently compute integral points on biquadratic curves. We illustrate this approach with a particular application of finding near-multiples of squares in Lucas sequences. As an example, we establish that among Fibonacci numbers only 2 and 34 are of the form 2m^2+2; only 1, 13, and 1597 are of the form m^2-3; and so on. As an auxiliary result, we also give an algorithm for solving a Diophantine equation k^2 = f(m,n)/g(m,n) in integers m,n,k, where f and g are homogeneous quadratic polynomials.

preprint2014arXiv

On the minimal teaching sets of two-dimensional threshold functions

It is known that a minimal teaching set of any threshold function on the twodimensional rectangular grid consists of 3 or 4 points. We derive exact formulae for the numbers of functions corresponding to these values and further refine them in the case of a minimal teaching set of size 3. We also prove that the average cardinality of the minimal teaching sets of threshold functions is asymptotically 7/2. We further present corollaries of these results concerning some special arrangements of lines in the plane.

preprint2013arXiv

On the number of permutations with bounded run lengths

In this work we obtain recurrent formulae for the number of permutations with either increasing or monotonic (i.e., both increasing and decreasing) runs of bounded length. Our formulae allow one to efficiently compute the number of such permutations. In particular, we use the formulae to find and correct a few miscalculations in the classic 1966 book by David, Kendall, and Barton. We further use our formulae to derive differential equations for the corresponding exponential generating functions. In the case of increasing runs, we solve these equations and obtain closed-form expressions for the generating functions.

preprint2012arXiv

BayesHammer: Bayesian clustering for error correction in single-cell sequencing

Error correction of sequenced reads remains a difficult task, especially in single-cell sequencing projects with extremely non-uniform coverage. While existing error correction tools designed for standard (multi-cell) sequencing data usually come up short in single-cell sequencing projects, algorithms actually used for single-cell error correction have been so far very simplistic. We introduce several novel algorithms based on Hamming graphs and Bayesian subclustering in our new error correction tool BayesHammer. While BayesHammer was designed for single-cell sequencing, we demonstrate that it also improves on existing error correction tools for multi-cell sequencing data while working much faster on real-life datasets. We benchmark BayesHammer on both $k$-mer counts and actual assembly results with the SPAdes genome assembler.

preprint2012arXiv

On pairwise distances and median score of three genomes under DCJ

In comparative genomics, the rearrangement distance between two genomes (equal the minimal number of genome rearrangements required to transform them into a single genome) is often used for measuring their evolutionary remoteness. Generalization of this measure to three genomes is known as the median score (while a resulting genome is called median genome). In contrast to the rearrangement distance between two genomes which can be computed in linear time, computing the median score for three genomes is NP-hard. This inspires a quest for simpler and faster approximations for the median score, the most natural of which appears to be the halved sum of pairwise distances which in fact represents a lower bound for the median score. In this work, we study relationship and interplay of pairwise distances between three genomes and their median score under the model of Double-Cut-and-Join (DCJ) rearrangements. Most remarkably we show that while a rearrangement may change the sum of pairwise distances by at most 2 (and thus change the lower bound by at most 1), even the most "powerful" rearrangements in this respect that increase the lower bound by 1 (by moving one genome farther away from each of the other two genomes), which we call strong, do not necessarily affect the median score. This observation implies that the two measures are not as well-correlated as one's intuition may suggest. We further prove that the median score attains the lower bound exactly on the triples of genomes that can be obtained from a single genome with strong rearrangements. While the sum of pairwise distances with the factor 2/3 represents an upper bound for the median score, its tightness remains unclear. Nonetheless, we show that the difference of the median score and its lower bound is not bounded by a constant.

preprint2011arXiv

On convergence of the Flint Hills series

It is not known whether the Flint Hills series $\sum_{n=1}^{\infty} \frac{1}{n^3\cdot\sin(n)^2}$ converges. We show that this question is closely related to the irrationality measure of $π$, denoted $μ(π)$. In particular, convergence of the Flint Hills series would imply $μ(π) \leq 2.5$ which is much stronger than the best currently known upper bound $μ(π)\leq 7.6063...$. This result easily generalizes to series of the form $\sum_{n=1}^{\infty} \frac{1}{n^u\cdot |\sin(n)|^v}$ where $u,v>0$. We use the currently known bound for $μ(π)$ to derive conditions on $u$ and $v$ that guarantee convergence of such series.

preprint2010arXiv

Limited Lifespan of Fragile Regions in Mammalian Evolution

An important question in genome evolution is whether there exist fragile regions (rearrangement hotspots) where chromosomal rearrangements are happening over and over again. Although nearly all recent studies supported the existence of fragile regions in mammalian genomes, the most comprehensive phylogenomic study of mammals (Ma et al. (2006) Genome Research 16, 1557-1565) raised some doubts about their existence. We demonstrate that fragile regions are subject to a "birth and death" process, implying that fragility has limited evolutionary lifespan. This finding implies that fragile regions migrate to different locations in different mammals, explaining why there exist only a few chromosomal breakpoints shared between different lineages. The birth and death of fragile regions phenomenon reinforces the hypothesis that rearrangements are promoted by matching segmental duplications and suggests putative locations of the currently active fragile regions in the human genome.

preprint2010arXiv

On the intersections of Fibonacci, Pell, and Lucas numbers

We describe how to compute the intersection of two Lucas sequences of the forms $\{U_n(P,\pm 1) \}_{n=0}^{\infty}$ or $\{V_n(P,\pm 1) \}_{n=0}^{\infty}$ with $P\in\mathbb{Z}$ that includes sequences of Fibonacci, Pell, Lucas, and Lucas-Pell numbers. We prove that such an intersection is finite except for the case $U_n(1,-1)$ and $U_n(3,1)$ and the case of two $V$-sequences when the product of their discriminants is a perfect square. Moreover, the intersection in these cases also forms a Lucas sequence. Our approach relies on solving homogeneous quadratic Diophantine equations and Thue equations. In particular, we prove that 0, 1, 2, and 5 are the only numbers that are both Fibonacci and Pell, and list similar results for many other pairs of Lucas sequences. We further extend our results to Lucas sequences with arbitrary initial terms.

preprint2010arXiv

On the number of two-dimensional threshold functions

A two-dimensional threshold function of k-valued logic can be viewed as coloring of the points of a k x k square lattice into two colors such that there exists a straight line separating points of different colors. For the number of such functions only asymptotic bounds are known. We give an exact formula for the number of two-dimensional threshold functions and derive more accurate asymptotics.

preprint2010arXiv

Weighted genomic distance can hardly impose a bound on the proportion of transpositions

Genomic distance between two genomes, i.e., the smallest number of genome rearrangements required to transform one genome into the other, is often used as a measure of evolutionary closeness of the genomes in comparative genomics studies. However, in models that include rearrangements of significantly different "power" such as reversals (that are "weak" and most frequent rearrangements) and transpositions (that are more "powerful" but rare), the genomic distance typically corresponds to a transformation with a large proportion of transpositions, which is not biologically adequate. Weighted genomic distance is a traditional approach to bounding the proportion of transpositions by assigning them a relative weight α > 1. A number of previous studies addressed the problem of computing weighted genomic distance with α \leq 2. Employing the model of multi-break rearrangements on circular genomes, that captures both reversals (modelled as 2-breaks) and transpositions (modelled as 3-breaks), we prove that for α \in (1,2], a minimum-weight transformation may entirely consist of transpositions, implying that the corresponding weighted genomic distance does not actually achieve its purpose of bounding the proportion of transpositions. We further prove that for α \in (1,2), the minimum-weight transformations do not depend on a particular choice of α from this interval. We give a complete characterization of such transformations and show that they coincide with the transformations that at the same time have the shortest length and make the smallest number of breakages in the genomes. Our results also provide a theoretical foundation for the empirical observation that for α < 2, transpositions are favored over reversals in the minimum-weight transformations.