Source author record

Marta Casanellas

Marta Casanellas appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

12works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

12 published item(s)

preprint2022arXiv

Designing weights for quartet-based methods when data is heterogeneous across lineages

Homogeneity across lineages is a common assumption in phylogenetics according to which nucleotide substitution rates remain constant in time and do not depend on lineages. This is a simplifying hypothesis which is often adopted to make the process of sequence evolution more tractable. However, its validity has been explored and put into question in several papers. On the other hand, dealing successfully with the general case (heterogeneity across lineages) is one of the key features of phylogenetic reconstruction methods based on algebraic tools. The goal of this paper is twofold. First, we present a new weighting system for quartets (ASAQ) based on algebraic and semi-algebraic tools, thus specially indicated to deal with data evolving under heterogeneus rates. This method combines the weights two previous methods by means of a test based on the positivity of the branch length estimated with the paralinear distance. ASAQ is statistically consistent when applied to GM data, considers rate and base composition heterogeneity among lineages and does not assume stationarity nor time-reversibility. Second, we test and compare the performance of several quartet-based methods for phylogenetic tree reconstruction (namely, Quartet Puzzling, Weight Optimization and Wilson's method) in combination with ASAQ weights and other weights based on algebraic and semi-algebraic methods or on the paralinear distance. These tests are applied to both simulated and real data and support Weight Optimization with ASAQ weights as a reliable and successful reconstruction method.

preprint2021arXiv

Robust estimation of tree structured models

Consider the problem of learning undirected graphical models on trees from corrupted data. Recently Katiyar et al. showed that it is possible to recover trees from noisy binary data up to a small equivalence class of possible trees. Their other paper on the Gaussian case follows a similar pattern. By framing this as a special phylogenetic recovery problem we largely generalize these two settings. Using the framework of linear latent tree models we discuss tree identifiability for binary data under a continuous corruption model. For the Ising and the Gaussian tree model we also provide a characterisation of when the Chow-Liu algorithm consistently learns the underlying tree from the noisy data.

preprint2020arXiv

An open set of $4\times4$ embeddable matrices whose principal logarithm is not a Markov generator

A Markov matrix is embeddable if it can represent a homogeneous continuous-time Markov process. It is well known that if a Markov matrix has real and pairwise-different eigenvalues, then the embeddability can be determined by checking whether its principal logarithm is a rate matrix or not. The same holds for Markov matrices close enough to the identity matrix or that rule a Markov process subjected to certain restrictions. In this paper we prove that this criterion cannot be generalized and we provide open sets of Markov matrices that are embeddable and whose principal logarithm is a not a rate matrix.

preprint2016arXiv

Phylogenetic mixtures and linear invariants for equal input models

The reconstruction of phylogenetic trees from molecular sequence data relies on modelling site substitutions by a Markov process, or a mixture of such processes. In general, allowing mixed processes can result in different tree topologies becoming indistinguishable from the data, even for infinitely long sequences. However, when the underlying Markov process supports linear phylogenetic invariants, then provided these are sufficiently informative, the identifiability of the tree topology can be restored. In this paper, we investigate a class of processes that support linear invariants once the stationary distribution is fixed, the `equal input model'. This model generalizes the `Felsenstein 1981' model (and thereby the Jukes--Cantor model) from four states to an arbitrary number of states (finite or infinite), and it can also be described by a `random cluster' process. We describe the structure and dimension of the vector spaces of phylogenetic mixtures and of linear invariants for any fixed phylogenetic tree (and for all trees -- the so called `model invariants'), on any number $n$ of leaves. We also provide a precise description of the space of mixtures and linear invariants for the special case of $n=4$ leaves. By combining techniques from discrete random processes and (multi-) linear algebra, our results build on a classic result that was first established by James Lake in 1987.

preprint2014arXiv

EM for phylogenetic topology reconstruction on non-homogeneous data

Background: The reconstruction of the phylogenetic tree topology of four taxa is, still nowadays, one of the main challenges in phylogenetics. Its difficulties lie in considering not too restrictive evolutionary models, and correctly dealing with the long-branch attraction problem. The correct reconstruction of 4-taxon trees is crucial for making quartet-based methods work and being able to recover large phylogenies. Results: In this paper we consider an expectation-maximization method for maximizing the likelihood of (time nonhomogeneous) evolutionary Markov models on trees. We study its success on reconstructing 4-taxon topologies and its performance as input method in quartet-based phylogenetic reconstruction methods such as QFIT and QuartetSuite. Our results show that the method proposed here outperforms neighbor-joining and the usual (time-homogeneous continuous-time) maximum likelihood methods on 4-leaved trees with among-lineage instantaneous rate heterogeneity, and perform similarly to usual continuous-time maximum-likelihood when data satisfies the assumptions of both methods. Conclusions: The method presented in this paper is well suited for reconstructing the topology of any number of taxa via quartet-based methods and is highly accurate, specially regarding largely divergent trees and time nonhomogeneous data.

preprint2014arXiv

Invariant versus classical quartet inference when evolution is heterogeneous across sites and lineages

One reason why classical phylogenetic reconstruction methods fail to correctly infer the underlying topology is because they assume oversimplified models. In this paper we propose a topology reconstruction method consistent with the most general Markov model of nucleotide substitution, which can also deal with data coming from mixtures on the same topology. It is based on an idea of Eriksson on using phylogenetic invariants and provides a system of weights that can be used as input of quartet-based methods. We study its performance on real data and on a wide range of simulated 4-taxon data (both time-homogeneous and nonhomogeneous, with or without among-site rate heterogeneity, and with different branch length settings). We compare it to the classical methods of neighbor-joining (with paralinear distance), maximum likelihood (with different underlying models), and maximum parsimony. Our results show that this method is accurate and robust, has a similar performance to ML when data satisfies the assumptions of both methods, and outperforms all methods when these are based on inappropriate substitution models or when both long and short branches are present. If alignments are long enough, then it also outperforms other methods when some of its assumptions are violated.

preprint2014arXiv

Local description of phylogenetic group-based models

Motivated by phylogenetics, our aim is to obtain a system of equations that define a phylogenetic variety on an open set containing the biologically meaningful points. In this paper we consider phylogenetic varieties defined via group-based models. For any finite abelian group $G$, we provide an explicit construction of $codim X$ phylogenetic invariants (polynomial equations) of degree at most $|G|$ that define the variety $X$ on a Zariski open set $U$. The set $U$ contains all biologically meaningful points when $G$ is the group of the Kimura 3-parameter model. In particular, our main result confirms a conjecture by the third author and, on the set $U$, a couple of conjectures by Bernd Sturmfels and Seth Sullivant.

preprint2012arXiv

Empar: EM-based algorithm for parameter estimation of Markov models on trees

The goal of branch length estimation in phylogenetic inference is to estimate the divergence time between a set of sequences based on compositional differences between them. A number of software is currently available facilitating branch lengths estimation for homogeneous and stationary evolutionary models. Homogeneity of the evolutionary process imposes fixed rates of evolution throughout the tree. In complex data problems this assumption is likely to put the results of the analyses in question. In this work we propose an algorithm for parameter and branch lengths inference in the discrete-time Markov processes on trees. This broad class of nonhomogeneous models comprises the general Markov model and all its submodels, including both stationary and nonstationary models. Here, we adapted the well-known Expectation-Maximization algorithm and present a detailed performance study of this approach for a selection of nonhomogeneous evolutionary models. We conducted an extensive performance assessment on multiple sequence alignments simulated under a variety of settings. We demonstrated high accuracy of the tool in parameter estimation and branch lengths recovery, proving the method to be a valuable tool for phylogenetic inference in real life problems. $\empar$ is an open-source C++ implementation of the methods introduced in this paper and is the first tool designed to handle nonhomogeneous data.

preprint2012arXiv

The space of phylogenetic mixtures for equivariant models

The selection of the most suitable evolutionary model to analyze the given molecular data is usually left to biologist's choice. In his famous book, J Felsenstein suggested that certain linear equations satisfied by the expected probabilities of patterns observed at the leaves of a phylogenetic tree could be used for model selection. It remained open the question regarding whether these equations were enough for characterizing the evolutionary model. Here we prove that, for equivariant models of evolution, the space of distributions satisfying these linear equations coincides with the space of distributions arising from mixtures of trees on a set of taxa. In other words, we prove that an alignment is produced from a mixture of phylogenetic trees under an equivariant evolutionary model if and only if its distribution of column patterns satisfies the linear equations mentioned above. Moreover, for each equivariant model and for any number of taxa, we provide a set of linearly independent equations defining this space of phylogenetic mixtures. This is a powerful tool that has already been successfully used in model selection. We also use the results obtained to study identifiability issues for phylogenetic mixtures.

preprint2011arXiv

Generating Markov evolutionary matrices for a given branch length

Under a markovian evolutionary process, the expected number of substitutions per site (also called branch length) that have occurred when a sequence has evolved from another according to a transition matrix $P$ can be approximated by $-1/4log det P.$ When the Markov process is assumed to be continuous in time, i.e. $P=\exp Qt$ it is easy to simulate this evolutionary process for a given branch length (this amounts to requiring $Q$ of a certain trace). For the more general case (what we call discrete-time models), it is not trivial to generate a substitution matrix $P$ of given determinant (i.e. corresponding to a process of given branch length). In this paper we solve this problem for the most well-known discrete-time models JC*, K80*, K81*, SSM and GMM. These models lie in the class of nonhomogeneous evolutionary models. For any of these models we provide concise algorithms to generate matrices $P$ of given determinant. Moreover, in the first four models, our results prove that any of these matrices can be generated in this way. Our techniques are mainly based on algebraic tools.

preprint2011arXiv

Stable Ulrich bundles

The existence of stable ACM vector bundles of high rank on algebraic varieties is a challenging problem. In this paper, we study stable Ulrich bundles (that is, stable ACM bundles whose corresponding module has the maximum number of generators) on nonsingular cubic surfaces $X \subset \mathbb{P}^3.$ We give necessary and sufficient conditions on the first Chern class $D$ for the existence of stable Ulrich bundles on $X$ of rank $r$ and $c_1=D$. When such bundles exist, we prove that that the corresponding moduli space of stable bundles is smooth and irreducible of dimension $D^2-2r^2+1$ and consists entirely of stable Ulrich bundles (see Theorem 1.1). As a consequence, we are also able to prove the existence of stable Ulrich bundles of any rank on nonsingular cubic threefolds in $\mathbb{P}^4$.