Source author record

Raul Rabadan

Raul Rabadan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

18works
15topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

18 published item(s)

preprint2020arXiv

Dose-response modeling in high-throughput cancer drug screenings: An end-to-end approach

Personalized cancer treatments based on the molecular profile of a patient's tumor are an emerging and exciting class of treatments in oncology. As genomic tumor profiling is becoming more common, targeted treatments to specific molecular alterations are gaining traction. To discover new potential therapeutics that may apply to broad classes of tumors matching some molecular pattern, experimentalists and pharmacologists rely on high-throughput, in-vitro screens of many compounds against many different cell lines. We propose a hierarchical Bayesian model of how cancer cell lines respond to drugs in these experiments and develop a method for fitting the model to real-world high-throughput screening data. Through a case study, the model is shown to capture nontrivial associations between molecular features and drug response, such as requiring both wild type TP53 and overexpression of MDM2 to be sensitive to Nutlin-3(a). In quantitative benchmarks, the model outperforms a standard approach in biology, with ~20% lower predictive error on held out data. When combined with a conditional randomization testing procedure, the model discovers biomarkers of therapeutic response that recapitulate known biology and suggest new avenues for investigation. All code for the paper is publicly available at https://github.com/tansey/deep-dose-response.

preprint2020arXiv

MREC: a fast and versatile framework for aligning and matching point clouds with applications to single cell molecular data

Comparing and aligning large datasets is a pervasive problem occurring across many different knowledge domains. We introduce and study MREC, a recursive decomposition algorithm for computing matchings between data sets. The basic idea is to partition the data, match the partitions, and then recursively match the points within each pair of identified partitions. The matching itself is done using black box matching procedures that are too expensive to run on the entire data set. Using an absolute measure of the quality of a matching, the framework supports optimization over parameters including partitioning procedures and matching algorithms. By design, MREC can be applied to extremely large data sets. We analyze the procedure to describe when we can expect it to work well and demonstrate its flexibility and power by applying it to a number of alignment problems arising in the analysis of single cell molecular data.

preprint2016arXiv

A Theory of Taxonomy

A taxonomy is a standardized framework to classify and organize items into categories. Hierarchical taxonomies are ubiquitous, ranging from the classification of organisms to the file system on a computer. Characterizing the typical distribution of items within taxonomic categories is an important question with applications in many disciplines. Ecologists have long sought to account for the patterns observed in species-abundance distributions (the number of individuals per species found in some sample), and computer scientists study the distribution of files per directory. Is there a universal statistical distribution describing how many items are typically found in each category in large taxonomies? Here, we analyze a wide array of large, real-world datasets -- including items lost and found on the New York City transit system, library books, and a bacterial microbiome -- and discover such an underlying commonality. A simple, non-parametric branching model that randomly categorizes items and takes as input only the total number of items and the total number of categories successfully reproduces the abundance distributions in these datasets. This result may shed light on patterns in species-abundance distributions long observed in ecology. The model also predicts the number of taxonomic categories that remain unrepresented in a finite sample.

preprint2016arXiv

Genomic data analysis in tree spaces

Recently, an elegant approach in phylogenetics was introduced by Billera-Holmes-Vogtmann that allows a systematic comparison of different evolutionary histories using the metric geometry of tree spaces. In many problem settings one encounters heavily populated phylogenetic trees, where the large number of leaves encumbers visualization and analysis in the relevant evolutionary moduli spaces. To address this issue, we introduce tree dimensionality reduction, a structured approach to reducing large phylogenetic trees to a distribution of smaller trees. We prove a stability theorem ensuring that small perturbations of the large trees are taken to small perturbations of the resulting distributions. We then present a series of four biologically motivated applications to the analysis of genomic data, spanning cancer and infectious disease. The first quantifies how chemotherapy can disrupt the evolution of common leukemias. The second examines a link between geometric information and the histologic grade in relapsed gliomas, where longer relapse branches were specific to high grade glioma. The third concerns genetic stability of xenograft models of cancer, where heterogeneity at the single cell level increased with later mouse passages. The last studies genetic diversity in seasonal influenza A virus. We apply tree dimensionality reduction to 24 years of longitudinally collected H3N2 hemagglutinin sequences, generating distributions of smaller trees spanning between three and five seasons. A negative correlation is observed between the influenza vaccine effectiveness during a season and the variance of the distributions produced using preceding seasons' sequence data. We also show how tree distributions relate to antigenic clusters and choice of influenza vaccine. Our formalism exposes links between viral genomic data and clinical observables such as vaccine selection and efficacy.

preprint2016arXiv

Inference of Ancestral Recombination Graphs through Topological Data Analysis

The recent explosion of genomic data has underscored the need for interpretable and comprehensive analyses that can capture complex phylogenetic relationships within and across species. Recombination, reassortment and horizontal gene transfer constitute examples of pervasive biological phenomena that cannot be captured by tree-like representations. Starting from hundreds of genomes, we are interested in the reconstruction of potential evolutionary histories leading to the observed data. Ancestral recombination graphs represent potential histories that explicitly accommodate recombination and mutation events across orthologous genomes. However, they are computationally costly to reconstruct, usually being infeasible for more than few tens of genomes. Recently, Topological Data Analysis (TDA) methods have been proposed as robust and scalable methods that can capture the genetic scale and frequency of recombination. We build upon previous TDA developments for detecting and quantifying recombination, and present a novel framework that can be applied to hundreds of genomes and can be interpreted in terms of minimal histories of mutation and recombination events, quantifying the scales and identifying the genomic locations of recombinations. We implement this framework in a software package, called TARGet, and apply it to several examples, including small migration between different populations, human recombination, and horizontal evolution in finches inhabiting the Galápagos Islands.

preprint2015arXiv

Multiscale Topology of Chromatin Folding

The three dimensional structure of DNA in the nucleus (chromatin) plays an important role in many cellular processes. Recent experimental advances have led to high-throughput methods of capturing information about chromatin conformation on genome-wide scales. New models are needed to quantitatively interpret this data at a global scale. Here we introduce the use of tools from topological data analysis to study chromatin conformation. We use persistent homology to identify and characterize conserved loops and voids in contact map data and identify scales of interaction. We demonstrate the utility of the approach on simulated data and then look data from both a bacterial genome and a human cell line. We identify substantial multiscale topology in these datasets.

preprint2015arXiv

Quantifying Reticulation in Phylogenetic Complexes Using Homology

Reticulate evolutionary processes result in phylogenetic histories that cannot be modeled using a tree topology. Here, we apply methods from topological data analysis to molecular sequence data with reticulations. Using a simple example, we demonstrate the correspondence between nontrivial higher homology and reticulate evolution. We discuss the sensitivity of the standard filtration and show cases where reticulate evolution can fail to be detected. We introduce an extension of the standard framework and define the median complex as a construction to recover signal of the frequency and scale of reticulate evolution by inferring and imputing putative ancestral states. Finally, we apply our methods to two datasets from phylogenetics. Our work expands on earlier ideas of using topology to extract important evolutionary features from genomic data.

preprint2014arXiv

Characterizing Scales of Genetic Recombination and Antibiotic Resistance in Pathogenic Bacteria Using Topological Data Analysis

Pathogenic bacteria present a large disease burden on human health. Control of these pathogens is hampered by rampant lateral gene transfer, whereby pathogenic strains may acquire genes conferring resistance to common antibiotics. Here we introduce tools from topological data analysis to characterize the frequency and scale of lateral gene transfer in bacteria, focusing on a set of pathogens of significant public health relevance. As a case study, we examine the spread of antibiotic resistance in Staphylococcus aureus. Finally, we consider the possible role of the human microbiome as a reservoir for antibiotic resistance genes.

preprint2014arXiv

Moduli Spaces of Phylogenetic Trees Describing Tumor Evolutionary Patterns

Cancers follow a clonal Darwinian evolution, with fitter subclones replacing more quiescent cells, ultimately giving rise to macroscopic disease. High-throughput genomics provides the opportunity to investigate these processes and determine specific genetic alterations driving disease progression. Genomic sampling of a patient's cancer provides a molecular history, represented by a phylogenetic tree. Cohorts of patients represent a forest of related phylogenetic structures. To extract clinically relevant information, one must represent and statistically compare these collections of trees. We propose a framework based on an application of the work by Billera, Holmes and Vogtmann on phylogenetic tree spaces to the case of unrooted trees of intra-individual cancer tissue samples. We observe that these tree spaces are globally nonpositively curved, allowing for statistical inference on populations of patient histories. A projective tree space is introduced, permitting visualizations of aggregate evolutionary behavior. Published data from three types of human malignancies are explored within our framework.

preprint2014arXiv

Parametric Inference using Persistence Diagrams: A Case Study in Population Genetics

Persistent homology computes topological invariants from point cloud data. Recent work has focused on developing statistical methods for data analysis in this framework. We show that, in certain models, parametric inference can be performed using statistics defined on the computed invariants. We develop this idea with a model from population genetics, the coalescent with recombination. We apply our model to an influenza dataset, identifying two scales of topological structure which have a distinct biological interpretation.

preprint2011arXiv

Identifying Hosts of Families of Viruses: A Machine Learning Approach

Identifying viral pathogens and characterizing their transmission is essential to developing effective public health measures in response to a pandemic. Phylogenetics, though currently the most popular tool used to characterize the likely host of a virus, can be ambiguous when studying species very distant to known species and when there is very little reliable sequence information available in the early stages of the pandemic. Motivated by an existing framework for representing biological sequence information, we learn sparse, tree-structured models, built from decision rules based on subsequences, to predict viral hosts from protein sequence data using popular discriminative machine learning tools. Furthermore, the predictive motifs robustly selected by the learning algorithm are found to show strong host-specificity and occur in highly conserved regions of the viral proteome.

preprint2011arXiv

Understanding the Origins of a Pandemic Virus

Understanding the origin of infectious diseases provides scientifically based rationales for implementing public health measures that may help to avoid or mitigate future epidemics. The recent ancestors of a pandemic virus provide invaluable information about the set of minimal genomic alterations that transformed a zoonotic agent into a full human pandemic. Since the first confirmed cases of the H1N1 pandemic virus in the spring of 2009, several hypotheses about the strain's origins have been proposed. However, how, where, and when it first infected humans is still far from clear. The only way to piece together this epidemiological puzzle relies on the collective effort of the international scientific community to increase genomic sequencing of influenza isolates, especially ones collected in the months prior to the origin of the pandemic.

preprint2010arXiv

Fractal-like Distributions over the Rational Numbers in High-throughput Biological and Clinical Data

Recent developments in extracting and processing biological and clinical data are allowing quantitative approaches to studying living systems. High-throughput sequencing, expression profiles, proteomics, and electronic health records are some examples of such technologies. Extracting meaningful information from those technologies requires careful analysis of the large volumes of data they produce. In this note, we present a set of distributions that commonly appear in the analysis of such data. These distributions present some interesting features: they are discontinuous in the rational numbers, but continuous in the irrational numbers, and possess a certain self-similar (fractal-like) structure. The first set of examples which we present here are drawn from a high-throughput sequencing experiment. Here, the self-similar distributions appear as part of the evaluation of the error rate of the sequencing technology and the identification of tumorogenic genomic alterations. The other examples are obtained from risk factor evaluation and analysis of relative disease prevalence and co-mordbidity as these appear in electronic clinical data. The distributions are also relevant to identification of subclonal populations in tumors and the study of the evolution of infectious diseases, and more precisely the study of quasi-species and intrahost diversity of viral populations.

preprint2008arXiv

Spectral Signatures of Photon-Particle Oscillations from Celestial Objects

We give detailed predictions for the spectral signatures arising from photon-particle oscillations in astrophysical objects. The calculations include quantum electrodynamic effects as well as those due to active relativistic plasma. We show that, by studying the spectra of compact sources, it may be possible to directly detect (pseudo-)scalar particles, such as the axion, with much greater sensitivity, by roughly three orders of magnitude, than is currently achievable by other methods. In particular, if such particles exist with masses m_a<0.01[eV] and coupling constant to the electromagnetic field, g>1e-13[1/GeV], then their oscillation signatures are likely to be lurking in the spectra of magnetars, pulsars, and quasars.

preprint2005arXiv

Non-perturbative orientifold transitions at the conifold

After orientifold projection, the conifold singularity in hypermultiplet moduli space of Calabi-Yau compactifications cannot be avoided by geometric deformations. We study the non-perturbative fate of this singularity in a local model involving O6-planes and D6-branes wrapping the deformed conifold in Type IIA string theory. We classify possible A-type orientifolds of the deformed conifold and find that they cannot all be continued to the small resolution. When passing through the singularity on the deformed side, the O-plane charge generally jumps by the class of the vanishing cycle. To decide which classical configurations are dynamically connected, we construct the quantum moduli space by lifting the orientifold to M-theory as well as by looking at the superpotential. We find a rich pattern of smooth and phase transitions depending on the total sixbrane charge. Non-BPS states from branes wrapped on non-supersymmetric bolts are responsible for a phase transition. We also clarify the nature of a Z_2 valued D0-brane charge in the 6-brane background. Along the way, we obtain a new metric of G_2 holonomy corresponding to an O6-plane on the three sphere of the deformed conifold.

preprint2003arXiv

(Re)constructing Dimensions

Compactifying a higher-dimensional theory defined in R^{1,3+n} on an n-dimensional manifold {\cal M} results in a spectrum of four-dimensional (bosonic) fields with masses m^2_i = λ_i, where - λ_i are the eigenvalues of the Laplacian on the compact manifold. The question we address in this paper is the inverse: given the masses of the Kaluza-Klein fields in four dimensions, what can we say about the size and shape (i.e. the topology and the metric) of the compact manifold? We present some examples of isospectral manifolds (i.e., different manifolds which give rise to the same Kaluza-Klein mass spectrum). Some of these examples are Ricci-flat, complex and Kähler and so they are isospectral backgrounds for string theory. Utilizing results from finite spectral geometry, we also discuss the accuracy of reconstructing the properties of the compact manifold (e.g., its dimension, volume, and curvature etc) from measuring the masses of only a finite number of Kaluza-Klein modes.

preprint2002arXiv

Complex structure moduli stability in toroidal compactifications

In this paper we present a classification of possible dynamics of closed string moduli within specific toroidal compactifications of Type II string theories due to the NS-NS tadpole terms in the reduced action. They appear as potential terms for the moduli when supersymmetry is broken due to the presence of D-branes. We particularise to specific constructions with two, four and six-dimensional tori, and study the stabilisation of the complex structure moduli at the disk level. We find that, depending on the cycle on the compact space where the brane is wrapped, there are three possible cases: i) there is a solution inside the complex structure moduli space, and the configuration is stable at the critical point, ii) the moduli fields are driven towards the boundary of the moduli space, iii) there is no stable solution at the minimum of the potential and the system decays into a set of branes.

preprint2002arXiv

M-theory lift of brane-antibrane systems and localised closed string tachyons

We discuss the lift of certain D6-antiD6-brane systems to M-theory. These are purely gravitational configurations with a bolt singularity. When reduced along a trivial circle, and for large bolt radius, the bolt is related to a non-supersymmetric orbifold type of singularity where some closed string tachyons are expected in the twisted sectors. This is a kind of open-closed string duality that relates open string tachyons on one side and localised tachyons in the other. We consider the evolution of the system of branes from M-theory point of view. This evolution gives rise to a brane-antibrane annihilation on the brane side. On the gravity side, the evolution is related to a reduction of the order of the orbifold and to a contraction of the bolt to a nut or flat space if the system has non-vanishing or vanishing charge, respectively. We also consider the inverse process of reducing a non-supersymmetric orbifold to a D6-brane system. For $C^2/Z_N\times Z_M$, the reduced system is a fractional D6-brane at an orbifold singularity $C/Z_M$.