Source author record

Valentina Boeva

Valentina Boeva appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
1topics
2close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2016arXiv

Dynamic read mapping and online consensus calling for better variant detection

Variant detection from high-throughput sequencing data is an essential step in identification of alleles involved in complex diseases and cancer. To deal with these massive data, elaborated sequence analysis pipelines are employed. A core component of such pipelines is a read mapping module whose accuracy strongly affects the quality of resulting variant calls. We propose a dynamic read mapping approach that significantly improves read alignment accuracy. The general idea of dynamic mapping is to continuously update the reference sequence on the basis of previously computed read alignments. Even though this concept already appeared in the literature, we believe that our work provides the first comprehensive analysis of this approach. To evaluate the benefit of dynamic mapping, we developed a software pipeline (http://github.com/karel-brinda/dymas) that mimics different dynamic mapping scenarios. The pipeline was applied to compare dynamic mapping with the conventional static mapping and, on the other hand, with the so-called iterative referencing - a computationally expensive procedure computing an optimal modification of the reference that maximizes the overall quality of all alignments. We conclude that in all alternatives, dynamic mapping results in a much better accuracy than static mapping, approaching the accuracy of iterative referencing. To correct the reference sequence in the course of dynamic mapping, we developed an online consensus caller named OCOCO (http://github.com/karel-brinda/ococo). OCOCO is the first consensus caller capable to process input reads in the online fashion. Finally, we provide conclusions about the feasibility of dynamic mapping and discuss main obstacles that have to be overcome to implement it. We also review a wide range of possible applications of dynamic mapping with a special emphasis on variant detection.

preprint2015arXiv

RNF: a general framework to evaluate NGS read mappers

Aligning reads to a reference sequence is a fundamental step in numerous bioinformatics pipelines. As a consequence, the sensitivity and precision of the mapping tool, applied with certain parameters to certain data, can critically affect the accuracy of produced results (e.g., in variant calling applications). Therefore, there has been an increasing demand of methods for comparing mappers and for measuring effects of their parameters. Read simulators combined with alignment evaluation tools provide the most straightforward way to evaluate and compare mappers. Simulation of reads is accompanied by information about their positions in the source genome. This information is then used to evaluate alignments produced by the mapper. Finally, reports containing statistics of successful read alignments are created. In default of standards for encoding read origins, every evaluation tool has to be made explicitly compatible with the simulator used to generate reads. In order to solve this obstacle, we have created a generic format RNF (Read Naming Format) for assigning read names with encoded information about original positions. Futhermore, we have developed an associated software package RNF containing two principal components. MIShmash applies one of popular read simulating tools (among DwgSim, Art, Mason, CuReSim etc.) and transforms the generated reads into RNF format. LAVEnder evaluates then a given read mapper using simulated reads in RNF format. A special attention is payed to mapping qualities that serve for parametrization of ROC curves, and to evaluation of the effect of read sample contamination.

preprint2014arXiv

Deciphering regulation in eukaryotic cell: from sequence to function

A transversal topic of my research has been the development and application of computational methods for DNA sequence analysis. The methods I have been developing aim at improving our understanding of the regulation processes happening in normal and cancer cells. This topic connects together the projects presented in this thesis. Two chapters of the thesis represent major areas of my research interests: (1) methods for deciphering transcriptional regulation and their application to answer specific biological questions, and (2) methods to study the genome structure and their application in cancer studies. The first chapter predominantly focuses on transcriptional regulation. Here I describe my contribution to the development of methodology for the discovery of transcription factor binding sites and the positioning of histone proteins. I also explain how sequence analysis, in combination with gene expression data, can allow the identification of direct target genes of a transcription factor under study, as well as the physical mechanisms of its action. As two examples, I provide the results of my study of transcriptional regulation by (i) oncogenic protein EWS-FLI1 in Ewing sarcoma and (ii) oncogenic transcription factor Spi-1/PU.1 in erythroleukemia. In the second chapter, I describe the sequence analysis methods aimed at the identification of the genomic rearrangements in species with existing reference genome. I explain how the developed methodology can be applied to detect the structure of cancer genomes. I provide an example of how such an analysis of tumor genomes can result in a discovery of a new phenomenon: chromothripsis, when hundreds of rearrangements occur in a single cellular catastrophe. The thesis is concluded by listing the major challenges in high-throughput sequencing analysis. I also discuss the current top questions demanding the integration of sequencing data.