Source author record

C. Donalek

C. Donalek appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

20works
11topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

20 published item(s)

preprint2016arXiv

Real-Time Data Mining of Massive Data Streams from Synoptic Sky Surveys

The nature of scientific and technological data collection is evolving rapidly: data volumes and rates grow exponentially, with increasing complexity and information content, and there has been a transition from static data sets to data streams that must be analyzed in real time. Interesting or anomalous phenomena must be quickly characterized and followed up with additional measurements via optimal deployment of limited assets. Modern astronomy presents a variety of such phenomena in the form of transient events in digital synoptic sky surveys, including cosmic explosions (supernovae, gamma ray bursts), relativistic phenomena (black hole formation, jets), potentially hazardous asteroids, etc. We have been developing a set of machine learning tools to detect, classify and plan a response to transient events for astronomy applications, using the Catalina Real-time Transient Survey (CRTS) as a scientific and methodological testbed. The ability to respond rapidly to the potentially most interesting events is a key bottleneck that limits the scientific returns from the current and anticipated synoptic sky surveys. Similar challenge arise in other contexts, from environmental monitoring using sensor networks to autonomous spacecraft systems. Given the exponential growth of data rates, and the time-critical response, we need a fully automated and robust approach. We describe the results obtained to date, and the possible future developments.

preprint2015arXiv

A serendipitous all sky survey for bright objects in the outer solar system

We use seven year's worth of observations from the Catalina Sky Survey and the Siding Spring Survey covering most of the northern and southern hemisphere at galactic latitudes higher than 20 degrees to search for serendipitously imaged moving objects in the outer solar system. These slowly moving objects would appear as stationary transients in these fast cadence asteroids surveys, so we develop methods to discover objects in the outer solar system using individual observations spaced by months, rather than spaced by hours, as is typically done. While we independently discover 8 known bright objects in the outer solar system, the faintest having $V=19.8\pm0.1$, no new objects are discovered. We find that the survey is nearly 100% efficient at detecting objects beyond 25 AU for $V\lesssim 19.1$ ($V\lesssim18.6$ in the southern hemisphere) and that the probability that there is one or more remaining outer solar system object of this brightness left to be discovered in the unsurveyed regions of the galactic plane is approximately 32%.

preprint2014arXiv

Automated Real-Time Classification and Decision Making in Massive Data Streams from Synoptic Sky Surveys

The nature of scientific and technological data collection is evolving rapidly: data volumes and rates grow exponentially, with increasing complexity and information content, and there has been a transition from static data sets to data streams that must be analyzed in real time. Interesting or anomalous phenomena must be quickly characterized and followed up with additional measurements via optimal deployment of limited assets. Modern astronomy presents a variety of such phenomena in the form of transient events in digital synoptic sky surveys, including cosmic explosions (supernovae, gamma ray bursts), relativistic phenomena (black hole formation, jets), potentially hazardous asteroids, etc. We have been developing a set of machine learning tools to detect, classify and plan a response to transient events for astronomy applications, using the Catalina Real-time Transient Survey (CRTS) as a scientific and methodological testbed. The ability to respond rapidly to the potentially most interesting events is a key bottleneck that limits the scientific returns from the current and anticipated synoptic sky surveys. Similar challenge arise in other contexts, from environmental monitoring using sensor networks to autonomous spacecraft systems. Given the exponential growth of data rates, and the time-critical response, we need a fully automated and robust approach. We describe the results obtained to date, and the possible future developments.

preprint2014arXiv

Cataclysmic Variables from the Catalina Real-time Transient Survey

We present 855 cataclysmic variable candidates detected by the Catalina Real-time Transient Survey (CRTS) of which at least 137 have been spectroscopically confirmed and 705 are new discoveries. The sources were identified from the analysis of five years of data, and come from an area covering three quarters of the sky. We study the amplitude distribution of the dwarf novae CVs discovered by CRTS during outburst, and find that in quiescence they are typically two magnitudes fainter compared to the spectroscopic CV sample identified by SDSS. However, almost all CRTS CVs in the SDSS footprint have ugriz photometry. We analyse the spatial distribution of the CVs and find evidence that many of the systems lie at scale heights beyond those expected for a Galactic thin disc population. We compare the outburst rates of newly discovered CRTS CVs with the previously known CV population, and find no evidence for a difference between them. However, we find that significant evidence for a systematic difference in orbital period distribution. We discuss the CVs found below the orbital period minimum and argue that many more are yet to be identified among the full CRTS CV sample. We cross-match the CVs with archival X-ray catalogs and find that most of the systems are dwarf novae rather than magnetic CVs.

preprint2014arXiv

Data Driven Discovery in Astrophysics

We review some aspects of the current state of data-intensive astronomy, its methods, and some outstanding data analysis challenges. Astronomy is at the forefront of "big data" science, with exponentially growing data volumes and data rates, and an ever-increasing complexity, now entering the Petascale regime. Telescopes and observatories from both ground and space, covering a full range of wavelengths, feed the data via processing pipelines into dedicated archives, where they can be accessed for scientific analysis. Most of the large archives are connected through the Virtual Observatory framework, that provides interoperability standards and services, and effectively constitutes a global data grid of astronomy. Making discoveries in this overabundance of data requires applications of novel, machine learning tools. We describe some of the recent examples of such applications.

preprint2014arXiv

The Catalina Surveys Periodic Variable Star Catalog

We present ~47,000 periodic variables found during the analysis of 5.4 million variable star candidates within a 20,000 square degree region covered by the Catalina Surveys Data Release-1 (CSDR1). Combining these variables with type-ab RR Lyrae from our previous work, we produce an on-line catalog containing periods, amplitudes, and classifications for ~61,000 periodic variables. By cross-matching these variables with those from prior surveys, we find that > 90% of the ~8,000 known periodic variables in the survey region are recovered. For these sources we find excellent agreement between our catalog and prior values of luminosity, period and amplitude, as well as classification. We investigate the rate of confusion between objects classified as contact binaries and type-c RR Lyrae (RRc's) based on periods, colours, amplitudes, metalicities, radial velocities and surface gravities. We find that no more than few percent of these variables in these classes are misidentified. By deriving distances for this clean sample of ~5,500 RRc's, we trace the path of the Sagittarius tidal streams within the Galactic halo. Selecting 146 outer-halo RRc's with SDSS radial velocities, we confirm the presence of a coherent halo structure that is inconsistent with current N-body simulations of the Sagittarius tidal stream. We also find numerous long-period variables that are very likely associated within the Sagittarius tidal streams system. Based on the examination of 31,000 contact binary light curves we find evidence for two subgroups exhibiting irregular lightcurves. One subgroup presents significant variations in mean brightness that are likely due to chromospheric activity. The other subgroup shows stable modulations over more than a thousand days and thereby provides evidence that the O'Connell effect is not due to stellar spots.

preprint2014arXiv

Ultra-short Period Binaries from the Catalina Surveys

We investigate the properties of 367 ultra-short period binary candidates selected from 31,000 sources recently identified from Catalina Surveys data. Based on light curve morphology, along with WISE, SDSS and GALEX multi-colour photometry, we identify two distinct groups of binaries with periods below the 0.22 day contact binary minimum. In contrast to most recent work, we spectroscopically confirm the existence of M-dwarf+M-dwarf contact binary systems. By measuring the radial velocity variations for five of the shortest-period systems, we find examples of rare cool-white dwarf+M-dwarf binaries. Only a few such systems are currently known. Unlike warmer white dwarf systems, their UV flux and their optical colours and spectra are dominated by the M-dwarf companion. We contrast our discoveries with previous photometrically-selected ultra-short period contact binary candidates, and highlight the ongoing need for confirmation using spectra and associated radial velocity measurements. Overall, our analysis increases the number of ultra-short period contact binary candidates by more than an order of magnitude.

preprint2013arXiv

Evidence for a Milky Way Tidal Stream Reaching Beyond 100 kpc

We present the analysis of 1,207 RR Lyrae found in photometry taken by the Catalina Survey's Mount Lemmon telescope. By combining accurate distances for these stars with measurements for ~14,000 type-AB RR Lyrae from the Catalina Schmid telescope, we reveal an extended association that reaches Galactocentric distances beyond 100 kpc and overlaps the Sagittarius streams system. This result confirms earlier evidence for the existence of an outer halo tidal stream resulting from a disrupted stellar system. By comparing the RR Lyrae source density with that expected based on halo models, we find the detection has ~8 sigma significance. We investigate the distances, radial velocities, metallicities, and period-amplitude distribution of the RR Lyrae. We find that both radial velocities and distances are inconsistent with current models of the Sagittarius stream. We also find tentative evidence for a division in source metallicities for the most distant sources. Following prior analyses, we compare the locations and distances of the RR Lyrae with photometrically selected candidate horizontal branch stars and find supporting evidence that this structure spans at least 60 deg of the sky. We investigate the prospects of an association between the stream and unusual globular cluster NGC 2419.

preprint2012arXiv

CLaSPS: a new methodology for Knowledge extraction from complex astronomical dataset

In this paper we present the Clustering-Labels-Score Patterns Spotter (CLaSPS), a new methodology for the determination of correlations among astronomical observables in complex datasets, based on the application of distinct unsupervised clustering techniques. The novelty in CLaSPS is the criterion used for the selection of the optimal clusterings, based on a quantitative measure of the degree of correlation between the cluster memberships and the distribution of a set of observables, the labels, not employed for the clustering. In this paper we discuss the applications of CLaSPS to two simple astronomical datasets, both composed of extragalactic sources with photometric observations at different wavelengths from large area surveys. The first dataset, CSC+, is composed of optical quasars spectroscopically selected in the SDSS data, observed in the X-rays by Chandra and with multi-wavelength observations in the near-infrared, optical and ultraviolet spectral intervals. One of the results of the application of CLaSPS to the CSC+ is the re-identification of a well-known correlation between the alphaOX parameter and the near ultraviolet color, in a subset of CSC+ sources with relatively small values of the near-ultraviolet colors. The other dataset consists of a sample of blazars for which photometric observations in the optical, mid and near infrared are available, complemented for a subset of the sources, by Fermi gamma-ray data. The main results of the application of CLaSPS to such datasets have been the discovery of a strong correlation between the multi-wavelength color distribution of blazars and their optical spectral classification in BL Lacs and Flat Spectrum Radio Quasars and a peculiar pattern followed by blazars in the WISE mid-infrared colors space. This pattern and its physical interpretation have been discussed in details in other papers by one of the authors.

preprint2012arXiv

Flashes in a Star Stream: Automated Classification of Astronomical Transient Events

An automated, rapid classification of transient events detected in the modern synoptic sky surveys is essential for their scientific utility and effective follow-up using scarce resources. This presents some unusual challenges: the data are sparse, heterogeneous and incomplete; evolving in time; and most of the relevant information comes not from the data stream itself, but from a variety of archival data and contextual information (spatial, temporal, and multi-wavelength). We are exploring a variety of novel techniques, mostly Bayesian, to respond to these challenges, using the ongoing CRTS sky survey as a testbed. The current surveys are already overwhelming our ability to effectively follow all of the potentially interesting events, and these challenges will grow by orders of magnitude over the next decade as the more ambitious sky surveys get under way. While we focus on an application in a specific domain (astrophysics), these challenges are more broadly relevant for event or anomaly detection and knowledge discovery in massive data streams.

preprint2012arXiv

Probing the Outer Galactic halo with RR Lyrae from the Catalina Surveys

We present the analysis of 12227 type-ab RR Lyrae found among the 200 million public lightcurves in the Catalina Surveys Data Release 1 (CSDR1). These stars span the largest volume of the Milky Way ever surveyed with RR Lyrae, covering ~20,000 square degrees of the sky (0 < RA < 360, -22 < Dec < 65 deg) to heliocentric distances of up to 60kpc. Each of the RR Lyrae are observed between 60 and 419 times over a six-year period. Using period finding and Fourier fitting techniques we determine periods and apparent magnitudes for each source. We find that the periods at generally accurate to sigma = 0.002% by comparison with 2842 previously known RR Lyrae and 100 RR Lyrae observed in overlapping survey fields. We photometrically calibrate the light curves using 445 Landolt standard stars and show that the resulting magnitudes are accurate to ~0.05 mags using SDSS data for ~1000 blue horizontal branch stars and 7788 of the RR Lyrae. By combining Catalina photometry with SDSS spectroscopy, we analyze the radial velocity and metallicity distributions for > 1500 of the RR Lyrae. Using the accurate distances derived for the RR Lyrae, we show the paths of the Sagittarius tidal streams crossing the sky at heliocentric distances from 20 to 60 kpc. By selecting samples of Galactic halo RR Lyrae, we compare their velocity, metallicity, and distance with predictions from a recent detailed N-body model of the Sagittarius system. We find that there are some significant differences between the distances and structures predicted and our observations.

preprint2012arXiv

Sky Surveys

Sky surveys represent a fundamental data basis for astronomy. We use them to map in a systematic way the universe and its constituents, and to discover new types of objects or phenomena. We review the subject, with an emphasis on the wide-field imaging surveys, placing them in a broader scientific and historical context. Surveys are the largest data generators in astronomy, propelled by the advances in information and computation technology, and have transformed the ways in which astronomy is done. We describe the variety and the general properties of surveys, the ways in which they may be quantified and compared, and offer some figures of merit that can be used to compare their scientific discovery potential. Surveys enable a very wide range of science; that is perhaps their key unifying characteristic. As new domains of the observable parameter space open up thanks to the advances in technology, surveys are often the initial step in their exploration. Science can be done with the survey data alone or a combination of different surveys, or with a targeted follow-up of potentially interesting selected sources. Surveys can be used to generate large, statistical samples of objects that can be studied as populations, or as tracers of larger structures. They can be also used to discover or generate samples of rare or unusual objects, and may lead to discoveries of some previously unknown types. We discuss a general framework of parameter spaces that can be used for an assessment and comparison of different surveys, and the strategies for their scientific exploration. As we move into the Petascale regime, an effective processing and scientific exploitation of such large data sets and data streams poses many challenges, some of which may be addressed in the framework of Virtual Observatory and Astroinformatics, with a broader application of data mining and knowledge discovery technologies.

preprint2011arXiv

Discovery, classification, and scientific exploration of transient events from the Catalina Real-time Transient Survey

Exploration of the time domain - variable and transient objects and phenomena - is rapidly becoming a vibrant research frontier, touching on essentially every field of astronomy and astrophysics, from the Solar system to cosmology. Time domain astronomy is being enabled by the advent of the new generation of synoptic sky surveys that cover large areas on the sky repeatedly, and generating massive data streams. Their scientific exploration poses many challenges, driven mainly by the need for a real-time discovery, classification, and follow-up of the interesting events. Here we describe the Catalina Real-Time Transient Survey (CRTS), that discovers and publishes transient events at optical wavelengths in real time, thus benefiting the entire community. We describe some of the scientific results to date, and then focus on the challenges of the automated classification and prioritization of transient events. CRTS represents a scientific and a technological testbed and precursor for the larger surveys in the future, including the Large Synoptic Survey Telescope (LSST) and the Square Kilometer Array (SKA).

preprint2011arXiv

Exploring the Time Domain With Synoptic Sky Surveys

Synoptic sky surveys are becoming the largest data generators in astronomy, and they are opening a new research frontier, that touches essentially every field of astronomy. Opening of the time domain to a systematic exploration will strengthen our understanding of a number of interesting known phenomena, and may lead to the discoveries of as yet unknown ones. We describe some lessons learned over the past decade, and offer some ideas that may guide strategic considerations in planning and execution of the future synoptic sky surveys.

preprint2011arXiv

Real Time Classification of Transient Events in Synoptic Sky Surveys

An automated, rapid classification of transient events detected in the modern synoptic sky surveys is essential for their scientific utility and effective follow-up using scarce resources. This problem will grow by orders of magnitude with the next generation of surveys. We are exploring a variety of novel automated classification techniques, mostly Bayesian, to respond to these challenges, using the ongoing CRTS sky survey as a testbed. We describe briefly some of the methods used.

preprint2011arXiv

The Catalina Real-time Transient Survey

The Catalina Real-time Transient Survey (CRTS) currently covers 33,000 deg^2 of the sky in search of transient astrophysical events, with time baselines ranging from 10 minutes to ~7 years. Data provided by the Catalina Sky Survey provides an unequaled baseline against which >4,000 unique optical transient events have been discovered and openly published in real-time. Here we highlight some of the discoveries of CRTS.

preprint2011arXiv

The Catalina Real-Time Transient Survey (CRTS)

Catalina Real-Time Transient Survey (CRTS) is a synoptic sky survey uses data streams from 3 wide-field telescopes in Arizona and Australia, covering the total area of ~30,000 deg2, down to the limiting magnitudes ~ 20 - 21 mag per exposure, with time baselines from 10 min to 6 years (and growing); there are now typically ~ 200 - 300 exposures per pointing, and coadded images reach deeper than 23 mag. The basic goal of CRTS is a systematic exploration and characterization of the faint, variable sky. The survey has detected ~ 3,000 high-amplitude transients to date, including ~ 1,000 supernovae, hundreds of CVs (the majority of them previously uncatalogued), and hundreds of blazars / OVV AGN, highly variable and flare stars, etc. CRTS has a complete open data philosophy: all transients are published immediately electronically, with no proprietary period at all, and all of the data (images, light curves) will be publicly available in the near future, thus benefiting the entire astronomical community. CRTS is a scientific and technological testbed and precursor for the grander synoptic sky surveys to come.

preprint2011arXiv

The DAME/VO-Neural Infrastructure: an Integrated Data Mining System Support for the Science Community

Astronomical data are gathered through a very large number of heterogeneous techniques and stored in very diversified and often incompatible data repositories. Moreover in the e-science environment, it is needed to integrate services across distributed, heterogeneous, dynamic "virtual organizations" formed by different resources within a single enterprise and/or external resource sharing and service provider relationships. The DAME/VONeural project, run jointly by the University Federico II, INAF (National Institute of Astrophysics) Astronomical Observatories of Napoli and the California Institute of Technology, aims at creating a single, sustainable, distributed e-infrastructure for data mining and exploration in massive data sets, to be offered to the astronomical (but not only) community as a web application. The framework makes use of distributed computing environments (e.g. S.Co.P.E.) and matches the international IVOA standards and requirements. The integration process is technically challenging due to the need of achieving a specific quality of service when running on top of different native platforms. In these terms, the result of the DAME/VO-Neural project effort will be a service-oriented architecture, obtained by using appropriate standards and incorporating Grid paradigms and restful Web services frameworks where needed, that will have as main target the integration of interdisciplinary distributed systems within and across organizational domains.

preprint2011arXiv

Towards an Automated Classification of Transient Events in Synoptic Sky Surveys

We describe the development of a system for an automated, iterative, real-time classification of transient events discovered in synoptic sky surveys. The system under development incorporates a number of Machine Learning techniques, mostly using Bayesian approaches, due to the sparse nature, heterogeneity, and variable incompleteness of the available data. The classifications are improved iteratively as the new measurements are obtained. One novel feature is the development of an automated follow-up recommendation engine, that suggest those measurements that would be the most advantageous in terms of resolving classification ambiguities and/or characterization of the astrophysically most interesting objects, given a set of available follow-up assets and their cost functions. This illustrates the symbiotic relationship of astronomy and applied computer science through the emerging discipline of AstroInformatics.

preprint2006arXiv

Some Pattern Recognition Challenges in Data-Intensive Astronomy

We review some of the recent developments and challenges posed by the data analysis in modern digital sky surveys, which are representative of the information-rich astronomy in the context of Virtual Observatory. Illustrative examples include the problems of an automated star-galaxy classification in complex and heterogeneous panoramic imaging data sets, and an automated, iterative, dynamical classification of transient events detected in synoptic sky surveys. These problems offer good opportunities for productive collaborations between astronomers and applied computer scientists and statisticians, and are representative of the kind of challenges now present in all data-intensive fields. We discuss briefly some emergent types of scalable scientific data analysis systems with a broad applicability.