Source author record

Tamás Budavári

Tamás Budavári appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

22works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

22 published item(s)

preprint2020arXiv

Computational Tools for the Spectroscopic Analysis of White Dwarfs

The spectroscopic features of white dwarfs are formed in the thin upper layer of their stellar photosphere. These features carry information about the white dwarf's surface temperature, surface gravity, and chemical composition (hereafter 'labels'). Existing methods to determine these labels rely on complex ab-initio theoretical models which are not always publicly available. Here we present two techniques to determine atmospheric labels from white dwarf spectra: a generative fitting pipeline that interpolates theoretical spectra with artificial neural networks, and a random forest regression model using parameters derived from absorption line features. We test and compare our methods using a large catalog of white dwarfs from the Sloan Digital Sky Survey (SDSS), achieving the same accuracy and negligible bias compared to previous studies. We package our techniques into an open-source Python module 'wdtools' that provides a computationally inexpensive way to determine stellar labels from white dwarf spectra observed from any facility. We will actively develop and update our tool as more theoretical models become publicly available. We discuss applications of our tool in its present form to identify interesting outlier white dwarf systems including those with magnetic fields, helium-rich atmospheres, and double-degenerate binaries.

preprint2020arXiv

Optimal Probabilistic Catalogue Matching for Radio Sources

Cross-matching catalogues from radio surveys to catalogues of sources at other wavelengths is extremely hard, because radio sources are often extended, often consist of several spatially separated components, and often no radio component is coincident with the optical/infrared host galaxy. Traditionally, the cross-matching is done by eye, but this does not scale to the millions of radio sources expected from the next generation of radio surveys. We present an innovative automated procedure, using Bayesian hypothesis testing, that models trial radio-source morphologies with putative positions of the host galaxy. This new algorithm differs from an earlier version by allowing more complex radio source morphologies, and performing a simultaneous fit over a large field. We show that this technique performs well in an unsupervised mode.

preprint2016arXiv

Exploring the SDSS Photometric Galaxies with Clustering Redshifts

We apply clustering-based redshift inference to all extended sources from the Sloan Digital Sky Survey photometric catalogue, down to magnitude r = 22. We map the relationships between colours and redshift, without assumption of the sources' spectral energy distributions (SED). We identify and locate star-forming, quiescent galaxies, and AGN, as well as colour changes due to spectral features, such as the 4000 Å break, redshifting through specific filters. Our mapping is globally in good agreement with colour-redshift tracks computed with SED templates, but reveals informative differences, such as the need for a lower fraction of M-type stars in certain templates. We compare our clustering-redshift estimates to photometric redshifts and find these two independent estimators to be in good agreement at each limiting magnitude considered. Finally, we present the global clustering-redshift distribution of all Sloan extended sources, showing objects up to z ~ 0.8. While the overall shape agrees with that inferred from photometric redshifts, the clustering redshift technique results in a smoother distribution, with no indication of structure in redshift space suggested by the photometric redshift estimates (likely artifacts imprinted by their spectroscopic training set). We also infer a higher fraction of high redshift objects. The mapping between the four observed colours and redshift can be used to estimate the redshift probability distribution function of individual galaxies. This work is an initial step towards producing a general mapping between redshift and all available observables in the photometric space, including brightness, size, concentration, and ellipticity.

preprint2016arXiv

Galaxy Redshifts from Discrete Optimization of Correlation Functions

We propose a new method of constraining the redshifts of individual extragalactic sources based on celestial coordinates and their ensemble statistics. Techniques from integer linear programming are utilized to optimize simultaneously for the angular two-point cross- and autocorrelation functions. Our novel formalism introduced here not only transforms the otherwise hopelessly expensive, brute-force combinatorial search into a linear system with integer constraints but also is readily implementable in off-the-shelf solvers. We adopt Gurobi, a commercial optimization solver, and use Python to build the cost function dynamically. The preliminary results on simulated data show potential for future applications to sky surveys by complementing and enhancing photometric redshift estimators. Our approach is the first application of integer linear programming to astronomical analysis.

preprint2016arXiv

Mapping the Similarities of Spectra: Global and Locally-biased Approaches to SDSS Galaxy Data

We apply a novel spectral graph technique, that of locally-biased semi-supervised eigenvectors, to study the diversity of galaxies. This technique permits us to characterize empirically the natural variations in observed spectra data, and we illustrate how this approach can be used in an exploratory manner to highlight both large-scale global as well as small-scale local structure in Sloan Digital Sky Survey (SDSS) data. We use this method in a way that simultaneously takes into account the measurements of spectral lines as well as the continuum shape. Unlike Principal Component Analysis, this method does not assume that the Euclidean distance between galaxy spectra is a good global measure of similarity between all spectra, but instead it only assumes that local difference information between similar spectra is reliable. Moreover, unlike other nonlinear dimensionality methods, this method can be used to characterize very finely both small-scale local as well as large-scale global properties of realistic noisy data. The power of the method is demonstrated on the SDSS Main Galaxy Sample by illustrating that the derived embeddings of spectra carry an unprecedented amount of information. By using a straightforward global or unsupervised variant, we observe that the main features correlate strongly with star formation rate and that they clearly separate active galactic nuclei. Computed parameters of the method can be used to describe line strengths and their interdependencies. By using a locally-biased or semi-supervised variant, we are able to focus on typical variations around specific objects of astronomical interest. We present several examples illustrating that this approach can enable new discoveries in the data as well as a detailed understanding of very fine local structure that would otherwise be overwhelmed by large-scale noise and global trends in the data.

preprint2016arXiv

Photometric redshifts for the SDSS Data Release 12

We present the methodology and data behind the photometric redshift database of the Sloan Digital Sky Survey Data Release 12 (SDSS DR12). We adopt a hybrid technique, empirically estimating the redshift via local regression on a spectroscopic training set, then fitting a spectrum template to obtain K-corrections and absolute magnitudes. The SDSS spectroscopic catalog was augmented with data from other, publicly available spectroscopic surveys to mitigate target selection effects. The training set is comprised of $1,976,978$ galaxies, and extends up to redshift $z\approx 0.8$, with a useful coverage of up to $z\approx 0.6$. We provide photometric redshifts and realistic error estimates for the $208,474,076$ galaxies of the SDSS primary photometric catalog. We achieve an average bias of $\overline{Δz_{\mathrm{norm}}} = 5.84 \times 10^{-5}$, a standard deviation of $σ\left(Δz_{\mathrm{norm}}\right)=0.0205$, and a $3σ$ outlier rate of $P_o=4.11\%$ when cross-validating on our training set. The published redshift error estimates and photometric error classes enable the selection of galaxies with high quality photometric redshifts. We also provide a supplementary error map that allows additional, sophisticated filtering of the data.

preprint2016arXiv

The Footprint Database and Web Services of the Herschel Space Observatory

Data from the Herschel Space Observatory is freely available to the public but no uniformly processed catalogue of the observations has been published so far. To date, the Herschel Science Archive does not contain the exact sky coverage (footprint) of individual observations and supports search for measurements based on bounding circles only. Drawing on previous experience in implementing footprint databases, we built the Herschel Footprint Database and Web Services for the Herschel Space Observatory to provide efficient search capabilities for typical astronomical queries. The database was designed with the following main goals in mind: (a) provide a unified data model for meta-data of all instruments and observational modes, (b) quickly find observations covering a selected object and its neighbourhood, (c) quickly find every observation in a larger area of the sky, (d) allow for finding solar system objects crossing observation fields. As a first step, we developed a unified data model of observations of all three Herschel instruments for all pointing and instrument modes. Then, using telescope pointing information and observational meta-data, we compiled a database of footprints. As opposed to methods using pixellation of the sphere, we represent sky coverage in an exact geometric form allowing for precise area calculations. For easier handling of Herschel observation footprints with rather complex shapes, two algorithms were implemented to reduce the outline. Furthermore, a new visualisation tool to plot footprints with various spherical projections was developed. Indexing of the footprints using Hierarchical Triangular Mesh makes it possible to quickly find observations based on sky coverage, time and meta-data. The database is accessible via a web site (http://herschel.vo.elte.hu) and also as a set of REST web service functions.

preprint2015arXiv

Matching Radio Catalogs with Realistic Geometry: Application to SWIRE and ATLAS

Crossmatching catalogs at different wavelengths is a difficult problem in astronomy, especially when the objects are not point-like. At radio wavelengths an object can have several components corresponding, for example, to a core and lobes. {Considering not all radio detections correspond to visible or infrared sources, matching these catalogs can be challenging.} Traditionally this is done by eye for better quality, which does not scale to the large data volumes expected from the next-generation of radio telescopes. We present a novel automated procedure, using Bayesian hypothesis testing, to achieve reliable associations by explicit modelling of a particular class of radio-source morphology. {The new algorithm not only assesses the likelihood of an association between data at two different wavelengths, but also tries to assess whether different radio sources are physically associated, are double-lobed radio galaxies, or just distinct nearby objects.} Application to the SWIRE and ATLAS CDF-S catalogs shows that this method performs well without human intervention.

preprint2014arXiv

CANDELS/GOODS-S, CDFS, ECDFS: Photometric Redshifts For Normal and for X-Ray-Detected Galaxies

We present photometric redshifts and associated probability distributions for all detected sources in the Extended Chandra Deep Field South (ECDFS). The work makes use of the most up-to-date data from the Cosmic Assembly Near-IR Deep Legacy Survey (CANDELS) and the Taiwan ECDFS Near-Infrared Survey (TENIS) in addition to other data. We also revisit multi-wavelength counterparts for published X-ray sources from the 4Ms-CDFS and 250ks-ECDFS surveys, finding reliable counterparts for 1207 out of 1259 sources ($\sim 96\%$). Data used for photometric redshifts include intermediate-band photometry deblended using the TFIT method, which is used for the first time in this work. Photometric redshifts for X-ray source counterparts are based on a new library of AGN/galaxy hybrid templates appropriate for the faint X-ray population in the CDFS. Photometric redshift accuracy for normal galaxies is 0.010 and for X-ray sources is 0.014, and outlier fractions are $4\%$ and $5.4\%$ respectively. The results within the CANDELS coverage area are even better as demonstrated both by spectroscopic comparison and by galaxy-pair statistics. Intermediate-band photometry, even if shallow, is valuable when combined with deep broad-band photometry. For best accuracy, templates must include emission lines.

preprint2014arXiv

Efficient Catalog Matching with Dropout Detection

Not only source catalogs are extracted from astronomy observations. Their sky coverage is always carefully recorded and used in statistical analyses, such as correlation and luminosity function studies. Here we present a novel method for catalog matching, which inherently builds on the coverage information for better performance and completeness. A modified version of the Zones Algorithm is introduced for matching partially overlapping observations, where irrelevant parts of the data are excluded up front for efficiency. Our design enables searches to focus on specific areas on the sky to further speed up the process. Another important advantage of the new method over traditional techniques is its ability to quickly detect dropouts, i.e., the missing components that are in the observed regions of the celestial sphere but did not reach the detection limit in some observations. These often provide invaluable insight into the spectral energy distribution of the matched sources but rarely available in traditional associations.

preprint2014arXiv

Efficient classification of billions of points into complex geographic regions using hierarchical triangular mesh

We present a case study about the spatial indexing and regional classification of billions of geographic coordinates from geo-tagged social network data using Hierarchical Triangular Mesh (HTM) implemented for Microsoft SQL Server. Due to the lack of certain features of the HTM library, we use it in conjunction with the GIS functions of SQL Server to significantly increase the efficiency of pre-filtering of spatial filter and join queries. For example, we implemented a new algorithm to compute the HTM tessellation of complex geographic regions and precomputed the intersections of HTM triangles and geographic regions for faster false-positive filtering. With full control over the index structure, HTM-based pre-filtering of simple containment searches outperforms SQL Server spatial indices by a factor of ten and HTM-based spatial joins run about a hundred times faster.

preprint2013arXiv

Graywulf: A platform for federated scientific databases and services

Many fields of science rely on relational database management systems to analyze, publish and share data. Since RDBMS are originally designed for, and their development directions are primarily driven by, business use cases they often lack features very important for scientific applications. Horizontal scalability is probably the most important missing feature which makes it challenging to adapt traditional relational database systems to the ever growing data sizes. Due to the limited support of array data types and metadata management, successful application of RDBMS in science usually requires the development of custom extensions. While some of these extensions are specific to the field of science, the majority of them could easily be generalized and reused in other disciplines. With the Graywulf project we intend to target several goals. We are building a generic platform that offers reusable components for efficient storage, transformation, statistical analysis and presentation of scientific data stored in Microsoft SQL Server. Graywulf also addresses the distributed computational issues arising from current RDBMS technologies. The current version supports load balancing of simple queries and parallel execution of partitioned queries over a set of mirrored databases. Uniform user access to the data is provided through a web based query interface and a data surface for software clients. Queries are formulated in a slightly modified syntax of SQL that offers a transparent view of the distributed data. The software library consists of several components that can be reused to develop complex scientific data warehouses: a system registry, administration tools to manage entire database server clusters, a sophisticated workflow execution framework, and a SQL parser library.

preprint2013arXiv

More than just halo mass: Modelling how the red galaxy fraction depends on multiscale density in a HOD framework

The fraction of galaxies with red colours depends sensitively on environment, and on the way in which environment is measured. To distinguish competing theories for the quenching of star formation, a robust and complete description of environment is required, to be applied to a large sample of galaxies. The environment of galaxies can be described using the density field of neighbours on multiple scales - the multiscale density field. We are using the Millennium simulation and a simple HOD prescription which describes the multiscale density field of Sloan Digital Sky Survey DR7 galaxies to investigate the dependence of the fraction of red galaxies on the environment. Using a volume limited sample where we have sufficient galaxies in narrow density bins, we have more dynamic range in halo mass and density for satellite galaxies than for central galaxies. Therefore we model the red fraction of central galaxies as a constant while we use a functional form to describe the red fraction of satellites as a function of halo mass which allows us to distinguish a sharp from a gradual transition. While it is clear that the data can only be explained by a gradual transition, an analysis of the multiscale density field on different scales suggests that colour segregation within the haloes is needed to explain the results. We also rule out a sharp transition for central galaxies, within the halo mass range sampled.

preprint2012arXiv

Astronomy and Computing: a New Journal for the Astronomical Computing Community

We introduce \emph{Astronomy and Computing}, a new journal for the growing population of people working in the domain where astronomy overlaps with computer science and information technology. The journal aims to provide a new communication channel within that community, which is not well served by current journals, and to help secure recognition of its true importance within modern astronomy. In this inaugural editorial, we describe the rationale for creating the journal, outline its scope and ambitions, and seek input from the community in defining in detail how the journal should work towards its high-level goals.

preprint2012arXiv

SkyQuery: An Implementation of a Parallel Probabilistic Join Engine for Cross-Identification of Multiple Astronomical Databases

Multi-wavelength astronomical studies require cross-identification of detections of the same celestial objects in multiple catalogs based on spherical coordinates and other properties. Because of the large data volumes and spherical geometry, the symmetric N-way association of astronomical detections is a computationally intensive problem, even when sophisticated indexing schemes are used to exclude obviously false candidates. Legacy astronomical catalogs already contain detections of more than a hundred million objects while the ongoing and future surveys will produce catalogs of billions of objects with multiple detections of each at different times. The varying statistical error of position measurements, moving and extended objects, and other physical properties make it necessary to perform the cross-identification using a mathematically correct, proper Bayesian probabilistic algorithm, capable of including various priors. One time, pair-wise cross-identification of these large catalogs is not sufficient for many astronomical scenarios. Consequently, a novel system is necessary that can cross-identify multiple catalogs on-demand, efficiently and reliably. In this paper, we present our solution based on a cluster of commodity servers and ordinary relational databases. The cross-identification problems are formulated in a language based on SQL, but extended with special clauses. These special queries are partitioned spatially by coordinate ranges and compiled into a complex workflow of ordinary SQL queries. Workflows are then executed in a parallel framework using a cluster of servers hosting identical mirrors of the same data sets.

preprint2012arXiv

Spatial Indexing of Large Multidimensional Databases

Scientific endeavors such as large astronomical surveys generate databases on the terabyte scale. These, usually multidimensional databases must be visualized and mined in order to find interesting objects or to extract meaningful and qualitatively new relationships. Many statistical algorithms required for these tasks run reasonably fast when operating on small sets of in-memory data, but take noticeable performance hits when operating on large databases that do not fit into memory. We utilize new software technologies to develop and evaluate fast multidimensional indexing schemes that inherently follow the underlying, highly non-uniform distribution of the data: they are layered uniform grid indices, hierarchical binary space partitioning, and sampled flat Voronoi tessellation of the data. Our working database is the 5-dimensional magnitude space of the Sloan Digital Sky Survey with more than 270 million data points, where we show that these techniques can dramatically speed up data mining operations such as finding similar objects by example, classifying objects or comparing extensive simulation sets with observations. We are also developing tools to interact with the multidimensional database and visualize the data at multiple resolutions in an adaptive manner.

preprint2011arXiv

A High Resolution Atlas of Composite SDSS Galaxy Spectra

In this work we present an atlas of composite spectra of galaxies based on the data of the Sloan Digital Sky Survey Data Release 7 (SDSS DR7). Galaxies are classified by colour, nuclear activity and star-formation activity to calculate average spectra of high signal-to-noise ratio and resolution (S/N = 132 - 4760 at {Dlambda = 1 A), using an algorithm that is robust against outliers. Besides composite spectra, we also compute the first five principal components of the distributions in each galaxy class to characterize the nature of variations of individual spectra around the averages. The continua of the composite spectra are fitted with BC03 stellar population synthesis models to extend the wavelength coverage beyond the coverage of the SDSS spectrographs. Common derived parameters of the composites are also calculated: integrated colours in the most popular filter systems, line strength measurements, and continuum absorption indices (including Lick indices). These derived parameters are compared with the distributions of parameters of individual galaxies and it is shown on many examples that the composites of the atlas cover much of the parameter space spanned by SDSS galaxies. By co-adding thousands of spectra, a total integration time of several months can be reached, which results in extremely low noise composites. The variations in redshift not only allow for extending the spectral coverage bluewards to the original wavelength limit of the SDSS spectrographs, but also make higher spectral resolution achievable. The composite spectrum atlas is available online at http://www.vo.elte.hu/compositeatlas.

preprint2011arXiv

Array Requirements for Scientific Applications and an Implementation for Microsoft SQL Server

This paper outlines certain scenarios from the fields of astrophysics and fluid dynamics simulations which require high performance data warehouses that support array data type. A common feature of all these use cases is that subsetting and preprocessing the data on the server side (as far as possible inside the database server process) is necessary to avoid the client-server overhead and to minimize IO utilization. Analyzing and summarizing the requirements of the various fields help software engineers to come up with a comprehensive design of an array extension to relational database systems that covers a wide range of scientific applications. We also present a working implementation of an array data type for Microsoft SQL Server 2008 to support large-scale scientific applications. We introduce the design of the array type, results from a performance evaluation, and discuss the lessons learned from this implementation. The library can be downloaded from our website at http://voservices.net/sqlarray/

preprint2011arXiv

Correlations between Nebular Emission and the Continuum Spectral Shape in SDSS Galaxies

We present a statistical study of the correlations and dimensionality of emission lines carried out on a sample of over 40,000 Sloan Digital Sky Survey (SDSS) galaxies. Using principal component analysis, we found that the equivalent widths of the 11 strongest lines can be well represented using three parameters. We also explore correlations of the emission pattern with the eigenspace representation of the continuum spectrum. The observed relations are used to provide an empirical prescription for expectation values and variances of emission-line strengths as a function of spectral shape. We show that this estimation of emission lines has a sufficient accuracy to make it suitable for photometric applications. The method has already proved useful in SDSS photometric redshift estimation.

preprint2010arXiv

Cross-Identification of Stars with Unknown Proper Motions

The cross-identification of sources in separate catalogs is one of the most basic tasks in observational astronomy. It is, however, surprisingly difficult and generally ill-defined. Recently Budavári & Szalay (2008) formulated the problem in the realm of probability theory, and laid down the statistical foundations of an extensible methodology. In this paper, we apply their Bayesian approach to stars that, we know, can move measurably on the sky, with detectable proper motion, and show how to associate their observations. We study models on a sample of stars in the Sloan Digital Sky Survey, which allow for an unknown proper motion per object, and demonstrate the improvements over the analytic static model. Our models and conclusions are directly applicable to upcoming surveys such as PanSTARRS, the Dark Energy Survey, Sky Mapper, and the LSST, whose data sets will contain hundreds of millions of stars observed multiple times over several years.

preprint2010arXiv

Redshift-Space Enhancement of Line-of-Sight Baryon Acoustic Oscillations in the SDSS Main-Galaxy Sample

We show that redshift-space distortions of galaxy correlations have a strong effect on correlation functions with distinct, localized features, like the signature of the baryon acoustic oscillations (BAO). Near the line of sight, the features become sharper as a result of redshift-space distortions. We demonstrate this effect by measuring the correlation function in Gaussian simulations and the Millennium Simulation. We also analyze the SDSS DR7 main-galaxy sample (MGS), splitting the sample into slices 2.5 degrees on the sky in various rotations. Measuring 2D correlation functions in each slice, we do see a sharp bump along the line of sight. Using Mexican-hat wavelets, we localize it to (110 +/- 10) Mpc/h. Averaging only along the line of sight, we estimate its significance at a particular wavelet scale and location at 2.2 sigma. In a flat angular weighting in the (pi,r_p) coordinate system, the noise level is suppressed, pushing the bump's significance to 4 sigma. We estimate that there is about a 0.2% chance of getting such a signal anywhere in the vicinity of the BAO scale from a power spectrum lacking a BAO feature. However, these estimates of the significances make some use of idealized Gaussian simulations, and thus are likely a bit optimistic.

preprint2009arXiv

The GALEX Arecibo SDSS Survey. I. Gas Fraction Scaling Relations of Massive Galaxies and First Data Release

We introduce the GALEX Arecibo SDSS Survey (GASS), an on-going large program that is gathering high quality HI-line spectra using the Arecibo radio telescope for an unbiased sample of ~1000 galaxies with stellar masses greater than 10^10 Msun and redshifts 0.025<z<0.05, selected from the SDSS spectroscopic and GALEX imaging surveys. The galaxies are observed until detected or until a low gas mass fraction limit (1.5-5%) is reached. This paper presents the first Data Release, consisting of ~20% of the final GASS sample. We use this data set to explore the main scaling relations of HI gas fraction with galaxy structure and NUV-r colour. A large fraction (~60%) of the galaxies in our sample are detected in HI. We find that the atomic gas fraction decreases strongly with stellar mass, stellar surface mass density and NUV-r colour, but is only weakly correlated with galaxy bulge-to-disk ratio (as measured by the concentration index of the r-band light). We also find that the fraction of galaxies with significant (more than a few percent) HI decreases sharply above a characteristic stellar surface mass density of 10^8.5 Msun kpc^-2. The fraction of gas-rich galaxies decreases much more smoothly with stellar mass. One of the key goals of GASS is to identify and quantify the incidence of galaxies that are transitioning between the blue, star-forming cloud and the red sequence of passively-evolving galaxies. Likely transition candidates can be identified as outliers from the mean scaling relations between gas fraction and other galaxy properties. [abridged]