Source author record

István Csabai

István Csabai appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

27works
12topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

27 published item(s)

preprint2025arXiv

Constraints on AvERA Cosmologies from Cosmic Chronometers and Type Ia Supernovae

We constrain AvERA cosmologies in comparison with the flat $Λ$CDM model using cosmic chronometer (CC) data and the Pantheon+ sample of type Ia supernovae (SNe Ia). The analysis includes fits to both CC and SN datasets using the \texttt{dynesty} dynamic nested sampling algorithm. For model comparison, we use the Bayesian model evidences and Anderson-Darling tests applied to the normalized residuals to assess consistency with a standard normal distribution. Best-fit parameters are derived within the redshift ranges $z \leq 2$ for CCs and $z \leq 2.3$ for SNe. For the baseline AvERA cosmology, we obtain best-fit values of the Hubble constant of ${H_0=68.32_{-3.27}^{+3.21}~\mathrm{km~s^{-1}~Mpc^{-1}}}$ from the CC analysis and ${H_0=71.99_{-1.03}^{+1.05}~\mathrm{km~s^{-1}~Mpc^{-1}}}$ from the SN analysis, each consistent within $1σ$ with the corresponding AvERA simulation value of $H(z=0)$. While both the CC and SN datasets yield higher Bayesian evidence for the flat $Λ$CDM model, they favor the AvERA cosmologies according to the Anderson-Darling test. We have identified signs of overfitting in each model, which suggests the possibility of overestimating the uncertainties in the Pantheon+ covariance matrix.

preprint2022arXiv

An empirical nonlinear power spectrum overdensity-response

Context. The overdensity inside a cosmological sub-volume and the tidal fields from its surroundings affect the matter distribution of the region. The resulting difference between the local and global power spectra is characterized by the response function. Aims. Our aim is to provide a new, simple, and accurate formula for the power spectrum overdensity response at highly nonlinear scales based on the results of cosmological simulations and paying special attention to the lognormal nature of the density field. Methods. We measured the dark matter power spectrum amplitude as a function of the overdensity ($δ_W$) in $N$-body simulation subsamples. We show that the response follows a power-law form in terms of $(1+δ_W)$, and we provide a new fit in terms of the variance, $σ(L)$, of a sub-volume of size $L$. Results. Our fit has a similar accuracy and a comparable complexity to second-order standard perturbation theory on large scales, but it is also valid for nonlinear (smaller) scales, where perturbation theory needs higher-order terms for a comparable precision. Furthermore, we show that the lognormal nature of the overdensity distribution causes a previously unidentified bias: the power spectrum amplitude for a subsample with an average density is typically underestimated by about $-2σ^2$. Although this bias falls to the sub-percent level above characteristic scales of $200Mpch^{-1}$, taking it into account improves the accuracy of estimating power spectra from zoom-in simulations and smaller high-resolution surveys embedded in larger low-resolution volumes.

preprint2021arXiv

Decomposition of stellar populations in CosmoDC2 galaxies using SCARLET and Deep Learning

We are presenting a novel, Deep Learning based approach to estimate the normalized broadband spectral energy distribution (SED) of different stellar populations in synthetic galaxies. In contrast to the non-parametric multiband source separation algorithm, SCARLET - where the SED and morphology are simultaneously fitted - in our study we provide a morphology-independent, statistical determination of the SEDs, where we only use the color distribution of the galaxy. We developed a neural network (sedNN) that accurately predicts the SEDs of the old, red and young, blue stellar populations of realistic synthetic galaxies from the color distribution of the galaxy-related pixels in simulated broadband images. We trained and tested the network on a subset of the recently published CosmoDC2 simulated galaxy catalog containing about 3,600 galaxies. The model performance was compared to the results of SCARLET, where we found that sedNN can predict the SEDs with 4-5% accuracy on average, which is about two times better than applying SCARLET. We also investigated the effect of this improvement on the flux determination accuracy of the bulge and disk. We found that using more accurate SEDs decreases the error in the flux determination of the components by approximately 30%.

preprint2021arXiv

The effect of emission lines on the performance of photometric redshift estimation algorithms

We investigate the effect of strong emission line galaxies on the performance of empirical photometric redshift estimation methods. In order to artificially control the contribution of photometric error and emission lines to total flux, we develop a PCA-based stochastic mock catalogue generation technique that allows for generating infinite signal-to-noise ratio model spectra with realistic emission lines on top of theoretical stellar continua. Instead of running the computationally expensive stellar population synthesis and nebular emission codes, our algorithm generates realistic spectra with a statistical approach, and - as an alternative to attempting to constrain the priors on input model parameters - works by matching output observational parameters. Hence, it can be used to match the luminosity, colour, emission line and photometric error distribution of any photometric sample with sufficient flux-calibrated spectroscopic follow-up. We test three simple empirical photometric estimation methods and compare the results with and without photometric noise and strong emission lines. While photometric noise clearly dominates the uncertainty of photometric redshift estimates, the key findings are that emission lines play a significant role in resolving colour space degeneracies and good spectroscopic coverage of the entire colour space is necessary to achieve good results with empirical photo-z methods. Template fitting methods, on the other hand, must use a template set with sufficient variation in emission line strengths and ratios, or even better, first estimate the redshift empirically and fit the colours with templates at the best-fit redshift to calculate the K-correction and various physical parameters.

preprint2021arXiv

The rich still get richer: Empirical comparison of preferential attachment via linking statistics in Bitcoin and Ethereum

Bitcoin and Ethereum transactions present one of the largest real-world complex networks that are publicly available for study, including a detailed picture of their time evolution. As such, they have received a considerable amount of attention from the network science community, beside analysis from an economic or cryptography perspective. Among these studies, in an analysis on the early instance of the Bitcoin network, we have shown the clear presence of the preferential attachment, or "rich-get-richer" phenomenon. Now, we revisit this question, using a recent version of the Bitcoin network that has grown almost 100-fold since our original analysis. Furthermore, we additionally carry out a comparison with Ethereum, the second most important cryptocurrency. Our results show that preferential attachment continues to be a key factor in the evolution of both the Bitcoin and Ethereum transactoin networks. To facilitate further analysis, we publish a recent version of both transaction networks, and an efficient software implementation that is able to evaluate linking statistics necessary for learn about preferential attachment on networks with several hundred million edges.

preprint2020arXiv

A common explanation of the Hubble tension and anomalous cold spots in the CMB

The standard cosmological paradigm narrates a reassuring story of a universe currently dominated by an enigmatic dark energy component. Disquietingly, its universal explaining power has recently been challenged by, above all, the $\sim4σ$ tension in the values of the Hubble constant. Another, less studied anomaly is the repeated observation of integrated Sachs-Wolfe imprints $\sim5\times$ stronger than expected in the $Λ$CDM model from R>100 $Mpc/h$ super-structures. Here we show that the inhomogeneous AvERA model of emerging curvature is capable of telling a plausible albeit radically different story that explains both observational anomalies without dark energy. We demonstrate that while stacked imprints of R>100 $Mpc/h$ supervoids in cosmic microwave background temperature maps can discriminate between the AvERA and $Λ$CDM models, their characteristic differences may remain hidden using alternative void definitions and stacking methodologies. Testing the extremes, we then also show that the CMB Cold Spot can plausibly be explained in the AvERA model as an ISW imprint. The coldest spot in the AvERA map is aligned with multiple low-$z$ supervoids with R>100 $Mpc/h$ and central underdensity $δ_{0}\approx-0.3$, resembling the observed large-scale galaxy density field in the Cold Spot area. We hence conclude that the anomalous imprint of supervoids may well be the canary in the coal mine, and existing observational evidence for dark energy should be re-interpreted to further test alternative models.

preprint2016arXiv

Photometric redshifts for the SDSS Data Release 12

We present the methodology and data behind the photometric redshift database of the Sloan Digital Sky Survey Data Release 12 (SDSS DR12). We adopt a hybrid technique, empirically estimating the redshift via local regression on a spectroscopic training set, then fitting a spectrum template to obtain K-corrections and absolute magnitudes. The SDSS spectroscopic catalog was augmented with data from other, publicly available spectroscopic surveys to mitigate target selection effects. The training set is comprised of $1,976,978$ galaxies, and extends up to redshift $z\approx 0.8$, with a useful coverage of up to $z\approx 0.6$. We provide photometric redshifts and realistic error estimates for the $208,474,076$ galaxies of the SDSS primary photometric catalog. We achieve an average bias of $\overline{Δz_{\mathrm{norm}}} = 5.84 \times 10^{-5}$, a standard deviation of $σ\left(Δz_{\mathrm{norm}}\right)=0.0205$, and a $3σ$ outlier rate of $P_o=4.11\%$ when cross-validating on our training set. The published redshift error estimates and photometric error classes enable the selection of galaxies with high quality photometric redshifts. We also provide a supplementary error map that allows additional, sophisticated filtering of the data.

preprint2016arXiv

Quantifying correlations between galaxy emission lines and stellar continua

We analyse the correlations between continuum properties and emission line equivalent widths of star-forming and active galaxies from the Sloan Digital Sky Survey. Since upcoming large sky surveys will make broad-band observations only, including strong emission lines into theoretical modelling of spectra will be essential to estimate physical properties of photometric galaxies. We show that emission line equivalent widths can be fairly well reconstructed from the stellar continuum using local multiple linear regression in the continuum principal component analysis (PCA) space. Line reconstruction is good for star-forming galaxies and reasonable for galaxies with active nuclei. We propose a practical method to combine stellar population synthesis models with empirical modelling of emission lines. The technique will help generate more accurate model spectra and mock catalogues of galaxies to fit observations of the new surveys. More accurate modelling of emission lines is also expected to improve template-based photometric redshift estimation methods. We also show that, by combining PCA coefficients from the pure continuum and the emission lines, automatic distinction between hosts of weak active galactic nuclei (AGNs) and quiescent star-forming galaxies can be made. The classification method is based on a training set consisting of high-confidence starburst galaxies and AGNs, and allows for the similar separation of active and star-forming galaxies as the empirical curve found by Kauffmann et al. We demonstrate the use of three important machine learning algorithms in the paper: k-nearest neighbour finding, k-means clustering and support vector machines.

preprint2016arXiv

Race, Religion and the City: Twitter Word Frequency Patterns Reveal Dominant Demographic Dimensions in the United States

Recently, numerous approaches have emerged in the social sciences to exploit the opportunities made possible by the vast amounts of data generated by online social networks (OSNs). Having access to information about users on such a scale opens up a range of possibilities, all without the limitations associated with often slow and expensive paper-based polls. A question that remains to be satisfactorily addressed, however, is how demography is represented in the OSN content? Here, we study language use in the US using a corpus of text compiled from over half a billion geo-tagged messages from the online microblogging platform Twitter. Our intention is to reveal the most important spatial patterns in language use in an unsupervised manner and relate them to demographics. Our approach is based on Latent Semantic Analysis (LSA) augmented with the Robust Principal Component Analysis (RPCA) methodology. We find spatially correlated patterns that can be interpreted based on the words associated with them. The main language features can be related to slang use, urbanization, travel, religion and ethnicity, the patterns of which are shown to correlate plausibly with traditional census data. Our findings thus validate the concept of demography being represented in OSN language use and show that the traits observed are inherently present in the word frequencies without any previous assumptions about the dataset. Thus, they could form the basis of further research focusing on the evaluation of demographic data estimation from other big data sources, or on the dynamical processes that result in the patterns found here.

preprint2014arXiv

Do the rich get richer? An empirical analysis of the BitCoin transaction network

The possibility to analyze everyday monetary transactions is limited by the scarcity of available data, as this kind of information is usually considered highly sensitive. Present econophysics models are usually employed on presumed random networks of interacting agents, and only macroscopic properties (e.g. the resulting wealth distribution) are compared to real-world data. In this paper, we analyze BitCoin, which is a novel digital currency system, where the complete list of transactions is publicly available. Using this dataset, we reconstruct the network of transactions, and extract the time and amount of each payment. We analyze the structure of the transaction network by measuring network characteristics over time, such as the degree distribution, degree correlations and clustering. We find that linear preferential attachment drives the growth of the network. We also study the dynamics taking place on the transaction network, i.e. the flow of money. We measure temporal patterns and the wealth accumulation. Investigating the microscopic statistics of money movement, we find that sublinear preferential attachment governs the evolution of the wealth distribution. We report a scaling relation between the degree and wealth associated to individual nodes.

preprint2014arXiv

Efficient classification of billions of points into complex geographic regions using hierarchical triangular mesh

We present a case study about the spatial indexing and regional classification of billions of geographic coordinates from geo-tagged social network data using Hierarchical Triangular Mesh (HTM) implemented for Microsoft SQL Server. Due to the lack of certain features of the HTM library, we use it in conjunction with the GIS functions of SQL Server to significantly increase the efficiency of pre-filtering of spatial filter and join queries. For example, we implemented a new algorithm to compute the HTM tessellation of complex geographic regions and precomputed the intersections of HTM triangles and geographic regions for faster false-positive filtering. With full control over the index structure, HTM-based pre-filtering of simple containment searches outperforms SQL Server spatial indices by a factor of ten and HTM-based spatial joins run about a hundred times faster.

preprint2014arXiv

Inferring the interplay of network structure and market effects in Bitcoin

A main focus in economics research is understanding the time series of prices of goods and assets. While statistical models using only the properties of the time series itself have been successful in many aspects, we expect to gain a better understanding of the phenomena involved if we can model the underlying system of interacting agents. In this article, we consider the history of Bitcoin, a novel digital currency system, for which the complete list of transactions is available for analysis. Using this dataset, we reconstruct the transaction network between users and analyze changes in the structure of the subgraph induced by the most active users. Our approach is based on the unsupervised identification of important features of the time variation of the network. Applying the widely used method of Principal Component Analysis to the matrix constructed from snapshots of the network at different times, we are able to show how structural changes in the network accompany significant changes in the exchange price of bitcoins.

preprint2014arXiv

Strong random correlations in networks of heterogeneous agents

Correlations and other collective phenomena in a schematic model of heterogeneous binary agents (individual spin-glass samples) are considered on the complete graph and also on 2d and 3d regular lattices. The system's stochastic dynamics is studied by numerical simulations. The dynamics is so slow that one can meaningfully speak of quasi-equilibrium states. Performing measurements of correlations in such a quasi-equilibrium state we find that they are random both as to their sign and absolute value, but on average they fall off very slowly with distance in all instances that we have studied. This means that the system is essentially non-local, small changes at one end may have a strong impact at the other. Correlations and other local quantities are extremely sensitive to the boundary conditions all across the system, although this sensitivity disappears upon averaging over the samples or partially averaging over the agents. The strong, random correlations tend to organize a large fraction of the agents into strongly correlated clusters that act together. If we think about this model as a distant metaphor of economic agents or bank networks, the systemic risk implications of this tendency are clear: any impact on even a single strongly correlated agent will spread, in an unforeseeable manner, to the whole system via the strong random correlations.

preprint2013arXiv

A multi-terabyte relational database for geo-tagged social network data

Despite their relatively low sampling factor, the freely available, randomly sampled status streams of Twitter are very useful sources of geographically embedded social network data. To statistically analyze the information Twitter provides via these streams, we have collected a year's worth of data and built a multi-terabyte relational database from it. The database is designed for fast data loading and to support a wide range of studies focusing on the statistics and geographic features of social networks, as well as on the linguistic analysis of tweets. In this paper we present the method of data collection, the database design, the data loading procedure and special treatment of geo-tagged and multi-lingual data. We also provide some SQL recipes for computing network statistics.

preprint2013arXiv

Graywulf: A platform for federated scientific databases and services

Many fields of science rely on relational database management systems to analyze, publish and share data. Since RDBMS are originally designed for, and their development directions are primarily driven by, business use cases they often lack features very important for scientific applications. Horizontal scalability is probably the most important missing feature which makes it challenging to adapt traditional relational database systems to the ever growing data sizes. Due to the limited support of array data types and metadata management, successful application of RDBMS in science usually requires the development of custom extensions. While some of these extensions are specific to the field of science, the majority of them could easily be generalized and reused in other disciplines. With the Graywulf project we intend to target several goals. We are building a generic platform that offers reusable components for efficient storage, transformation, statistical analysis and presentation of scientific data stored in Microsoft SQL Server. Graywulf also addresses the distributed computational issues arising from current RDBMS technologies. The current version supports load balancing of simple queries and parallel execution of partitioned queries over a set of mirrored databases. Uniform user access to the data is provided through a web based query interface and a data surface for software clients. Queries are formulated in a slightly modified syntax of SQL that offers a transparent view of the distributed data. The software library consists of several components that can be reused to develop complex scientific data warehouses: a system registry, administration tools to manage entire database server clusters, a sophisticated workflow execution framework, and a SQL parser library.

preprint2013arXiv

Measuring the dimension of partially embedded networks

Scaling phenomena have been intensively studied during the past decade in the context of complex networks. As part of these works, recently novel methods have appeared to measure the dimension of abstract and spatially embedded networks. In this paper we propose a new dimension measurement method for networks, which does not require global knowledge on the embedding of the nodes, instead it exploits link-wise information (link lengths, link delays or other physical quantities). Our method can be regarded as a generalization of the spectral dimension, that grasps the network's large-scale structure through local observations made by a random walker while traversing the links. We apply the presented method to synthetic and real-world networks, including road maps, the Internet infrastructure and the Gowalla geosocial network. We analyze the theoretically and empirically designated case when the length distribution of the links has the form P(r) ~ 1/r. We show that while previous dimension concepts are not applicable in this case, the new dimension measure still exhibits scaling with two distinct scaling regimes. Our observations suggest that the link length distribution is not sufficient in itself to entirely control the dimensionality of complex networks, and we show that the proposed measure provides information that complements other known measures.

preprint2013arXiv

Photo-Met: a non-parametric method for estimating stellar metallicity from photometric observations

Getting spectra at good signal-to-noise ratios takes orders of magnitudes more time than photometric observations. Building on the technique developed for photometric redshift estimation of galaxies, we develop and demonstrate a non-parametric photometric method for estimating the chemical composition of galactic stars. We investigate the efficiency of our method using spectroscopically determined stellar metallicities from SDSS DR7. The technique is generic in the sense that it is not restricted to certain stellar types or stellar parameter ranges and makes it possible to obtain metallicities and error estimates for a much larger sample than spectroscopic surveys would allow. We find that our method performs well, especially for brighter stars and higher metallicities and, in contrast to many other techniques, we are able to reliably estimate the error of the predicted metallicities.

preprint2013arXiv

Refined position angle measurements for galaxies of the SDSS Stripe 82 co-added dataset

Position angle measurements of Sloan Digital Sky Survey (SDSS) galaxies, as measured by the surface brightness profile fitting code of the SDSS photometric pipeline (Lupton 2001), are known to be strongly biased, especially in the case of almost face-on and highly inclined galaxies. To address this issue we developed a reliable algorithm which determines position angles by means of isophote fitting. In this paper we present our algorithm and a catalogue of position angles for 26397 SDSS galaxies taken from the deep co-added Stripe 82 (equatorial stripe) images.

preprint2013arXiv

Using Robust PCA to estimate regional characteristics of language use from geo-tagged Twitter messages

Principal component analysis (PCA) and related techniques have been successfully employed in natural language processing. Text mining applications in the age of the online social media (OSM) face new challenges due to properties specific to these use cases (e.g. spelling issues specific to texts posted by users, the presence of spammers and bots, service announcements, etc.). In this paper, we employ a Robust PCA technique to separate typical outliers and highly localized topics from the low-dimensional structure present in language use in online social networks. Our focus is on identifying geospatial features among the messages posted by the users of the Twitter microblogging service. Using a dataset which consists of over 200 million geolocated tweets collected over the course of a year, we investigate whether the information present in word usage frequencies can be used to identify regional features of language use and topics of interest. Using the PCA pursuit method, we are able to identify important low-dimensional features, which constitute smoothly varying functions of the geographic location.

preprint2012arXiv

Plane-Sweep Incremental Algorithm: Computing Delaunay Tessellations of Large Datasets

We present the plane-sweep incremental algorithm, a hybrid approach for computing Delaunay tessellations of large point sets whose size exceeds the computer's main memory. This approach unites the simplicity of the incremental algorithms with the comparatively low memory requirements of plane-sweep approaches. The procedure is to first sort the point set along the first principal component and then to sequentially insert the points into the tessellation, essentially simulating a sweeping plane. The part of the tessellation that has been passed by the sweeping plane can be evicted from memory and written to disk, limiting the memory requirement of the program to the "thickness" of the data set along its first principal component. We implemented the algorithm and used it to compute the Delaunay tessellation and Voronoi partition of the Sloan Digital Sky Survey magnitude space consisting of 287 million points.

preprint2012arXiv

Revealing a strongly reddened, faint active galactic nucleus population by stacking deep co-added images

More than half of the sources identified by recent radio sky surveys have not been detected by wide-field optical surveys. We present a study based on our co-added image stacking technique, in which our aim is to detect the optical emission from unresolved, isolated radio sources of the Very Large Array (VLA) Faint Images of the Radio Sky at Twenty-cm (FIRST) survey that have no identified optical counterparts in the Sloan Digital Sky Survey (SDSS) Stripe 82 co-added data set. From the FIRST catalogue, 2116 such radio point sources were selected, and cut-out images, centred on the FIRST coordinates, were generated from the Stripe 82 images. The already co-added cut-outs were stacked once again to obtain images of high signal-to-noise ratio, in the hope that optical emission from the radio sources would become detectable. Multiple stacks were generated, based on the radio luminosity of the point sources. The resulting stacked images show central peaks similar to point sources. The peaks have very red colours with steep optical spectral energy distributions. We have found that the optical spectral index alpha_nu falls in the range -2.9 < alpha_nu < -2.2, depending only weakly on the radio flux. The total integration times of the stacks are between 270 and 300 h, and the corresponding 5 sigma detection limit is estimated to be about m_r = 26.6 mag. We argue that the detected light is mainly from the central regions of dust-reddened Type 1 active galactic nuclei. Dust-reddened quasars might represent an early phase of quasar evolution, and thus they can also give us an insight into the formation of massive galaxies. The data used in the paper are available on-line at http://www.vo.elte.hu/doublestacking.

preprint2012arXiv

SkyQuery: An Implementation of a Parallel Probabilistic Join Engine for Cross-Identification of Multiple Astronomical Databases

Multi-wavelength astronomical studies require cross-identification of detections of the same celestial objects in multiple catalogs based on spherical coordinates and other properties. Because of the large data volumes and spherical geometry, the symmetric N-way association of astronomical detections is a computationally intensive problem, even when sophisticated indexing schemes are used to exclude obviously false candidates. Legacy astronomical catalogs already contain detections of more than a hundred million objects while the ongoing and future surveys will produce catalogs of billions of objects with multiple detections of each at different times. The varying statistical error of position measurements, moving and extended objects, and other physical properties make it necessary to perform the cross-identification using a mathematically correct, proper Bayesian probabilistic algorithm, capable of including various priors. One time, pair-wise cross-identification of these large catalogs is not sufficient for many astronomical scenarios. Consequently, a novel system is necessary that can cross-identify multiple catalogs on-demand, efficiently and reliably. In this paper, we present our solution based on a cluster of commodity servers and ordinary relational databases. The cross-identification problems are formulated in a language based on SQL, but extended with special clauses. These special queries are partitioned spatially by coordinate ranges and compiled into a complex workflow of ordinary SQL queries. Workflows are then executed in a parallel framework using a cluster of servers hosting identical mirrors of the same data sets.

preprint2012arXiv

Spatial Indexing of Large Multidimensional Databases

Scientific endeavors such as large astronomical surveys generate databases on the terabyte scale. These, usually multidimensional databases must be visualized and mined in order to find interesting objects or to extract meaningful and qualitatively new relationships. Many statistical algorithms required for these tasks run reasonably fast when operating on small sets of in-memory data, but take noticeable performance hits when operating on large databases that do not fit into memory. We utilize new software technologies to develop and evaluate fast multidimensional indexing schemes that inherently follow the underlying, highly non-uniform distribution of the data: they are layered uniform grid indices, hierarchical binary space partitioning, and sampled flat Voronoi tessellation of the data. Our working database is the 5-dimensional magnitude space of the Sloan Digital Sky Survey with more than 270 million data points, where we show that these techniques can dramatically speed up data mining operations such as finding similar objects by example, classifying objects or comparing extensive simulation sets with observations. We are also developing tools to interact with the multidimensional database and visualize the data at multiple resolutions in an adaptive manner.

preprint2011arXiv

A High Resolution Atlas of Composite SDSS Galaxy Spectra

In this work we present an atlas of composite spectra of galaxies based on the data of the Sloan Digital Sky Survey Data Release 7 (SDSS DR7). Galaxies are classified by colour, nuclear activity and star-formation activity to calculate average spectra of high signal-to-noise ratio and resolution (S/N = 132 - 4760 at {Dlambda = 1 A), using an algorithm that is robust against outliers. Besides composite spectra, we also compute the first five principal components of the distributions in each galaxy class to characterize the nature of variations of individual spectra around the averages. The continua of the composite spectra are fitted with BC03 stellar population synthesis models to extend the wavelength coverage beyond the coverage of the SDSS spectrographs. Common derived parameters of the composites are also calculated: integrated colours in the most popular filter systems, line strength measurements, and continuum absorption indices (including Lick indices). These derived parameters are compared with the distributions of parameters of individual galaxies and it is shown on many examples that the composites of the atlas cover much of the parameter space spanned by SDSS galaxies. By co-adding thousands of spectra, a total integration time of several months can be reached, which results in extremely low noise composites. The variations in redshift not only allow for extending the spectral coverage bluewards to the original wavelength limit of the SDSS spectrographs, but also make higher spectral resolution achievable. The composite spectrum atlas is available online at http://www.vo.elte.hu/compositeatlas.

preprint2011arXiv

Array Requirements for Scientific Applications and an Implementation for Microsoft SQL Server

This paper outlines certain scenarios from the fields of astrophysics and fluid dynamics simulations which require high performance data warehouses that support array data type. A common feature of all these use cases is that subsetting and preprocessing the data on the server side (as far as possible inside the database server process) is necessary to avoid the client-server overhead and to minimize IO utilization. Analyzing and summarizing the requirements of the various fields help software engineers to come up with a comprehensive design of an array extension to relational database systems that covers a wide range of scientific applications. We also present a working implementation of an array data type for Microsoft SQL Server 2008 to support large-scale scientific applications. We introduce the design of the array type, results from a performance evaluation, and discuss the lessons learned from this implementation. The library can be downloaded from our website at http://voservices.net/sqlarray/

preprint2011arXiv

Correlations between Nebular Emission and the Continuum Spectral Shape in SDSS Galaxies

We present a statistical study of the correlations and dimensionality of emission lines carried out on a sample of over 40,000 Sloan Digital Sky Survey (SDSS) galaxies. Using principal component analysis, we found that the equivalent widths of the 11 strongest lines can be well represented using three parameters. We also explore correlations of the emission pattern with the eigenspace representation of the continuum spectrum. The observed relations are used to provide an empirical prescription for expectation values and variances of emission-line strengths as a function of spectral shape. We show that this estimation of emission lines has a sufficient accuracy to make it suitable for photometric applications. The method has already proved useful in SDSS photometric redshift estimation.

preprint2010arXiv

Cross-Identification of Stars with Unknown Proper Motions

The cross-identification of sources in separate catalogs is one of the most basic tasks in observational astronomy. It is, however, surprisingly difficult and generally ill-defined. Recently Budavári & Szalay (2008) formulated the problem in the realm of probability theory, and laid down the statistical foundations of an extensible methodology. In this paper, we apply their Bayesian approach to stars that, we know, can move measurably on the sky, with detectable proper motion, and show how to associate their observations. We study models on a sample of stars in the Sloan Digital Sky Survey, which allow for an unknown proper motion per object, and demonstrate the improvements over the analytic static model. Our models and conclusions are directly applicable to upcoming surveys such as PanSTARRS, the Dark Energy Survey, Sky Mapper, and the LSST, whose data sets will contain hundreds of millions of stars observed multiple times over several years.