Source author record

Joseph W. Richards

Joseph W. Richards appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

20works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

20 published item(s)

preprint2014arXiv

Mid-Infrared Period--Luminosity Relations of RR Lyrae Stars Derived from the AllWISE Data Release

We use photometry from the recent AllWISE Data Release of the Wide-field In-frared Survey Explorer (WISE) of 129 calibration stars, combined with prior distances obtained from the established $M_V-$[Fe/H] relation and Hubble Space Telescope trigonometric parallax, to derive mid-infrared period--luminosity relations for RR Lyrae pulsating variable stars. We derive relations in the W1, W2, and W3 wavebands (3.4, 4.6, and 12 μm, respectively), and for each of the two main RR Lyrae sub-types (RRab and RRc). We report an error on the period--luminosity relation slope for RRab stars of 0.2. We also fit posterior distances for the calibration catalog and find a median fractional distance error of 0.8 per cent.

preprint2014arXiv

The Highly-Eccentric Detached Eclipsing Binaries in ACVS and MACC

Next-generation synoptic photometric surveys will yield unprecedented (for the astronomical community) volumes of data and the processes of discovery and rare-object identification are, by necessity, becoming more autonomous. Such autonomous searches can be used to find objects of interest applicable to a wide range of outstanding problems in astronomy, and in this paper we present the methods and results of a largely autonomous search for highly eccentric detached eclipsing binary systems in the Machine-learned ASAS Classification Catalog. 106 detached eclipsing binaries with eccentricities greater than 0.1 are presented, most of which are identified here for the first time. We also present new radial-velocity curves and absolute parameters for 6 of those systems with the long-term goal of increasing the number of highly eccentric systems with orbital solutions, thereby facilitating further studies of the tidal circularization process in binary stars.

preprint2013arXiv

A Mid-infrared Study of RR Lyrae Stars with the WISE All-Sky Data Release

We present a group of 3740 previously identified RR Lyrae variables well-observed with the Wide-field Infrared Survey Explorer (WISE). We explore how the shape of the generic RR Lyrae mid-infrared light curve evolves in period-space, comparing light curves in mid-infrared and optical bands. We find that optical light curves exhibit high amplitudes and a large spectrum of light curve shapes, while mid-infrared light curves have low amplitudes and uniform light curve shapes. From the period-space analysis, we hope to improve the classification methods of RR Lyrae variables and enable reliable discovery of these pulsators in the WISE catalog and future mid-infrared surveys such as the James Webb Space Telescope (JWST). We provide mid-infrared templates for typical RR Lyrae stars and demonstrate how these templates can be applied to improve estimates of mid-infrared RR Lyrae mean magnitude, which is used for distance measurement. This method of template fitting is particularly beneficial for improving observational efficiency. For example, using light curves with observational noise of 0.05 mag, we obtain the same level of accuracy in mean magnitude estimates for light curves randomly sampled at 12 data points with template fitting as with light curves randomly sampled at 20 data points with harmonic modelling.

preprint2013arXiv

Discovery and Early Multi-Wavelength Measurements of the Energetic Type Ic Supernova PTF12gzk: A Massive-Star Explosion in a Dwarf Host Galaxy

We present the discovery and extensive early-time observations of the Type Ic supernova (SN) PTF12gzk. Our finely sampled light curves show a rise of 0.8mag within 2.5hr. Power-law fits [f(t)\sim(t-t_0)^n] to these data constrain the explosion date to within one day. We cannot rule out the expected quadratic fireball model, but higher values of n are possible as well for larger areas in the fit parameter space. Our bolometric light curve and a dense spectral sequence are used to estimate the physical parameters of the exploding star and of the explosion. We show that the photometric evolution of PTF12gzk is slower than that of most SNe Ic, and its high ejecta velocities (~30,000km/s four days after explosion) are closer to the observed velocities of broad-lined SNe Ic associated with gamma-ray bursts (GRBs) than to the observed velocities in normal Type Ic SNe. The high velocities are sustained through the SN early evolution, and are similar to those of GRB-SNe when the SN reach peak magnitude. By comparison with the spectroscopically similar SN 2004aw, we suggest that the observed properties of PTF12gzk indicate an initial progenitor mass of 25-35 solar mass and a large (5-10E51 erg) kinetic energy, close to the regime of GRB-SN properties. The host-galaxy characteristics are consistent with GRB-SN hosts, and not with normal SN Ic hosts as well, yet this SN does not show the broad lines over extended periods of time that are typical of broad-line Type Ic SNe.

preprint2012arXiv

A Bayesian Approach to Calibrating Period-Luminosity Relations of RR Lyrae Stars in the Mid-Infrared

A Bayesian approach to calibrating period-luminosity (PL) relations has substantial benefits over generic least-squares fits. In particular, the Bayesian approach takes into account the full prior distribution of the model parameters, such as the a priori distances, and refits these parameters as part of the process of settling on the most highly-constrained final fit. Additionally, the Bayesian approach can naturally ingest data from multiple wavebands and simultaneously fit the parameters of PL relations for each waveband in a procedure that constrains the parameter posterior distributions so as to minimize the scatter of the final fits appropriately in all wavebands. Here we describe the generalized approach to Bayesian model fitting and then specialize to a detailed description of applying Bayesian linear model fitting to the mid-infrared PL relations of RR Lyrae variable stars. For this example application we quantify the improvement afforded by using a Bayesian model fit. We also compare distances previously predicted in our example application to recently published parallax distances measured with the Hubble Space Telescope and find their agreement to be a vindication of our methodology. Our intent with this article is to spread awareness of the benefits and applicability of this Bayesian approach and encourage future PL relation investigations to consider employing this powerful analysis method.

preprint2012arXiv

Construction of a Calibrated Probabilistic Classification Catalog: Application to 50k Variable Sources in the All-Sky Automated Survey

With growing data volumes from synoptic surveys, astronomers must become more abstracted from the discovery and introspection processes. Given the scarcity of follow-up resources, there is a particularly sharp onus on the frameworks that replace these human roles to provide accurate and well-calibrated probabilistic classification catalogs. Such catalogs inform the subsequent follow-up, allowing consumers to optimize the selection of specific sources for further study and permitting rigorous treatment of purities and efficiencies for population studies. Here, we describe a process to produce a probabilistic classification catalog of variability with machine learning from a multi-epoch photometric survey. In addition to producing accurate classifications, we show how to estimate calibrated class probabilities, and motivate the importance of probability calibration. We also introduce a methodology for feature-based anomaly detection, which allows discovery of objects in the survey that do not fit within the predefined class taxonomy. Finally, we apply these methods to sources observed by the All Sky Automated Survey (ASAS), and unveil the Machine-learned ASAS Classification Catalog (MACC), which is a 28-class probabilistic classification catalog of 50,124 ASAS sources. We estimate that MACC achieves a sub-20% classification error rate, and demonstrate that the class posterior probabilities are reasonably calibrated. MACC classifications compare favorably to the classifications of several previous domain-specific ASAS papers and to the ASAS Catalog of Variable Stars, which had classified only 24% of those sources into one of 12 science classes. The MACC is publicly available at http://www.bigmacc.info.

preprint2012arXiv

Optimizing Automated Classification of Periodic Variable Stars in New Synoptic Surveys

Efficient and automated classification of periodic variable stars is becoming increasingly important as the scale of astronomical surveys grows. Several recent papers have used methods from machine learning and statistics to construct classifiers on databases of labeled, multi--epoch sources with the intention of using these classifiers to automatically infer the classes of unlabeled sources from new surveys. However, the same source observed with two different synoptic surveys will generally yield different derived metrics (features) from the light curve. Since such features are used in classifiers, this survey-dependent mismatch in feature space will typically lead to degraded classifier performance. In this paper we show how and why feature distributions change using OGLE and \textit{Hipparcos} light curves. To overcome survey systematics, we apply a method, \textit{noisification}, which attempts to empirically match distributions of features between the labeled sources used to construct the classifier and the unlabeled sources we wish to classify. Results from simulated and real--world light curves show that noisification can significantly improve classifier performance. In a three--class problem using light curves from \textit{Hipparcos} and OGLE, noisification reduces the classifier error rate from 27.0% to 7.0%. We recommend that noisification be used for upcoming surveys such as Gaia and LSST and describe some of the promises and challenges of applying noisification to these surveys.

preprint2012arXiv

Prototype selection for parameter estimation in complex models

Parameter estimation in astrophysics often requires the use of complex physical models. In this paper we study the problem of estimating the parameters that describe star formation history (SFH) in galaxies. Here, high-dimensional spectral data from galaxies are appropriately modeled as linear combinations of physical components, called simple stellar populations (SSPs), plus some nonlinear distortions. Theoretical data for each SSP is produced for a fixed parameter vector via computer modeling. Though the parameters that define each SSP are continuous, optimizing the signal model over a large set of SSPs on a fine parameter grid is computationally infeasible and inefficient. The goal of this study is to estimate the set of parameters that describes the SFH of each galaxy. These target parameters, such as the average ages and chemical compositions of the galaxy's stellar populations, are derived from the SSP parameters and the component weights in the signal model. Here, we introduce a principled approach of choosing a small basis of SSP prototypes for SFH parameter estimation. The basic idea is to quantize the vector space and effective support of the model components. In addition to greater computational efficiency, we achieve better estimates of the SFH target parameters. In simulations, our proposed quantization method obtains a substantial improvement in estimating the target parameters over the common method of employing a parameter grid. Sparse coding techniques are not appropriate for this problem without proper constraints, while constrained sparse coding methods perform poorly for parameter estimation because their objective is signal reconstruction, not estimation of the target parameters.

preprint2012arXiv

The XMM Cluster Survey: The Stellar Mass Assembly of Fossil Galaxies

This paper presents both the result of a search for fossil systems (FSs) within the XMM Cluster Survey and the Sloan Digital Sky Survey and the results of a study of the stellar mass assembly and stellar populations of their fossil galaxies. In total, 17 groups and clusters are identified at z < 0.25 with large magnitude gaps between the first and fourth brightest galaxies. All the information necessary to classify these systems as fossils is provided. For both groups and clusters, the total and fractional luminosity of the brightest galaxy is positively correlated with the magnitude gap. The brightest galaxies in FSs (called fossil galaxies) have stellar populations and star formation histories which are similar to normal brightest cluster galaxies (BCGs). However, at fixed group/cluster mass, the stellar masses of the fossil galaxies are larger compared to normal BCGs, a fact that holds true over a wide range of group/cluster masses. Moreover, the fossil galaxies are found to contain a significant fraction of the total optical luminosity of the group/cluster within 0.5R200, as much as 85%, compared to the non-fossils, which can have as little as 10%. Our results suggest that FSs formed early and in the highest density regions of the universe and that fossil galaxies represent the end products of galaxy mergers in groups and clusters. The online FS catalog can be found at http://www.astro.ljmu.ac.uk/~xcs/Harrison2012/XCSFSCat.html.

preprint2012arXiv

Using Machine Learning for Discovery in Synoptic Survey Imaging

Modern time-domain surveys continuously monitor large swaths of the sky to look for astronomical variability. Astrophysical discovery in such data sets is complicated by the fact that detections of real transient and variable sources are highly outnumbered by bogus detections caused by imperfect subtractions, atmospheric effects and detector artefacts. In this work we present a machine learning (ML) framework for discovery of variability in time-domain imaging surveys. Our ML methods provide probabilistic statements, in near real time, about the degree to which each newly observed source is astrophysically relevant source of variable brightness. We provide details about each of the analysis steps involved, including compilation of the training and testing sets, construction of descriptive image-based and contextual features, and optimization of the feature subset and model tuning parameters. Using a validation set of nearly 30,000 objects from the Palomar Transient Factory, we demonstrate a missed detection rate of at most 7.7% at our chosen false-positive rate of 1% for an optimized ML classifier of 23 features, selected to avoid feature correlation and over-fitting from an initial library of 42 attributes. Importantly, we show that our classification methodology is insensitive to mis-labelled training data up to a contamination of nearly 10%, making it easier to compile sufficient training sets for accurate performance in future surveys. This ML framework, if so adopted, should enable the maximization of scientific gain from future synoptic survey and enable fast follow-up decisions on the vast amounts of streaming data produced by such experiments.

preprint2011arXiv

Active Learning to Overcome Sample Selection Bias: Application to Photometric Variable Star Classification

Despite the great promise of machine-learning algorithms to classify and predict astrophysical parameters for the vast numbers of astrophysical sources and transients observed in large-scale surveys, the peculiarities of the training data often manifest as strongly biased predictions on the data of interest. Typically, training sets are derived from historical surveys of brighter, more nearby objects than those from more extensive, deeper surveys (testing data). This sample selection bias can cause catastrophic errors in predictions on the testing data because a) standard assumptions for machine-learned model selection procedures break down and b) dense regions of testing space might be completely devoid of training data. We explore possible remedies to sample selection bias, including importance weighting (IW), co-training (CT), and active learning (AL). We argue that AL---where the data whose inclusion in the training set would most improve predictions on the testing set are queried for manual follow-up---is an effective approach and is appropriate for many astronomical applications. For a variable star classification problem on a well-studied set of stars from Hipparcos and OGLE, AL is the optimal method in terms of error rate on the testing data, beating the off-the-shelf classifier by 3.4% and the other proposed methods by at least 3.0%. To aid with manual labeling of variable stars, we developed a web interface which allows for easy light curve visualization and querying of external databases. Finally, we apply active learning to classify variable stars in the ASAS survey, finding dramatic improvement in our agreement with the ACVS catalog, from 65.5% to 79.5%, and a significant increase in the classifier's average confidence for the testing set, from 14.6% to 42.9%, after a few AL iterations.

preprint2011arXiv

Constraints on the Progenitor System of the Type Ia Supernova SN 2011fe/PTF11kly

Type Ia supernovae (SNe) serve as a fundamental pillar of modern cosmology, owing to their large luminosity and a well-defined relationship between light-curve shape and peak brightness. The precision distance measurements enabled by SNe Ia first revealed the accelerating expansion of the universe, now widely believed (though hardly understood) to require the presence of a mysterious "dark" energy. General consensus holds that Type Ia SNe result from thermonuclear explosions of a white dwarf (WD) in a binary system; however, little is known of the precise nature of the companion star and the physical properties of the progenitor system. Here we make use of extensive historical imaging obtained at the location of SN 2011fe/PTF11kly, the closest SN Ia discovered in the digital imaging era, to constrain the visible-light luminosity of the progenitor to be 10-100 times fainter than previous limits on other SN Ia progenitors. This directly rules out luminous red giants and the vast majority of helium stars as the mass-donating companion to the exploding white dwarf. Any evolved red companion must have been born with mass less than 3.5 times the mass of the Sun. These observations favour a scenario where the exploding WD of SN 2011fe/PTF11kly, accreted matter either from another WD, or by Roche-lobe overflow from a subgiant or main-sequence companion star.

preprint2011arXiv

Data Mining and Machine-Learning in Time-Domain Discovery & Classification

The changing heavens have played a central role in the scientific effort of astronomers for centuries. Galileo's synoptic observations of the moons of Jupiter and the phases of Venus starting in 1610, provided strong refutation of Ptolemaic cosmology. In more modern times, the discovery of a relationship between period and luminosity in some pulsational variable stars led to the inference of the size of the Milky Way, the distance scale to the nearest galaxies, and the expansion of the Universe. Distant explosions of supernovae were used to uncover the existence of dark energy and provide a precise numerical account of dark matter. Indeed, time-domain observations of transient events and variable stars, as a technique, influences a broad diversity of pursuits in the entire astronomy endeavor. While, at a fundamental level, the nature of the scientific pursuit remains unchanged, the advent of astronomy as a data-driven discipline presents fundamental challenges to the way in which the scientific process must now be conducted. Digital images (and data cubes) are not only getting larger, there are more of them. On logistical grounds, this taxes storage and transport systems. But it also implies that the intimate connection that astronomers have always enjoyed with their data---from collection to processing to analysis to inference---necessarily must evolve. The pathway to scientific inference is now influenced (if not driven by) modern automation processes, computing, data-mining and machine learning. The emerging reliance on computation and machine learning is a general one, but the time-domain aspect of the data and the objects of interest presents some unique challenges, which we describe and explore in this chapter.

preprint2011arXiv

Mid-infrared Period-Luminosity Relations of RR Lyrae Stars Derived from the WISE Preliminary Data Release

Interstellar dust presents a significant challenge to extending parallax-determined distances of optically observed pulsational variables to larger volumes. Distance ladder work at mid-infrared wavebands, where dust effects are negligible and metallicity correlations are minimized, have been largely focused on few-epoch Cepheid studies. Here we present the first determination of mid-infrared period-luminosity (PL) relations of RR Lyrae stars from phase-resolved imaging using the preliminary data release of the Wide-Field Infrared Survey Explorer (WISE). We present a novel statistical framework to predict posterior distances of 76 well-observed RR Lyrae that uses the optically constructed prior distance moduli while simultaneously imposing a power-law PL relation to WISE-determined mean magnitudes. We find that the absolute magnitude in the bluest WISE filter is M_W1 = (-0.421+-0.014) - (1.681+-0.147)*log(P/0.50118 day), with no evidence for a correlation with metallicity. Combining the results from the three bluest WISE filters, we find that a typical star in our sample has a distance measurement uncertainty of 0.97% (statistical) plus 1.17% (systematic). We do not fundamentalize the periods of RRc stars to improve their fit to the relations. Taking the Hipparcos-derived mean V-band magnitudes, we use the distance posteriors to determine a new optical metallicity-luminosity relation which we present in Section 5. The results of this analysis will soon be tested by HST parallax measurements and, eventually, with the Gaia astrometric mission.

preprint2011arXiv

On Machine-Learned Classification of Variable Stars with Sparse and Noisy Time-Series Data

With the coming data deluge from synoptic surveys, there is a growing need for frameworks that can quickly and automatically produce calibrated classification probabilities for newly-observed variables based on a small number of time-series measurements. In this paper, we introduce a methodology for variable-star classification, drawing from modern machine-learning techniques. We describe how to homogenize the information gleaned from light curves by selection and computation of real-numbered metrics ("feature"), detail methods to robustly estimate periodic light-curve features, introduce tree-ensemble methods for accurate variable star classification, and show how to rigorously evaluate the classification results using cross validation. On a 25-class data set of 1542 well-studied variable stars, we achieve a 22.8% overall classification error using the random forest classifier; this represents a 24% improvement over the best previous classifier on these data. This methodology is effective for identifying samples of specific science classes: for pulsational variables used in Milky Way tomography we obtain a discovery efficiency of 98.2% and for eclipsing systems we find an efficiency of 99.1%, both at 95% purity. We show that the random forest (RF) classifier is superior to other machine-learned methods in terms of accuracy, speed, and relative immunity to features with no useful class information; the RF classifier can also be used to estimate the importance of each feature in classification. Additionally, we present the first astronomical use of hierarchical classification methods to incorporate a known class taxonomy in the classifier, which further reduces the catastrophic error rate to 7.8%. Excluding low-amplitude sources, our overall error rate improves to 14%, with a catastrophic error rate of 3.5%.

preprint2011arXiv

Rapid, Machine-Learned Resource Allocation: Application to High-redshift GRB Follow-up

As the number of observed Gamma-Ray Bursts (GRBs) continues to grow, follow-up resources need to be used more efficiently in order to maximize science output from limited telescope time. As such, it is becoming increasingly important to rapidly identify bursts of interest as soon as possible after the event, before the afterglows fade beyond detectability. Studying the most distant (highest redshift) events, for instance, remains a primary goal for many in the field. Here we present our Random forest Automated Triage Estimator for GRB redshifts (RATE GRB-z) for rapid identification of high-redshift candidates using early-time metrics from the three telescopes onboard Swift. While the basic RATE methodology is generalizable to a number of resource allocation problems, here we demonstrate its utility for telescope-constrained follow-up efforts with the primary goal to identify and study high-z GRBs. For each new GRB, RATE GRB-z provides a recommendation - based on the available telescope time - of whether the event warrants additional follow-up resources. We train RATE GRB-z using a set consisting of 135 Swift bursts with known redshifts, only 18 of which are z > 4. Cross-validated performance metrics on this training data suggest that ~56% of high-z bursts can be captured from following up the top 20% of the ranked candidates, and ~84% of high-z bursts are identified after following up the top ~40% of candidates. We further use the method to rank 200+ Swift bursts with unknown redshifts according to their likelihood of being high-z.

preprint2011arXiv

Semi-supervised Learning for Photometric Supernova Classification

We present a semi-supervised method for photometric supernova typing. Our approach is to first use the nonlinear dimension reduction technique diffusion map to detect structure in a database of supernova light curves and subsequently employ random forest classification on a spectroscopically confirmed training set to learn a model that can predict the type of each newly observed supernova. We demonstrate that this is an effective method for supernova typing. As supernova numbers increase, our semi-supervised method efficiently utilizes this information to improve classification, a property not enjoyed by template based methods. Applied to supernova data simulated by Kessler et al. (2010b) to mimic those of the Dark Energy Survey, our methods achieve (cross-validated) 95% Type Ia purity and 87% Type Ia efficiency on the spectroscopic sample, but only 50% Type Ia purity and 50% efficiency on the photometric sample due to their spectroscopic follow-up strategy. To improve the performance on the photometric sample, we search for better spectroscopic follow-up procedures by studying the sensitivity of our machine learned supernova classification on the specific strategy used to obtain training sets. With a fixed amount of spectroscopic follow-up time, we find that deeper magnitude-limited spectroscopic surveys are better for producing training sets. For supernova Ia (II-P) typing, we obtain a 44% (1%) increase in purity to 72% (87%) and 30% (162%) increase in efficiency to 65% (84%) of the sample using a 25th (24.5th) magnitude-limited survey instead of the shallower spectroscopic sample used in the original simulations. When redshift information is available, we incorporate it into our analysis using a novel method of altering the diffusion map representation of the supernovae. Incorporating host redshifts leads to a 5% improvement in Type Ia purity and 13% improvement in Type Ia efficiency.

preprint2010arXiv

Results from the Supernova Photometric Classification Challenge

We report results from the Supernova Photometric Classification Challenge (SNPCC), a publicly released mix of simulated supernovae (SNe), with types (Ia, Ibc, and II) selected in proportion to their expected rate. The simulation was realized in the griz filters of the Dark Energy Survey (DES) with realistic observing conditions (sky noise, point-spread function and atmospheric transparency) based on years of recorded conditions at the DES site. Simulations of non-Ia type SNe are based on spectroscopically confirmed light curves that include unpublished non-Ia samples donated from the Carnegie Supernova Project (CSP), the Supernova Legacy Survey (SNLS), and the Sloan Digital Sky Survey-II (SDSS-II). A spectroscopically confirmed subset was provided for training. We challenged scientists to run their classification algorithms and report a type and photo-z for each SN. Participants from 10 groups contributed 13 entries for the sample that included a host-galaxy photo-z for each SN, and 9 entries for the sample that had no redshift information. Several different classification strategies resulted in similar performance, and for all entries the performance was significantly better for the training subset than for the unconfirmed sample. For the spectroscopically unconfirmed subset, the entry with the highest average figure of merit for classifying SNe~Ia has an efficiency of 0.96 and an SN~Ia purity of 0.79. As a public resource for the future development of photometric SN classification and photo-z estimators, we have released updated simulations with improvements based on our experience from the SNPCC, added samples corresponding to the Large Synoptic Survey Telescope (LSST) and the SDSS, and provided the answer keys so that developers can evaluate their own analysis.

preprint2009arXiv

Accurate parameter estimation for star formation history in galaxies using SDSS spectra

To further our knowledge of the complex physical process of galaxy formation, it is essential that we characterize the formation and evolution of large databases of galaxies. The spectral synthesis STARLIGHT code of Cid Fernandes et al. (2004) was designed for this purpose. Results of STARLIGHT are highly dependent on the choice of input basis of simple stellar population (SSP) spectra. Speed of the code, which uses random walks through the parameter space, scales as the square of the number of basis spectra, making it computationally necessary to choose a small number of SSPs that are coarsely sampled in age and metallicity. In this paper, we develop methods based on diffusion map (Lafon & Lee, 2006) that, for the first time, choose appropriate bases of prototype SSP spectra from a large set of SSP spectra designed to approximate the continuous grid of age and metallicity of SSPs of which galaxies are truly composed. We show that our techniques achieve better accuracy of physical parameter estimation for simulated galaxies. Specifically, we show that our methods significantly decrease the age-metallicity degeneracy that is common in galaxy population synthesis methods. We analyze a sample of 3046 galaxies in SDSS DR6 and compare the parameter estimates obtained from different basis choices.

preprint2008arXiv

Exploiting Low-Dimensional Structure in Astronomical Spectra

Dimension-reduction techniques can greatly improve statistical inference in astronomy. A standard approach is to use Principal Components Analysis (PCA). In this work we apply a recently-developed technique, diffusion maps, to astronomical spectra for data parameterization and dimensionality reduction, and develop a robust, eigenmode-based framework for regression. We show how our framework provides a computationally efficient means by which to predict redshifts of galaxies, and thus could inform more expensive redshift estimators such as template cross-correlation. It also provides a natural means by which to identify outliers (e.g., misclassified spectra, spectra with anomalous features). We analyze 3835 SDSS spectra and show how our framework yields a more than 95% reduction in dimensionality. Finally, we show that the prediction error of the diffusion map-based regression approach is markedly smaller than that of a similar approach based on PCA, clearly demonstrating the superiority of diffusion maps over PCA for this regression task.