Source author record

Andrew J. Connolly

Andrew J. Connolly appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

24works
12topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

24 published item(s)

preprint2026arXiv

Hyrax: An Extensible Framework for Rapid ML Experimentation and Unsupervised Discovery in the Era of Rubin, Roman, and Euclid

The NSF-DOE Vera C. Rubin Observatory, Roman Space Telescope, Euclid, and other next-generation surveys will deliver imaging, spectroscopic, and time-domain data at scales that increasingly shift the bottleneck in astronomical machine learning (ML) projects from model design to infrastructure. We present Hyrax, an open-source, modular, GPU-enabled Python framework that supports the full ML lifecycle in astronomy: from data acquisition and training to inference and experiment comparison, with capabilities including multimodal dataset support, integrated vector databases for similarity search, and interactive two- and three-dimensional latent-space exploration for unsupervised discovery. We demonstrate Hyrax's versatility through five representative applications on real survey data: (i) unsupervised representation learning on $\sim 4\times10^5$ Rubin Legacy Survey of Space and Time (LSST) Data Preview 1 (DP1) galaxies, surfacing new merger and low-surface-brightness candidates missing from reference Euclid and Dark Energy Survey catalogs, while also isolating imaging artifacts -- all without labeled training data; (ii) hybrid density-based clustering for identifying cluster-scale gravitational lens candidates in DP1 data; (iii) multimodal early-time transient classification in the Zwicky Transient Facility leveraging light curves, spectra, images, and metadata; (iv) supervised false-positive filtering in shift-and-stack searches for distant solar system objects in the Dark Energy Camera Ecliptic Exploration Project survey; and (v) supervised detection of semi-resolved dwarf galaxies in Hyper Suprime-Cam and LSST-like imaging using synthetic source injection. Together, these results demonstrate that Hyrax provides astronomy-specific ML infrastructure that enables systematic discovery and rapid methodological iteration across next-generation astronomical surveys.

preprint2026arXiv

You Only Stack Once (YOSO): A Motion-Filtered, Deep-Learning Framework for Detecting Faint Moving Sources

We present You Only Stack Once (YOSO), an automated pipeline designed to detect faint, slow-moving Solar System objects in wide-field astronomical surveys. The pipeline integrates a novel Gaussian Motion Filter (GMoF) that operates at the pixel level to enhance signal-to-noise for objects exhibiting a range of apparent rates of motion. Unlike conventional shift-and-stack methods, which rely on discrete velocity trials, GMoF amplifies trails while suppressing random noise and static background features. Applied to a subset of DEEP observations from the Dark Energy Camera, YOSO recovered 45 out of 73 previously detected objects, as well as 11 new TNOs. It also discovered 216 objects in the near Solar System. Although alternative shift-and-stack methods are sensitive to objects about 0.88 magnitudes fainter, YOSO's false positive rate is extremely low, since it detects only sources that exhibit a trail and are consistent with a point source when shifted at the right rate. We show how this method can be deployed on large surveys like LSST, and adapted for other domains that require motion-based signal enhancement, including exoplanet imaging through Angular Differential Imaging (ADI), and near-Earth object (NEO) detection for missions like NEO Surveyor. YOSO thus provides a versatile, scalable approach for extracting faint, motion-dependent signals in the era of data-intensive astronomy.

preprint2022arXiv

From Data to Software to Science with the Rubin Observatory LSST

The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) dataset will dramatically alter our understanding of the Universe, from the origins of the Solar System to the nature of dark matter and dark energy. Much of this research will depend on the existence of robust, tested, and scalable algorithms, software, and services. Identifying and developing such tools ahead of time has the potential to significantly accelerate the delivery of early science from LSST. Developing these collaboratively, and making them broadly available, can enable more inclusive and equitable collaboration on LSST science. To facilitate such opportunities, a community workshop entitled "From Data to Software to Science with the Rubin Observatory LSST" was organized by the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) and partners, and held at the Flatiron Institute in New York, March 28-30th 2022. The workshop included over 50 in-person attendees invited from over 300 applications. It identified seven key software areas of need: (i) scalable cross-matching and distributed joining of catalogs, (ii) robust photometric redshift determination, (iii) software for determination of selection functions, (iv) frameworks for scalable time-series analyses, (v) services for image access and reprocessing at scale, (vi) object image access (cutouts) and analysis at scale, and (vii) scalable job execution systems. This white paper summarizes the discussions of this workshop. It considers the motivating science use cases, identified cross-cutting algorithms, software, and services, their high-level technical specifications, and the principles of inclusive collaborations needed to develop them. We provide it as a useful roadmap of needs, as well as to spur action and collaboration between groups and individuals looking to develop reusable software for early LSST science.

preprint2022arXiv

Machine Learning and Cosmology

Methods based on machine learning have recently made substantial inroads in many corners of cosmology. Through this process, new computational tools, new perspectives on data collection, model development, analysis, and discovery, as well as new communities and educational pathways have emerged. Despite rapid progress, substantial potential at the intersection of cosmology and machine learning remains untapped. In this white paper, we summarize current and ongoing developments relating to the application of machine learning within cosmology and provide a set of recommendations aimed at maximizing the scientific impact of these burgeoning tools over the coming decade through both technical development as well as the fostering of emerging communities.

preprint2022arXiv

MUSSES2020J: The Earliest Discovery of a Fast Blue Ultraluminous Transient at Redshift 1.063

In this Letter, we report the discovery of an ultraluminous fast-evolving transient in rest-frame UV wavelengths, MUSSES2020J, soon after its occurrence by using the Hyper Suprime-Cam (HSC) mounted on the 8.2 m Subaru telescope. The rise time of about 5 days with an extremely high UV peak luminosity shares similarities to a handful of fast blue optical transients whose peak luminosities are comparable with the most luminous supernovae while their timescales are significantly shorter (hereafter "fast blue ultraluminous transient," FBUT). In addition, MUSSES2020J is located near the center of a normal low-mass galaxy at a redshift of 1.063, suggesting a possible connection between the energy source of MUSSES2020J and the central part of the host galaxy. Possible physical mechanisms powering this extreme transient such as a wind-driven tidal disruption event and an interaction between supernova and circumstellar material are qualitatively discussed based on the first multiband early-phase light curve of FBUTs, although whether the scenarios can quantitatively explain the early photometric behavior of MUSSES2020J requires systematical theoretical investigations. Thanks to the ultrahigh luminosity in UV and blue optical wavelengths of these extreme transients, a promising number of FBUTs from the local to the high-z universe can be discovered through deep wide-field optical surveys in the near future.

preprint2021arXiv

Optimization of the Observing Cadence for the Rubin Observatory Legacy Survey of Space and Time: a pioneering process of community-focused experimental design

Vera C. Rubin Observatory is a ground-based astronomical facility under construction, a joint project of the National Science Foundation and the U.S. Department of Energy, designed to conduct a multi-purpose 10-year optical survey of the southern hemisphere sky: the Legacy Survey of Space and Time. Significant flexibility in survey strategy remains within the constraints imposed by the core science goals of probing dark energy and dark matter, cataloging the Solar System, exploring the transient optical sky, and mapping the Milky Way. The survey's massive data throughput will be transformational for many other astrophysics domains and Rubin's data access policy sets the stage for a huge potential users' community. To ensure that the survey science potential is maximized while serving as broad a community as possible, Rubin Observatory has involved the scientific community at large in the process of setting and refining the details of the observing strategy. The motivation, history, and decision-making process of this strategy optimization are detailed in this paper, giving context to the science-driven proposals and recommendations for the survey strategy included in this Focus Issue.

preprint2021arXiv

The LSST DESC DC2 Simulated Sky Survey

We describe the simulated sky survey underlying the second data challenge (DC2) carried out in preparation for analysis of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) by the LSST Dark Energy Science Collaboration (LSST DESC). Significant connections across multiple science domains will be a hallmark of LSST; the DC2 program represents a unique modeling effort that stresses this interconnectivity in a way that has not been attempted before. This effort encompasses a full end-to-end approach: starting from a large N-body simulation, through setting up LSST-like observations including realistic cadences, through image simulations, and finally processing with Rubin's LSST Science Pipelines. This last step ensures that we generate data products resembling those to be delivered by the Rubin Observatory as closely as is currently possible. The simulated DC2 sky survey covers six optical bands in a wide-fast-deep (WFD) area of approximately 300 deg^2 as well as a deep drilling field (DDF) of approximately 1 deg^2. We simulate 5 years of the planned 10-year survey. The DC2 sky survey has multiple purposes. First, the LSST DESC working groups can use the dataset to develop a range of DESC analysis pipelines to prepare for the advent of actual data. Second, it serves as a realistic testbed for the image processing software under development for LSST by the Rubin Observatory. In particular, simulated data provide a controlled way to investigate certain image-level systematic effects. Finally, the DC2 sky survey enables the exploration of new scientific ideas in both static and time-domain cosmology.

preprint2020arXiv

Applying Information Theory to Design Optimal Filters for Photometric Redshifts

In this paper we apply ideas from information theory to create a method for the design of optimal filters for photometric redshift estimation. We show the method applied to a series of simple example filters in order to motivate an intuition for how photometric redshift estimators respond to the properties of photometric passbands. We then design a realistic set of six filters covering optical wavelengths that optimize photometric redshifts for $z <= 2.3$ and $i < 25.3$. We create a simulated catalog for these optimal filters and use our filters with a photometric redshift estimation code to show that we can improve the standard deviation of the photometric redshift error by 7.1% overall and improve outliers 9.9% over the standard filters proposed for the Large Synoptic Survey Telescope (LSST). We compare features of our optimal filters to LSST and find that the LSST filters incorporate key features for optimal photometric redshift estimation. Finally, we describe how information theory can be applied to a range of optimization problems in astronomy.

preprint2020arXiv

Dimensionality Reduction of SDSS Spectra with Variational Autoencoders

High resolution galaxy spectra contain much information about galactic physics, but the high dimensionality of these spectra makes it difficult to fully utilize the information they contain. We apply variational autoencoders (VAEs), a non-linear dimensionality reduction technique, to a sample of spectra from the Sloan Digital Sky Survey. In contrast to Principal Component Analysis (PCA), a widely used technique, VAEs can capture non-linear relationships between latent parameters and the data. We find that a VAE can reconstruct the SDSS spectra well with only six latent parameters, outperforming PCA with the same number of components. Different galaxy classes are naturally separated in this latent space, without class labels having been given to the VAE. The VAE latent space is interpretable because the VAE can be used to make synthetic spectra at any point in latent space. For example, making synthetic spectra along tracks in latent space yields sequences of realistic spectra that interpolate between two different types of galaxies. Using the latent space to find outliers may yield interesting spectra: in our small sample, we immediately find unusual data artifacts and stars misclassified as galaxies. In this exploratory work, we show that VAEs create compact, interpretable latent spaces that capture non-linear features of the data. While a VAE takes substantial time to train (~1 day for 48000 spectra), once trained, VAEs can enable the fast exploration of large astronomical data sets.

preprint2020arXiv

Photometric Redshifts with the LSST II: The Impact of Near-Infrared and Near-Ultraviolet Photometry

Accurate photometric redshift (photo-$z$) estimates are essential to the cosmological science goals of the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). In this work we use simulated photometry for mock galaxy catalogs to explore how LSST photo-$z$ estimates can be improved by the addition of near-infrared (NIR) and/or ultraviolet (UV) photometry from the Euclid, WFIRST, and/or CASTOR space telescopes. Generally, we find that deeper optical photometry can reduce the standard deviation of the photo-$z$ estimates more than adding NIR or UV filters, but that additional filters are the only way to significantly lower the fraction of galaxies with catastrophically under- or over-estimated photo-$z$. For Euclid, we find that the addition of ${JH}$ $5σ$ photometric detections can reduce the standard deviation for galaxies with $z>1$ ($z>0.3$) by ${\sim}20\%$ (${\sim}10\%$), and the fraction of outliers by ${\sim}40\%$ (${\sim}25\%$). For WFIRST, we show how the addition of deep ${YJHK}$ photometry could reduce the standard deviation by ${\gtrsim}50\%$ at $z>1.5$ and drastically reduce the fraction of outliers to just ${\sim}2\%$ overall. For CASTOR, we find that the addition of its ${UV}$ and $u$-band photometry could reduce the standard deviation by ${\sim}30\%$ and the fraction of outliers by ${\sim}50\%$ for galaxies with $z<0.5$. We also evaluate the photo-$z$ results within sky areas that overlap with both the NIR and UV surveys, and when spectroscopic training sets built from the surveys' small-area deep fields are used.

preprint2014arXiv

Determining Frequentist Confidence Limits Using a Directed Parameter Space Search

We consider the problem of inferring constraints on a high-dimensional parameter space with a computationally expensive likelihood function. We propose a machine learning algorithm that maps out the Frequentist confidence limit on parameter space by intelligently targeting likelihood evaluations so as to quickly and accurately characterize the likelihood surface in both low- and high-likelihood regions. We compare our algorithm to Bayesian credible limits derived by the well-tested Markov Chain Monte Carlo (MCMC) algorithm using both multi-modal toy likelihood functions and the 7-year WMAP cosmic microwave background likelihood function. We find that our algorithm correctly identifies the location, general size, and general shape of high-likelihood regions in parameter space while being more robust against multi-modality than MCMC.

preprint2014arXiv

Introduction to astroML: Machine Learning for Astrophysics

Astronomy and astrophysics are witnessing dramatic increases in data volume as detectors, telescopes and computers become ever more powerful. During the last decade, sky surveys across the electromagnetic spectrum have collected hundreds of terabytes of astronomical data for hundreds of millions of sources. Over the next decade, the data volume will enter the petabyte domain, and provide accurate measurements for billions of sources. Astronomy and physics students are not traditionally trained to handle such voluminous and complex data sets. In this paper we describe astroML; an initiative, based on Python and scikit-learn, to develop a compendium of machine learning tools designed to address the statistical needs of the next generation of students and astronomical surveys. We introduce astroML and present a number of example applications that are enabled by this package.

preprint2013arXiv

Variability-based AGN selection using image subtraction in the SDSS and LSST era

With upcoming all sky surveys such as LSST poised to generate a deep digital movie of the optical sky, variability-based AGN selection will enable the construction of highly-complete catalogs with minimum contamination. In this study, we generate $g$-band difference images and construct light curves for QSO/AGN candidates listed in SDSS Stripe 82 public catalogs compiled from different methods, including spectroscopy, optical colors, variability, and X-ray detection. Image differencing excels at identifying variable sources embedded in complex or blended emission regions such as Type II AGNs and other low-luminosity AGNs that may be omitted from traditional photometric or spectroscopic catalogs. To separate QSOs/AGNs from other sources using our difference image light curves, we explore several light curve statistics and parameterize optical variability by the characteristic damping timescale ($τ$) and variability amplitude. By virtue of distinguishable variability parameters of AGNs, we are able to select them with high completeness of 93.4% and efficiency (i.e., purity) of 71.3%. Based on optical variability, we also select highly variable blazar candidates, whose infrared colors are consistent with known blazars. One third of them are also radio detected. With the X-ray selected AGN candidates, we probe the optical variability of X-ray detected optically-extended sources using their difference image light curves for the first time. A combination of optical variability and X-ray detection enables us to select various types of host-dominated AGNs. Contrary to the AGN unification model prediction, two Type II AGN candidates (out of 6) show detectable variability on long-term timescales like typical Type I AGNs. This study will provide a baseline for future optical variability studies of extended sources.

preprint2012arXiv

The DEEP2 Galaxy Redshift Survey: Design, Observations, Data Reduction, and Redshifts

We describe the design and data sample from the DEEP2 Galaxy Redshift Survey, the densest and largest precision-redshift survey of galaxies at z ~ 1 completed to date. The survey has conducted a comprehensive census of massive galaxies, their properties, environments, and large-scale structure down to absolute magnitude M_B = -20 at z ~ 1 via ~90 nights of observation on the DEIMOS spectrograph at Keck Observatory. DEEP2 covers an area of 2.8 deg^2 divided into four separate fields, observed to a limiting apparent magnitude of R_AB=24.1. Objects with z < 0.7 are rejected based on BRI photometry in three of the four DEEP2 fields, allowing galaxies with z > 0.7 to be targeted ~2.5 times more efficiently than in a purely magnitude-limited sample. Approximately sixty percent of eligible targets are chosen for spectroscopy, yielding nearly 53,000 spectra and more than 38,000 reliable redshift measurements. Most of the targets which fail to yield secure redshifts are blue objects that lie beyond z ~ 1.45. The DEIMOS 1200-line/mm grating used for the survey delivers high spectral resolution (R~6000), accurate and secure redshifts, and unique internal kinematic information. Extensive ancillary data are available in the DEEP2 fields, particularly in the Extended Groth Strip, which has evolved into one of the richest multiwavelength regions on the sky. DEEP2 surpasses other deep precision-redshift surveys at z ~ 1 in terms of galaxy numbers, redshift accuracy, sample number density, and amount of spectral information. We also provide an overview of the scientific highlights of the DEEP2 survey thus far. This paper is intended as a handbook for users of the DEEP2 Data Release 4, which includes all DEEP2 spectra and redshifts, as well as for the publicly-available DEEP2 DEIMOS data reduction pipelines. [Abridged]

preprint2011arXiv

Classification of Stellar Spectra with LLE

We investigate the use of dimensionality reduction techniques for the classification of stellar spectra selected from the SDSS. Using local linear embedding (LLE), a technique that preserves the local (and possibly non-linear) structure within high dimensional data sets, we show that the majority of stellar spectra can be represented as a one dimensional sequence within a three dimensional space. The position along this sequence is highly correlated with spectral temperature. Deviations from this "stellar locus" are indicative of spectra with strong emission lines (including misclassified galaxies) or broad absorption lines (e.g. Carbon stars). Based on this analysis, we propose a hierarchical classification scheme using LLE that progressively identifies and classifies stellar spectra in a manner that requires no feature extraction and that can reproduce the classic MK classifications to an accuracy of one type.

preprint2011arXiv

Milky Way Tomography IV: Dissecting Dust

We use SDSS photometry of 73 million stars to simultaneously obtain best-fit main-sequence stellar energy distribution (SED) and amount of dust extinction along the line of sight towards each star. Using a subsample of 23 million stars with 2MASS photometry, whose addition enables more robust results, we show that SDSS photometry alone is sufficient to break degeneracies between intrinsic stellar color and dust amount when the shape of extinction curve is fixed. When using both SDSS and 2MASS photometry, the ratio of the total to selective absorption, $R_V$, can be determined with an uncertainty of about 0.1 for most stars in high-extinction regions. These fits enable detailed studies of the dust properties and its spatial distribution, and of the stellar spatial distribution at low Galactic latitudes. Our results are in good agreement with the extinction normalization given by the Schlegel et al. (1998, SFD) dust maps at high northern Galactic latitudes, but indicate that the SFD extinction map appears to be consistently overestimated by about 20% in the southern sky, in agreement with Schlafly et al. (2010). The constraints on the shape of the dust extinction curve across the SDSS and 2MASS bandpasses support the models by Fitzpatrick (1999) and Cardelli et al. (1989). For the latter, we find an $R_V=3.0\pm0.1$(random) $\pm0.1$(systematic) over most of the high-latitude sky. At low Galactic latitudes (|b|<5), we demonstrate that the SFD map cannot be reliably used to correct for extinction as most stars are embedded in dust, rather than behind it. We introduce a method for efficient selection of candidate red giant stars in the disk, dubbed "dusty parallax relation", which utilizes a correlation between distance and the extinction along the line of sight. We make these best-fit parameters, as well as all the input SDSS and 2MASS data, publicly available in a user-friendly format.

preprint2011arXiv

Pixel-z: Studying Substructure and Stellar Populations in Galaxies out to z~3 using Pixel Colors I. Systematics

We perform a pixel-by-pixel analysis of 467 galaxies in the GOODS-VIMOS survey to study systematic effects in extracting properties of stellar populations (age, dust, metallicity and SFR) from pixel colors using the pixel-z method. The systematics studied include the effect of the input stellar population synthesis model, passband limitations and differences between individual SED fits to pixels and global SED-fitting to a galaxy's colors. We find that with optical-only colors, the systematic errors due to differences among the models are well constrained. The largest impact on the age and SFR e-folding time estimates in the pixels arises from differences between the Maraston models and the Bruzual&Charlot models, when optical colors are used. This results in systematic differences larger than the 2σ uncertainties in over 10 percent of all pixels in the galaxy sample. The effect of restricting the available passbands is more severe. In 26 percent of pixels in the full sample, passband limitations result in systematic biases in the age estimates which are larger than the 2σ uncertainties. Systematic effects from model differences are reexamined using Near-IR colors for a subsample of 46 galaxies in the GOODS-NICMOS survey. For z > 1, the observed optical/NIR colors span the rest frame UV-optical SED, and the use of different models does not significantly bias the estimates of the stellar population parameters compared to using optical-only colors. We then illustrate how pixel-z can be applied robustly to make detailed studies of substructure in high redshift galaxies such as (a) radial gradients of age, SFR, sSFR and dust and (b) the distribution of these properties within subcomponents such as spiral arms and clumps. Finally, we show preliminary results from the CANDELS survey illustrating how the new HST/WFC3 data can be exploited to probe substructure in z~1-3 galaxies.

preprint2010arXiv

Cross-Identification of Stars with Unknown Proper Motions

The cross-identification of sources in separate catalogs is one of the most basic tasks in observational astronomy. It is, however, surprisingly difficult and generally ill-defined. Recently Budavári & Szalay (2008) formulated the problem in the realm of probability theory, and laid down the statistical foundations of an extensible methodology. In this paper, we apply their Bayesian approach to stars that, we know, can move measurably on the sky, with detectable proper motion, and show how to associate their observations. We study models on a sample of stars in the Sloan Digital Sky Survey, which allow for an unknown proper motion per object, and demonstrate the improvements over the analytic static model. Our models and conclusions are directly applicable to upcoming surveys such as PanSTARRS, the Dark Energy Survey, Sky Mapper, and the LSST, whose data sets will contain hundreds of millions of stars observed multiple times over several years.

preprint2010arXiv

Three-Point Correlation Functions of SDSS Galaxies: Constraining Galaxy-Mass Bias

We constrain the linear and quadratic bias parameters from the configuration dependence of the three-point correlation function (3PCF) in both redshift and projected space, utilizing measurements of spectroscopic galaxies in the Sloan Digital Sky Survey (SDSS) Main Galaxy Sample. We show that bright galaxies (M_r < -21.5) are biased tracers of mass, measured at a significance of 4.5 sigma in redshift space and 2.5 sigma in projected space by using a thorough error analysis in the quasi-linear regime (9-27 Mpc/h). Measurements on a fainter galaxy sample are consistent with an unbiased model. We demonstrate that a linear bias model appears sufficient to explain the galaxy-mass bias of our samples, although a model using both linear and quadratic terms results in a better fit. In contrast, the bias values obtained from the linear model appear in better agreement with the data by inspection of the relative bias, and yield implied values of sigma_8 that are more consistent with current constraints. We investigate the covariance of the 3PCF, which itself is a measurement of galaxy clustering. We assess the accuracy of our error estimates by comparing results from mock galaxy catalogs to jackknife re-sampling methods. We identify significant differences in the structure of the covariance. However, the impact of these discrepancies appears to be mitigated by an eigenmode analysis that can account for the noisy, unresolved modes. Our results demonstrate that using this technique is sufficient to remove potential systematics even when using less-than-ideal methods to estimate errors.

preprint2010arXiv

Three-Point Correlation Functions of SDSS Galaxies: Luminosity and Color Dependence in Redshift and Projected Space

The three-point correlation function (3PCF) provides an important view into the clustering of galaxies that is not available to its lower order cousin, the two-point correlation function (2PCF). Higher order statistics, such as the 3PCF, are necessary to probe the non-Gaussian structure and shape information expected in these distributions. We measure the clustering of spectroscopic galaxies in the Main Galaxy Sample of the Sloan Digital Sky Survey (SDSS), focusing on the shape or configuration dependence of the reduced 3PCF in both redshift and projected space. This work constitutes the largest number of galaxies ever used to investigate the reduced 3PCF, using over 220,000 galaxies in three volume-limited samples. We find significant configuration dependence of the reduced 3PCF at 3-27 Mpc/h, in agreement with LCDM predictions and in disagreement with the hierarchical ansatz. Below 6 Mpc/h, the redshift space reduced 3PCF shows a smaller amplitude and weak configuration dependence in comparison with projected measurements suggesting that redshift distortions, and not galaxy bias, can make the reduced 3PCF appear consistent with the hierarchical ansatz. The reduced 3PCF shows a weaker dependence on luminosity than the 2PCF, with no significant dependence on scales above 9 Mpc/h. On scales less than 9 Mpc/h, the reduced 3PCF appears more affected by galaxy color than luminosty. We demonstrate the extreme sensitivity of the 3PCF to systematic effects such as sky completeness and binning scheme, along with the difficulty of resolving the errors. Some comparable analyses make assumptions that do not consistently account for these effects.

preprint2009arXiv

Clustering of Low-Redshift (z <= 2.2) Quasars from the Sloan Digital Sky Survey

We present measurements of the quasar two-point correlation function, ξ_{Q}, over the redshift range z=0.3-2.2 based upon data from the SDSS. Using a homogeneous sample of 30,239 quasars with spectroscopic redshifts from the DR5 Quasar Catalogue, our study represents the largest sample used for this type of investigation to date. With this redshift range and an areal coverage of approx 4,000 deg^2, we sample over 25 h^-3 Gpc^3 (comoving) assuming the current LCDM cosmology. Over this redshift range, we find that the redshift-space correlation function, xi(s), is adequately fit by a single power-law, with s_{0}=5.95+/-0.45 h^-1 Mpc and γ_{s}=1.16+0.11-0.16 when fit over s=1-25 h^-1 Mpc. Using the projected correlation function we calculate the real-space correlation length, r_{0}=5.45+0.35-0.45 h^-1 Mpc and γ=1.90+0.04-0.03, over scales of rp=1-130 h^-1 Mpc. Dividing the sample into redshift slices, we find very little, if any, evidence for the evolution of quasar clustering, with the redshift-space correlation length staying roughly constant at s_{0} ~ 6-7 h^-1 Mpc at z<2.2 (and only increasing at redshifts greater than this). Comparing our clustering measurements to those reported for X-ray selected AGN at z=0.5-1, we find reasonable agreement in some cases but significantly lower correlation lengths in others. We find that the linear bias evolves from b~1.4 at z=0.5 to b~3 at z=2.2, with b(z=1.27)=2.06+/-0.03 for the full sample. We compare our data to analytical models and infer that quasars inhabit dark matter haloes of constant mass M ~2 x 10^12 h^-1 M_Sol from redshifts z~2.5 (the peak of quasar activity) to z~0. [ABRIDGED]

preprint2008arXiv

Quasar Clustering from SDSS DR5: Dependences on Physical Properties

Using a homogenous sample of 38,208 quasars with a sky coverage of $4000 {\rm deg^2}$ drawn from the SDSS Data Release Five quasar catalog, we study the dependence of quasar clustering on luminosity, virial black hole mass, quasar color, and radio loudness. At $z<2.5$, quasar clustering depends weakly on luminosity and virial black hole mass, with typical uncertainty levels $\sim 10%$ for the measured correlation lengths. These weak dependences are consistent with models in which substantial scatter between quasar luminosity, virial black hole mass and the host dark matter halo mass has diluted any clustering difference, where halo mass is assumed to be the relevant quantity that best correlates with clustering strength. However, the most luminous and most massive quasars are more strongly clustered (at the $\sim 2σ$ level) than the remainder of the sample, which we attribute to the rapid increase of the bias factor at the high-mass end of host halos. We do not observe a strong dependence of clustering strength on quasar colors within our sample. On the other hand, radio-loud quasars are more strongly clustered than are radio-quiet quasars matched in redshift and optical luminosity (or virial black hole mass), consistent with local observations of radio galaxies and radio-loud type 2 AGN. Thus radio-loud quasars reside in more massive and denser environments in the biased halo clustering picture. Using the Sheth et al.(2001) formula for the linear halo bias, the estimated host halo mass for radio-loud quasars is $\sim 10^{13} h^{-1}M_\odot$, compared to $\sim 2\times 10^{12} h^{-1}M_\odot$ for radio-quiet quasar hosts at $z\sim 1.5$.

preprint2004arXiv

Three-point Correlation Functions of SDSS Galaxies in Redshift Space: Morphology, Color, and Luminosity Dependence

We present measurements of the redshift--space three-point correlation function of galaxies in the Sloan Digital Sky Survey (SDSS). For the first time, we analyze the dependence of this statistic on galaxy morphology, color and luminosity. In order to control systematics due to selection effects, we used $r$--band, volume-limited samples of galaxies, constructed from the magnitude-limited SDSS data ($14.5<r<17.5$), and further divided the samples into two morphological types (early and late) or two color populations (red and blue). The three-point correlation function of SDSS galaxies follow the hierarchical relation well and the reduced three-point amplitudes in redshift--space are almost scale-independent ($Q_z=0.5\sim1.0$). In addition, their dependence on the morphology, color and luminosity is not statistically significant. Given the robust morphological, color and luminosity dependences of the two-point correlation function, this implies that galaxy biasing is complex on weakly non-linear to non-linear scales. We show that simple deterministic linear relation with the underlying mass could not explain our measurements on these scales.

preprint1999arXiv

Creating Spectral Templates from Multicolor Redshift Surveys

We present a novel method capable of creating optimal eigenspectra from multicolor redshift surveys for photometric redshift estimation. Our iterative training algorithm modifies the templates to represent the photometric measurements better. We present a short description of our algorithm here. We show that the corrected templates give more precise photometric redshifts, essentially a ``free'' feature, since we were not fitting for the redshifts themselves.