Source author record

David W. Hogg

David W. Hogg appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

120works
24topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

120 published item(s)

preprint2022arXiv

An empirical model of the Gaia DR3 selection function

Interpreting and modelling astronomical catalogues requires an understanding of the catalogues' completeness or selection function: objects of what properties had a chance to end up in the catalogue. Here we set out to empirically quantify the completeness of the overall Gaia DR3 catalogue. This task is not straightforward because Gaia is the all-sky optical survey with the highest angular resolution to date and no consistent ``ground truth'' exists to allow direct comparisons. However, well-characterised deeper imaging enables an empirical assessment of Gaia's $G$-band completeness across parts of the sky. On this basis, we devised a simple analytical completeness model of Gaia as a function of the observed $G$ magnitude and position over the sky, which accounts for both the effects of crowding and the complex Gaia scanning law. Our model only depends on a single quantity: the median magnitude $M_{10}$ in a patch of the sky of catalogued sources with $\texttt{astrometric_matched_transits}$ $\leq 10$. $M_{10}$ reflects elementary completeness decisions in the Gaia pipeline and is computable from the Gaia DR3 catalogue itself and therefore applicable across the whole sky. We calibrate our model using the Dark Energy Camera Plane Survey (DECaPS) and test its predictions against Hubble Space Telescope observations of globular clusters. We find that our model predicts Gaia's completeness values to a few per cent across the sky. We make the model available as a part of the $\texttt{gaiasf}$ Python package built and maintained by the GaiaUnlimited project: $\texttt{https://github.com/gaia-unlimited/gaiaunlimited}$

preprint2022arXiv

GREAT3 results I: systematic errors in shear estimation and the impact of real galaxy morphology

We present first results from the third GRavitational lEnsing Accuracy Testing (GREAT3) challenge, the third in a sequence of challenges for testing methods of inferring weak gravitational lensing shear distortions from simulated galaxy images. GREAT3 was divided into experiments to test three specific questions, and included simulated space- and ground-based data with constant or cosmologically-varying shear fields. The simplest (control) experiment included parametric galaxies with a realistic distribution of signal-to-noise, size, and ellipticity, and a complex point spread function (PSF). The other experiments tested the additional impact of realistic galaxy morphology, multiple exposure imaging, and the uncertainty about a spatially-varying PSF; the last two questions will be explored in Paper II. The 24 participating teams competed to estimate lensing shears to within systematic error tolerances for upcoming Stage-IV dark energy surveys, making 1525 submissions overall. GREAT3 saw considerable variety and innovation in the types of methods applied. Several teams now meet or exceed the targets in many of the tests conducted (to within the statistical errors). We conclude that the presence of realistic galaxy morphology in simulations changes shear calibration biases by $\sim 1$ per cent for a wide range of methods. Other effects such as truncation biases due to finite galaxy postage stamps, and the impact of galaxy type as measured by the Sérsic index, are quantified for the first time. Our results generalize previous studies regarding sensitivities to galaxy size and signal-to-noise, and to PSF properties such as seeing and defocus. Almost all methods' results support the simple model in which additive shear biases depend linearly on PSF ellipticity.

preprint2022arXiv

Magnitudes, distance moduli, bolometric corrections, and so much more

This pedagogical document about stellar photometry - aimed at those for whom astronomical arcana seem arcane - endeavours to explain the concepts of magnitudes, color indices, absolute magnitudes, distance moduli, extinctions, attenuations, color excesses, K corrections, and bolometric corrections. I include some discussion of observational technique, and some discussion of epistemology, but the primary focus here is on the theoretical or interpretive connections between the observational astronomical quantities and the physical properties of the observational targets.

preprint2022arXiv

Mapping Interstellar Dust with Gaussian Processes

Interstellar dust corrupts nearly every stellar observation, and accounting for it is crucial to measuring physical properties of stars. We model the dust distribution as a spatially varying latent field with a Gaussian process (GP) and develop a likelihood model and inference method that scales to millions of astronomical observations. Modeling interstellar dust is complicated by two factors. The first is integrated observations. The data come from a vantage point on Earth and each observation is an integral of the unobserved function along our line of sight, resulting in a complex likelihood and a more difficult inference problem than in classical GP inference. The second complication is scale; stellar catalogs have millions of observations. To address these challenges we develop ziggy, a scalable approach to GP inference with integrated observations based on stochastic variational inference. We study ziggy on synthetic data and the Ananke dataset, a high-fidelity mechanistic model of the Milky Way with millions of stars. ziggy reliably infers the spatial dust map with well-calibrated posterior uncertainties.

preprint2022arXiv

Stellar Abundance Maps of the Milky Way Disk

To understand the formation of the Milky Way's prominent bar it is important to know whether stars in the bar differ in the chemical element composition of their birth material as compared to disk stars. This requires stellar abundance measurements for large samples across the Milky Way's body. Such samples, e.g. luminous red giant stars observed by SDSS's Apogee survey, will inevitably span a range of stellar parameters; as a consequence, both modelling imperfections and stellar evolution may preclude consistent and precise estimates of their chemical composition at a level of purported bar signatures, which has left current analyses of a chemically distinct bar inconclusive. Here, we develop a new self-calibration approach to eliminate both modelling and astrophysical abundance systematics among red giant branch (RGB) stars of different luminosities (and hence surface gravity $\log g$). We apply our method to $48,853$ luminous Apogee DR16 RGB stars to construct spatial abundance maps of $20$ chemical elements near the Milky Way's mid-plane, covering Galactocentric radii of $0\,{\rm kpc}<R_{\rm GC}<20\,\rm kpc$. Our results indicate that there are no abundance variations whose geometry matches that of the bar, and that the mean abundance gradients vary smoothly and monotonically with Galactocentric radius. We confirm that the high-$α$ disk is chemically homogeneous, without spatial gradients. Furthermore, we present the most precise [Fe/H] vs. $R_{\rm GC}$ gradient to date with a slope of $-0.057\pm0.001\rm~dex\,kpc^{-1}$ out to approximately $15$ kpc.

preprint2022arXiv

The EXPRES Stellar Signals Project II. State of the Field in Disentangling Photospheric Velocities

Measured spectral shifts due to intrinsic stellar variability (e.g., pulsations, granulation) and activity (e.g., spots, plages) are the largest source of error for extreme precision radial velocity (EPRV) exoplanet detection. Several methods are designed to disentangle stellar signals from true center-of-mass shifts due to planets. The EXPRES Stellar Signals Project (ESSP) presents a self-consistent comparison of 22 different methods tested on the same extreme-precision spectroscopic data from EXPRES. Methods derived new activity indicators, constructed models for mapping an indicator to the needed RV correction, or separated out shape- and shift-driven RV components. Since no ground truth is known when using real data, relative method performance is assessed using the total and nightly scatter of returned RVs and agreement between the results of different methods. Nearly all submitted methods return a lower RV RMS than classic linear decorrelation, but no method is yet consistently reducing the RV RMS to sub-meter-per-second levels. There is a concerning lack of agreement between the RVs returned by different methods. These results suggest that continued progress in this field necessitates increased interpretability of methods, high-cadence data to capture stellar signals at all timescales, and continued tests like the ESSP using consistent data sets with more advanced metrics for method performance. Future comparisons should make use of various well-characterized data sets -- such as solar data or data with known injected planetary and/or stellar signals -- to better understand method performance and whether planetary signals are preserved.

preprint2022arXiv

The Thresher: Lucky Imaging without the Waste

In traditional lucky imaging (TLI), many consecutive images of the same scene are taken with a high frame-rate camera, and all but the sharpest images are discarded before constructing the final shift-and-add image. Here we present an alternative image analysis pipeline -- The Thresher -- for these kinds of data, based on online multi-frame blind deconvolution. It makes use of all available data to obtain a best estimate of the astronomical scene in the context of reasonable computational limits; it does not require prior estimates of the point-spread functions in the images, or knowledge of point sources in the scene that could provide such estimates. Most importantly, the scene it aims to return is the optimum of a justified scalar objective based on the likelihood function. Because it uses the full set of images in the stack, The Thresher outperforms TLI in signal-to-noise; as it accounts for the individual-frame PSFs, it does this without loss of angular resolution. We demonstrate the effectiveness of our algorithm on both simulated data and real Electron-Multiplying CCD images obtained at the Danish 1.54m telescope (hosted by ESO, La Silla). We also explore the current limitations of the algorithm, and find that for the choice of image model presented here, non-linearities in flux are introduced into the returned scene. Ongoing development of the software can be viewed at https://github.com/jah1994/TheThresher.

preprint2022arXiv

The unpopular Package: a Data-driven Approach to De-trend TESS Full Frame Image Light Curves

The majority of observed pixels on the Transiting Exoplanet Survey Satellite (TESS) are delivered in the form of full frame images (FFI). However, the FFIs contain systematic effects such as pointing jitter and scattered light from the Earth and Moon that must be removed before downstream analysis. We present unpopular, an open-source Python package to de-trend TESS FFI light curves based on the causal pixel model method. Under the assumption that shared flux variations across multiple distant pixels are likely to be systematics, unpopular removes these common (i.e., popular) trends by modeling the systematics in a given pixel's light curve as a linear combination of light curves from many other distant pixels. To prevent overfitting we employ ridge regression and a train-and-test framework where the data points being de-trended are separated from those used to obtain the model coefficients. We also allow for simultaneous fitting with a polynomial model to capture any long-term astrophysical trends. We validate our method by de-trending different sources (e.g., supernova, tidal disruption event, exoplanet-hosting star, fast rotating star) and comparing our light curves to those obtained by other pipelines when appropriate. We also show that unpopular is able to preserve sector-length astrophysical signals, allowing for the extraction of multi-sector light curves from the FFI data. The unpopular source code and tutorials are freely available online.

preprint2021arXiv

An unsupervised method for identifying $X$-enriched stars directly from spectra: Li in LAMOST

Stars with peculiar element abundances are important markers of chemical enrichment mechanisms. We present a simple method, tangent space projection (TSP), for the detection of $X$-enriched stars, for arbitrary elements $X$, even from blended lines. Our method does not require stellar labels, but instead directly estimates the counterfactual unrenriched spectrum from other unlabelled spectra. As a case study, we apply this method to the $6708~$Å Li doublet in LAMOST DR5, identifying 8,428 Li-enriched stars seamlessly across evolutionary state. We comment on the explanation for Li-enrichement for different subpopulations, including planet accretion, nonstandard mixing, and youth.

preprint2021arXiv

Dimensionality reduction, regularization, and generalization in overparameterized regressions

Overparameterization in deep learning is powerful: Very large models fit the training data perfectly and yet often generalize well. This realization brought back the study of linear models for regression, including ordinary least squares (OLS), which, like deep learning, shows a "double-descent" behavior: (1) The risk (expected out-of-sample prediction error) can grow arbitrarily when the number of parameters $p$ approaches the number of samples $n$, and (2) the risk decreases with $p$ for $p>n$, sometimes achieving a lower value than the lowest risk for $p<n$. The divergence of the risk for OLS can be avoided with regularization. In this work, we show that for some data models it can also be avoided with a PCA-based dimensionality reduction (PCA-OLS, also known as principal component regression). We provide non-asymptotic bounds for the risk of PCA-OLS by considering the alignments of the population and empirical principal components. We show that dimensionality reduction improves robustness while OLS is arbitrarily susceptible to adversarial attacks, particularly in the overparameterized regime. We compare PCA-OLS theoretically and empirically with a wide range of projection-based methods, including random projections, partial least squares (PLS), and certain classes of linear two-layer neural networks. These comparisons are made for different data generation models to assess the sensitivity to signal-to-noise and the alignment of regression coefficients with the features. We find that methods in which the projection depends on the training data can outperform methods where the projections are chosen independently of the training data, even those with oracle knowledge of population quantities, another seemingly paradoxical phenomenon that has been identified previously. This suggests that overparameterization may not be necessary for good generalization.

preprint2021arXiv

Snails Across Scales: Local and Global Phase-Mixing Structures as Probes of the Past and Future Milky Way

Signatures of vertical disequilibrium have been observed across the Milky Way's disk. These signatures manifest locally as unmixed phase-spirals in $z$--$v_z$ space ("snails-in-phase") and globally as nonzero mean $z$ and $v_z$ which wraps around as a physical spiral across the $x$--$y$ plane ("snails-in-space"). We explore the connection between these local and global spirals through the example of a satellite perturbing a test-particle Milky Way (MW)-like disk. We anticipate our results to broadly apply to any vertical perturbation. Using a $z$--$v_z$ asymmetry metric we demonstrate that in test-particle simulations: (a) multiple local phase-spiral morphologies appear when stars are binned by azimuthal action $J_ϕ$, excited by a single event (in our case, a satellite disk-crossing); (b) these distinct phase-spirals are traced back to distinct disk locations; and (c) they are excited at distinct times. Thus, local phase-spirals offer a global view of the MW's perturbation history from multiple perspectives. Using a toy model for a Sagittarius (Sgr)-like satellite crossing the disk, we show that the full interaction takes place on timescales comparable to orbital periods of disk stars within $R \lesssim 10$ kpc. Hence such perturbations have widespread influence which peaks in distinct regions of the disk at different times. This leads us to examine the ongoing MW-Sgr interaction. While Sgr has not yet crossed the disk (currently, $z_{Sgr} \approx -6$ kpc, $v_{z,Sgr} \approx 210$ km/s), we demonstrate that the peak of the impact has already passed. Sgr's pull over the past 150 Myr creates a global $v_z$ signature with amplitude $\propto M_{Sgr}$, which might be detectable in future spectroscopic surveys.

preprint2020arXiv

Close Binary Companions to APOGEE DR16 Stars: 20,000 Binary-star Systems Across the Color-Magnitude Diagram

Many problems in contemporary astrophysics---from understanding the formation of black holes to untangling the chemical evolution of galaxies---rely on knowledge about binary stars. This, in turn, depends on discovery and characterization of binary companions for large numbers of different kinds of stars in different chemical and dynamical environments. Current stellar spectroscopic surveys observe hundreds of thousands to millions of stars with (typically) few observational epochs, which allows binary discovery but makes orbital characterization challenging. We use a custom Monte Carlo sampler (The Joker) to perform discovery and characterization of binary systems through radial-velocities, in the regime of sparse, noisy, and poorly sampled multi-epoch data. We use it to generate posterior samplings in Keplerian parameters for 232,531 sources released in APOGEE Data Release 16. Our final catalog contains 19,635 high-confidence close-binary (P < few years, a < few AU) systems that show interesting relationships between binary occurrence rate and location in the color-magnitude diagram. We find notable faint companions at high masses (black-hole candidates), at low masses (substellar candidates), and at very close separations (mass-transfer candidates). We also use the posterior samplings in a (toy) hierarchical inference to measure the long-period binary-star eccentricity distribution. We release the full set of posterior samplings for the entire parent sample of 232,531 stars. This set of samplings involves no heuristic "discovery" threshold and therefore can be used for myriad statistical purposes, including hierarchical inferences about binary-star populations and sub-threshold searches.

preprint2020arXiv

Data Analysis Recipes: Products of multivariate Gaussians in Bayesian inferences

A product of two Gaussians (or normal distributions) is another Gaussian. That's a valuable and useful fact! Here we use it to derive a refactoring of a common product of multivariate Gaussians: The product of a Gaussian likelihood times a Gaussian prior, where some or all of those parameters enter the likelihood only in the mean and only linearly. That is, a linear, Gaussian, Bayesian model. This product of a likelihood times a prior pdf can be refactored into a product of a marginalized likelihood (or a Bayesian evidence) times a posterior pdf, where (in this case) both of these are also Gaussian. The means and variance tensors of the refactored Gaussians are straightforward to obtain as closed-form expressions; here we deliver these expressions, with discussion. The closed-form expressions can be used to speed up and improve the precision of inferences that contain linear parameters with Gaussian priors. We connect these methods to inferences that arise frequently in physics and astronomy. If all you want is the answer, the question is posed and answered at the beginning of Section 3. We show two toy examples, in the form of worked exercises, in Section 4. The solutions, discussion, and exercises in this Note are aimed at someone who is already familiar with the basic ideas of Bayesian inference and probability.

preprint2020arXiv

Forward modeling the orbits of companions to pulsating stars from their light travel time variations

Mutual gravitation between a pulsating star and an orbital companion leads to a time-dependent variation in path length for starlight traveling to Earth. These variations can be used for coherently pulsating stars, such as the δ Scuti variables, to constrain the masses and orbits of their companions. Observing these variations for δ Scuti stars has previously relied on subdividing the light curve and measuring the average pulsation phase in equally sized subdivisions, which leads to under-sampling near periapsis. We introduce a new approach that simultaneously forward-models each sample in the light curve and show that this method improves upon current sensitivity limits - especially in the case of highly eccentric and short-period binaries. We find that this approach is sensitive enough to observe Jupiter mass planets around δ Scuti stars under ideal conditions, and use gravity-mode pulsations in the subdwarf B star KIC 7668647 to detect its companion without radial velocity data. We further provide robust detection limits as a function of the SNR of the pulsation mode and determine that the minimum detectable light travel time amplitude for a typical Kepler δ Scuti is around 2 s. This new method significantly enhances the application of light travel time variations to detecting short period binaries with pulsating components, and pulsating A-type exoplanet host stars, especially as a tool for eliminating false positives.

preprint2020arXiv

High-resolution spectroscopy of the GD-1 stellar stream localizes the perturber near the orbital plane of Sagittarius

The $100^\circ$-long thin stellar stream in the Milky Way halo, GD-1, has an ensemble of features that may be due to dynamical interactions. Using high-resolution MMT/Hectochelle spectroscopy we show that a spur of GD-1-like stars outside of the main stream are kinematically and chemically consistent with the main stream. In the spur, as in the main stream, GD-1 has a low intrinsic radial velocity dispersion, $σ_{V_r}\lesssim1\,\rm km\,s^{-1}$, is metal-poor, $\rm [Fe/H]\approx-2.3$, with little $\rm [Fe/H]$ spread and some variation in $\rm [α/Fe]$ abundances, which point to a common globular cluster progenitor. At a fixed location along the stream, the median radial velocity offset between the spur and the main stream is smaller than $0.5\,\rm km\,s^{-1}$, comparable to the measurement uncertainty. A flyby of a massive, compact object can change orbits of stars in a stellar stream and produce features like the spur observed in GD-1. In this scenario, the radial velocity of the GD-1 spur relative to the stream constrains the orbit of the perturber and its current on-sky position to $\approx5,000\,\rm deg^2$. The family of acceptable perturber orbits overlaps the stellar and dark-matter debris of the Sagittarius dwarf galaxy in present-day position and velocity. This suggests that GD-1 may have been perturbed by a globular cluster or an extremely compact dark-matter subhalo formerly associated with Sagittarius.

preprint2020arXiv

How to obtain the redshift distribution from probabilistic redshift estimates

A trustworthy estimate of the redshift distribution $n(z)$ is crucial for using weak gravitational lensing and large-scale structure of galaxy catalogs to study cosmology. Spectroscopic redshifts for the dim and numerous galaxies of next-generation weak-lensing surveys are expected to be unavailable, making photometric redshift (photo-$z$) probability density functions (PDFs) the next-best alternative for comprehensively encapsulating the nontrivial systematics affecting photo-$z$ point estimation. The established stacked estimator of $n(z)$ avoids reducing photo-$z$ PDFs to point estimates but yields a systematically biased estimate of $n(z)$ that worsens with decreasing signal-to-noise, the very regime where photo-$z$ PDFs are most necessary. We introduce Cosmological Hierarchical Inference with Probabilistic Photometric Redshifts (CHIPPR), a statistically rigorous probabilistic graphical model of redshift-dependent photometry, which correctly propagates the redshift uncertainty information beyond the best-fit estimator of $n(z)$ produced by traditional procedures and is provably the only self-consistent way to recover $n(z)$ from photo-$z$ PDFs. We present the $\texttt{chippr}$ prototype code, noting that the mathematically justifiable approach incurs computational expense. The CHIPPR approach is applicable to any one-point statistic of any random variable, provided the prior probability density used to produce the posteriors is explicitly known; if the prior is implicit, as may be the case for popular photo-$z$ techniques, then the resulting posterior PDFs cannot be used for scientific inference. We therefore recommend that the photo-$z$ community focus on developing methodologies that enable the recovery of photo-$z$ likelihoods with support over all redshifts, either directly or via a known prior probability density.

preprint2020arXiv

Principled point-source detection in collections of astronomical images

We review the well-known matched filter method for the detection of point sources in astronomical images. This is shown to be optimal (that is, to saturate the Cramer--Rao bound) under stated conditions that are very strong: an isolated source in background-dominated imaging with perfectly known background level, point-spread function, and noise models. We show that the matched filter produces a maximum-likelihood estimate of the brightness of a purported point source, and this leads to a simple way to combine multiple images---taken through the same bandpass filter but with different noise levels and point-spread functions---to produce an optimal point source detection map. We then extend the approach to images taken through different bandpass filters, introducing the SED-matched filter, which allows us to combine images taken through different filters, but requires us to specify the colors of the objects we wish to detect. We show that this approach is superior to some methods traditionally employed, and that other traditional methods can be seen as instances of SED-matched filtering with implied (and often unreasonable) priors. We present a Bayesian formulation, including a flux prior that leads to a closed-form expression with low computational cost.

preprint2020arXiv

Temperatures and Metallicities of M dwarfs in the APOGEE Survey

M dwarfs have enormous potential for our understanding of structure and formation on both Galactic and exoplanetary scales through their properties and compositions. However, current atmosphere models have limited ability to reproduce spectral features in stars at the coolest temperatures ($T_{\rm eff} < 4200\,$K) and to fully exploit the information content of current and upcoming large-scale spectroscopic surveys. Here we present a catalog of spectroscopic temperatures, metallicities, and spectral types for 5875 M dwarfs in the Apache Point Observatory Galactic Evolution Experiment (APOGEE) and Gaia DR2 surveys using The Cannon: a flexible, data-driven spectral-modeling and parameter-inference framework demonstrated to estimate stellar-parameter labels ($T_{\rm eff}$, log$g$, [Fe/H], and detailed abundances) to high precision. Using a training sample of 87 M dwarfs with optically derived labels spanning $2860 < T_{\rm eff} < 4130\,$K calibrated with bolometric temperatures, and $-0.5 < $[Fe/H]$ < 0.5\,$dex calibrated with FGK binary metallicities, we train a two-parameter model with predictive accuracy (in cross-validation) to $77\,$K and $0.09\,$dex respectively. We also train a one-dimensional spectral classification model using 51 M dwarfs with Sloan Digital Sky Survey optical spectral types ranging from M0 to M6, to predictive accuracy of 0.7 types. We find Cannon temperatures to be in agreement to within $60\,$K compared to a subsample of 1702 sources with color-derived temperatures, and Cannon metallicities to be in agreement to within $0.08\,$dex metallicity compared to a subsample of 15 FGK+M or M+M binaries. Finally, our comparison between Cannon and APOGEE pipeline (ASPCAP DR14) labels finds that ASPCAP is systematically biased toward reporting higher temperatures and lower metallicities for M dwarfs.

preprint2020arXiv

The power of co-ordinate transformations in dynamical interpretations of Galactic structure

$Gaia$ DR2 has provided an unprecedented wealth of information about the positions and motions of stars in our Galaxy, and has highlighted the degree of disequilibria in the disc. As we collect data over a wider area of the disc it becomes increasingly appealing to start analysing stellar actions and angles, which specifically label orbit space, instead of their current phase space location. Conceptually, while $\bar{x}$ and $\bar{v}$ tell us about the potential and local interactions, grouping in action puts together stars that have similar frequencies and hence similar responses to dynamical effects occurring over several orbits. Grouping in actions and angles refines this further to isolate stars which are travelling together through space and hence have shared histories. Mixing these coordinate systems can confuse the interpretation. For example, it has been suggested that by moving stars to their guiding radius, the Milky Way spiral structure is visible as ridge-like overdensities in the $Gaia$ data \citep{Khoperskov+19b}. However, in this work, we show that these features are in fact the known kinematic moving groups, both in the $L_z-ϕ$ and the $v_{\mathrm{R}}-v_ϕ$ planes. Using simulations we show how this distinction will become even more important as we move to a global view of the Milky Way. As an example, we show that the radial velocity wave seen in the Galactic disc in $Gaia$ and APOGEE should become stronger in the action-angle frame, and that it can be reproduced by transient spiral structure.

preprint2020arXiv

The Sixteenth Data Release of the Sloan Digital Sky Surveys: First Release from the APOGEE-2 Southern Survey and Full Release of eBOSS Spectra

This paper documents the sixteenth data release (DR16) from the Sloan Digital Sky Surveys; the fourth and penultimate from the fourth phase (SDSS-IV). This is the first release of data from the southern hemisphere survey of the Apache Point Observatory Galactic Evolution Experiment 2 (APOGEE-2); new data from APOGEE-2 North are also included. DR16 is also notable as the final data release for the main cosmological program of the Extended Baryon Oscillation Spectroscopic Survey (eBOSS), and all raw and reduced spectra from that project are released here. DR16 also includes all the data from the Time Domain Spectroscopic Survey (TDSS) and new data from the SPectroscopic IDentification of ERosita Survey (SPIDERS) programs, both of which were co-observed on eBOSS plates. DR16 has no new data from the Mapping Nearby Galaxies at Apache Point Observatory (MaNGA) survey (or the MaNGA Stellar Library "MaStar"). We also preview future SDSS-V operations (due to start in 2020), and summarize plans for the final SDSS-IV data release (DR17).

preprint2020arXiv

The Strength of the Dynamical Spiral Perturbation in the Galactic Disk

The mean Galactocentric radial velocities $\langle v_{R}\rangle(R,φ)$ of luminous red giant stars within the mid-plane of the Milky Way reveal a spiral signature, which could plausibly reflect the response to a non-axisymmetric perturbation of the gravitational potential in the Galactic disk. We apply a simple steady-state toy model of a logarithmic spiral to interpret these observations, and find a good qualitative and quantitative match. Presuming that the amplitude of the gravitational potential perturbation is proportionate to that in the disk's surface mass density, we estimate the surface mass density amplitude to be $Σ_{\rm max} (R_{\odot})\approx 5.5\,\rm M_{\odot}\,pc^{-2}$ at the solar radius when choosing a fixed pattern speed of $Ω_{\mathrm p}=12\,\rm km\,s^{-1}\,kpc^{-1}$. Combined with the local disk density, this implies a surface mass density contrast between the arm and inter-arm regions of approximately $\pm 10\%$ at the solar radius, with an increases towards larger radii. Our model constrains the pitch angle of the dynamical spiral arms to be approximately $12^{\circ}$.

preprint2016arXiv

A 14 $h^{-3}$ Gpc$^3$ study of cosmic homogeneity using BOSS DR12 quasar sample

The BOSS quasar sample is used to study cosmic homogeneity with a 3D survey in the redshift range $2.2<z<2.8$. We measure the count-in-sphere, $N(<\! r)$, i.e. the average number of objects around a given object, and its logarithmic derivative, the fractal correlation dimension, $D_2(r)$. For a homogeneous distribution $N(<\! r) \propto r^3$ and $D_2(r)=3$. Due to the uncertainty on tracer density evolution, 3D surveys can only probe homogeneity up to a redshift dependence, i.e. they probe so-called "spatial isotropy". Our data demonstrate spatial isotropy of the quasar distribution in the redshift range $2.2<z<2.8$ in a model-independent way, independent of any FLRW fiducial cosmology, resulting in $3-\langle D_2 \rangle < 1.7 \times 10^{-3}$ (2 $σ$) over the range $250<r<1200 \, h^{-1}$Mpc for the quasar distribution. If we assume that quasars do not have a bias much less than unity, this implies spatial isotropy of the matter distribution on large scales. Then, combining with the Copernican principle, we finally get homogeneity of the matter distribution on large scales. Alternatively, using a flat $Λ$CDM fiducial cosmology with CMB-derived parameters, and measuring the quasar bias relative to this $Λ$CDM model, our data provide a consistency check of the model, in terms of how homogeneous the Universe is on different scales. $D_2(r)$ is found to be compatible with our $Λ$CDM model on the whole $10<r<1200 \, h^{-1}$Mpc range. For the matter distribution we obtain $3-\langle D_2 \rangle < 5 \times 10^{-5}$ (2 $σ$) over the range $250<r<1200 \, h^{-1}$Mpc, consistent with homogeneity on large scales.

preprint2016arXiv

A Causal, Data-Driven Approach to Modeling the Kepler Data

Astronomical observations are affected by several kinds of noise, each with its own causal source; there is photon noise, stochastic source variability, and residuals coming from imperfect calibration of the detector or telescope. The precision of NASA Kepler photometry for exoplanet science---the most precise photometric measurements of stars ever made---appears to be limited by unknown or untracked variations in spacecraft pointing and temperature, and unmodeled stellar variability. Here we present the Causal Pixel Model (CPM) for Kepler data, a data-driven model intended to capture variability but preserve transit signals. The CPM works at the pixel level so that it can capture very fine-grained information about the variation of the spacecraft. The CPM predicts each target pixel value from a large number of pixels of other stars sharing the instrument variabilities while not containing any information on possible transits in the target star. In addition, we use the target star's future and past (auto-regression). By appropriately separating, for each data point, the data into training and test sets, we ensure that information about any transit will be perfectly isolated from the model. The method has four hyper-parameters (the number of predictor stars, the auto-regressive window size, and two L2-regularization amplitudes for model components), which we set by cross-validation. We determine a generic set of hyper-parameters that works well for most of the stars and apply the method to a corresponding set of target stars. We find that we can consistently outperform (for the purposes of exoplanet detection) the Kepler Pre-search Data Conditioning (PDC) method for exoplanet discovery.

preprint2016arXiv

Campaign 9 of the $K2$ Mission: Observational Parameters, Scientific Drivers, and Community Involvement for a Simultaneous Space- and Ground-based Microlensing Survey

$K2$'s Campaign 9 ($K2$C9) will conduct a $\sim$3.7 deg$^{2}$ survey toward the Galactic bulge from 7/April through 1/July of 2016 that will leverage the spatial separation between $K2$ and the Earth to facilitate measurement of the microlens parallax $π_{\rm E}$ for $\gtrsim$127 microlensing events. These will include several that are planetary in nature as well as many short-timescale microlensing events, which are potentially indicative of free-floating planets (FFPs). These satellite parallax measurements will in turn allow for the direct measurement of the masses of and distances to the lensing systems. In this white paper we provide an overview of the $K2$C9 space- and ground-based microlensing survey. Specifically, we detail the demographic questions that can be addressed by this program, including the frequency of FFPs and the Galactic distribution of exoplanets, the observational parameters of $K2$C9, and the array of resources dedicated to concurrent observations. Finally, we outline the avenues through which the larger community can become involved, and generally encourage participation in $K2$C9, which constitutes an important pathfinding mission and community exercise in anticipation of $WFIRST$.

preprint2016arXiv

Chaotic Dispersal of Tidal Debris

Several long, dynamically cold stellar streams have been observed around the Milky Way Galaxy, presumably formed from the tidal disruption of globular clusters. In integrable potentials---where all orbits are regular---tidal debris phase-mixes close to the orbit of the progenitor system. However, the Milky Way's dark matter halo is expected not to be fully integrable; an appreciable fraction of orbits will be chaotic. This paper examines the influence of chaos on the phase-space morphology of cold tidal streams. Streams even in weakly chaotic regions look very different from those in regular regions. We find that streams can be sensitive to chaos on a much shorter time-scale than any standard prediction (from the Lyapunov or frequency-diffusion times). For example, on a weakly chaotic orbit with a chaotic timescale predicted to be $>$1000 orbital periods ($>$1000 Gyr), the resulting stellar stream is, after just a few 10's of orbits, substantially more diffuse than any formed on a nearby but regular orbit. We find that the enhanced diffusion of the stream stars can be understood by looking at the variance in orbital frequencies of orbit ensembles centered around the parent (progenitor) orbit. Our results suggest that long, cold streams around our Galaxy must exist only on regular (or very nearly regular) orbits; they potentially provide a map of the regular regions of the Milky Way potential. This suggests a promising new direction for the use of tidal streams to constrain the distribution of dark matter around our Galaxy.

preprint2016arXiv

Chemical tagging can work: Identification of stellar phase-space structures purely by chemical-abundance similarity

Chemical tagging promises to use detailed abundance measurements to identify spatially separated stars that were in fact born together (in the same molecular cloud), long ago. This idea has not yielded much practical success, presumably because of the noise and incompleteness in chemical-abundance measurements. We have succeeded in substantially improving spectroscopic measurements with The Cannon, which has now delivered 15 individual abundances for ~100,000 stars observed as part of the APOGEE spectroscopic survey, with precisions around 0.04 dex. We test the chemical-tagging hypothesis by looking at clusters in abundance space and confirming that they are clustered in phase space. We identify (by the k-means algorithm) overdensities of stars in the 15-dimensional chemical-abundance space delivered by The Cannon, and plot the associated stars in phase space. We use only abundance-space information (no positional information) to identify stellar groups. We find that clusters in abundance space are indeed clusters in phase space. We recover some known phase-space clusters and find other interesting structures. This is the first-ever project to identify phase-space structures at survey-scale by blind search purely in abundance space; it verifies the precision of the abundance measurements delivered by The Cannon; the prospects for future data sets appear very good.

preprint2016arXiv

Constructing Polynomial Spectral Models for Stars

Stellar spectra depend on the stellar parameters and on dozens of photospheric elemental abundances. Simultaneous fitting of these $\mathcal{N}\sim 10-40$ model labels to observed spectra has been deemed unfeasible, because the number of ab initio spectral model grid calculations scales exponentially with $\mathcal{N}$. We suggest instead the construction of a polynomial spectral model (PSM) of order $\mathcal{O}$ for the model flux at each wavelength. Building this approximation requires a minimum of only ${\mathcal{N}+\mathcal{O}\choose\mathcal{O}}$ calculations: e.g. a quadratic spectral model ($\mathcal{O}=2$) to fit $\mathcal{N}=20$ labels simultaneously, can be constructed from as few as $231$ ab initio spectral model calculations; in practice, a somewhat larger number ($\sim 300-1000$) of randomly chosen models lead to a better performing PSM. Such a PSM can be a good approximation only over a portion of label space, which will vary case by case. Yet, taking the APOGEE survey as an example, a single quadratic PSM provides a remarkably good approximation to the exact ab initio spectral models across much of this survey: for random labels within that survey the PSM approximates the flux to within $10^{-3}$, and recovers the abundances to within $\sim 0.02$ dex rms of the exact models. This enormous speed-up enables the simultaneous many-label fitting of spectra with computationally expensive ab initio models for stellar spectra, such as non-LTE models. A PSM also enables the simultaneous fitting of observational parameters, such as the spectrum's continuum or line-spread function.

preprint2016arXiv

Do fast stellar centroiding methods saturate the Cramér-Rao lower bound?

One of the most demanding tasks in astronomical image processing---in terms of precision---is the centroiding of stars. Upcoming large surveys are going to take images of billions of point sources, including many faint stars, with short exposure times. Real-time estimation of the centroids of stars is crucial for real-time PSF estimation, and maximal precision is required for measurements of proper motion. The fundamental Cramér-Rao lower bound sets a limit on the root-mean-squared-error achievable by optimal estimators. In this work, we aim to compare the performance of various centroiding methods, in terms of saturating the bound, when they are applied to relatively low signal-to-noise ratio unsaturated stars assuming zero-mean constant Gaussian noise. In order to make this comparison, we present the ratio of the root-mean-squared-errors of these estimators to their corresponding Cramér-Rao bound as a function of the signal-to-noise ratio and the full-width at half-maximum of faint stars. We discuss two general circumstances in centroiding of faint stars: (i) when we have a good estimate of the PSF, (ii) when we do not know the PSF. In the case that we know the PSF, we show that a fast polynomial centroiding after smoothing the image by the PSF can be as efficient as the maximum-likelihood estimator at saturating the bound. In the case that we do not know the PSF, we demonstrate that although polynomial centroiding is not as optimal as PSF profile fitting, it comes very close to saturating the Cramér-Rao lower bound in a wide range of conditions. We also show that the moment-based method of center-of-light never comes close to saturating the bound, and thus it does not deliver reliable estimates of centroids.

preprint2016arXiv

Hydrogen Emission from the Ionized Gaseous Halos of Low Redshift Galaxies

Using a sample of nearly half million galaxies, intersected by over 7 million lines of sight from the Sloan Digital Sky Survey Data Release 12, we trace H$α$ + [N{\small II}] emission from a galactocentric projected radius, $r_p$, of 5 kpc to more than 100 kpc. The emission flux surface brightness is $\propto r_p^{-1.9 \pm 0.4}$. We obtain consistent results using only the H$α$ or [N{\small II}] flux. We measure a stronger signal for the bluer half of the target sample than for the redder half on small scales, $r_p <$ 20 kpc. We obtain a $3σ$ detection of H$α$ + [N{\small II}] emission in the 50 to 100 kpc $r_p$ bin. The mean emission flux within this bin is $(1.10 \pm 0.35) \times 10^{-20}$ erg cm$^{-2}$ s$^{-1}$ Å$^{-1}$, which corresponds to $1.87 \times 10^{-20}$ erg cm$^{-2}$ s$^{-1}$ arcsec$^{-2}$ or 0.0033 Rayleigh. This detection is 34 times fainter than a previous strict limit obtained using deep narrow-band imaging. The faintness of the signal demonstrates why it has been so difficult to trace recombination radiation out to large radii around galaxies. This signal, combined with published estimates of n$_{\rm H}$, lead us to estimate the temperature of the gas to be 12,000 K, consistent with independent empirical estimates based on metal ion absorption lines and expectations from numerical simulations.

preprint2016arXiv

State of the Field: Extreme Precision Radial Velocities

The Second Workshop on Extreme Precision Radial Velocities defined circa 2015 the state of the art Doppler precision and identified the critical path challenges for reaching 10 cm/s measurement precision. The presentations and discussion of key issues for instrumentation and data analysis and the workshop recommendations for achieving this precision are summarized here. Beginning with the HARPS spectrograph, technological advances for precision radial velocity measurements have focused on building extremely stable instruments. To reach still higher precision, future spectrometers will need to produce even higher fidelity spectra. This should be possible with improved environmental control, greater stability in the illumination of the spectrometer optics, better detectors, more precise wavelength calibration, and broader bandwidth spectra. Key data analysis challenges for the precision radial velocity community include distinguishing center of mass Keplerian motion from photospheric velocities, and the proper treatment of telluric contamination. Success here is coupled to the instrument design, but also requires the implementation of robust statistical and modeling techniques. Center of mass velocities produce Doppler shifts that affect every line identically, while photospheric velocities produce line profile asymmetries with wavelength and temporal dependencies that are different from Keplerian signals. Exoplanets are an important subfield of astronomy and there has been an impressive rate of discovery over the past two decades. Higher precision radial velocity measurements are required to serve as a discovery technique for potentially habitable worlds and to characterize detections from transit missions. The future of exoplanet science has very different trajectories depending on the precision that can ultimately be achieved with Doppler measurements.

preprint2016arXiv

The Cannon 2: A data-driven model of stellar spectra for detailed chemical abundance analyses

We have shown that data-driven models are effective for inferring physical attributes of stars (labels; Teff, logg, [M/H]) from spectra, even when the signal-to-noise ratio is low. Here we explore whether this is possible when the dimensionality of the label space is large (Teff, logg, and 15 abundances: C, N, O, Na, Mg, Al, Si, S, K, Ca, Ti, V, Mn, Fe, Ni) and the model is non-linear in its response to abundance and parameter changes. We adopt ideas from compressed sensing to limit overall model complexity while retaining model freedom. The model is trained with a set of 12,681 red-giant stars with high signal-to-noise spectroscopic observations and stellar parameters and abundances taken from the APOGEE Survey. We find that we can successfully train and use a model with 17 stellar labels. Validation shows that the model does a good job of inferring all 17 labels (typical abundance precision is 0.04 dex), even when we degrade the signal-to-noise by discarding ~50% of the observing time. The model dependencies make sense: the spectral derivatives with respect to abundances correlate with known atomic lines, and we identify elements belonging to atomic lines that were previously unknown. We recover (anti-)correlations in abundance labels for globular cluster stars, consistent with the literature. However we find the intrinsic spread in globular cluster abundances is 3--4 times smaller than previously reported. We deliver 17 labels with associated errors for 87,563 red giant stars, as well as open-source code to extend this work to other spectroscopic surveys.

preprint2016arXiv

The Panchromatic Hubble Andromeda Treasury XV. The BEAST: Bayesian Extinction and Stellar Tool

We present the Bayesian Extinction And Stellar Tool (BEAST), a probabilistic approach to modeling the dust extinguished photometric spectral energy distribution of an individual star while accounting for observational uncertainties common to large resolved star surveys. Given a set of photometric measurements and an observational uncertainty model, the BEAST infers the physical properties of the stellar source using stellar evolution and atmosphere models and constrains the line of sight extinction using a newly developed mixture model that encompasses the full range of dust extinction curves seen in the Local Group. The BEAST is specifically formulated for use with large multi-band surveys of resolved stellar populations. Our approach accounts for measurement uncertainties and any covariance between them due to stellar crowding (both systematic biases and uncertainties in the bias) and absolute flux calibration, thereby incorporating the full information content of the measurement. We illustrate the accuracy and precision possible with the BEAST using data from the Panchromatic Hubble Andromeda Treasury. While the BEAST has been developed for this survey, it can be easily applied to similar existing and planned resolved star surveys.

preprint2016arXiv

The population of long-period transiting exoplanets

The Kepler Mission has discovered thousands of exoplanets and revolutionized our understanding of their population. This large, homogeneous catalog of discoveries has enabled rigorous studies of the occurrence rate of exoplanets and planetary systems as a function of their physical properties. However, transit surveys like Kepler are most sensitive to planets with orbital periods much shorter than the orbital periods of Jupiter and Saturn, the most massive planets in our Solar System. To address this deficiency, we perform a fully automated search for long-period exoplanets with only one or two transits in the archival Kepler light curves. When applied to the $\sim 40,000$ brightest Sun-like target stars, this search produces 16 long-period exoplanet candidates. Of these candidates, 6 are novel discoveries and 5 are in systems with inner short-period transiting planets. Since our method involves no human intervention, we empirically characterize the detection efficiency of our search. Based on these results, we measure the average occurrence rate of exoplanets smaller than Jupiter with orbital periods in the range 2-25 years to be $2.0\pm0.7$ planets per Sun-like star.

preprint2015arXiv

A systematic search for transiting planets in the K2 data

Photometry of stars from the K2 extension of NASA's Kepler mission is afflicted by systematic effects caused by small (few-pixel) drifts in the telescope pointing and other spacecraft issues. We present a method for searching K2 light curves for evidence of exoplanets by simultaneously fitting for these systematics and the transit signals of interest. This method is more computationally expensive than standard search algorithms but we demonstrate that it can be efficiently implemented and used to discover transit signals. We apply this method to the full Campaign 1 dataset and report a list of 36 planet candidates transiting 31 stars, along with an analysis of the pipeline performance and detection efficiency based on artificial signal injections and recoveries. For all planet candidates, we present posterior distributions on the properties of each system based strictly on the transit observables.

preprint2015arXiv

Action-space clustering of tidal streams to infer the Galactic potential

We present a new method for constraining the Milky Way halo gravitational potential by simultaneously fitting multiple tidal streams. This method requires full three-dimensional positions and velocities for all stars to be fit, but does not require identification of any specific stream or determination of stream membership for any star. We exploit the principle that the action distribution of stream stars is most clustered when the potential used to calculate the actions is closest to the true potential. Clustering is quantified with the Kullback-Leibler Divergence (KLD), which also provides conditional uncertainties for our parameter estimates. We show, for toy Gaia-like data in a spherical isochrone potential, that maximizing the KLD of the action distribution relative to a smoother distribution recovers the true values of the potential parameters. The precision depends on the observational errors and the number of streams in the sample; using KIII giants as tracers, we measure the enclosed mass at the average radius of the sample stars accurate to 3% and precise to 20-40%. Recovery of the scale radius is precise to 25%, and is biased 50% high by the small galactocentric distance range of stars in our mock sample (1-25 kpc, or about three scale radii, with mean 6.5 kpc). About 15 streams, with at least 100 stars per stream, are needed to obtain upper and lower bounds on the enclosed mass and scale radius when observational errors are taken into account; 20-25 streams are required to stabilize the size of the confidence interval. If radial velocities are provided for stars out to 100 kpc (10 scale radii), all parameters can be determined with 10% accuracy and 20% precision (1.3% accuracy in the case of the enclosed mass), underlining the need for ground-based spectroscopic follow-up to complete the radial velocity catalog for faint halo stars observed by Gaia.

preprint2015arXiv

Constructing A Flexible Likelihood Function For Spectroscopic Inference

We present a modular, extensible likelihood framework for spectroscopic inference based on synthetic model spectra. The subtraction of an imperfect model from a continuously sampled spectrum introduces covariance between adjacent datapoints (pixels) into the residual spectrum. For the high signal-to-noise data with large spectral range that is commonly employed in stellar astrophysics, that covariant structure can lead to dramatically underestimated parameter uncertainties (and, in some cases, biases). We construct a likelihood function that accounts for the structure of the covariance matrix, utilizing the machinery of Gaussian process kernels. This framework specifically address the common problem of mismatches in model spectral line strengths (with respect to data) due to intrinsic model imperfections (e.g., in the atomic/molecular databases or opacity prescriptions) by developing a novel local covariance kernel formalism that identifies and self-consistently downweights pathological spectral line "outliers." By fitting many spectra in a hierarchical manner, these local kernels provide a mechanism to learn about and build data-driven corrections to synthetic spectral libraries. An open-source software implementation of this approach is available at http://iancze.github.io/Starfish, including a sophisticated probabilistic scheme for spectral interpolation when using model libraries that are sparsely sampled in the stellar parameters. We demonstrate some salient features of the framework by fitting the high resolution $V$-band spectrum of WASP-14, an F5 dwarf with a transiting exoplanet, and the moderate resolution $K$-band spectrum of Gliese 51, an M5 field dwarf.

preprint2015arXiv

Fast Direct Methods for Gaussian Processes

A number of problems in probability and statistics can be addressed using the multivariate normal (Gaussian) distribution. In the one-dimensional case, computing the probability for a given mean and variance simply requires the evaluation of the corresponding Gaussian density. In the $n$-dimensional setting, however, it requires the inversion of an $n \times n$ covariance matrix, $C$, as well as the evaluation of its determinant, $\det(C)$. In many cases, such as regression using Gaussian processes, the covariance matrix is of the form $C = σ^2 I + K$, where $K$ is computed using a specified covariance kernel which depends on the data and additional parameters (hyperparameters). The matrix $C$ is typically dense, causing standard direct methods for inversion and determinant evaluation to require $\mathcal O(n^3)$ work. This cost is prohibitive for large-scale modeling. Here, we show that for the most commonly used covariance functions, the matrix $C$ can be hierarchically factored into a product of block low-rank updates of the identity matrix, yielding an $\mathcal O (n\log^2 n) $ algorithm for inversion. More importantly, we show that this factorization enables the evaluation of the determinant $\det(C)$, permitting the direct calculation of probabilities in high dimensions under fairly broad assumptions on the kernel defining $K$. Our fast algorithm brings many problems in marginalization and the adaptation of hyperparameters within practical reach using a single CPU core. The combination of nearly optimal scaling in terms of problem size with high-performance computing resources will permit the modeling of previously intractable problems. We illustrate the performance of the scheme on standard covariance kernels.

preprint2015arXiv

Finding, characterizing and classifying variable sources in multi-epoch sky surveys: QSOs and RR Lyrae in PS1 3$π$ data

In area and depth, the Pan-STARRS1 (PS1) 3$π$ survey is unique among many-epoch, multi-band surveys and has enormous potential for all-sky identification of variable sources. PS1 has observed the sky typically seven times in each of its five bands ($grizy$) over 3.5 years, but unlike SDSS not simultaneously across the bands. Here we develop a new approach for quantifying statistical properties of non-simultaneous, sparse, multi-color lightcurves through light-curve structure functions, effectively turning PS1 into a $\sim 35$-epoch survey. We use this approach to estimate variability amplitudes and timescales $(ω_r, τ)$ for all point-sources brighter than $r_{\mathrm{P1}}=21.5$ mag in the survey. With PS1 data on SDSS Stripe 82 as ``ground truth", we use a Random Forest Classifier to identify QSOs and RR Lyrae based on their variability and their mean PS1 and WISE colors. We find that, aside from the Galactic plane, QSO and RR Lyrae samples of purity $\sim$75\% and completeness $\sim$92\% can be selected. On this basis we have identified a sample of $\sim 1,000,000$ QSO candidates, as well as an unprecedentedly large and deep sample of $\sim$150,000 RR Lyrae candidates with distances from $\sim$10 kpc to $\sim$120 kpc. Within the Draco dwarf spheroidal, we demonstrate a distance precision of 6\% for RR Lyrae candidates. We provide a catalog of all likely variable point sources and likely QSOs in PS1, a total of $25.8\times 10^6$ sources.

preprint2015arXiv

Globular Cluster Streams as Galactic High-Precision Scales - The Poster Child Palomar 5

Using the example of the tidal stream of the Milky Way globular cluster Palomar 5 (Pal 5), we demonstrate how observational data on streams can be efficiently reduced in dimensionality and modeled in a Bayesian framework. Our approach combines detection of stream overdensities by a Difference-of-Gaussians process with fast streakline models, a continuous likelihood function built from these models, and inference with MCMC. By generating $\approx10^7$ model streams, we show that the geometry of the Pal 5 debris yields powerful constraints on the solar position and motion, the Milky Way and Pal 5 itself. All 10 model parameters were allowed to vary over large ranges without additional prior information. Using only SDSS data and a few radial velocities from the literature, we find that the distance of the Sun from the Galactic Center is $8.30\pm0.25$ kpc, and the transverse velocity is $253\pm16$ km/s. Both estimates are in excellent agreement with independent measurements of these quantities. Assuming a standard disk and bulge model, we determine the Galactic mass within Pal 5's apogalactic radius of 19 kpc to be $(2.1\pm0.4)\times10^{11}$ M$_\odot$. Moreover, we find the potential of the dark halo with a flattening of $q_z = 0.95^{+0.16}_{-0.12}$ to be essentially spherical within the radial range that is effectively probed by Pal 5. We also determine Pal 5's mass, distance and proper motion independently from other methods, which enables us to perform vital cross-checks. We conclude that with more observational data and by using additional prior information, the precision of this method can be significantly increased.

preprint2015arXiv

Spectroscopic determination of masses (and implied ages) for red giants

The mass of a star is arguably its most fundamental parameter. For red giant stars, tracers luminous enough to be observed across the Galaxy, mass implies a stellar evolution age. It has proven to be extremely difficult to infer ages and masses directly from red giant spectra using existing methods. From the KEPLER and APOGEE surveys, samples of several thousand stars exist with high-quality spectra and asteroseismic masses. Here we show that from these data we can build a data-driven spectral model using The Cannon, which can determine stellar masses to $\sim$ 0.07 dex from APOGEE DR12 spectra of red giants; these imply age estimates accurate to $\sim$ 0.2 dex (40 percent). We show that The Cannon constrains these ages foremost from spectral regions with CN absorption lines, elements whose surface abundances reflect mass-dependent dredge-up. We deliver an unprecedented catalog of 80,000 giants (including 20,000 red-clump stars) with mass and age estimates, spanning the entire disk (from the Galactic center to R $\sim$ 20 kpc). We show that the age information in the spectra is not simply a corollary of the birth-material abundances [Fe/H] and [$α$/Fe], and that even within a mono-abundance population of stars, there are age variations that vary sensibly with Galactic position. Such stellar age constraints across the Milky Way open up new avenues in Galactic archeology.

preprint2015arXiv

Stellar and Planetary Properties of K2 Campaign 1 Candidates and Validation of 17 Planets, Including a Planet Receiving Earth-like Insolation

The extended Kepler mission, K2, is now providing photometry of new fields every three months in a search for transiting planets. In a recent study, Foreman-Mackey and collaborators presented a list of 36 planet candidates orbiting 31 stars in K2 Campaign 1. In this contribution, we present stellar and planetary properties for all systems. We combine ground-based seeing-limited survey data and adaptive optics imaging with an automated transit analysis scheme to validate 21 candidates as planets, 17 for the first time, and identify 6 candidates as likely false positives. Of particular interest is K2-18 (EPIC 201912552), a bright (K=8.9) M2.8 dwarf hosting a 2.23 \pm 0.25 R_Earth planet with T_eq = 272 \pm 15 K and an orbital period of 33 days. We also present two new open-source software packages which enable this analysis. The first, isochrones, is a flexible tool for fitting theoretical stellar models to observational data to determine stellar properties using a nested sampling scheme to capture the multimodal nature of the posterior distributions of the physical parameters of stars that may plausibly be evolved. The second is vespa, a new general-purpose procedure to calculate false positive probabilities and statistically validate transiting exoplanets.

preprint2015arXiv

The Cannon: A data-driven approach to stellar label determination

New spectroscopic surveys offer the promise of consistent stellar parameters and abundances ('stellar labels') for hundreds of thousands of stars in the Milky Way: this poses a formidable spectral modeling challenge. In many cases, there is a sub-set of reference objects for which the stellar labels are known with high(er) fidelity. We take advantage of this with The Cannon, a new data-driven approach for determining stellar labels from spectroscopic data. The Cannon learns from the 'known' labels of reference stars how the continuum-normalized spectra depend on these labels by fitting a flexible model at each wavelength; then, The Cannon uses this model to derive labels for the remaining survey stars. We illustrate The Cannon by training the model on only 542 stars in 19 clusters as reference objects, with Teff, log g and [Fe/H] as the labels, and then applying it to the spectra of 56,000 stars from APOGEE DR10. The Cannon is very accurate. Its stellar labels compare well to the stars for which APOGEE pipeline (ASPCAP) labels are provided in DR10, with rms differences that are basically identical to the stated ASPCAP uncertainties. Beyond the reference labels, The Cannon makes no use of stellar models nor any line-list, but needs a set of reference objects that span label-space. The Cannon performs well at lower signal-to-noise, as it delivers comparably good labels even at one ninth the APOGEE observing time. We discuss the limitations of The Cannon and its future potential, particularly, to bring different spectroscopic surveys onto a consistent scale of stellar labels.

preprint2015arXiv

The Eleventh and Twelfth Data Releases of the Sloan Digital Sky Survey: Final Data from SDSS-III

The third generation of the Sloan Digital Sky Survey (SDSS-III) took data from 2008 to 2014 using the original SDSS wide-field imager, the original and an upgraded multi-object fiber-fed optical spectrograph, a new near-infrared high-resolution spectrograph, and a novel optical interferometer. All the data from SDSS-III are now made public. In particular, this paper describes Data Release 11 (DR11) including all data acquired through 2013 July, and Data Release 12 (DR12) adding data acquired through 2014 July (including all data included in previous data releases), marking the end of SDSS-III observing. Relative to our previous public release (DR10), DR12 adds one million new spectra of galaxies and quasars from the Baryon Oscillation Spectroscopic Survey (BOSS) over an additional 3000 sq. deg of sky, more than triples the number of H-band spectra of stars as part of the Apache Point Observatory (APO) Galactic Evolution Experiment (APOGEE), and includes repeated accurate radial velocity measurements of 5500 stars from the Multi-Object APO Radial Velocity Exoplanet Large-area Survey (MARVELS). The APOGEE outputs now include measured abundances of 15 different elements for each star. In total, SDSS-III added 2350 sq. deg of ugriz imaging; 155,520 spectra of 138,099 stars as part of the Sloan Exploration of Galactic Understanding and Evolution 2 (SEGUE-2) survey; 2,497,484 BOSS spectra of 1,372,737 galaxies, 294,512 quasars, and 247,216 stars over 9376 sq. deg; 618,080 APOGEE spectra of 156,593 stars; and 197,040 MARVELS spectra of 5,513 stars. Since its first light in 1998, SDSS has imaged over 1/3 of the Celestial sphere in five bands and obtained over five million astronomical spectra.

preprint2015arXiv

The High-Mass Stellar Initial Mass Function in M31 Clusters

We have undertaken the largest systematic study of the high-mass stellar initial mass function (IMF) to date using the optical color-magnitude diagrams (CMDs) of 85 resolved, young (4 Myr < t < 25 Myr), intermediate mass star clusters (10^3-10^4 Msun), observed as part of the Panchromatic Hubble Andromeda Treasury (PHAT) program. We fit each cluster's CMD to measure its mass function (MF) slope for stars >2 Msun. For the ensemble of clusters, the distribution of stellar MF slopes is best described by $Γ=+1.45^{+0.03}_{-0.06}$ with a very small intrinsic scatter. The data also imply no significant dependencies of the MF slope on cluster age, mass, and size, providing direct observational evidence that the measured MF represents the IMF. This analysis implies that the high-mass IMF slope in M31 clusters is universal with a slope ($Γ=+1.45^{+0.03}_{-0.06}$) that is steeper than the canonical Kroupa (+1.30) and Salpeter (+1.35) values. Using our inference model on select Milky Way (MW) and LMC high-mass IMF studies from the literature, we find $Γ_{\rm MW} \sim+1.15\pm0.1$ and $Γ_{\rm LMC} \sim+1.3\pm0.1$, both with intrinsic scatter of ~0.3-0.4 dex. Thus, while the high-mass IMF in the Local Group may be universal, systematics in literature IMF studies preclude any definitive conclusions; homogenous investigations of the high-mass IMF in the local universe are needed to overcome this limitation. Consequently, the present study represents the most robust measurement of the high-mass IMF slope to date. We have grafted the M31 high-mass IMF slope onto widely used sub-solar mass Kroupa and Chabrier IMFs and show that commonly used UV- and Halpha-based star formation rates should be increased by a factor of ~1.3-1.5 and the number of stars with masses >8 Msun are ~25% fewer than expected for a Salpeter/Kroupa IMF. [abridged]

preprint2015arXiv

The Panchromatic Hubble Andromeda Treasury VIII: A Wide-Area, High-Resolution Map of Dust Extinction in M31

We map the distribution of dust in M31 at 25pc resolution, using stellar photometry from the Panchromatic Hubble Andromeda Treasury. We develop a new mapping technique that models the NIR color-magnitude diagram (CMD) of red giant branch (RGB) stars. The model CMDs combine an unreddened foreground of RGB stars with a reddened background population viewed through a log-normal column density distribution of dust. Fits to the model constrain the median extinction, the width of the extinction distribution, and the fraction of reddened stars. The resulting extinction map has >4 times better resolution than maps of dust emission, while providing a more direct measurement of the dust column. There is superb morphological agreement between the new map and maps of the extinction inferred from dust emission by Draine et al. 2014. However, the widely-used Draine & Li (2007) dust models overpredict the observed extinction by a factor of ~2.5, suggesting that M31's true dust mass is lower and that dust grains are significantly more emissive than assumed in Draine et al. (2014). The discrepancy we identify is consistent with similar findings in the Milky Way by the Planck Collaboration (2015), but has a more complex dependence on parameters from the Draine & Li (2007) dust models. We also show that the discrepancy with the Draine et al. (2014) map is lowest where the interstellar radiation field has a harder spectrum than average. We discuss possible improvements to the CMD dust mapping technique, and explore further applications.

preprint2014arXiv

10 Simple Rules for the Care and Feeding of Scientific Data

This article offers a short guide to the steps scientists can take to ensure that their data and associated analyses continue to be of value and to be recognized. In just the past few years, hundreds of scholarly papers and reports have been written on questions of data sharing, data provenance, research reproducibility, licensing, attribution, privacy, and more, but our goal here is not to review that literature. Instead, we present a short guide intended for researchers who want to know why it is important to "care for and feed" data, with some practical advice on how to do that.

preprint2014arXiv

Exoplanet population inference and the abundance of Earth analogs from noisy, incomplete catalogs

No true extrasolar Earth analog is known. Hundreds of planets have been found around Sun-like stars that are either Earth-sized but on shorter periods, or else on year-long orbits but somewhat larger. Under strong assumptions, exoplanet catalogs have been used to make an extrapolated estimate of the rate at which Sun-like stars host Earth analogs. These studies are complicated by the fact that every catalog is censored by non-trivial selection effects and detection efficiencies, and every property (period, radius, etc.) is measured noisily. Here we present a general hierarchical probabilistic framework for making justified inferences about the population of exoplanets, taking into account survey completeness and, for the first time, observational uncertainties. We are able to make fewer assumptions about the distribution than previous studies; we only require that the occurrence rate density be a smooth function of period and radius (employing a Gaussian process). By applying our method to synthetic catalogs, we demonstrate that it produces more accurate estimates of the whole population than standard procedures based on weighting by inverse detection efficiency. We apply the method to an existing catalog of small planet candidates around G dwarf stars (Petigura et al. 2013). We confirm a previous result that the radius distribution changes slope near Earth's radius. We find that the rate density of Earth analogs is about 0.02 (per star per natural logarithmic bin in period and radius) with large uncertainty. This number is much smaller than previous estimates made with the same data but stronger assumptions.

preprint2014arXiv

Hierarchical probabilistic inference of cosmic shear

Point estimators for the shearing of galaxy images induced by gravitational lensing involve a complex inverse problem in the presence of noise, pixelization, and model uncertainties. We present a probabilistic forward modeling approach to gravitational lensing inference that has the potential to mitigate the biased inferences in most common point estimators and is practical for upcoming lensing surveys. The first part of our statistical framework requires specification of a likelihood function for the pixel data in an imaging survey given parameterized models for the galaxies in the images. We derive the lensing shear posterior by marginalizing over all intrinsic galaxy properties that contribute to the pixel data (i.e., not limited to galaxy ellipticities) and learn the distributions for the intrinsic galaxy properties via hierarchical inference with a suitably flexible conditional probabilitiy distribution specification. We use importance sampling to separate the modeling of small imaging areas from the global shear inference, thereby rendering our algorithm computationally tractable for large surveys. With simple numerical examples we demonstrate the improvements in accuracy from our importance sampling approach, as well as the significance of the conditional distribution specification for the intrinsic galaxy properties when the data are generated from an unknown number of distinct galaxy populations with different morphological characteristics.

preprint2014arXiv

IGM Constraints from the SDSS-III/BOSS DR9 Ly-alpha Forest Flux Probability Distribution Function

The Ly$α$ forest transmission probability distribution function (PDF) is an established probe of the intergalactic medium (IGM) astrophysics, especially the temperature-density relationship of the IGM. We measure the transmission PDF from 3393 Baryon Oscillations Spectroscopic Survey (BOSS) quasars from SDSS Data Release 9, and compare with mock spectra that include careful modeling of the noise, continuum, and astrophysical uncertainties. The BOSS transmission PDFs, measured at $\langle z \rangle = [2.3,2.6,3.0]$, are compared with PDFs created from mock spectra drawn from a suite of hydrodynamical simulations that sample the IGM temperature-density relationship, $γ$, and temperature at mean-density, $T_0$, where $T(Δ) = T_0 Δ^{γ-1}$. We find that a significant population of partial Lyman-limit systems with a column-density distribution slope of $β_\mathrm{pLLS} \sim -2$ are required to explain the data at the low-transmission end of transmission PDF, while uncertainties in the mean Ly$α$ forest transmission affect the high-transmission end. After modelling the LLSs and marginalizing over mean-transmission uncertainties, we find that $γ=1.6$ best describes the data over our entire redshift range, although constraints on $T_0$ are affected by systematic uncertainties. Within our model framework, isothermal or inverted temperature-density relationships ($γ\leq 1$) are disfavored at a significance of over 4$σ$, although this could be somewhat weakened by cosmological and astrophysical uncertainties that we did not model.

preprint2014arXiv

Inferring the gravitational potential of the Milky Way with a few precisely measured stars

The dark matter halo of the Milky Way is expected to be triaxial and filled with substructure. It is hoped that streams or shells of stars produced by tidal disruption of stellar systems will provide precise measures of the gravitational potential to test these predictions. We develop a method for inferring the Galactic potential with tidal streams based on the idea that the stream stars were once close in phase space. Our method can flexibly adapt to any form for the Galactic potential: it works in phase-space rather than action-space and hence relies neither on our ability to derive actions nor on the integrability of the potential. Our model is probabilistic, with a likelihood function and priors on the parameters. The method can properly account for finite observational uncertainties and missing data dimensions. We test our method on synthetic datasets generated from N-body simulations of satellite disruption in a static, multi-component Milky Way including a triaxial dark matter halo with observational uncertainties chosen to mimic current and near-future surveys of various stars. We find that with just four well-measured stream stars, we can infer properties of a triaxial potential with precisions of order 5-7 percent. Without proper motions we obtain 15 percent constraints on potential parameters and precisions around 25 percent for recovering missing phase-space coordinates. These results are encouraging for the eventual goal of using flexible, time-dependent potential models combined with larger data sets to unravel the detailed shape of the dark matter distribution around the Milky Way.

preprint2014arXiv

Milky Way Mass and Potential Recovery Using Tidal Streams in a Realistic Halo

We present a new method for determining the Galactic gravitational potential based on forward modeling of tidal stellar streams. We use this method to test the performance of smooth and static analytic potentials in representing realistic dark matter halos, which have substructure and are continually evolving by accretion. Our FAST-FORWARD method uses a Markov Chain Monte Carlo algorithm to compare, in 6D phase space, an "observed" stream to models created in trial analytic potentials. We analyze a large sample of streams evolved in the Via Lactea II (VL2) simulation, which represents a realistic Galactic halo potential. The recovered potential parameters are in agreement with the best fit to the global, present-day VL2 potential. However, merely assuming an analytic potential limits the dark matter halo mass measurement to an accuracy of 5 to 20%, depending on the choice of analytic parametrization. Collectively, mass estimates using streams from our sample reach this fundamental limit, but individually they can be highly biased. Individual streams can both under- and overestimate the mass, and the bias is progressively worse for those with smaller perigalacticons, motivating the search for tidal streams at galactocentric distances larger than 70 kpc. We estimate that the assumption of a static and smooth dark matter potential in modeling of the GD-1 and Pal5-like streams introduces an error of up to 50% in the Milky Way mass estimates.

preprint2014arXiv

S4: A Spatial-Spectral model for Speckle Suppression

High dynamic-range imagers aim to block out or null light from a very bright primary star to make it possible to detect and measure far fainter companions; in real systems a small fraction of the primary light is scattered, diffracted, and unocculted. We introduce S4, a flexible data-driven model for the unocculted (and highly speckled) light in the P1640 spectroscopic coronograph. The model uses Principal Components Analysis (PCA) to capture the spatial structure and wavelength dependence of the speckles but not the signal produced by any companion. Consequently, the residual typically includes the companion signal. The companion can thus be found by filtering this error signal with a fixed companion model. The approach is sensitive to companions that are of order a percent of the brightness of the speckles, or up to $10^{-7}$ times the brightness of the primary star. This outperforms existing methods by a factor of 2-3 and is close to the shot-noise physical limit.

preprint2014arXiv

The nature of massive black hole binary candidates: II. Spectral energy distribution atlas

Recoiling supermassive black holes (SMBHs) are considered one plausible physical mechanism to explain high velocity shifts between narrow and broad emission lines sometimes observed in quasar spectra. If the sphere of influence of the recoiling SMBH is such that only the accretion disc is bound, the dusty torus would be left behind, hence the SED should then present distinctive features (i.e. a mid-infrared deficit). Here we present results from fitting the Spectral Energy Distributions (SEDs) of 32 Type-1 AGN with high velocity shifts between broad and narrow lines. The aim is to find peculiar properties in the multi-wavelength SEDs of such objects by comparing their physical parameters (torus and disc luminosity, intrinsic reddening, and size of the 12$μ$m emitter) with those estimated from a control sample of $\sim1000$ \emph{typical} quasars selected from the Sloan Digital Sky Survey in the same redshift range. We find that all sources, with the possible exception of J1154+0134, analysed here present a significant amount of 12~$μ$m emission. This is in contrast with a scenario of a SMBH displaced from the center of the galaxy, as expected for an undergoing recoil event.

preprint2014arXiv

The Probabilities of Orbital-Companion Models for Stellar Radial Velocity Data

The fully marginalized likelihood, or Bayesian evidence, is of great importance in probabilistic data analysis, because it is involved in calculating the posterior probability of a model or re-weighting a mixture of models conditioned on data. It is, however, extremely challenging to compute. This paper presents a geometric-path Monte Carlo method, inspired by multi-canonical Monte Carlo to evaluate the fully marginalized likelihood. We show that the algorithm is very fast and easy to implement and produces a justified uncertainty estimate on the fully marginalized likelihood. The algorithm performs efficiently on a trial problem and multi-companion model fitting for radial velocity data. For the trial problem, the algorithm returns the correct fully marginalized likelihood, and the estimated uncertainty is also consistent with the standard deviation of results from multiple runs. We apply the algorithm to the problem of fitting radial velocity data from HIP 88048 ($ν$ Oph) and Gliese 581. We evaluate the fully marginalized likelihood of 1, 2, 3, and 4-companion models given data from HIP 88048 and various choices of prior distributions. We consider prior distributions with three different minimum radial velocity amplitude $K_{\mathrm{min}}$. Under all three priors, the 2-companion model has the largest marginalized likelihood, but the detailed values depend strongly on $K_{\mathrm{min}}$. We also evaluate the fully marginalized likelihood of 3, 4, 5, and 6-planet model given data from Gliese 581 and find that the fully marginalized likelihood of the 5-planet model is too close to that of the 6-planet model for us to confidently decide between them.

preprint2014arXiv

The Tenth Data Release of the Sloan Digital Sky Survey: First Spectroscopic Data from the SDSS-III Apache Point Observatory Galactic Evolution Experiment

The Sloan Digital Sky Survey (SDSS) has been in operation since 2000 April. This paper presents the tenth public data release (DR10) from its current incarnation, SDSS-III. This data release includes the first spectroscopic data from the Apache Point Observatory Galaxy Evolution Experiment (APOGEE), along with spectroscopic data from the Baryon Oscillation Spectroscopic Survey (BOSS) taken through 2012 July. The APOGEE instrument is a near-infrared R~22,500 300-fiber spectrograph covering 1.514--1.696 microns. The APOGEE survey is studying the chemical abundances and radial velocities of roughly 100,000 red giant star candidates in the bulge, bar, disk, and halo of the Milky Way. DR10 includes 178,397 spectra of 57,454 stars, each typically observed three or more times, from APOGEE. Derived quantities from these spectra (radial velocities, effective temperatures, surface gravities, and metallicities) are also included.DR10 also roughly doubles the number of BOSS spectra over those included in the ninth data release. DR10 includes a total of 1,507,954 BOSS spectra, comprising 927,844 galaxy spectra; 182,009 quasar spectra; and 159,327 stellar spectra, selected over 6373.2 square degrees.

preprint2014arXiv

Towards building a Crowd-Sourced Sky Map

We describe a system that builds a high dynamic-range and wide-angle image of the night sky by combining a large set of input images. The method makes use of pixel-rank information in the individual input images to improve a "consensus" pixel rank in the combined image. Because it only makes use of ranks and the complexity of the algorithm is linear in the number of images, the method is useful for large sets of uncalibrated images that might have undergone unknown non-linear tone mapping transformations for visualization or aesthetic reasons. We apply the method to images of the night sky (of unknown provenance) discovered on the Web. The method permits discovery of astronomical objects or features that are not visible in any of the input images taken individually. More importantly, however, it permits scientific exploitation of a huge source of astronomical images that would not be available to astronomical research without our automatic system.

preprint2013arXiv

A New Approach to Identifying the Most Powerful Gravitational Lensing Telescopes

The best gravitational lenses for detecting distant galaxies are those with the largest mass concentrations and the most advantageous configurations of that mass along the line of sight. Our new method for finding such gravitational telescopes uses optical data to identify projected concentrations of luminous red galaxies (LRGs). LRGs are biased tracers of the underlying mass distribution, so lines of sight with the highest total luminosity in LRGs are likely to contain the largest total mass. We apply this selection technique to the Sloan Digital Sky Survey and identify the 200 fields with the highest total LRG luminosities projected within a 3.5' radius over the redshift range 0.1 < z < 0.7. The redshift and angular distributions of LRGs in these fields trace the concentrations of non-LRG galaxies. These fields are diverse; 22.5% contain one known galaxy cluster and 56.0% contain multiple known clusters previously identified in the literature. Thus, our results confirm that these LRGs trace massive structures and that our selection technique identifies fields with large total masses. These fields contain 2-3 times higher total LRG luminosities than most known strong-lensing clusters and will be among the best gravitational lensing fields for the purpose of detecting the highest redshift galaxies.

preprint2013arXiv

Baryon Acoustic Oscillations in the Ly-α forest of BOSS quasars

We report a detection of the baryon acoustic oscillation (BAO) feature in the three-dimensional correlation function of the transmitted flux fraction in the \Lya forest of high-redshift quasars. The study uses 48,640 quasars in the redshift range $2.1\le z \le 3.5$ from the Baryon Oscillation Spectroscopic Survey (BOSS) of the third generation of the Sloan Digital Sky Survey (SDSS-III). At a mean redshift $z=2.3$, we measure the monopole and quadrupole components of the correlation function for separations in the range $20\hMpc<r<200\hMpc$. A peak in the correlation function is seen at a separation equal to $(1.01\pm0.03)$ times the distance expected for the BAO peak within a concordance $Λ$CDM cosmology. This first detection of the BAO peak at high redshift, when the universe was strongly matter dominated, results in constraints on the angular diameter distance $\da$ and the expansion rate $H$ at $z=2.3$ that, combined with priors on $H_0$ and the baryon density, require the existence of dark energy. Combined with constraints derived from Cosmic Microwave Background (CMB) observations, this result implies $H(z=2.3)=(224\pm8){\rm km\,s^{-1}Mpc^{-1}}$, indicating that the time derivative of the cosmological scale parameter $\dot{a}=H(z=2.3)/(1+z)$ is significantly greater than that measured with BAO at $z\sim0.5$. This demonstrates that the expansion was decelerating in the range $0.7<z<2.3$, as expected from the matter domination during this epoch. Combined with measurements of $H_0$, one sees the pattern of deceleration followed by acceleration characteristic of a dark-energy dominated universe.

preprint2013arXiv

emcee: The MCMC Hammer

We introduce a stable, well tested Python implementation of the affine-invariant ensemble sampler for Markov chain Monte Carlo (MCMC) proposed by Goodman & Weare (2010). The code is open source and has already been used in several published projects in the astrophysics literature. The algorithm behind emcee has several advantages over traditional MCMC sampling methods and it has excellent performance as measured by the autocorrelation time (or function calls per independent sample). One major advantage of the algorithm is that it requires hand-tuning of only 1 or 2 parameters compared to $\sim N^2$ for a traditional algorithm in an N-dimensional parameter space. In this document, we describe the algorithm and the details of our implementation and API. Exploiting the parallelism of the ensemble method, emcee permits any user to take advantage of multiple CPU cores without extra effort. The code is available online at http://dan.iel.fm/emcee under the MIT License.

preprint2013arXiv

Maximizing Kepler science return per telemetered pixel: Detailed models of the focal plane in the two-wheel era

Kepler's immense photometric precision to date was maintained through satellite stability and precise pointing. In this white paper, we argue that image modeling--fitting the Kepler-downlinked raw pixel data--can vastly improve the precision of Kepler in pointing-degraded two-wheel mode. We argue that a non-trivial modeling effort may permit continuance of photometry at 10-ppm-level precision. We demonstrate some baby steps towards precise models in both data-driven (flexible) and physics-driven (interpretably parameterized) modes. We demonstrate that the expected drift or jitter in positions in the two-weel era will help with constraining calibration parameters. In particular, we show that we can infer the device flat-field at higher than pixel resolution; that is, we can infer pixel-to-pixel variations in intra-pixel sensitivity. These results are relevant to almost any scientific goal for the repurposed mission; image modeling ought to be a part of any two-wheel repurpose for the satellite. We make other recommendations for Kepler operations, but fundamentally advocate that the project stick with its core mission of finding and characterizing Earth analogs. [abridged]

preprint2013arXiv

Maximizing Kepler science return per telemetered pixel: Searching the habitable zones of the brightest stars

In today's mailing, Hogg et al. propose image modeling techniques to maintain 10-ppm-level precision photometry in Kepler data with only two working reaction wheels. While these results are relevant to many scientific goals for the repurposed mission, all modeling efforts so far have used a toy model of the Kepler telescope. Because the two-wheel performance of Kepler remains to be determined, we advocate for the consideration of an alternate strategy for a >1 year program that maximizes the science return from the "low-torque" fields across the ecliptic plane. Assuming we can reach the precision of the original Kepler mission, we expect to detect 800 new planet candidates in the first year of such a mission. Our proposed strategy has benefits for transit timing variation and transit duration variation studies, especially when considered in concert with the future TESS mission. We also expect to help address the first key science goal of Kepler: the frequency of planets in the habitable zone as a function of spectral type.

preprint2013arXiv

Probabilistic Catalogs for Crowded Stellar Fields

We present and implement a probabilistic (Bayesian) method for producing catalogs from images of stellar fields. The method is capable of inferring the number of sources N in the image and can also handle the challenges introduced by noise, overlapping sources, and an unknown point spread function (PSF). The luminosity function of the stars can also be inferred even when the precise luminosity of each star is uncertain, via the use of a hierarchical Bayesian model. The computational feasibility of the method is demonstrated on two simulated images with different numbers of stars. We find that our method successfully recovers the input parameter values along with principled uncertainties even when the field is crowded. We also compare our results with those obtained from the SExtractor software. While the two approaches largely agree about the fluxes of the bright stars, the Bayesian approach provides more accurate inferences about the faint stars and the number of stars, particularly in the crowded case.

preprint2013arXiv

Reconnaissance of the HR 8799 Exosolar System I: Near IR Spectroscopy

We obtained spectra, in the wavelength range λ= 995 - 1769 nm, of all four known planets orbiting the star HR 8799. Using the suite of instrumentation known as Project 1640 on the Palomar 5-m Hale Telescope, we acquired data at two epochs. This allowed for multiple imaging detections of the companions and multiple extractions of low-resolution (R ~ 35) spectra. Data reduction employed two different methods of speckle suppression and spectrum extraction, both yielding results that agree. The spectra do not directly correspond to those of any known objects, although similarities with L and T-dwarfs are present, as well as some characteristics similar to planets such as Saturn. We tentatively identify the presence of CH_4 along with NH_3 and/or C_2H_2, and possibly CO_2 or HCN in varying amounts in each component of the system. Other studies suggested red colors for these faint companions, and our data confirm those observations. Cloudy models, based on previous photometric observations, may provide the best explanation for the new data presented here. Notable in our data is that these presumably co-eval objects of similar luminosity have significantly different spectra; the diversity of planets may be greater than previously thought. The techniques and methods employed in this paper represent a new capability to observe and rapidly characterize exoplanetary systems in a routine manner over a broad range of planet masses and separations. These are the first simultaneous spectroscopic observations of multiple planets in a planetary system other than our own.

preprint2013arXiv

Searching for comets on the World Wide Web: The orbit of 17P/Holmes from the behavior of photographers

We performed an image search for "Comet Holmes," using the Yahoo Web search engine, on 2010 April 1. Thousands of images were returned. We astrometrically calibrated---and therefore vetted---the images using the Astrometry.net system. The calibrated image pointings form a set of data points to which we can fit a test-particle orbit in the Solar System, marginalizing over image dates and detecting outliers. The approach is Bayesian and the model is, in essence, a model of how comet astrophotographers point their instruments. In this work, we do not measure the position of the comet within each image, but rather use the celestial position of the whole image to infer the orbit. We find very strong probabilistic constraints on the orbit, although slightly off the JPL ephemeris, probably due to limitations of our model. Hyperparameters of the model constrain the reliability of date meta-data and where in the image astrophotographers place the comet; we find that ~70 percent of the meta-data are correct and that the comet typically appears in the central third of the image footprint. This project demonstrates that discoveries and measurements can be made using data of extreme heterogeneity and unknown provenance. As the size and diversity of astronomical data sets continues to grow, approaches like ours will become more essential. This project also demonstrates that the Web is an enormous repository of astronomical information; and that if an object has been given a name and photographed thousands of times by observers who post their images on the Web, we can (re-)discover it and infer its dynamical properties.

preprint2013arXiv

SYNMAG Photometry: A Fast Tool for Catalog-Level Matched Colors of Extended Sources

Obtaining reliable, matched photometry for galaxies imaged by different observatories represents a key challenge in the era of wide-field surveys spanning more than several hundred square degrees. Methods such as flux fitting, profile fitting, and PSF homogenization followed by matched-aperture photometry are all computationally expensive. We present an alternative solution called "synthetic aperture photometry" that exploits galaxy profile fits in one band to efficiently model the observed, PSF-convolved light profile in other bands and predict the flux in arbitrarily sized apertures. Because aperture magnitudes are the most widely tabulated flux measurements in survey catalogs, producing synthetic aperture magnitudes (SYNMAGs) enables very fast matched photometry at the catalog level, without reprocessing imaging data. We make our code public and apply it to obtain matched photometry between SDSS ugriz and UKIDSS YJHK imaging, recovering red-sequence colors and photometric redshifts with a scatter and accuracy as good as if not better than FWHM-homogenized photometry from the GAMA Survey. Finally, we list some specific measurements that upcoming surveys could make available to facilitate and ease the use of SYNMAGs.

preprint2013arXiv

The nature of massive black hole binary candidates: I. Spectral properties and evolution

Theoretically, bound binaries of massive black holes are expected as the natural outcome of mergers of massive galaxies. From the observational side, however, massive black hole binaries remain elusive. Velocity shifts between narrow and broad emission lines in quasar spectra are considered a promising observational tool to search for spatially unresolved, dynamically bound binaries. In this series of papers we investigate the nature of such candidates through analyses of their spectra, images and multi-wavelength spectral energy distributions. Here we investigate the properties of the optical spectra, including the evolution of the broad line profiles, of all the sources identified in our previous study. We find a diverse phenomenology of broad and narrow line luminosities, widths, shapes, ionization conditions and time variability, which we can broadly ascribe to 4 classes based on the shape of the broad line profiles: 1) Objects with bell-shaped broad lines with big velocity shifts (>1000 km/s) compared to their narrow lines show a variety of broad line widths and luminosities, modest flux variations over a few years, and no significant change in the broad line peak wavelength. 2) Objects with double-peaked broad emission lines tend to show very luminous and broadened lines, and little time variability. 3) Objects with asymmetric broad emission lines show a broad range of broad line luminosities and significant variability of the line profiles. 4) The remaining sources tend to show moderate to low broad line luminosities, and can be ascribed to diverse phenomena. We discuss the implications of our findings in the context of massive black hole binary searches.

preprint2013arXiv

The PRIsm MUlti-object Survey (PRIMUS). II. Data Reduction and Redshift Fitting

The PRIsm MUti-object Survey (PRIMUS) is a spectroscopic galaxy redshift survey to z~1 completed with a low-dispersion prism and slitmasks allowing for simultaneous observations of ~2,500 objects over 0.18 square degrees. The final PRIMUS catalog includes ~130,000 robust redshifts over 9.1 sq. deg. In this paper, we summarize the PRIMUS observational strategy and present the data reduction details used to measure redshifts, redshift precision, and survey completeness. The survey motivation, observational techniques, fields, target selection, slitmask design, and observations are presented in Coil et al 2010. Comparisons to existing higher-resolution spectroscopic measurements show a typical precision of sigma_z/(1+z)=0.005. PRIMUS, both in area and number of redshifts, is the largest faint galaxy redshift survey completed to date and is allowing for precise measurements of the relationship between AGNs and their hosts, the effects of environment on galaxy evolution, and the build up of galactic systems over the latter half of cosmic history.

preprint2012arXiv

A data-driven model for spectra: Finding double redshifts in the Sloan Digital Sky Survey

We present a data-driven method - heteroscedastic matrix factorization, a kind of probabilistic factor analysis - for modeling or performing dimensionality reduction on observed spectra or other high-dimensional data with known but non-uniform observational uncertainties. The method uses an iterative inverse-variance-weighted least-squares minimization procedure to generate a best set of basis functions. The method is similar to principal components analysis, but with the substantial advantage that it uses measurement uncertainties in a responsible way and accounts naturally for poorly measured and missing data; it models the variance in the noise-deconvolved data space. A regularization can be applied, in the form of a smoothness prior (inspired by Gaussian processes) or a non-negative constraint, without making the method prohibitively slow. Because the method optimizes a justified scalar (related to the likelihood), the basis provides a better fit to the data in a probabilistic sense than any PCA basis. We test the method on SDSS spectra, concentrating on spectra known to contain two redshift components: These are spectra of gravitational lens candidates and massive black-hole binaries. We apply a hypothesis test to compare one-redshift and two-redshift models for these spectra, utilizing the data-driven model trained on a random subset of all SDSS spectra. This test confirms 129 of the 131 lens candidates in our sample and all of the known binary candidates, and turns up very few false positives.

preprint2012arXiv

Data analysis recipes: Probability calculus for inference

In this pedagogical text aimed at those wanting to start thinking about or brush up on probabilistic inference, I review the rules by which probability distribution functions can (and cannot) be combined. I connect these rules to the operations performed in probabilistic data analysis. Dimensional analysis is emphasized as a valuable tool for helping to construct non-wrong probabilistic statements. The applications of probability calculus in constructing likelihoods, marginalized likelihoods, posterior probabilities, and posterior predictions are all discussed.

preprint2012arXiv

Designing Imaging Surveys for a Retrospective Relative Photometric Calibration

In this paper, we investigate the impact of survey strategy on the performance of self-calibration when the goal is to produce accurate photometric catalogs from wide-field imaging surveys. This self-calibration technique utilizes multiple measurements of sources at different focal-plane positions to constrain instruments' large-scale response (flat-field) from survey science data alone. We create an artificial sky of sources and synthetically observe it under four basic survey strategies, creating an end-to-end simulation of an imaging survey for each. These catalog-level simulations include realistic measurement uncertainties and a complex focal-plane dependence of the instrument response. In the self-calibration step, we simultaneously fit for all the star fluxes and the parameters of a position-dependent flat-field. For realism, we deliberately fit with a wrong noise model and a flat-field functional basis that does not include the model that generated the synthetic data. We demonstrate that with a favorable survey strategy, a complex instrument response can be precisely self-calibrated. We show that returning the same sources to very different focal-plane positions is the key property of any survey strategy designed for accurate retrospective calibration of this type. The results of this work suggest the following advice for those considering the design of large-scale imaging surveys: Do not use a regular, repeated tiling of the sky; instead return the same sources to very different focal-plane positions.

preprint2012arXiv

Galaxy growth by merging in the nearby universe

We measure the mass growth rate by merging for a wide range of galaxy types. We present the small-scale (0.014 < r < 11 h70^{-1} Mpc) projected cross-correlation functions w(rp) of galaxy subsamples from the spectroscopic sample of the NYU VAGC (5 \times 10^5 galaxies of redshifts 0.03 < z < 0.15) with galaxy subsamples from the SDSS imaging (4 \times 10^7 galaxies). We use smooth fits to de-project the two-dimensional functions w(rp) to obtain smooth three-dimensional real-space cross-correlation functions ξ(r) for each of several spectroscopic subsamples with each of several imaging subsamples. Because close pairs are expected to merge, the three-space functions and dynamical evolution time estimates provide galaxy accretion rates. We find that the accretion onto massive blue galaxies and onto red galaxies is dominated by red companions, and that onto small-mass blue galaxies, red and blue galaxies make comparable contributions. We integrate over all types of companions and find that at fixed stellar mass, the total fractional accretion rates onto red galaxies (\sim 1.5 h70 percent per Gyr) is greater than that onto blue galaxies (\sim 0.5 h70 percent per Gyr). Although these rates are very low, they are almost certainly over-estimates because we have assumed that all close pairs merge as quickly as dynamical friction permits.

preprint2012arXiv

Photometric redshifts and quasar probabilities from a single, data-driven generative model

We describe a technique for simultaneously classifying and estimating the redshift of quasars. It can separate quasars from stars in arbitrary redshift ranges, estimate full posterior distribution functions for the redshift, and naturally incorporate flux uncertainties, missing data, and multi-wavelength photometry. We build models of quasars in flux-redshift space by applying the extreme deconvolution technique to estimate the underlying density. By integrating this density over redshift one can obtain quasar flux-densities in different redshift ranges. This approach allows for efficient, consistent, and fast classification and photometric redshift estimation. This is achieved by combining the speed obtained by choosing simple analytical forms as the basis of our density model with the flexibility of non-parametric models through the use of many simple components with many parameters. We show that this technique is competitive with the best photometric quasar classification techniques---which are limited to fixed, broad redshift ranges and high signal-to-noise ratio data---and with the best photometric redshift techniques when applied to broadband optical data. We demonstrate that the inclusion of UV and NIR data significantly improves photometric quasar--star separation and essentially resolves all of the redshift degeneracies for quasars inherent to the ugriz filter system, even when included data have a low signal-to-noise ratio. For quasars spectroscopically confirmed by the SDSS 84 and 97 percent of the objects with GALEX UV and UKIDSS NIR data have photometric redshifts within 0.1 and 0.3, respectively, of the spectroscopic redshift; this amounts to about a factor of three improvement over ugriz-only photometric redshifts. Our code to calculate quasar probabilities and redshift probability distributions is publicly available.

preprint2012arXiv

Replacing standard galaxy profiles with mixtures of Gaussians

Exponential, de Vaucouleurs, and Sérsic profiles are simple and successful models for fitting two-dimensional images of galaxies. One numerical issue encountered in this kind of fitting is the pixel rendering and convolution (or correlation) of the models with the telescope point-spread function (PSF); these operations are slow, and easy to get slightly wrong at small radii. Here we exploit the realization that these models can be approximated to arbitrary accuracy with a mixture (linear superposition) of two-dimensional Gaussians (MoGs). MoGs are fast to render and fast to affine-transform. Most importantly, if you have a MoG model for the pixel-convolved PSF, the PSF-convolved, affine-transformed galaxy models are themselves MoGs and therefore very fast to compute, integrate, and render precisely. We present worked examples that can be directly used in image fitting; we are using them ourselves. The MoG profiles we provide can be swapped in to replace the standard models in any image-fitting code; they sped up model fitting in our projects by an order of magnitude; they ought to make any code faster at essentially no cost in precision.

preprint2012arXiv

Star-Galaxy Classification in Multi-Band Optical Imaging

Ground-based optical surveys such as PanSTARRS, DES, and LSST, will produce large catalogs to limiting magnitudes of r > 24. Star-galaxy separation poses a major challenge to such surveys because galaxies---even very compact galaxies---outnumber halo stars at these depths. We investigate photometric classification techniques on stars and galaxies with intrinsic FWHM < 0.2 arcsec. We consider unsupervised spectral energy distribution template fitting and supervised, data-driven Support Vector Machines (SVM). For template fitting, we use a Maximum Likelihood (ML) method and a new Hierarchical Bayesian (HB) method, which learns the prior distribution of template probabilities from the data. SVM requires training data to classify unknown sources; ML and HB don't. We consider i.) a best-case scenario (SVM_best) where the training data is (unrealistically) a random sampling of the data in both signal-to-noise and demographics, and ii.) a more realistic scenario where training is done on higher signal-to-noise data (SVM_real) at brighter apparent magnitudes. Testing with COSMOS ugriz data we find that HB outperforms ML, delivering ~80% completeness, with purity of ~60-90% for both stars and galaxies, respectively. We find no algorithm delivers perfect performance, and that studies of metal-poor main-sequence turnoff stars may be challenged by poor star-galaxy separation. Using the Receiver Operating Characteristic curve, we find a best-to-worst ranking of SVM_best, HB, ML, and SVM_real. We conclude, therefore, that a well trained SVM will outperform template-fitting methods. However, a normally trained SVM performs worse. Thus, Hierarchical Bayesian template fitting may prove to be the optimal classification method in future surveys.

preprint2012arXiv

The Baryon Oscillation Spectroscopic Survey of SDSS-III

The Baryon Oscillation Spectroscopic Survey (BOSS) is designed to measure the scale of baryon acoustic oscillations (BAO) in the clustering of matter over a larger volume than the combined efforts of all previous spectroscopic surveys of large scale structure. BOSS uses 1.5 million luminous galaxies as faint as i=19.9 over 10,000 square degrees to measure BAO to redshifts z<0.7. Observations of neutral hydrogen in the Lyman alpha forest in more than 150,000 quasar spectra (g<22) will constrain BAO over the redshift range 2.15<z<3.5. Early results from BOSS include the first detection of the large-scale three-dimensional clustering of the Lyman alpha forest and a strong detection from the Data Release 9 data set of the BAO in the clustering of massive galaxies at an effective redshift z = 0.57. We project that BOSS will yield measurements of the angular diameter distance D_A to an accuracy of 1.0% at redshifts z=0.3 and z=0.57 and measurements of H(z) to 1.8% and 1.7% at the same redshifts. Forecasts for Lyman alpha forest constraints predict a measurement of an overall dilation factor that scales the highly degenerate D_A(z) and H^{-1}(z) parameters to an accuracy of 1.9% at z~2.5 when the survey is complete. Here, we provide an overview of the selection of spectroscopic targets, planning of observations, and analysis of data and data quality of BOSS.

preprint2012arXiv

The Extreme Small Scales: Do Satellite Galaxies Trace Dark Matter?

We investigate the radial distribution of galaxies within their host dark matter halos by modeling their small-scale clustering, as measured in the Sloan Digital Sky Survey. Specifically, we model the Jiang et al. (2011) measurements of the galaxy two-point correlation function down to very small projected separations (10 < r < 400 kpc/h), in a wide range of luminosity threshold samples (absolute r-band magnitudes of -18 up to -23). We use a halo occupation distribution (HOD) framework with free parameters that specify both the number and spatial distribution of galaxies within their host dark matter halos. We assume that the first galaxy in each halo lives at the halo center and that additional satellite galaxies follow a radial density profile similar to the dark matter Navarro-Frenk-White (NFW) profile, except that the concentration and inner slope are allowed to vary. We find that in low luminosity samples, satellite galaxies have radial profiles that are consistent with NFW. M_r < -20 and brighter satellite galaxies have radial profiles with significantly steeper inner slopes than NFW (we find inner logarithmic slopes ranging from -1.6 to -2.1, as opposed to -1 for NFW). We define a useful metric of concentration, M_(1/10), which is the fraction of satellite galaxies (or mass) that are enclosed within one tenth of the virial radius of a halo. We find that M_(1/10) for low luminosity satellite galaxies agrees with NFW, whereas for luminous galaxies it is 2.5-4 times higher, demonstrating that these galaxies are substantially more centrally concentrated within their dark matter halos than the dark matter itself. Our results therefore suggest that the processes that govern the spatial distribution of galaxies, once they have merged into larger halos, must be luminosity dependent, such that luminous galaxies become poor tracers of the underlying dark matter.

preprint2012arXiv

The Milky Way has no thick disk

Different stellar sub-populations of the Milky Way's stellar disk are known to have different vertical scale heights, their thickness increasing with age. Using SEGUE spectroscopic survey data, we have recently shown that mono-abundance sub-populations, defined in the [α/Fe]-[Fe/H] space, are well described by single exponential spatial-density profiles in both the radial and the vertical direction; therefore any star of a given abundance is clearly associated with a sub-population of scale height h_z. Here, we work out how to determine the stellar surface-mass density contributions at the solar radius R_0 of each such sub-population, accounting for the survey selection function, and for the fraction of the stellar population mass that is reflected in the spectroscopic target stars given populations of different abundances and their presumed age distributions. Taken together, this enables us to derive Σ_{R_0}(h_z), the surface-mass contributions of stellar populations with scale height h_z. Surprisingly, we find no hint of a thin-thick disk bi-modality in this mass-weighted scale-height distribution, but a smoothly decreasing function, approximately Σ_{R_0}(h_z)\propto \exp(-h_z), from h_z ~ 200 pc to h_z ~ 1 kpc. As h_z is ultimately the structurally defining property of a thin or thick disk, this shows clearly that the Milky Way has a continuous and monotonic distribution of disk thicknesses: there is no 'thick disk' sensibly characterized as a distinct component. We discuss how our result is consistent with evidence for seeming bi-modality in purely geometric disk decompositions, or chemical abundances analyses. We constrain the total visible stellar surface-mass density at the Solar radius to be Σ^*_{R_0} = 30 +/- 1 M_\odot pc^{-2}.

preprint2012arXiv

The Milky Way's circular velocity curve between 4 and 14 kpc from APOGEE data

We measure the Milky Way's rotation curve over the Galactocentric range 4 kpc <~ R <~ 14 kpc from the first year of data from the Apache Point Observatory Galactic Evolution Experiment (APOGEE). We model the line-of-sight velocities of 3,365 stars in fourteen fields with b = 0 deg between 30 deg < l < 210 deg out to distances of 10 kpc using an axisymmetric kinematical model that includes a correction for the asymmetric drift of the warm tracer population (σ_R ~ 35 km/s). We determine the local value of the circular velocity to be V_c(R_0) = 218 +/- 6 km/s and find that the rotation curve is approximately flat with a local derivative between -3.0 km/s/kpc and 0.4 km/s/kpc. We also measure the Sun's position and velocity in the Galactocentric rest frame, finding the distance to the Galactic center to be 8 kpc < R_0 < 9 kpc, radial velocity V_{R,sun} = -10 +/- 1 km/s, and rotational velocity V_{ϕ,sun} = 242^{+10}_{-3} km/s, in good agreement with local measurements of the Sun's radial velocity and with the observed proper motion of Sgr A*. We investigate various systematic uncertainties and find that these are limited to offsets at the percent level, ~2 km/s in V_c. Marginalizing over all the systematics that we consider, we find that V_c(R_0) < 235 km/s at >99% confidence. We find an offset between the Sun's rotational velocity and the local circular velocity of 26 +/- 3 km/s, which is larger than the locally-measured solar motion of 12 km/s. This larger offset reconciles our value for V_c with recent claims that V_c >~ 240 km/s. Combining our results with other data, we find that the Milky Way's dark-halo mass within the virial radius is ~8x10^{11} M_sun.

preprint2012arXiv

The Ninth Data Release of the Sloan Digital Sky Survey: First Spectroscopic Data from the SDSS-III Baryon Oscillation Spectroscopic Survey

The Sloan Digital Sky Survey III (SDSS-III) presents the first spectroscopic data from the Baryon Oscillation Spectroscopic Survey (BOSS). This ninth data release (DR9) of the SDSS project includes 535,995 new galaxy spectra (median z=0.52), 102,100 new quasar spectra (median z=2.32), and 90,897 new stellar spectra, along with the data presented in previous data releases. These spectra were obtained with the new BOSS spectrograph and were taken between 2009 December and 2011 July. In addition, the stellar parameters pipeline, which determines radial velocities, surface temperatures, surface gravities, and metallicities of stars, has been updated and refined with improvements in temperature estimates for stars with T_eff<5000 K and in metallicity estimates for stars with [Fe/H]>-0.5. DR9 includes new stellar parameters for all stars presented in DR8, including stars from SDSS-I and II, as well as those observed as part of the SDSS-III Sloan Extension for Galactic Understanding and Exploration-2 (SEGUE-2). The astrometry error introduced in the DR8 imaging catalogs has been corrected in the DR9 data products. The next data release for SDSS-III will be in Summer 2013, which will present the first data from the Apache Point Observatory Galactic Evolution Experiment (APOGEE) along with another year of data from BOSS, followed by the final SDSS-III data release in December 2014.

preprint2012arXiv

The Panchromatic Hubble Andromeda Treasury IV. A Probabilistic Approach to Inferring the High Mass Stellar Initial Mass Function and Other Power-law Functions

We present a probabilistic approach for inferring the parameters of the present day power-law stellar mass function (MF) of a resolved young star cluster. This technique (a) fully exploits the information content of a given dataset; (b) accounts for observational uncertainties in a straightforward way; (c) assigns meaningful uncertainties to the inferred parameters; (d) avoids the pitfalls associated with binning data; and (e) is applicable to virtually any resolved young cluster, laying the groundwork for a systematic study of the high mass stellar MF (M > 1 Msun). Using simulated clusters and Markov chain Monte Carlo sampling of the probability distribution functions, we show that estimates of the MF slope, α, are unbiased and that the uncertainty, Δα, depends primarily on the number of observed stars and stellar mass range they span, assuming that the uncertainties on individual masses and the completeness are well-characterized. Using idealized mock data, we compute the lower limit precision on α and provide an analytic approximation for Δα as a function of the observed number of stars and mass range. We find that ~ 3/4 of quoted literature uncertainties are smaller than the theoretical lower limit. By correcting these uncertainties to the theoretical lower limits, we find the literature studies yield <α>=2.46 with a 1-σ dispersion of 0.35 dex. We verify that it is impossible for a power-law MF to obtain meaningful constraints on the upper mass limit of the IMF. We show that avoiding substantial biases in the MF slope requires: (1) including the MF as a prior when deriving individual stellar mass estimates; (2) modeling the uncertainties in the individual stellar masses; and (3) fully characterizing and then explicitly modeling the completeness for stars of a given mass. (abridged)

preprint2012arXiv

The Sloan Digital Sky Survey quasar catalog: ninth data release

We present the Data Release 9 Quasar (DR9Q) catalog from the Baryon Oscillation Spectroscopic Survey (BOSS) of the Sloan Digital Sky Survey III. The catalog includes all BOSS objects that were targeted as quasar candidates during the survey, are spectrocopically confirmed as quasars via visual inspection, have luminosities Mi[z=2]<-20.5 (in a $Λ$CDM cosmology with H0 = 70 km/s/Mpc, $Ω_{\rm M}$ = 0.3, and $Ω_Λ$ = 0.7) and either display at least one emission line with full width at half maximum (FWHM) larger than 500 km/s or, if not, have interesting/complex absorption features. It includes as well, known quasars (mostly from SDSS-I and II) that were reobserved by BOSS. This catalog contains 87,822 quasars (78,086 are new discoveries) detected over 3,275 deg$^{2}$ with robust identification and redshift measured by a combination of principal component eigenspectra newly derived from a training set of 8,632 spectra from SDSS-DR7. The number of quasars with $z>2.15$ (61,931) is ~2.8 times larger than the number of z>2.15 quasars previously known. Redshifts and FWHMs are provided for the strongest emission lines (CIV, CIII], MgII). The catalog identifies 7,533 broad absorption line quasars and gives their characteristics. For each object the catalog presents five-band (u,g,r,i,z) CCD-based photometry with typical accuracy of 0.03 mag, and information on the morphology and selection method. The catalog also contains X-ray, ultraviolet, near-infrared, and radio emission properties of the quasars, when available, from other large-area surveys.

preprint2012arXiv

The spatial structure of mono-abundance sub-populations of the Milky Way disk

The spatial, kinematic, and elemental-abundance structure of the Milky Way's stellar disk is complex, and has been difficult to dissect with local spectroscopic or global photometric data. Here, we develop and apply a rigorous density modeling approach for Galactic spectroscopic surveys that enables investigation of the global spatial structure of stellar sub-populations in narrow bins of [α/Fe] and [Fe/H], using 23,767 G-type dwarfs from SDSS/SEGUE. We fit models for the number density of each such mono-abundance component, properly accounting for the complex spectroscopic SEGUE sampling of the underlying stellar population. We find that each mono-abundance sub-population has a simple spatial structure that can be described by a single exponential in both the vertical and radial direction, with continuously increasing scale heights (~200 pc to 1 kpc) and decreasing scale lengths (>4.5 kpc to 2 kpc) for increasingly older sub-populations, as indicated by their lower metallicities and [α/Fe] enhancements. That the abundance-selected sub-components with the largest scale heights have the shortest scale lengths is in sharp contrast with purely geometric `thick--thin disk' decompositions. To the extent that [α/Fe] is an adequate proxy for age, our results directly show that older disk sub-populations are more centrally concentrated, which implies inside-out formation of galactic disks. The fact that the largest scale-height sub-components are most centrally concentrated in the Milky Way is an almost inevitable consequence of explaining the vertical structure of the disk through internal evolution. Whether the simple spatial structure of the mono-abundance sub-components, and the striking correlations between age, scale length, and scale height can be plausibly explained by satellite accretion or other external heating remains to be seen.

preprint2012arXiv

The vertical motions of mono-abundance sub-populations in the Milky Way disk

We present the vertical kinematics of stars in the Milky Way's stellar disk inferred from SDSS/SEGUE G-dwarf data, deriving the vertical velocity dispersion, σ_z, as a function of vertical height |z| and Galactocentric radius R for a set of 'mono-abundance' sub-populations of stars with very similar elemental abundances [α/Fe] and [Fe/H]. We find that all components exhibit nearly isothermal kinematics in |z|, and a slow outward decrease of the vertical velocity dispersion: σ_z (z,R|[α/Fe],[Fe/H]) ~ σ_z ([α/Fe],[Fe/H]) x \exp (-(R-R_0)/7 kpc}). The characteristic velocity dispersions of these components vary from ~ 15 km/s for chemically young, metal-rich stars, to >~ 50 km/s for metal poor stars. The mean σ_z gradient away from the mid plane is only 0.3 +/- 0.2 km/s/kpc. We find a continuum of vertical kinetic temperatures (~σ^2_z) as function of ([α/Fe],[Fe/H]), which contribute to the stellar surface mass density as Σ_{R_0}(σ^2_z) ~ \exp(-σ^2_z). The existence of isothermal mono-abundance populations with intermediate dispersions reject the notion of a thin-thick disk dichotomy. This continuum of disks argues against models where the thicker disk portions arise from massive satellite infall or heating; scenarios where either the oldest disk portion was born hot, or where internal evolution plays a major role, seem the most viable. The wide range of σ_z ([α/Fe],[Fe/H]) combined with a constant σ_z(z) for each abundance bin provides an independent check on the precision of the SEGUE abundances: δ_[α/Fe] ~ 0.07 dex and δ_[Fe/H] ~ 0.15 dex. The radial decline of the vertical dispersion presumably reflects the decrease in disk surface-mass density. This measurement constitutes a first step toward a purely dynamical estimate of the mass profile the disk in our Galaxy. [abridged]

preprint2011arXiv

A systematic search for massive black hole binaries in SDSS spectroscopic sample

We present the results of a systematic search for massive black hole binaries in the Sloan Digital Sky Survey spectroscopic database. We focus on bound binaries, under the assumption that one of the black holes is active. In this framework, the broad lines associated to the accreting black hole are expected to show systematic velocity shifts with respect to the narrow lines, which trace the rest-frame of the galaxy. For a sample of 54586 quasars and 3929 galaxies at redshifts 0.1<z<1.5 we brute-force model each spectrum as a mixture of two quasars at two different redshifts. The spectral model is a data-driven dimensionality reduction of the SDSS quasar spectra based on a matrix factorization. We identified 32 objects with peculiar spectra. Nine of them can be interpreted as black hole binaries. This doubles the number of known black hole binary candidates. We also report on the discovery of a new class of extreme double-peaked emitters with exceptionally broad and faint Balmer lines. For all the interesting sources, we present detailed analysis of the spectra, and discuss possible interpretations.

preprint2011arXiv

An Affine-Invariant Sampler for Exoplanet Fitting and Discovery in Radial Velocity Data

Markov Chain Monte Carlo (MCMC) proves to be powerful for Bayesian inference and in particular for exoplanet radial velocity fitting because MCMC provides more statistical information and makes better use of data than common approaches like chi-square fitting. However, the non-linear density functions encountered in these problems can make MCMC time-consuming. In this paper, we apply an ensemble sampler respecting affine invariance to orbital parameter extraction from radial velocity data. This new sampler has only one free parameter, and it does not require much tuning for good performance, which is important for automatization. The autocorrelation time of this sampler is approximately the same for all parameters and far smaller than Metropolis-Hastings, which means it requires many fewer function calls to produce the same number of independent samples. The affine-invariant sampler speeds up MCMC by hundreds of times compared with Metropolis-Hastings in the same computing situation. This novel sampler would be ideal for projects involving large datasets such as statistical investigations of planet distribution. The biggest obstacle to ensemble samplers is the existence of multiple local optima; we present a clustering technique to deal with local optima by clustering based on the likelihood of the walkers in the ensemble. We demonstrate the effectiveness of the sampler on real radial velocity data.

preprint2011arXiv

Clumpy Streams from Clumpy Halos: Detecting Missing Satellites with Cold Stellar Structures

Dynamically cold stellar streams are ideal probes of the gravitational field of the Milky Way. This paper re-examines the question of how such streams might be used to test for the presence of "missing satellites" -the many thousands of dark-matter subhalos with masses 10^5-10^7Msolar which are seen to orbit within Galactic-scale dark-matter halos in simulations of structure formation in LCDM cosmologies. Analytical estimates of the frequency and energy scales of stream encounters indicate that these missing satellites should have a negligible effect on hot debris structures, such as the tails from the Sagittarius dwarf galaxy. However, long cold streams, such as the structure known as GD-1 or those from the globular cluster Palomar 5 (Pal 5) are expected to suffer many tens of direct impacts from missing satellites during their lifetimes. Numerical experiments confirm that these impacts create gaps in the debris' orbital energy distribution, which will evolve into degree- and sub-degree- scale fluctuations in surface density over the age of the debris. Maps of Pal 5's own stream contain surface density fluctuations on these scales. The presence and frequency of these inhomogeneities suggests the existence of a population of missing satellites in numbers predicted in the standard LCDM cosmologies.

preprint2011arXiv

Extreme deconvolution: Inferring complete distribution functions from noisy, heterogeneous and incomplete observations

We generalize the well-known mixtures of Gaussians approach to density estimation and the accompanying Expectation--Maximization technique for finding the maximum likelihood parameters of the mixture to the case where each data point carries an individual $d$-dimensional uncertainty covariance and has unique missing data properties. This algorithm reconstructs the error-deconvolved or "underlying" distribution function common to all samples, even when the individual data points are samples from different distributions, obtained by convolving the underlying distribution with the heteroskedastic uncertainty distribution of the data point and projecting out the missing data directions. We show how this basic algorithm can be extended with conjugate priors on all of the model parameters and a "split-and-merge" procedure designed to avoid local maxima of the likelihood. We demonstrate the full method by applying it to the problem of inferring the three-dimensional velocity distribution of stars near the Sun from noisy two-dimensional, transverse velocity measurements from the Hipparcos satellite.

preprint2011arXiv

SDSS-III: Massive Spectroscopic Surveys of the Distant Universe, the Milky Way Galaxy, and Extra-Solar Planetary Systems

Building on the legacy of the Sloan Digital Sky Survey (SDSS-I and II), SDSS-III is a program of four spectroscopic surveys on three scientific themes: dark energy and cosmological parameters, the history and structure of the Milky Way, and the population of giant planets around other stars. In keeping with SDSS tradition, SDSS-III will provide regular public releases of all its data, beginning with SDSS DR8 (which occurred in Jan 2011). This paper presents an overview of the four SDSS-III surveys. BOSS will measure redshifts of 1.5 million massive galaxies and Lya forest spectra of 150,000 quasars, using the BAO feature of large scale structure to obtain percent-level determinations of the distance scale and Hubble expansion rate at z<0.7 and at z~2.5. SEGUE-2, which is now completed, measured medium-resolution (R=1800) optical spectra of 118,000 stars in a variety of target categories, probing chemical evolution, stellar kinematics and substructure, and the mass profile of the dark matter halo from the solar neighborhood to distances of 100 kpc. APOGEE will obtain high-resolution (R~30,000), high signal-to-noise (S/N>100 per resolution element), H-band (1.51-1.70 micron) spectra of 10^5 evolved, late-type stars, measuring separate abundances for ~15 elements per star and creating the first high-precision spectroscopic survey of all Galactic stellar populations (bulge, bar, disks, halo) with a uniform set of stellar tracers and spectral diagnostics. MARVELS will monitor radial velocities of more than 8000 FGK stars with the sensitivity and cadence (10-40 m/s, ~24 visits per star) needed to detect giant planets with periods up to two years, providing an unprecedented data set for understanding the formation and dynamical evolution of giant planet systems. (Abridged)

preprint2011arXiv

Statistics of gamma-ray point sources below the Fermi detection limit

An analytic relation between the statistics of photons in pixels and the number counts of multi-photon point sources is used to constrain the distribution of gamma-ray point sources below the Fermi detection limit at energies above 1 GeV and at latitudes below and above 30 degrees. The derived source-count distribution is consistent with the distribution found by the Fermi collaboration based on the first Fermi point source catalogue. In particular, we find that the contribution of resolved and unresolved active galactic nuclei (AGN) to the total gamma-ray flux is below 20% - 25%. In the best fit model, the AGN-like point source fraction is 17% +- 2%. Using the fact that the Galactic emission varies across the sky while the extra-galactic diffuse emission is isotropic, we put a lower limit of 51% on Galactic diffuse emission and an upper limit of 32% on the contribution from extra-galactic weak sources, such as star-forming galaxies. Possible systematic uncertainties are discussed.

preprint2011arXiv

The Color Variability of Quasars

We quantify quasar color-variability using an unprecedented variability database - ugriz photometry of 9093 quasars from SDSS Stripe 82, observed over 8 years at ~60 epochs each. We confirm previous reports that quasars become bluer when brightening. We find a redshift dependence of this blueing in a given set of bands (e.g. g and r), but show that it is the result of the flux contribution from less-variable or delayed emission lines in the different SDSS bands at different redshifts. After correcting for this effect, quasar color-variability is remarkably uniform, and independent not only of redshift, but also of quasar luminosity and black hole mass. The color variations of individual quasars, as they vary in brightness on year timescales, are much more pronounced than the ranges in color seen in samples of quasars across many orders of magnitude in luminosity. This indicates distinct physical mechanisms behind quasar variability and the observed range of quasar luminosities at a given black hole mass - quasar variations cannot be explained by changes in the mean accretion rate. We do find some dependence of the color variability on the characteristics of the flux variations themselves, with fast, low-amplitude, brightness variations producing more color variability. The observed behavior could arise if quasar variability results from flares or ephemeral hot spots in an accretion disc.

preprint2011arXiv

The Eighth Data Release of the Sloan Digital Sky Survey: First Data from SDSS-III

The Sloan Digital Sky Survey (SDSS) started a new phase in August 2008, with new instrumentation and new surveys focused on Galactic structure and chemical evolution, measurements of the baryon oscillation feature in the clustering of galaxies and the quasar Ly alpha forest, and a radial velocity search for planets around ~8000 stars. This paper describes the first data release of SDSS-III (and the eighth counting from the beginning of the SDSS). The release includes five-band imaging of roughly 5200 deg^2 in the Southern Galactic Cap, bringing the total footprint of the SDSS imaging to 14,555 deg^2, or over a third of the Celestial Sphere. All the imaging data have been reprocessed with an improved sky-subtraction algorithm and a final, self-consistent photometric recalibration and flat-field determination. This release also includes all data from the second phase of the Sloan Extension for Galactic Understanding and Evolution (SEGUE-2), consisting of spectroscopy of approximately 118,000 stars at both high and low Galactic latitudes. All the more than half a million stellar spectra obtained with the SDSS spectrograph have been reprocessed through an improved stellar parameters pipeline, which has better determination of metallicity for high metallicity stars.

preprint2011arXiv

The PRIsm MUlti-Object Survey (PRIMUS) I: Survey Overview and Characteristics

We present the PRIsm MUlti-object Survey (PRIMUS), a spectroscopic faint galaxy redshift survey to z~1. PRIMUS uses a low-dispersion prism and slitmasks to observe ~2,500 objects at once in a 0.18 deg^2 field of view, using the IMACS camera on the Magellan I Baade 6.5m telescope at Las Campanas Observatory. PRIMUS covers a total of 9.1 deg^2 of sky to a depth of i_AB~23.5 in seven different deep, multi-wavelength fields that have coverage from GALEX, Spitzer and either XMM or Chandra, as well as multiple-band optical and near-IR coverage. PRIMUS includes ~130,000 robust redshifts of unique objects with a redshift precision of dz/(1+z)~0.005. The redshift distribution peaks at z=0.6 and extends to z=1.2 for galaxies and z=5 for broad-line AGN. The motivation, observational techniques, fields, target selection, slitmask design, and observations are presented here, with a brief summary of the redshift precision; a companion paper presents the data reduction, redshift fitting, redshift confidence, and survey completeness. PRIMUS is the largest faint galaxy survey undertaken to date. The high targeting fraction (~80%) and large survey size will allow for precise measures of galaxy properties and large-scale structure to z~1.

preprint2011arXiv

The SDSS-III Baryon Oscillation Spectroscopic Survey: Quasar Target Selection for Data Release Nine

The SDSS-III Baryon Oscillation Spectroscopic Survey (BOSS), a five-year spectroscopic survey of 10,000 deg^2, achieved first light in late 2009. One of the key goals of BOSS is to measure the signature of baryon acoustic oscillations in the distribution of Ly-alpha absorption from the spectra of a sample of ~150,000 z>2.2 quasars. Along with measuring the angular diameter distance at z\approx2.5, BOSS will provide the first direct measurement of the expansion rate of the Universe at z > 2. One of the biggest challenges in achieving this goal is an efficient target selection algorithm for quasars over 2.2 < z < 3.5, where their colors overlap those of stars. During the first year of the BOSS survey, quasar target selection methods were developed and tested to meet the requirement of delivering at least 15 quasars deg^-2 in this redshift range, out of 40 targets deg^-2. To achieve these surface densities, the magnitude limit of the quasar targets was set at g <= 22.0 or r<=21.85. While detection of the BAO signature in the Ly-alpha absorption in quasar spectra does not require a uniform target selection, many other astrophysical studies do. We therefore defined a uniformly-selected subsample of 20 targets deg^-2, for which the selection efficiency is just over 50%. This "CORE" subsample will be fixed for Years Two through Five of the survey. In this paper we describe the evolution and implementation of the BOSS quasar target selection algorithms during the first two years of BOSS operations. We analyze the spectra obtained during the first year. 11,263 new z>2.2 quasars were spectroscopically confirmed by BOSS. Our current algorithms select an average of 15 z > 2.2 quasars deg^-2 from 40 targets deg^-2 using single-epoch SDSS imaging. Multi-epoch optical data and data at other wavelengths can further improve the efficiency and completeness of BOSS quasar target selection. [Abridged]

preprint2011arXiv

Think Outside the Color Box: Probabilistic Target Selection and the SDSS-XDQSO Quasar Targeting Catalog

We present the SDSS-XDQSO quasar targeting catalog for efficient flux-based quasar target selection down to the faint limit of the Sloan Digital Sky Survey (SDSS) catalog, even at medium redshifts (2.5 <~ z <~ 3) where the stellar contamination is significant. We build models of the distributions of stars and quasars in flux space down to the flux limit by applying the extreme-deconvolution method to estimate the underlying density. We convolve this density with the flux uncertainties when evaluating the probability that an object is a quasar. This approach results in a targeting algorithm that is more principled, more efficient, and faster than other similar methods. We apply the algorithm to derive low-redshift (z < 2.2), medium-redshift (2.2 <= z <= 3.5), and high-redshift (z > 3.5) quasar probabilities for all 160,904,060 point sources with dereddened i-band magnitude between 17.75 and 22.45 mag in the 14,555 deg^2 of imaging from SDSS Data Release 8. The catalog can be used to define a uniformly selected and efficient low- or medium-redshift quasar survey, such as that needed for the SDSS-III's Baryon Oscillation Spectroscopic Survey project. We show that the XDQSO technique performs as well as the current best photometric quasar-selection technique at low redshift, and outperforms all other flux-based methods for selecting the medium-redshift quasars of our primary interest. We make code to reproduce the XDQSO quasar target selection publicly available.

preprint2010arXiv

Are the Ultra-Faint Dwarf Galaxies Just Cusps?

We develop a technique to investigate the possibility that some of the recently discovered ultra-faint dwarf satellites of the Milky Way might be cusp caustics rather than gravitationally self-bound systems. Such cusps can form when a stream of stars folds, creating a region where the projected 2-D surface density is enhanced. In this work, we construct a Poisson maximum likelihood test to compare the cusp and exponential models of any substructure on an equal footing. We apply the test to the Hercules dwarf (d ~ 113 kpc, M_V ~ -6.2, e ~ 0.67). The flattened exponential model is strongly favored over the cusp model in the case of Hercules, ruling out at high confidence that Hercules is a cusp catastrophe. This test can be applied to any of the Milky Way dwarfs, and more generally to the entire stellar halo population, to search for the cusp catastrophes that might be expected in an accreted stellar halo.

preprint2010arXiv

Constraining the Milky Way potential with a 6-D phase-space map of the GD-1 stellar stream

The narrow GD-1 stream of stars, spanning 60 deg on the sky at a distance of ~10 kpc from the Sun and ~15 kpc from the Galactic center, is presumed to be debris from a tidally disrupted star cluster that traces out a test-particle orbit in the Milky Way halo. We combine SDSS photometry, USNO-B astrometry, and SDSS and Calar Alto spectroscopy to construct a complete, empirical 6-dimensional phase-space map of the stream. We find that an eccentric orbit in a flattened isothermal potential describes this phase-space map well. Even after marginalizing over the stream orbital parameters and the distance from the Sun to the Galactic center, the orbital fit to GD-1 places strong constraints on the circular velocity at the Sun's radius V_c=224 \pm 13 km/s and total potential flattening q_Φ=0.87^{+0.07}_{-0.04}. When we drop any informative priors on V_c the GD-1 constraint becomes V_c=221 \pm 18 km/s. Our 6-D map of GD-1 therefore yields the best current constraint on V_c and the only strong constraint on q_Φat Galactocentric radii near R~15 kpc. Much, if not all, of the total potential flattening may be attributed to the mass in the stellar disk, so the GD-1 constraints on the flattening of the halo itself are weak: q_{Φ,halo}>0.89 at 90% confidence. The greatest uncertainty in the 6-D map and the orbital analysis stems from the photometric distances, which will be obviated by Gaia.

preprint2010arXiv

Data analysis recipes: Fitting a model to data

We go through the many considerations involved in fitting a model to data, using as an example the fit of a straight line to a set of points in a two-dimensional plane. Standard weighted least-squares fitting is only appropriate when there is a dimension along which the data points have negligible uncertainties, and another along which all the uncertainties can be described by Gaussians of known variance; these conditions are rarely met in practice. We consider cases of general, heterogeneous, and arbitrarily covariant two-dimensional uncertainties, and situations in which there are bad data (large outliers), unknown uncertainties, and unknown but expected intrinsic scatter in the linear relationship being fit. Above all we emphasize the importance of having a "generative model" for the data, even an approximate one. Once there is a generative model, the subsequent fitting is non-arbitrary because the model permits direct computation of the likelihood of the parameters or the posterior probability distribution. Construction of a posterior probability distribution is indispensible if there are "nuisance parameters" to marginalize away.

preprint2010arXiv

Dynamical inference from a kinematic snapshot: The force law in the Solar System

If a dynamical system is long-lived and non-resonant (that is, if there is a set of tracers that have evolved independently through many orbital times), and if the system is observed at any non-special time, it is possible to infer the dynamical properties of the system (such as the gravitational force or acceleration law) from a snapshot of the positions and velocities of the tracer population at a single moment in time. In this paper we describe a general inference technique that solves this problem while allowing (1) the unknown distribution function of the tracer population to be simultaneously inferred and marginalized over, and (2) prior information about the gravitational field and distribution function to be taken into account. As an example, we consider the simplest problem of this kind: We infer the force law in the Solar System using only an instantaneous kinematic snapshot (valid at 2009 April 1.0) for the eight major planets. We consider purely radial acceleration laws of the form a_r = -A [r/r_0]^{-α}, where r is the distance from the Sun. Using a probabilistic inference technique, we infer 1.989 < α< 2.052 (95 percent interval), largely independent of any assumptions about the distribution of energies and eccentricities in the system beyond the assumption that the system is phase-mixed. Generalizations of the methods used here will permit, among other things, inference of Milky Way dynamics from Gaia-like observations.

preprint2010arXiv

Inferring the eccentricity distribution

Standard maximum-likelihood estimators for binary-star and exoplanet eccentricities are biased high, in the sense that the estimated eccentricity tends to be larger than the true eccentricity. As with most non-trivial observables, a simple histogram of estimated eccentricities is not a good estimate of the true eccentricity distribution. Here we develop and test a hierarchical probabilistic method for performing the relevant meta-analysis, that is, inferring the true eccentricity distribution, taking as input the likelihood functions for the individual-star eccentricities, or samplings of the posterior probability distributions for the eccentricities (under a given, uninformative prior). The method is a simple implementation of a hierarchical Bayesian model; it can also be seen as a kind of heteroscedastic deconvolution. It can be applied to any quantity measured with finite precision--other orbital parameters, or indeed any astronomical measurements of any kind, including magnitudes, parallaxes, or photometric redshifts--so long as the measurements have been communicated as a likelihood function or a posterior sampling.

preprint2010arXiv

Stellar Population Variations in the Milky Way's Stellar Halo

If the stellar halos of disk galaxies are built up from the disruption of dwarf galaxies, models predict highly structured variations in the stellar populations within these halos. We test this prediction by studying the ratio of blue horizontal branch stars (BHB stars; more abundant in old, metal-poor populations) to main-sequence turn-off stars (MSTO stars; a feature of all populations) in the stellar halo of the Milky Way using data from the Sloan Digital Sky Survey. We develop and apply an improved technique to select BHB stars using ugr color information alone, yielding a sample of ~9000 g<18 candidates where ~70% of them are BHB stars. We map the BHB/MSTO ratio across ~1/4 of the sky at the distance resolution permitted by the absolute magnitude distribution of MSTO stars. We find large variations of BHB/MSTO star ratio in the stellar halo. Previously identified, stream-like halo structures have distinctive BHB/MSTO ratios, indicating different ages/metallicities. Some halo features, e.g., the low-latitude structure, appear to be almost completely devoid of BHB stars, whereas other structures appear to be rich in BHB stars. The Sagittarius tidal stream shows an apparent variation in BHB/MSTO ratio along its extent, which we interpret in terms of population gradients within the progenitor dwarf galaxy. Our detection of coherent stellar population variations between different stellar halo substructures provides yet more support to cosmologically motivated models for stellar halo growth.

preprint2010arXiv

Telescopes don't make catalogues!

Astronomical instruments make intensity measurements; any precise astronomical experiment ought to involve modeling those measurements. People make catalogues, but because a catalogue requires hard decisions about calibration and detection, no catalogue can contain all of the information in the raw pixels relevant to most scientific investigations. Here we advocate making catalogue-like data outputs that permit investigators to test hypotheses with almost the power of the original image pixels. The key is to provide users with approximations to likelihood tests against the raw image pixels. We advocate three options, in order of increasing difficulty: The first is to define catalogue entries and associated uncertainties such that the catalogue contains the parameters of an approximate description of the image-level likelihood function. The second is to produce a K-catalogue sampling in "catalogue space" that samples a posterior probability distribution of catalogues given the data. The third is to expose a web service or equivalent that can re-compute on demand the full image-level likelihood for any user-supplied catalogue.

preprint2010arXiv

The Dual Origin of Stellar Halos II: Chemical Abundances as Tracers of Formation History

Fully cosmological, high resolution N-Body + SPH simulations are used to investigate the chemical abundance trends of stars in simulated stellar halos as a function of their origin. These simulations employ a physically motivated supernova feedback recipe, as well as metal enrichment, metal cooling and metal diffusion. As presented in an earlier paper, the simulated galaxies in this study are surrounded by stellar halos whose inner regions contain both stars accreted from satellite galaxies and stars formed in situ in the central regions of the main galaxies and later displaced by mergers into their inner halos. The abundance patterns ([Fe/H] and [O/Fe]) of halo stars located within 10 kpc of a solar-like observer are analyzed. We find that for galaxies which have not experienced a recent major merger, in situ stars at the high [Fe/H] end of the metallicity distribution function are more [alpha/Fe]-rich than accreted stars at similar [Fe/H]. This dichotomy in the [O/Fe] of halo stars at a given [Fe/H] results from the different potential wells within which in situ and accreted halo stars form. These results qualitatively match recent observations of local Milky Way halo stars. It may thus be possible for observers to uncover the relative contribution of different physical processes to the formation of stellar halos by observing such trends in the halo populations of the Milky Way, and other local L* galaxies.

preprint2010arXiv

The velocity distribution of nearby stars from Hipparcos data II. The nature of the low-velocity moving groups

The velocity distribution of nearby stars contains many "moving groups" that are inconsistent with the standard assumption of an axisymmetric, time-independent, and steady-state Galaxy. We study the age and metallicity properties of the low-velocity moving groups based on the reconstruction of the local velocity distribution in Paper I of this series. We perform stringent, conservative hypothesis testing to establish for each of these moving groups whether it could conceivably consist of a coeval population of stars. We conclude that they do not: the moving groups are not trivially associated with their eponymous open clusters nor with any other inhomogeneous star formation event. Concerning a possible dynamical origin of the moving groups, we test whether any of the moving groups has a higher or lower metallicity than the background population of thin disk stars, as would generically be the case if the moving groups are associated with resonances of the bar or spiral structure. We find clear evidence that the Hyades moving group has higher than average metallicity and weak evidence that the Sirius moving group has lower than average metallicity, which could indicate that these two groups are related to the inner Lindblad resonance of the spiral structure. Further we find weak evidence that the Hercules moving group has higher than average metallicity, as would be the case if it is associated with the bar's outer Lindblad resonance. The Pleiades moving group shows no clear metallicity anomaly, arguing against a common dynamical origin for the Hyades and Pleiades groups. Overall, however, the moving groups are barely distinguishable from the background population of stars, raising the likelihood that the moving groups are associated with transient perturbations. [abridged]

preprint2009arXiv

Astrometry.net: Blind astrometric calibration of arbitrary astronomical images

We have built a reliable and robust system that takes as input an astronomical image, and returns as output the pointing, scale, and orientation of that image (the astrometric calibration or WCS information). The system requires no first guess, and works with the information in the image pixels alone; that is, the problem is a generalization of the "lost in space" problem in which nothing--not even the image scale--is known. After robust source detection is performed in the input image, asterisms (sets of four or five stars) are geometrically hashed and compared to pre-indexed hashes to generate hypotheses about the astrometric calibration. A hypothesis is only accepted as true if it passes a Bayesian decision theory test against a background hypothesis. With indices built from the USNO-B Catalog and designed for uniformity of coverage and redundancy, the success rate is 99.9% for contemporary near-ultraviolet and visual imaging survey data, with no false positives. The failure rate is consistent with the incompleteness of the USNO-B Catalog; augmentation with indices built from the 2MASS Catalog brings the completeness to 100% with no false positives. We are using this system to generate consistent and standards-compliant meta-data for digital and digitized imaging from plate repositories, automated observatories, individual scientific investigators, and hobbyists. This is the first step in a program of making it possible to trust calibration meta-data for astronomical data of arbitrary provenance.

preprint2009arXiv

Cosmic transparency: A test with the baryon acoustic feature and type Ia supernovae

Conservation of the phase-space density of photons plus Lorentz invariance requires that the cosmological luminosity distance be larger than the angular diameter distance by a factor of $(1+z)^2$, where $z$ is the redshift. Because this is a fundamental symmetry, this prediction--known sometimes as the "Etherington relation" or the "Tolman test"--is independent of world model, or even the assumptions of homogeneity and isotropy. It depends, however, on Lorentz invariance and transparency. Transparency can be affected by intergalactic dust or interactions between photons and the dark sector. Baryon acoustic feature and type Ia supernovae measures of the expansion history are differently sensitive to the angular diameter and luminosity distances and can therefore be used in conjunction to limit cosmic transparency. At the present day, the comparison only limits the change $Δτ$ in the optical depth from redshift 0.20 to 0.35 at visible wavelengths to $Δτ< 0.13$ at 95% confidence. In a model with a constant comoving number density $n$ of scatterers of constant proper cross-section $σ$, this limit implies $n σ< 2\times10^{-4} h \Mpc^{-1}$. These limits depend weakly on the cosmological world model. Within the next few years, the limits could extend to redshifts $z\approx2.5$ and improve to $n σ<1.1 \times10^{-5} h \Mpc^{-1}$. Cosmic variance will eventually limit the sensitivity of any test using the BAF at the $n σ\sim 4\times10^{-7} h \Mpc^{-1}$ level. Comparison with other measures of the transparency is provided; no other measure in the visible is as free of astrophysical assumptions.

preprint2009arXiv

Measuring the undetectable: Proper motions and parallaxes of very faint sources

The near future of astrophysics involves many large solid-angle, multi-epoch, multi-band imaging surveys. These surveys will, at their faint limits, have data on large numbers of sources that are too faint to be detected at any individual epoch. Here we show that it is possible to measure in multi-epoch data not only the fluxes and positions, but also the parallaxes and proper motions of sources that are too faint to be detected at any individual epoch. The method involves fitting a model of a moving point source simultaneously to all imaging, taking account of the noise and point-spread function in each image. By this method it is possible to measure the proper motion of a point source with an uncertainty close to the minimum possible uncertainty given the information in the data, which is limited by the point-spread function, the distribution of observation times (epochs), and the total signal-to-noise in the combined data. We demonstrate our technique on multi-epoch Sloan Digital Sky Survey imaging of the SDSS Southern Stripe. We show that we can distinguish very red brown dwarfs by their proper motions from very high-redshift quasars more than $1.6\mag$ fainter than with traditional technique on these SDSS data, and with better better fidelity than by multi-band imaging alone. We re-discover all 10 known brown dwarfs in our sample and present 9 new candidate brown dwarfs, identified on the basis of high proper motion.

preprint2009arXiv

The Aromatic Features in Very Faint Dwarf Galaxies

We present optical and mid-infrared photometry of a statistically complete sample of 29 very faint dwarf galaxies (M_r > -15 mag) selected from the SDSS spectroscopic sample and observed in the mid-infrared with Spitzer IRAC. This sample contains nearby (redshift z<0.005) galaxies three magnitudes fainter than previously studied samples. We compare our sample with other star-forming galaxies that have been observed with both IRAC and SDSS. We examine the relationship of the infrared color, sensitive to PAH abundance, with star-formation rates, gas-phase metallicities and radiation hardness, all estimated from optical emission lines. Consistent with studies of more luminous dwarfs, we find that the very faint dwarf galaxies show much weaker PAH emission than more luminous galaxies with similar specific star-formation rates. Unlike more luminous galaxies, we find that the very faint dwarf galaxies show no significant dependence at all of PAH emission on star-formation rate, metallicity, or radiation hardness, despite the fact that the sample spans a significant range in all of these quantities. When the very faint dwarfs in our sample are compared with more luminous (M_r ~ -18 mag) dwarfs, we find that PAH emission depends on metallicity and radiation hardness. These two parameters are correlated; we look at the PAH-metallicity relation at fixed radiation hardness and the PAH-hardness relation at fixed metallicity. This test shows that the PAH emission in dwarf galaxies depends most directly on metallicity.

preprint2009arXiv

The kinematic origin of the cosmological redshift

A common belief about big-bang cosmology is that the cosmological redshift cannot be properly viewed as a Doppler shift (that is, as evidence for a recession velocity), but must be viewed in terms of the stretching of space. We argue that, contrary to this view, the most natural interpretation of the redshift is as a Doppler shift, or rather as the accumulation of many infinitesimal Doppler shifts. The stretching-of-space interpretation obscures a central idea of relativity, namely that it is always valid to choose a coordinate system that is locally Minkowskian. We show that an observed frequency shift in any spacetime can be interpreted either as a kinematic (Doppler) shift or a gravitational shift by imagining a suitable family of observers along the photon's path. In the context of the expanding universe the kinematic interpretation corresponds to a family of comoving observers and hence is more natural.

preprint2009arXiv

What bandwidth do I need for my image?

Computer representations of real numbers are necessarily discrete, with some finite resolution, discreteness, quantization, or minimum representable difference. We perform astrometric and photometric measurements on stars and co-add multiple observations of faint sources to demonstrate that essentially all of the scientific information in an optical astronomical image can be preserved or transmitted when the minimum representable difference is a factor of two finer than the root-variance of the per-pixel noise. Adopting a representation this coarse reduces bandwidth for data acquisition, transmission, or storage, or permits better use of the system dynamic range, without sacrificing any information for down-stream data analysis, including information on sources fainter than the minimum representable difference itself.

preprint2008arXiv

Automated detection of galaxy-scale gravitational lenses in high resolution imaging data

Lens modeling is the key to successful and meaningful automated strong galaxy-scale gravitational lens detection. We have implemented a lens-modeling "robot" that treats every bright red galaxy (BRG) in a large imaging survey as a potential gravitational lens system. Using a simple model optimized for "typical" galaxy-scale lenses, we generate four assessments of model quality that are used in an automated classification. The robot infers the lens classification parameter H that a human would have assigned; the inference is performed using a probability distribution generated from a human-classified training set, including realistic simulated lenses and known false positives drawn from the HST/EGS survey. We compute the expected purity, completeness and rejection rate, and find that these can be optimized for a particular application by changing the prior probability distribution for H, equivalent to defining the robot's "character." Adopting a realistic prior based on the known abundance of lenses, we find that a lens sample may be generated that is ~100% pure, but only ~20% complete. This shortfall is due primarily to the over-simplicity of the lens model. With a more optimistic robot, ~90% completeness can be achieved while rejecting ~90% of the candidate objects. The remaining candidates must be classified by human inspectors. We are able to classify lens candidates by eye at a rate of a few seconds per system, suggesting that a future 1000 square degree imaging survey containing 10^7 BRGs, and some 10^4 lenses, could be successfully, and reproducibly, searched in a modest amount of time. [Abridged]

preprint2008arXiv

Cleaning the USNO-B Catalog through automatic detection of optical artifacts

The USNO-B Catalog contains spurious entries that are caused by diffraction spikes and circular reflection halos around bright stars in the original imaging data. These spurious entries appear in the Catalog as if they were real stars; they are confusing for some scientific tasks. The spurious entries can be identified by simple computer vision techniques because they produce repeatable patterns on the sky. Some techniques employed here are variants of the Hough transform, one of which is sensitive to (two-dimensional) overdensities of faint stars in thin right-angle cross patterns centered on bright ($<13 \mag$) stars, and one of which is sensitive to thin annular overdensities centered on very bright ($<7 \mag$) stars. After enforcing conservative statistical requirements on spurious-entry identifications, we find that of the 1,042,618,261 entries in the USNO-B Catalog, 24,148,382 of them ($2.3 \percent$) are identified as spurious by diffraction-spike criteria and 196,133 ($0.02 \percent$) are identified as spurious by reflection-halo criteria. The spurious entries are often detected in more than 2 bands and are not overwhelmingly outliers in any photometric properties; they therefore cannot be rejected easily on other grounds, i.e., without the use of computer vision techniques. We demonstrate our method, and return to the community in electronic form a table of spurious entries in the Catalog.

preprint2006arXiv

What aspects of galaxy environment matter?

We determine what aspects of the density field surrounding galaxies most affect their properties. For Sloan Digital Sky Survey galaxies, we measure the group environment, meaning the host group luminosity and the distance from the group center (hereafter, ``groupocentric distance''). For comparison, we measure the surrounding density field on scales ranging from 100 kpc/h to 10 Mpc/h. We use the relationship between color and group environment to test the null hypothesis that only the group environment matters, searching for a residual dependence of properties on the surrounding density. Generally, red galaxies are slightly more clustered on small scales (about 100--300 kpc/h) than the null hypothesis predicts, possibly indicating that substructure within groups has some importance. At large scales (> 1 Mpc/h), the actual projected correlation functions of galaxies are biased at less than the 5% level with respect to the null hypothesis predictions. We exclude strongly the converse null hypothesis, that only the surrounding density (on any scale) matters. These results generally encourage the use of the halo model description of galaxy bias, which models the galaxy distribution as a function of host halo mass alone. We compare these results to proposed galaxy formation scenarios within the Cold Dark Matter cosmological model.

preprint2005arXiv

A New Milky Way Dwarf Galaxy in Ursa Major

In this Letter, we report the discovery of a new dwarf satellite to the Milky Way, located at ($α_{2000}, δ_{2000}$) $=$ (158.72,51.92) in the constellation of Ursa Major. This object was detected as an overdensity of red, resolved stars in Sloan Digital Sky Survey data. The color-magnitude diagram of the Ursa Major dwarf looks remarkably similar to that of Sextans, the lowest surface brightness Milky Way companion known, but with approximately an order of magnitude fewer stars. Deeper follow-up imaging confirms this object has an old and metal-poor stellar population and is $\sim$ 100 kpc away. We roughly estimate M$_V =$ -6.75 and $r_{1/2} =$ 250 pc for this dwarf. Its luminosity is several times fainter than the faintest known Milky Way dwarf. However, its physical size is typical for dSphs. Even though its absolute magnitude and size are presently quite uncertain, Ursa Major is likely the lowest luminosity and lowest surface brightness galaxy yet known.

preprint2002arXiv

The luminosity density of red galaxies

A complete sample of $7.7\times 10^4$ galaxies with five-band imaging and spectroscopic redshifts from the Sloan Digital Sky Survey is used to determine the fraction of the optical luminosity density of the Local Universe (redshifts $0.02<z<0.22$) emitted by red galaxies. The distribution in the space of rest-frame color, central surface brightness, and concentration is shown to be highly clustered and bimodal; galaxies fall primarily into one of two distinct classes. One class is red, concentrated and high in surface brightness; the other is bluer, less concentrated, and lower in central surface brightness. Elliptical and bulge-dominated galaxies preferentially belong to the red class. Even with a very restrictive definition of the red class that includes limits on color, surface brightness and concentration, the class comprises roughly one fifth of the number density of galaxies more luminous than $0.05 \Lstar$ and produces two fifths of the total cosmic galaxy luminosity density at $0.7 \mathrm{μm}$. The natural interpretation is that a large fraction of the stellar mass density of the Local Universe is in very old stellar populations.

preprint2001arXiv

A photometricity and extinction monitor at the Apache Point Observatory

An unsupervised software ``robot'' that automatically and robustly reduces and analyzes CCD observations of photometric standard stars is described. The robot measures extinction coefficients and other photometric parameters in real time and, more carefully, on the next day. It also reduces and analyzes data from an all-sky $10 μm$ camera to detect clouds; photometric data taken during cloudy periods are automatically rejected. The robot reports its findings back to observers and data analysts via the World-Wide Web. It can be used to assess photometricity, and to build data on site conditions. The robot's automated and uniform site monitoring represents a minimum standard for any observing site with queue scheduling, a public data archive, or likely participation in any future National Virtual Observatory.

preprint1997arXiv

Counts and Colors of Faint Galaxies in the U and R Bands

Ground-based counts and colors of faint galaxies in the U and R bands in one field at high Galactic latitude are presented. Integrated over flux, a total of 1.2x10^5 sources per square degree are found to U=25.5 mag and 6.3x10^5 sources per square degree to R=27 mag, with d log N/dm ~ 0.5 in the U band and d log N/dm ~ 0.3 in the R band. Consistent with these number-magnitude curves, sources become bluer with increasing magnitude to median U-R=0.6 mag at 24<U<25 mag and U-R=1.2 mag at 25 < R < 26 mag. Because the Lyman break redshifts into the U band at z~3, at least 1.2x10^5 sources per square degree must be at redshifts z<3. Measurable U-band fluxes of 73 percent of the 6.3x10^5 sources per square degree suggest that the majority of these also lie at z < 3. These results require an enormous space density of objects in any cosmological model.

preprint1997arXiv

Near Infrared Imaging of the Hubble Deep Field with The Keck Telescope

Two deep K-band ($2.2 μm$) images, with point-source detection limits of $K=25.2$ mag (one sigma), taken with the Keck Telescope in subfields of the Hubble Deep Field, are presented and analyzed. A sample of objects to K=24 mag is constructed and $V_{606}-I_{814}$ and $I_{814}-K$ colors are measured. By stacking visually selected objects, mean $I_{814}-K$ colors can be measured to very faint levels; the mean $I_{814}-K$ color is constant with apparent magnitude down to $V_{606}=28$ mag.

preprint1995arXiv

The Discovery of Two Giant Arcs in the Rich Cluster A2219 with the Keck Telescope

We report the discovery with the Keck telescope of two new multiply-imaged arcs in the luminous X-ray cluster A2219 ($z=0.225$). The brighter arc in the field is red and we use spectroscopic and photometric information to identify it as a $z \sim 1$ moderately star-forming system. The brightness of this arc implies that it is formed from two merging images of the background source, and we identify possible candidates for the third image of this source. The second giant arc in this cluster is blue, and while fainter than the red arc it has a similarly large angular extent (32 arcsec). This arc comprises three images of a single nucleated source -- the relative parities of the three images are discernible in our best resolution images. The presence of several bright multiply imaged arcs in a single cluster allows detailed modelling of the cluster mass distribution, especially when redshift information is available. We present a lensing model of the cluster which explains the properties of the various arcs, and we contrast this model with the optical and X-ray information available on the cluster. We uncover significant differences between the distributions of mass and X-ray gas in the cluster. We suggest that such discrepancies may indicate an on-going merger event in the cluster core, possibly associated with a group around the second brightest cluster member. The preponderance of similar merger signatures in a large fraction of the moderate redshift clusters would indicate their dynamical immaturity.

preprint1993arXiv

The Gravitational Lens System B1422$+$231: Dark Matter, Superluminal Expansion and the Hubble Constant

A gravitational lens model of the radio quasar B1422+231 is presented which can account for the image arrangement and approximately for the relative magnifications. The locations of the principal lensing mass and a more distant secondary mass concentration were predicted and subsequently luminous galaxies were found at these locations. This argues against the existence of substantial numbers of ``dark'' galaxies. The model suggests that if the compact radio source is intrinsically superluminal then the observed component motions may be as large as ~100c with image B moving in the opposite direction to images A and C. The prospects for a measuring the Hubble constant from a model incorporating lens galaxy locations, compact radio source expansion speeds and radio time delays, if and when these are measured, are briefly assessed.