Source author record

Howard H. Chang

Howard H. Chang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2022arXiv

A Bayesian framework for incorporating exposure uncertainty into health analyses with application to air pollution and stillbirth

Studies of the relationships between environmental exposures and adverse health outcomes often rely on a two-stage statistical modeling approach, where exposure is modeled/predicted in the first stage and used as input to a separately fit health outcome analysis in the second stage. Uncertainty in these predictions is frequently ignored, or accounted for in an overly simplistic manner, when estimating the associations of interest. Working in the Bayesian setting, we propose a flexible kernel density estimation (KDE) approach for fully utilizing posterior output from the first stage modeling/prediction to make accurate inference on the association between exposure and health in the second stage, derive the full conditional distributions needed for efficient model fitting, detail its connections with existing approaches, and compare its performance through simulation. Our KDE approach is shown to generally have improved performance across several settings and model comparison metrics. Using competing approaches, we investigate the association between lagged daily ambient fine particulate matter levels and stillbirth counts in New Jersey (2011-2015), observing an increase in risk with elevated exposure three days prior to delivery. The newly developed methods are available in the R package KDExp.

preprint2019arXiv

A comparison of statistical and machine learning methods for creating national daily maps of ambient PM$_{2.5}$ concentration

A typical problem in air pollution epidemiology is exposure assessment for individuals for which health data are available. Due to the sparsity of monitoring sites and the limited temporal frequency with which measurements of air pollutants concentrations are collected (for most pollutants, once every 3 or 6 days), epidemiologists have been moving away from characterizing ambient air pollution exposure solely using measurements. In the last few years, substantial research efforts have been placed in developing statistical methods or machine learning techniques to generate estimates of air pollution at finer spatial and temporal scales (daily, usually) with complete coverage. Some of these methods include: geostatistical techniques, such as kriging; spatial statistical models that use the information contained in air quality model outputs (statistical downscaling models); linear regression modeling approaches that leverage the information in GIS covariates (land use regression); or machine learning methods that mine the information contained in relevant variables (neural network and deep learning approaches). Although some of these exposure modeling approaches have been used in several air pollution epidemiological studies, it is not clear how much the predicted exposures generated by these methods differ, and which method generates more reliable estimates. In this paper, we aim to address this gap by evaluating a variety of exposure modeling approaches, comparing their predictive performance and computational difficulty. Using PM$_{2.5}$ in year 2011 over the continental U.S. as case study, we examine the methods' performances across seasons, rural vs urban settings, and levels of PM$_{2.5}$ concentrations (low, medium, high).

preprint2016arXiv

Dynamic communicability and epidemic spread: a case study on an empirical dynamic contact network

We analyze a recently proposed temporal centrality measure applied to an empirical network based on person-to-person contacts in an emergency department of a busy urban hospital. We show that temporal centrality identifies a distinct set of top-spreaders than centrality based on the time-aggregated binarized contact matrix, so that taken together, the accuracy of capturing top-spreaders improves significantly. However, with respect to predicting epidemic outcome, the temporal measure does not necessarily outperform less complex measures. Our results also show that other temporal markers such as duration observed and the time of first appearance in the the network can be used in a simple predictive model to generate predictions that capture the trend of the observed data remarkably well.

preprint2016arXiv

Weighted SAMGSR: combining significance analysis of microarray-gene set reduction algorithm with pathway topology-based weights to select relevant genes

Introduction It has been demonstrated that a pathway-based feature selection method which incorporates biological information within pathways into the process of feature selection usually outperform a gene-based feature selection algorithm in terms of predictive accuracy, stability, and biological interpretation. Significance analysis of microarray-gene set reduction algorithm (SAMGSR), an extension to a gene set analysis method with further reduction of the selected pathways to their respective core subsets, can be regarded as a pathway-based feature selection method. Results and Discussion In SAMGSR, whether a gene is selected is mainly determined by its expression difference between the phenotypes, and partially by the number of pathways to which this gene belongs, but ignoring the topology information among pathways. In this study, we propose a weighted version of the SAMGSR algorithm by constructing weights based on the connectivity among genes and then incorporating these weights in the test statistic. Conclusions Using both simulated and real-world data, we evaluate the performance of the proposed SAMGSR extension and demonstrate that gene connectivity is indeed informative for feature selection.

preprint2015arXiv

Feature selection for longitudinal microarray data by adapting a pathway analysis method

Introduction: Feature selection and gene set analysis are of increasing interest in bioinformatics. While these two approaches have been developed for different purposes, we describe how some gene set analysis methods can be used to conduct feature selection. Here we adapt the gene set analysis method, significance analysis of microarray gene set reduction (SAMGSR), for feature selection, and propose two extensions-simple SAMGSR and two-level SAMGSR to identify relevant features for longitudinal microarray data. Results and Discussion: When applied to a real-world application, both simple and two-level SAMGSR work comparably well. Using simulated data, we further demonstrate that both SAMGSR extensions have the ability to identify the true relevant genes. If the relevant genes are not highly correlated with the irrelevant ones, the final models given by the two SAMGSR extensions are parsimonious as well. Conclusions: By adapting SAMGSR for feature selection and applying the proposed algorithms on a longitudinal gene expression dataset, we demonstrate that a gene set analysis method can be used for the purpose of feature selection. We believe this work paves the way for more research to bridge feature selection and gene set analysis with the development of novel algorithms.