Source author record

Joseph S. Koopmeiners

Joseph S. Koopmeiners appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2025arXiv

Jointly modeling multiple endpoints for efficient treatment effect estimation in randomized controlled trials

Randomized controlled trials are the gold standard for evaluating the efficacy of an intervention. However, there is often a trade-off between selecting the most scientifically relevant primary endpoint versus a less relevant, but more powerful, endpoint. For example, in the context of tobacco regulatory science many trials evaluate cigarettes per day as the primary endpoint instead of abstinence from smoking due to limited power. Additionally, it is often of interest to consider subgroup analyses to answer additional questions; such analyses are rarely adequately powered. In practice, trials often collect multiple endpoints. Heuristically, if multiple endpoints demonstrate a similar treatment effect we would be more confident in the results of this trial. However, there is limited research on leveraging information from secondary endpoints besides using composite endpoints which can be difficult to interpret. In this paper, we develop an estimator for the treatment effect on the primary endpoint based on a joint model for primary and secondary efficacy endpoints. This estimator gains efficiency over the standard treatment effect estimator when the model is correctly specified but is robust to model misspecification via model averaging. We illustrate our approach by estimating the effect of very low nicotine content cigarettes on the proportion of Black people who smoke who achieve abstinence and find our approach reduces the standard error by 27%.

preprint2020arXiv

Bayesian Spatial Models for Voxel-wise Prostate Cancer Classification Using Multi-parametric MRI Data

Multi-parametric magnetic resonance imaging (mpMRI) plays an increasingly important role in the diagnosis of prostate cancer. Various computer-aided detection algorithms have been proposed for automated prostate cancer detection by combining information from various mpMRI data components. However, there exist other features of mpMRI, including the spatial correlation between voxels and between-patient heterogeneity in the mpMRI parameters, that have not been fully explored in the literature but could potentially improve cancer detection if leveraged appropriately. This paper proposes novel voxel-wise Bayesian classifiers for prostate cancer that account for the spatial correlation and between-patient heterogeneity in mpMRI. Modeling the spatial correlation is challenging due to the extreme high dimensionality of the data, and we consider three computationally efficient approaches using Nearest Neighbor Gaussian Process (NNGP), knot-based reduced-rank approximation, and a conditional autoregressive (CAR) model, respectively. The between-patient heterogeneity is accounted for by adding a subject-specific random intercept on the mpMRI parameter model. Simulation results show that properly modeling the spatial correlation and between-patient heterogeneity improves classification accuracy. Application to in vivo data illustrates that classification is improved by spatial modeling using NNGP and reduced-rank approximation but not the CAR model, while modeling the between-patient heterogeneity does not further improve our classifier. Among our proposed models, the NNGP-based model is recommended considering its robust classification accuracy and high computational efficiency.

preprint2020arXiv

Borrowing from Supplemental Sources to Estimate Causal Effects from a Primary Data Source

The increasing multiplicity of data sources offers exciting possibilities in estimating the effects of a treatment, intervention, or exposure, particularly if observational and experimental sources could be used simultaneously. Borrowing between sources can potentially result in more efficient estimators, but it must be done in a principled manner to mitigate increased bias and Type I error. Furthermore, when the effect of treatment is confounded, as in observational sources or in clinical trials with noncompliance, causal effect estimators are needed to simultaneously adjust for confounding and to estimate effects across data sources. We consider the problem of estimating causal effects from a primary source and borrowing from any number of supplemental sources. We propose using regression-based estimators that borrow based on assuming exchangeability of the regression coefficients and parameters between data sources. Borrowing is accomplished with multisource exchangeability models and Bayesian model averaging. We show via simulation that a Bayesian linear model and Bayesian additive regression trees both have desirable properties and borrow under appropriate circumstances. We apply the estimators to recently completed trials of very low nicotine content cigarettes investigating their impact on smoking behavior.

preprint2020arXiv

Statistical design considerations for trials that study multiple indications

Breakthroughs in cancer biology have defined new research programs emphasizing the development of therapies that target specific pathways in tumor cells. Innovations in clinical trial design have followed with master protocols defined by inclusive eligibility criteria and evaluations of multiple therapies and/or histologies. Consequently, characterization of subpopulation heterogeneity has become central to the formulation and selection of a study design. However, this transition to master protocols has led to challenges in identifying the optimal trial design and proper calibration of hyperparameters. We often evaluate a range of null and alternative scenarios, however there has been little guidance on how to synthesize the potentially disparate recommendations for what may be optimal. This may lead to the selection of suboptimal designs and statistical methods that do not fully accommodate the subpopulation heterogeneity. This article proposes novel optimization criteria for calibrating and evaluating candidate statistical designs of master protocols in the presence of the potential for treatment effect heterogeneity among enrolled patient subpopulations. The framework is applied to demonstrate the statistical properties of conventional study designs when treatments offer heterogeneous benefit as well as identify optimal designs devised to monitor the potential for heterogeneity among patients with differing clinical indications using Bayesian modeling.

preprint2015arXiv

Who's good this year? Comparing the Information Content of Games in the Four Major US Sports

In the four major North American professional sports (baseball, basketball, football, and hockey), the primary purpose of the regular season is to determine which teams most deserve to advance to the playoffs. Interestingly, while the ultimate goal of identifying the best teams is the same, the number of regular season games played differs dramatically between the sports, ranging from 16 (football) to 82 (basketball and hockey) to 162 (baseball). Though length of season is partially determined by many factors including travel logistics, rest requirements, playoff structure and television contracts, it is hard to reconcile the 10-fold difference in the number of games between, for example, the NFL and MLB unless football games are somehow more "informative" than baseball games. In this paper, we aim to quantify the amount of information games yield about the relative strength of the teams involved. Our strategy is to assess how well simple paired comparison models fitted from $X%$ of the games within a season predict the outcomes of the remaining $(100-X)%$ of games, for multiple values of $X$. We compare the resulting predictive accuracy curves between seasons within the same sport and across all four sports, and find dramatic differences in the amount of information yielded by individual game results in the four major U.S. sports.

preprint2014arXiv

The Randomized CRM: An Approach to Overcoming the Long-Memory Property of the CRM

The primary object of a phase I clinical trial is to determine the maximum tolerated dose (MTD). Typically, the MTD is identified using a dose-escalation study, where initial subjects are treated at the lowest dose level and subsequent subjects are treated at progressively higher dose levels until the MTD is identified. The continual reassessment method (CRM) is a popular model-based dose-escalation design, which utilizes a formal model for the relationship between dose and toxicity to guide dose-finding. Recently, it was shown that the CRM has a tendency to get "stuck" on a dose-level, with little escalation or de-escalation in the late stages of the trial, due to the long-memory property of the CRM. We propose the randomized CRM (rCRM), which introduces random escalation and de-escalation into the standard CRM dose-finding algorithm, as an approach to overcoming the long-memory property of the CRM. We discuss two approaches to random escalation and de-escalation and compare the operating characteristics of the rCRM to the standard CRM by simulation. Our simulation results show that the rCRM identifies the true MTD at a similar rate and results in a similar number of DLTs compared to the standard CRM, while reducing the trial-to-trial variability in the number of cohorts treated at the true MTD.

preprint2012arXiv

Asymptotic properties of the sequential empirical ROC, PPV and NPV curves under case-control sampling

The receiver operating characteristic (ROC) curve, the positive predictive value (PPV) curve and the negative predictive value (NPV) curve are three measures of performance for a continuous diagnostic biomarker. The ROC, PPV and NPV curves are often estimated empirically to avoid assumptions about the distributional form of the biomarkers. Recently, there has been a push to incorporate group sequential methods into the design of diagnostic biomarker studies. A thorough understanding of the asymptotic properties of the sequential empirical ROC, PPV and NPV curves will provide more flexibility when designing group sequential diagnostic biomarker studies. In this paper, we derive asymptotic theory for the sequential empirical ROC, PPV and NPV curves under case-control sampling using sequential empirical process theory. We show that the sequential empirical ROC, PPV and NPV curves converge to the sum of independent Kiefer processes and show how these results can be used to derive asymptotic results for summaries of the sequential empirical ROC, PPV and NPV curves.