Source author record

Jay Bartroff

Jay Bartroff appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

18works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

18 published item(s)

preprint2022arXiv

Optimal and fast confidence intervals for hypergeometric successes

We present an efficient method of calculating exact confidence intervals for the hypergeometric parameter representing the number of "successes," or "special items," in the population. The method inverts minimum-width acceptance intervals after shifting them to make their endpoints nondecreasing while preserving their level. The resulting set of confidence intervals achieves minimum possible average size, and even in comparison with confidence sets not required to be intervals it attains the minimum possible cardinality most of the time, and always within $1$. The method compares favorably with existing methods not only in the size of the intervals but also in the time required to compute them. The available \textsf{R} package \texttt{hyperMCI} implements the proposed method.

preprint2020arXiv

Asymptotically optimal sequential FDR and pFDR control with (or without) prior information on the number of signals

We investigate asymptotically optimal multiple testing procedures for streams of sequential data in the context of prior information on the number of false null hypotheses ("signals"). We show that the "gap" and "gap-intersection" procedures, recently proposed and shown by Song and Fellouris (2017, Electron. J. Statist.) to be asymptotically optimal for controlling type 1 and 2 familywise error rates (FWEs), are also asymptotically optimal for controlling FDR/FNR when their critical values are appropriately adjusted. Generalizing this result, we show that these procedures, again with appropriately adjusted critical values, are asymptotically optimal for controlling any multiple testing error metric that is bounded between multiples of FWE in a certain sense. This class of metrics includes FDR/FNR but also pFDR/pFNR, the per-comparison and per-family error rates, and the false positive rate. Our analysis includes asymptotic regimes in which the number of null hypotheses approaches $\infty$ as the type 1 and 2 error metrics approach $0$.

preprint2016arXiv

Multiple Hypothesis Tests Controlling Generalized Error Rates for Sequential Data

The $γ$-FDP and $k$-FWER multiple testing error metrics, which are tail probabilities of the respective error statistics, have become popular recently as less-stringent alternatives to the FDR and FWER. We propose general and flexible stepup and stepdown procedures for testing multiple hypotheses about sequential (or streaming) data that simultaneously control both the type I and II versions of $γ$-FDP, or $k$-FWER. The error control holds regardless of the dependence between data streams, which may be of arbitrary size and shape. All that is needed is a test statistic for each data stream that controls the conventional type I and II error probabilities, and no information or assumptions are required about the joint distribution of the statistics or data streams. The procedures can be used with sequential, group sequential, truncated, or other sampling schemes. We give recommendations for the procedures' implementation including closed-form expressions for the needed critical values in some commonly-encountered testing situations. The proposed sequential procedures are compared with each other and with comparable fixed sample size procedures in the context of strongly positively correlated Gaussian data streams. For this setting we conclude that both the stepup and stepdown sequential procedures provide substantial savings over the fixed sample procedures in terms of expected sample size, and the stepup procedure performs slightly but consistently better than the stepdown for $γ$-FDP control, with the relationship reversed for $k$-FWER control.

preprint2015arXiv

A Rejection Principle for Sequential Tests of Multiple Hypotheses Controlling Familywise Error Rates

We present a unifying approach to multiple testing procedures for sequential (or streaming) data by giving sufficient conditions for a sequential multiple testing procedure to control the familywise error rate (FWER), extending to the sequential domain the work of Goeman and Solari (2010) who accomplished this for fixed sample size procedures. Together we call these conditions the "rejection principle for sequential tests," which we then apply to some existing sequential multiple testing procedures to give simplified understanding of their FWER control. Next the principle is applied to derive two new sequential multiple testing procedures with provable FWER control, one for testing hypotheses in order and another for closed testing. Examples of these new procedures are given by applying them to a chromosome aberration data set and to finding the maximum safe dose of a treatment.

preprint2014arXiv

A New Approach to Designing Phase I-II Cancer Trials for Cytotoxic Chemotherapies

Recently there has been much work on early phase cancer designs that incorporate both toxicity and efficacy data, called Phase I-II designs because they combine elements of both phases. However, they do not explicitly address the Phase II hypothesis test of $H_0: p\le p_0$, where $p$ is the probability of efficacy at the estimated maximum tolerated dose (MTD) $\widehatη$ from Phase I and $p_0$ is the baseline efficacy rate. Standard practice for Phase II remains to treat $p$ as a fixed, unknown parameter and to use Simon's 2-stage design with all patients dosed at $\widehatη$. We propose a Phase I-II design that addresses the uncertainty in the estimate $p=p(\widehatη)$ in $H_0$ by using sequential generalized likelihood theory. Combining this with a Phase I design that incorporates efficacy data, the Phase I-II design provides a common framework that can be used all the way from the first dose of Phase I through the final accept/reject decision about $H_0$ at the end of Phase II, utilizing both toxicity and efficacy data throughout. Efficient group sequential testing is used in Phase II that allows for early stopping to show treatment effect or futility. The proposed Phase I-II design thus removes the artificial barrier between Phase I and Phase II, and fulfills the objectives of searching for the MTD and testing if the treatment has an acceptable response rate to enter into a Phase III trial.

preprint2014arXiv

Sequential Tests of Multiple Hypotheses Controlling Type I and II Familywise Error Rates

This paper addresses the following general scenario: A scientist wishes to perform a battery of experiments, each generating a sequential stream of data, to investigate some phenomenon. The scientist would like to control the overall error rate in order to draw statistically-valid conclusions from each experiment, while being as efficient as possible. The between-stream data may differ in distribution and dimension but also may be highly correlated, even duplicated exactly in some cases. Treating each experiment as a hypothesis test and adopting the familywise error rate (FWER) metric, we give a procedure that sequentially tests each hypothesis while controlling both the type I and II FWERs regardless of the between-stream correlation, and only requires arbitrary sequential test statistics that control the error rates for a given stream in isolation. The proposed procedure, which we call the sequential Holm procedure because of its inspiration from Holm's (1979) seminal fixed-sample procedure, shows simultaneous savings in expected sample size and less conservative error control relative to fixed sample, sequential Bonferroni, and other recently proposed sequential procedures in a simulation study.

preprint2013arXiv

Two General Methods for Population Pharmacokinetic Modeling: Non-Parametric Adaptive Grid and Non-Parametric Bayesian

Population pharmacokinetic (PK) modeling methods can be statistically classified as either parametric or nonparametric (NP). Each classification can be divided into maximum likelihood (ML) or Bayesian (B) approaches. In this paper we discuss the nonparametric case using both maximum likelihood and Bayesian approaches. We present two nonparametric methods for estimating the unknown joint population distribution of model parameter values in a pharmacokinetic/pharmacodynamic (PK/PD) dataset. The first method is the NP Adaptive Grid (NPAG). The second is the NP Bayesian (NPB) algorithm with a stick-breaking process to construct a Dirichlet prior. Our objective is to compare the performance of these two methods using a simulated PK/PD dataset. Our results showed excellent performance of NPAG and NPB in a realistically simulated PK study. This simulation allowed us to have benchmarks in the form of the true population parameters to compare with the estimates produced by the two methods, while incorporating challenges like unbalanced sample times and sample numbers as well as the ability to include the covariate of patient weight. We conclude that both NPML and NPB can be used in realistic PK/PD population analysis problems. The advantages of one versus the other are discussed in the paper. NPAG and NPB are implemented in R and freely available for download within the Pmetrics package from www.lapk.org.

preprint2011arXiv

A New Characterization of Elfving's Method for High Dimensional Computation

We give a new characterization of Elfving's (1952) method for computing c-optimal designs in k dimensions which gives explicit formulae for the k unknown optimal weights and k unknown signs in Elfving's characterization. This eliminates the need to search over these parameters to compute c-optimal designs, and thus reduces the computational burden from solving a family of optimization problems to solving a single optimization problem for the optimal finite support set. We give two illustrative examples: a high dimensional polynomial regression model and a logistic regression model, the latter showing that the method can be used for locally optimal designs in nonlinear models as well.

preprint2011arXiv

A Proof of the Bomber Problem's Spend-It-All Conjecture

The Bomber Problem concerns optimal sequential allocation of partially effective ammunition $x$ while under attack from enemies arriving according to a Poisson process over a time interval of length $t$. In the doubly-continuous setting, in certain regions of $(x,t)$-space we are able to solve the integral equation defining the optimal survival probability and find the optimal allocation function $K(x,t)$ exactly in these regions. As a consequence, we complete the proof of the "spend-it-all" conjecture of Bartroff et al. (2010b) which gives the boundary of the region where $K(x,t)=x$.

preprint2011arXiv

Efficient adaptive designs with mid-course sample size adjustment in clinical trials

Adaptive designs have been proposed for clinical trials in which the nuisance parameters or alternative of interest are unknown or likely to be misspecified before the trial. Whereas most previous works on adaptive designs and mid-course sample size re-estimation have focused on two-stage or group sequential designs in the normal case, we consider here a new approach that involves at most three stages and is developed in the general framework of multiparameter exponential families. Not only does this approach maintain the prescribed type I error probability, but it also provides a simple but asymptotically efficient sequential test whose finite-sample performance, measured in terms of the expected sample size and power functions, is shown to be comparable to the optimal sequential design, determined by dynamic programming, in the simplified normal mean case with known variance and prespecified alternative, and superior to the existing two-stage designs and also to adaptive group sequential designs when the alternative or nuisance parameters are unknown or misspecified.

preprint2011arXiv

Generalized Likelihood Ratio Statistics and Uncertainty Adjustments in Efficient Adaptive Design of Clinical Trials

A new approach to adaptive design of clinical trials is proposed in a general multiparameter exponential family setting, based on generalized likelihood ratio statistics and optimal sequential testing theory. These designs are easy to implement, maintain the prescribed Type I error probability, and are asymptotically efficient. Practical issues involved in clinical trials allowing mid-course adaptation and the large literature on this subject are discussed, and comparisons between the proposed and existing designs are presented in extensive simulation studies of their finite-sample performance, measured in terms of the expected sample size and power functions.

preprint2011arXiv

Incorporating Individual and Collective Ethics into Phase I Cancer Trial Designs

A general framework is proposed for Bayesian model-based designs of Phase I cancer trials, in which a general criterion for coherence (Cheung, 2005) of a design is also developed. This framework can incorporate both "individual" and "collective" ethics into the design of the trial. We propose a new design which minimizes a risk function composed of two terms, with one representing the individual risk of the current dose and the other representing the collective risk. The performance of this design, which is measured in terms of the accuracy of the estimated target dose at the end of the trial, the toxicity and overdose rates, and certain loss functions reflecting the individual and collective ethics, is studied and compared with existing Bayesian model-based designs and is shown to have better performance than existing designs.

preprint2011arXiv

Multistage tests of multiple hypotheses

Conventional multiple hypothesis tests use step-up, step-down, or closed testing methods to control the overall error rates. We will discuss marrying these methods with adaptive multistage sampling rules and stopping rules to perform efficient multiple hypothesis testing in sequential experimental designs. The result is a multistage step-down procedure that adaptively tests multiple hypotheses while preserving the family-wise error rate, and extends Holm's (1979) step-down procedure to the sequential setting, yielding substantial savings in sample size with small loss in power.

preprint2011arXiv

On Optimal Allocation of a Continuous Resource Using an Iterative Approach and Total Positivity

We study a class of optimal allocation problems, including the well-known Bomber Problem, with the following common probabilistic structure. An aircraft equipped with an amount~$x$ of ammunition is intercepted by enemy airplanes arriving according to a homogenous Poisson process over a fixed time duration~$t$. Upon encountering an enemy, the aircraft has the choice of spending any amount~$0\le y\le x$ of its ammunition, resulting in the aircraft's survival with probability equal to some known increasing function of $y$. Two different goals have been considered in the literature concerning the optimal amount~$K(x,t)$ of ammunition spent: (i)~Maximizing the probability of surviving for time~$t$, which is the so-called Bomber Problem, and (ii) maximizing the number of enemy airplanes shot down during time~$t$, which we call the Fighter Problem. Several authors have attempted to settle the following conjectures about the monotonicity of $K(x,t)$: [A] $K(x,t)$ is decreasing in $t$, [B] $K(x,t)$ is increasing in $x$, and [C] the amount~$x-K(x,t)$ held back is increasing in $x$. [A] and [C] have been shown for the Bomber Problem with discrete ammunition, while [B] is still an open question. In this paper we consider both time and ammunition continuous, and for the Bomber Problem prove [A] and [C], while for the Fighter we prove [A] and [C] for one special case and [B] and [C] for another. These proofs involve showing that the optimal survival probability and optimal number shot down are totally positive of order 2 ($\mbox{TP}_2$) in the Bomber and Fighter Problems, respectively. The $\mbox{TP}_2$ property is shown by constructing convergent sequences of approximating functions through an iterative operation which preserves $\mbox{TP}_2$ and other properties.

preprint2011arXiv

Optimal Multistage Sampling in a Boundary-Crossing Problem

Brownian motion with known positive drift is sampled in stages until it crosses a positive boundary $a$. A family of multistage samplers that control the expected overshoot over the boundary by varying the stage size at each stage is shown to be optimal for large $a$, minimizing a linear combination of overshoot and number of stages. Applications to hypothesis testing are discussed.

preprint2011arXiv

The Fighter Problem: Optimal Allocation of a Discrete Commodity

The Fighter problem with discrete ammunition is studied. An aircraft (fighter) equipped with $n$ anti-aircraft missiles is intercepted by enemy airplanes, the appearance of which follows a homogeneous Poisson process with known intensity. If $j$ of the $n$ missiles are spent at an encounter they destroy an enemy plane with probability $a(j)$, where $a(0) = 0 $ and $\{a(j)\}$ is a known, strictly increasing concave sequence, e.g., $a(j) = 1-q^j, \; \, 0 < q < 1$. If the enemy is not destroyed, the enemy shoots the fighter down with known probability $1-u$, where $0 \le u \le 1$. The goal of the fighter is to shoot down as many enemy airplanes as possible during a given time period $[0, T]$. Let $K (n, t)$ be the smallest optimal number of missiles to be used at a present encounter, when the fighter has flying time $t$ remaining and $n$ missiles remaining. Three seemingly obvious properties of $K(n, t)$ have been conjectured: [A] The closer to the destination, the more of the $n$ missiles one should use, [B] the more missiles one has, the more one should use, and [C] the more missiles one has, the more one should save for possible future encounters. We show that [C] holds for all $0 \le u \le 1$, that [A] and [B] hold for the "Invincible Fighter" ($u=1$), and that [A] holds but [B] fails for the "Frail Fighter" ($u=0$), the latter through a surprising counterexample.

preprint2010arXiv

Approximate Dynamic Programming and Its Applications to the Design of Phase I Cancer Trials

Optimal design of a Phase I cancer trial can be formulated as a stochastic optimization problem. By making use of recent advances in approximate dynamic programming to tackle the problem, we develop an approximation of the Bayesian optimal design. The resulting design is a convex combination of a "treatment" design, such as Babb et al.'s (1998) escalation with overdose control, and a "learning" design, such as Haines et al.'s (2003) $c$-optimal design, thus directly addressing the treatment versus experimentation dilemma inherent in Phase I trials and providing a simple and intuitive design for clinical use. Computational details are given and the proposed design is compared to existing designs in a simulation study. The design can also be readily modified to include a first stage that cautiously escalates doses similarly to traditional nonparametric step-up/down schemes, while validating the Bayesian parametric model for the efficient model-based design in the second stage.

preprint2010arXiv

The Spend-It-All Region and Small Time Results for the Continuous Bomber Problem

A problem of optimally allocating partially effective ammunition $x$ to be used on randomly arriving enemies in order to maximize an aircraft's probability of surviving for time~$t$, known as the Bomber Problem, was first posed by \citet{Klinger68}. They conjectured a set of apparently obvious monotonicity properties of the optimal allocation function $K(x,t)$. Although some of these conjectures, and versions thereof, have been proved or disproved by other authors since then, the remaining central question, that $K(x,t)$ is nondecreasing in~$x$, remains unsettled. After reviewing the problem and summarizing the state of these conjectures, in the setting where $x$ is continuous we prove the existence of a ``spend-it-all'' region in which $K(x,t)=x$ and find its boundary, inside of which the long-standing, unproven conjecture of monotonicity of~$K(\cdot,t)$ holds. A new approach is then taken of directly estimating~$K(x,t)$ for small~$t$, providing a complete small-$t$ asymptotic description of~$K(x,t)$ and the optimal probability of survival.