Source author record

Dmitri V. Zaykin

Dmitri V. Zaykin appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
2topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2020arXiv

A new measure for the analysis of epidemiological associations: Cannabis use disorder examples

Analyses of population-based surveys are instrumental to research on prevention and treatment of mental and substance use disorders. Population-based data provides descriptive characteristics of multiple determinants of public health and are typically available to researchers as an annual data release. To provide trends in national estimates or to update the existing ones, a meta-analytical approach to year-by-year data is typically employed with ORs as effect sizes. However, if the estimated ORs exhibit different patterns over time, some normalization of ORs may be warranted. We propose a new normalized measure of effect size and derive an asymptotic distribution for the respective test statistic. The normalization constant is based on the maximum range of the standardized log(OR), for which we establish a connection to the Laplace Limit Constant. Furthermore, we propose to employ standardized log(OR) in a novel way to obtain accurate posterior inference. Through simulation studies, we show that our new statistic is more powerful than the traditional one for testing the hypothesis OR=1. We then applied it to the United States population estimates of co-occurrence of side effect problem-experiences (SEPE) among newly incident cannabis users, based on the the National Survey on Drug Use and Health (NSDUH), 2004-2014.

preprint2016arXiv

Assessment of P-value variability in the current replicability crisis

Increased availability of data and accessibility of computational tools in recent years have created unprecedented opportunities for scientific research driven by statistical analysis. Inherent limitations of statistics impose constrains on reliability of conclusions drawn from data but misuse of statistical methods is a growing concern. Significance, hypothesis testing and the accompanying P-values are being scrutinized as representing most widely applied and abused practices. One line of critique is that P-values are inherently unfit to fulfill their ostensible role as measures of scientific hypothesis's credibility. It has also been suggested that while P-values may have their role as summary measures of effect, researchers underappreciate the degree of randomness in the P-value. High variability of P-values would suggest that having obtained a small P-value in one study, one is, nevertheless, likely to obtain a much larger P-value in a similarly powered replication study. Thus, "replicability of P-value" is itself questionable. To characterize P-value variability one can use prediction intervals whose endpoints reflect the likely spread of P-values that could have been obtained by a replication study. Unfortunately, the intervals currently in use, the P-intervals, are based on unrealistic implicit assumptions. Namely, P-intervals are constructed with the assumptions that imply substantial chances of encountering large values of effect size in an observational study, which leads to bias. As an alternative to P-intervals, we develop a method that gives researchers flexibility by providing them with the means to control these assumptions. Unlike endpoints of P-intervals, endpoints of our intervals are directly interpreted as probabilistic bounds for replication P-values and are resistant to selection bias contingent upon approximate prior knowledge of the effect size distribution.

preprint2016arXiv

The more you test, the more you find: Smallest P-values become increasingly enriched with real findings as more tests are conducted

Increasing accessibility of data to researchers makes it possible to conduct massive amounts of statistical testing. Rather than follow a carefully crafted set of scientific hypotheses with statistical analysis, researchers can now test many possible relations and let P-values or other statistical summaries generate hypotheses for them. Genetic epidemiology field is an illustrative case in this paradigm shift. Driven by technological advances, testing a handful of genetic variants in relation to a health outcome has been abandoned in favor of agnostic screening of the entire genome, followed by selection of top hits, e.g., by selection of genetic variants with the smallest association P-values. At the same time, nearly total lack of replication of claimed associations that has been shaming the field turned to a flow of reports whose findings have been robustly replicating. Researchers may have adopted better statistical practices by learning from past failures, but we suggest that a steep increase in the amount of statistical testing itself is an important factor. Regardless of whether statistical significance has been reached, an increased number of tested hypotheses leads to enrichment of smallest P-values with genuine associations. In this study, we quantify how the expected proportion of genuine signals (EPGS) among top hits changes with an increasing number of tests. When the rate of occurrence of genuine signals does not decrease too sharply to zero as more tests are performed, the smallest P-values are increasingly more likely to represent genuine associations in studies with more tests.