Source author record

Jingchen Liu

Jingchen Liu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

20works
12topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

20 published item(s)

preprint2020arXiv

ProcData: An R Package for Process Data Analysis

Process data refer to data recorded in the log files of computer-based items. These data, represented as timestamped action sequences, keep track of respondents' response processes of solving the items. Process data analysis aims at enhancing educational assessment accuracy and serving other assessment purposes by utilizing the rich information contained in response processes. The R package ProcData presented in this article is designed to provide tools for processing, describing, and analyzing process data. We define an S3 class "proc" for organizing process data and extend generic methods summary and print for class "proc". Two feature extraction methods for process data are implemented in the package for compressing information in the irregular response processes into regular numeric vectors. ProcData also provides functions for fitting and making predictions from a neural-network-based sequence model. These functions call relevant functions in package keras for constructing and training neural networks. In addition, several response process generators and a real dataset of response processes of the climate control item in the 2012 Programme for International Student Assessment are included in the package.

preprint2020arXiv

Statistical Analysis of Multi-Relational Network Recovery

In this paper, we develop asymptotic theories for a class of latent variable models for large-scale multi-relational networks. In particular, we establish consistency results and asymptotic error bounds for the (penalized) maximum likelihood estimators when the size of the network tends to infinity. The basic technique is to develop a non-asymptotic error bound for the maximum likelihood estimators through large deviations analysis of random fields. We also show that these estimators are nearly optimal in terms of minimax risk.

preprint2020arXiv

Subtask Analysis of Process Data Through a Predictive Model

Response process data collected from human-computer interactive items contain rich information about respondents' behavioral patterns and cognitive processes. Their irregular formats as well as their large sizes make standard statistical tools difficult to apply. This paper develops a computationally efficient method for exploratory analysis of such process data. The new approach segments a lengthy individual process into a sequence of short subprocesses to achieve complexity reduction, easy clustering and meaningful interpretation. Each subprocess is considered a subtask. The segmentation is based on sequential action predictability using a parsimonious predictive model combined with the Shannon entropy. Simulation studies are conducted to assess performance of the new methods. We use the process data from PIAAC 2012 to demonstrate how exploratory analysis of process data can be done with the new approach.

preprint2016arXiv

A Fused Latent and Graphical Model for Multivariate Binary Data

We consider modeling, inference, and computation for analyzing multivariate binary data. We propose a new model that consists of a low dimensional latent variable component and a sparse graphical component. Our study is motivated by analysis of item response data in cognitive assessment and has applications to many disciplines where item response data are collected. Standard approaches to item response data in cognitive assessment adopt the multidimensional item response theory (IRT) models. However, human cognition is typically a complicated process and thus may not be adequately described by just a few factors. Consequently, a low-dimensional latent factor model, such as the multidimensional IRT models, is often insufficient to capture the structure of the data. The proposed model adds a sparse graphical component that captures the remaining ad hoc dependence. It reduces to a multidimensional IRT model when the graphical component becomes degenerate. Model selection and parameter estimation are carried out simultaneously through construction of a pseudo-likelihood function and properly chosen penalty terms. The convexity of the pseudo-likelihood function allows us to develop an efficient algorithm, while the penalty terms generate a low-dimensional latent component and a sparse graphical structure. Desirable theoretical properties are established under suitable regularity conditions. The method is applied to the revised Eysenck's personality questionnaire, revealing its usefulness in item analysis. Simulation results are reported that show the new method works well in practical situations.

preprint2016arXiv

A Multilevel Approach towards Unbiased Sampling of Random Elliptic Partial Differential Equations

Partial differential equation is a powerful tool to characterize various physics systems. In practice, measurement errors are often present and probability models are employed to account for such uncertainties. In this paper, we present a Monte Carlo scheme that yields unbiased estimators for expectations of random elliptic partial differential equations. This algorithm combines multilevel Monte Carlo [Giles, 2008] and a randomization scheme proposed by [Rhee and Glynn, 2012, Rhee and Glynn, 2013]. Furthermore, to obtain an estimator with both finite variance and finite expected computational cost, we employ higher order approximations.

preprint2016arXiv

Chernoff Index for Cox Test of Separate Parametric Families

The asymptotic efficiency of a generalized likelihood ratio test proposed by Cox is studied under the large deviations framework for error probabilities developed by Chernoff. In particular, two separate parametric families of hypotheses are considered [Cox, 1961, 1962]. The significance level is set such that the maximal type I and type II error probabilities for the generalized likelihood ratio test decay exponentially fast with the same rate. We derive the analytic form of such a rate that is also known as the Chernoff index [Chernoff, 1952], a relative efficiency measure when there is no preference between the null and the alternative hypotheses. We further extend the analysis to approximate error probabilities when the two families are not completely separated. Discussions are provided concerning the implications of the present result on model selection.

preprint2016arXiv

Sequential Hypothesis Test with Online Usage-Constrained Sensor Selection

This work investigates the sequential hypothesis testing problem with online sensor selection and sensor usage constraints. That is, in a sensor network, the fusion center sequentially acquires samples by selecting one "most informative" sensor at each time until a reliable decision can be made. In particular, the sensor selection is carried out in the online fashion since it depends on all the previous samples at each time. Our goal is to develop the sequential test (i.e., stopping rule and decision function) and sensor selection strategy that minimize the expected sample size subject to the constraints on the error probabilities and sensor usages. To this end, we first recast the usage-constrained formulation into a Bayesian optimal stopping problem with different sampling costs for the usage-contrained sensors. The Bayesian problem is then studied under both finite- and infinite-horizon setups, based on which, the optimal solution to the original usage-constrained problem can be readily established. Moreover, by capitalizing on the structures of the optimal solution, a lower bound is obtained for the optimal expected sample size. In addition, we also propose algorithms to approximately evaluate the parameters in the optimal sequential test so that the sensor usage and error probability constraints are satisfied. Finally, numerical experiments are provided to illustrate the theoretical findings, and compare with the existing methods.

preprint2014arXiv

Efficient rare event simulation for failure problems in random media

In this paper we study rare events associated to solutions of elliptic partial differential equations with spatially varying random coefficients. The random coefficients follow the lognormal distribution, which is determined by a Gaussian process. This model is employed to study the failure problem of elastic materials in random media in which the failure is characterized by that the strain field exceeds a high threshold. We propose an efficient importance sampling scheme to compute small failure probabilities in the high threshold limit. The change of measure in our scheme is parametrized by two density functions. The efficiency of the importance sampling scheme is validated by numerical examples.

preprint2014arXiv

On the conditional distributions and the efficient simulations of exponential integrals of Gaussian random fields

In this paper, we consider the extreme behavior of a Gaussian random field $f(t)$ living on a compact set $T$. In particular, we are interested in tail events associated with the integral $\int_Te^{f(t)}\,dt$. We construct a (non-Gaussian) random field whose distribution can be explicitly stated. This field approximates the conditional Gaussian random field $f$ (given that $\int_Te^{f(t)}\,dt$ exceeds a large value) in total variation. Based on this approximation, we show that the tail event of $\int_Te^{f(t)}\,dt$ is asymptotically equivalent to the tail event of $\sup_Tγ(t)$ where $γ(t)$ is a Gaussian process and it is an affine function of $f(t)$ and its derivative field. In addition to the asymptotic description of the conditional field, we construct an efficient Monte Carlo estimator that runs in polynomial time of $\log b$ to compute the probability $P(\int_Te^{f(t)}\,dt>b)$ with a prescribed relative accuracy.

preprint2014arXiv

Total variation approximations and conditional limit theorems for multivariate regularly varying random walks conditioned on ruin

We study a new technique for the asymptotic analysis of heavy-tailed systems conditioned on large deviations events. We illustrate our approach in the context of ruin events of multidimensional regularly varying random walks. Our approach is to study the Markov process described by the random walk conditioned on hitting a rare target set. We construct a Markov chain whose transition kernel can be evaluated directly from the increment distribution of the associated random walk. This process is shown to approximate the conditional process of interest in total variation. Then, by analyzing the approximating process, we are able to obtain asymptotic conditional joint distributions and a conditional functional central limit theorem of several objects such as the time until ruin, the whole random walk prior to ruin, and the overshoot on the target set. These types of joint conditional limit theorems have been obtained previously in the literature only in the one dimensional case. In addition to using different techniques, our results include features that are qualitatively different from the one dimensional case. For instance, the asymptotic conditional law of the time to ruin is no longer purely Pareto as in the multidimensional case.

preprint2013arXiv

Extreme Analysis of a Non-convex and Nonlinear Functional of Gaussian Processes -- On the Tail Asymptotics of Random Ordinary Differential Equations

In this paper, we consider a stochastic system described by a differential equation admitting a spatially varying random coefficient. The differential equation has been employed to model various static physics systems such as elastic deformation, water flow, electric-magnetic fields, temperature distribution, etc. A random coefficient is introduced to account for the system's uncertainty and/or imperfect measurements. This random coefficient is described by a Gaussian process (the input process) and thus the solution to the differential equation (under certain boundary conditions) is a complexed functional of the input Gaussian process. In this paper, we focus the analysis on the one-dimensional case and derive asymptotic approximations of the tail probabilities of the solution to the equation that has various physics interpretations under different contexts. This analysis rests on the literature of the extreme analysis of Gaussian processes (such as the tail approximations of the supremum) and extends the analysis to more complexed functionals.

preprint2013arXiv

Rare-event Simulation and Efficient Discretization for the Supremum of Gaussian Random Fields

In this paper, we consider a classic problem concerning the high excursion probabilities of a Gaussian random field $f$ living on a compact set $T$. We develop efficient computational methods for the tail probabilities $P(\sup_T f(t) > b)$ and the conditional expectations $E(Γ(f) | \sup_T f(t) > b)$ as $b\rightarrow \infty$. For each $\varepsilon$ positive, we present Monte Carlo algorithms that run in \emph{constant} time and compute the interesting quantities with $\varepsilon$ relative error for arbitrarily large $b$. The efficiency results are applicable to a large class of Hölder continuous Gaussian random fields. Besides computations, the proposed change of measure and its analysis techniques have several theoretical and practical indications in the asymptotic analysis of extremes of Gaussian random fields.

preprint2013arXiv

Theory of self-learning $Q$-matrix

Cognitive assessment is a growing area in psychological and educational measurement, where tests are given to assess mastery/deficiency of attributes or skills. A key issue is the correct identification of attributes associated with items in a test. In this paper, we set up a mathematical framework under which theoretical properties may be discussed. We establish sufficient conditions to ensure that the attributes required by each item are learnable from the data.

preprint2012arXiv

Efficient Monte Carlo for high excursions of Gaussian random fields

Our focus is on the design and analysis of efficient Monte Carlo methods for computing tail probabilities for the suprema of Gaussian random fields, along with conditional expectations of functionals of the fields given the existence of excursions above high levels, b. Naïve Monte Carlo takes an exponential, in b, computational cost to estimate these probabilities and conditional expectations for a prescribed relative accuracy. In contrast, our Monte Carlo procedures achieve, at worst, polynomial complexity in b, assuming only that the mean and covariance functions are Hölder continuous. We also explain how to fine tune the construction of our procedures in the presence of additional regularity, such as homogeneity and smoothness, in order to further improve the efficiency.

preprint2012arXiv

On the Stationary Distribution of Iterative Imputations

Iterative imputation, in which variables are imputed one at a time each given a model predicting from all the others, is a popular technique that can be convenient and flexible, as it replaces a potentially difficult multivariate modeling problem with relatively simple univariate regressions. In this paper, we begin to characterize the stationary distributions of iterative imputations and their statistical properties. More precisely, when the conditional models are compatible (defined in the text), we give a set of sufficient conditions under which the imputation distribution converges in total variation to the posterior distribution of a Bayesian model. When the conditional models are incompatible but are valid, we show that the combined imputation estimator is consistent.

preprint2012arXiv

Tail approximations of integrals of Gaussian random fields

This paper develops asymptotic approximations of $P(\int_Te^{f(t)}\,dt>b)$ as $b\rightarrow\infty$ for a homogeneous smooth Gaussian random field, $f$, living on a compact $d$-dimensional Jordan measurable set $T$. The integral of an exponent of a Gaussian random field is an important random variable for many generic models in spatial point processes, portfolio risk analysis, asset pricing and so forth. The analysis technique consists of two steps: 1. evaluate the tail probability $P(\int_Ξe^{f(t)}\,dt>b)$ over a small domain $Ξ$ depending on $b$, where $\operatorname {mes}(Ξ)\rightarrow0$ as $b\rightarrow \infty$ and $\operatorname {mes}(\cdot)$ is the Lebesgue measure; 2. with $Ξ$ appropriately chosen, we show that $P(\int_Te^{f(t)}\,dt>b)=(1+o(1))\operatorname{mes}(T)\times \operatorname{mes}^{-1}(Ξ)P(\int_Ξe^{f(t)}\,dt>b)$.

preprint2011arXiv

Learning Item-Attribute Relationship in Q-Matrix Based Diagnostic Classification Models

Recent surge of interests in cognitive assessment has led to the developments of novel statistical models for diagnostic classification. Central to many such models is the well-known Q-matrix, which specifies the item-attribute relationship. This paper proposes a principled estimation procedure for the Q-matrix and related model parameters. Desirable theoretic properties are established through large sample analysis. The proposed method also provides a platform under which important statistical issues, such as hypothesis testing and model selection, can be addressed.

preprint2011arXiv

Some Asymptotic Results of Gaussian Random Fields with Varying Mean Functions and the Associated Processes

In this paper, we derive tail approximations of integrals of exponential functions of Gaussian random fields with varying mean functions and approximations of the associated point processes. This study is motivated naturally by multiple applications such as hypothesis testing for spatial models, study of the distribution of Bayesian marginal likelihood and Bayes factor, and financial applications.

preprint2010arXiv

Efficient Simulation and Conditional Functional Limit Theorems for Ruinous Heavy-tailed Random Walks

The contribution of this paper is to introduce change of measure based techniques for the rare-event analysis of heavy-tailed stochastic processes. Our changes-of-measure are parameterized by a family of distributions admitting a mixture form. We exploit our methodology to achieve two types of results. First, we construct Monte Carlo estimators that are strongly efficient (i.e. have bounded relative mean squared error as the event of interest becomes rare). These estimators are used to estimate both rare-event probabilities of interest and associated conditional expectations. We emphasize that our techniques allow us to control the expected termination time of the Monte Carlo algorithm even if the conditional expected stopping time (under the original distribution) given the event of interest is infinity -- a situation that sometimes occurs in heavy-tailed settings. Second, the mixture family serves as a good approximation (in total variation) of the conditional distribution of the whole process given the rare event of interest. The convenient form of the mixture family allows us to obtain, as a corollary, functional conditional central limit theorems that extend classical results in the literature. We illustrate our methodology in the context of the ruin probability $P(\sup_n S_n >b)$, where $S_n$ is a random walk with heavy-tailed increments that have negative drift. Our techniques are based on the use of Lyapunov inequalities for variance control and termination time. The conditional limit theorems combine the application of Lyapunov bounds with coupling arguments.