Source author record

Ting Yan

Ting Yan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

14works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

14 published item(s)

preprint2026arXiv

Optimal estimators and tests for reciprocal effects

The $p_1$ model plays a fundamental role in modeling directed networks, where the reciprocal effect parameter $ρ$ is of special interest in practice. However, due to nonlinear factors in this model, how to estimate $ρ$ efficiently is a long-standing open problem. We tackle the problem by the cycle count approach. The challenge is, due to the nonlinear factors in the model, for any given type of generalized cycles, the expected count is a complicated function of many parameters in the model, so it is unclear how to use cycle counts to estimate $ρ$. However, somewhat surprisingly, we discover that, among many types of generalized cycles with the same length, we can carefully pick a pair of them such that in the ratio between the expected cycle counts of the two types, the non-linear factors cancel out nicely with each other, and as a result, the ratio equals to $\mathrm{exp}(ρ)$ exactly. Therefore, though the expected count of cycles of any type is not tractable, the ratio between the expected cycle counts of a (carefully chosen) pair of generalized cycles may have an utterly simple form. We study to what extent such pairs exist, and use our discovery to derive both an estimate for $ρ$ and a testing procedure for testing $ρ= ρ_0$. In a setting where we allow a wide range of reciprocal effects and a wide variety of network sparsity and degree heterogeneity, we show that our estimator achieves the optimal rate and our test achieves the optimal phase transition. Technically, first, motivated by what we observe on real networks, we do not want to impose strong conditions on reciprocal effects, network sparsity, and degree heterogeneity. Second, our proposed statistic is a type of $U$-statistic, the analysis of which involves complex combinatorics and is error-prone. For these reasons, our analysis is long and delicate.

preprint2026arXiv

Triple-dyad ratio estimation for the $p_1$ model

Although the $p_1$ model was proposed 40 years ago, little progress has been made to address asymptotic theories in this model, that is, neither consistency of the maximum likelihood estimator (MLE) nor other parameter estimation with statistical guarantees is understood. This problem has been acknowledged as a long-standing open problem. To address it, we propose a novel parametric estimation method based on the ratios of the sum of a sequence of triple-dyad indicators to another one, where a triple-dyad indicator means the product of three dyad indicators. Our proposed estimators, called \emph{triple-dyad ratio estimator}, have explicit expressions and can be scaled to very large networks with millions of nodes. We establish the consistency and asymptotic normality of the triple-dyad ratio estimator when the number of nodes reaches infinity. Based on the asymptotic results, we develop a test statistic for evaluating whether is a reciprocity effect in directed networks. The estimators for the density and reciprocity parameters contain bias terms, where analytical bias correction formulas are proposed to make valid inference. Numerical studies demonstrate the findings of our theories and show that the estimator is comparable to the MLE in large networks.

preprint2023arXiv

A degree-corrected Cox model for dynamic networks

Continuous time network data have been successfully modeled by multivariate counting processes, in which the intensity function is characterized by covariate information. However, degree heterogeneity has not been incorporated into the model which may lead to large biases for the estimation of homophily effects. In this paper, we propose a degree-corrected Cox network model to simultaneously analyze the dynamic degree heterogeneity and homophily effects for continuous time directed network data. Since each node has individual-specific in- and out-degree effects in the model, the dimension of the time-varying parameter vector grows with the number of nodes, which makes the estimation problem non-standard. We develop a local estimating equations approach to estimate unknown time-varying parameters, and establish consistency and asymptotic normality of the proposed estimators by using the powerful martingale process theories. We further propose test statistics to test for trend and degree heterogeneity in dynamic networks. Simulation studies are provided to assess the finite sample performance of the proposed method and a real data analysis is used to illustrate its practical utility.

preprint2022arXiv

Asymptotic theory in network models with covariates and a growing number of node parameters

We propose a general model that jointly characterizes degree heterogeneity and homophily in weighted, undirected networks. We present a moment estimation method using node degrees and homophily statistics. We establish consistency and asymptotic normality of our estimator using novel analysis. We apply our general framework to three applications, including both exponential family and non-exponential family models. Comprehensive numerical studies and a data example also demonstrate the usefulness of our method.

preprint2020arXiv

Asymptotic Theory for Differentially Private Generalized $β$-models with Parameters Increasing

Modelling edge weights play a crucial role in the analysis of network data, which reveals the extent of relationships among individuals. Due to the diversity of weight information, sharing these data has become a complicated challenge in a privacy-preserving way. In this paper, we consider the case of the non-denoising process to achieve the trade-off between privacy and weight information in the generalized $β$-model. Under the edge differential privacy with a discrete Laplace mechanism, the Z-estimators from estimating equations for the model parameters are shown to be consistent and asymptotically normally distributed. The simulations and a real data example are given to further support the theoretical results.

preprint2019arXiv

Directed Networks with a Differentially Private Bi-degree Sequence

Although a lot of approaches are developed to release network data with a differentially privacy guarantee, inference using noisy data in many network models is still unknown or not properly explored. In this paper, we release the bi-degree sequences of directed networks using the Laplace mechanism and use the $p_0$ model for inferring the degree parameters. The $p_0$ model is an exponential random graph model with the bi-degree sequence as its exclusively sufficient statistic. We show that the estimator of the parameter without the denoised process is asymptotically consistent and normally distributed. This is contrast sharply with some known results that valid inference such as the existence and consistency of the estimator needs the denoised process. Along the way, a new phenomenon is revealed in which an additional variance factor appears in the asymptotic variance of the estimator when the noise becomes large. Further, we propose an efficient algorithm for finding the closet point lying in the set of all graphical bi-degree sequences under the global $L_1$ optimization problem. Numerical studies demonstrate our theoretical findings.

preprint2016arXiv

Asymptotics in directed exponential random graph models with an increasing bi-degree sequence

Although asymptotic analyses of undirected network models based on degree sequences have started to appear in recent literature, it remains an open problem to study statistical properties of directed network models. In this paper, we provide for the first time a rigorous analysis of directed exponential random graph models using the in-degrees and out-degrees as sufficient statistics with binary as well as continuous weighted edges. We establish the uniform consistency and the asymptotic normality for the maximum likelihood estimate, when the number of parameters grows and only one realized observation of the graph is available. One key technique in the proofs is to approximate the inverse of the Fisher information matrix using a simple matrix with high accuracy. Numerical studies confirm our theoretical findings.

preprint2015arXiv

Invisible Active Galactic Nuclei. II Radio Morphologies & Five New HI 21 cm Absorption Line Detections

We have selected a sample of 80 candidates for obscured radio-loud active galactic nuclei and presented their basic optical/near-infrared (NIR) properties in Paper 1. In this paper, we present both high-resolution radio continuum images for all of these sources and HI 21cm absorption spectroscopy for a few selected sources in this sample. A-configuration 4.9 and 8.5 GHz VLA continuum observations find that 52 sources are compact or have substantial compact components with size <0.5" and flux density >0.1 Jy at 4.9 GHz. The most compact 36 sources were then observed with the VLBA at 1.4 GHz. One definite and 10 candidate Compact Symmetric Objects (CSOs) are newly identified, a detection rate of CSOs ~3 times higher than the detection rate previously found in purely flux-limited samples. Based on possessing compact components with high flux densities, 60 of these sources are good candidates for absorption-line searches. Twenty seven sources were observed for HI 21cm absorption at their photometric or spectroscopic redshifts with only 6 detections made (one detection is tentative). However, five of these were from a small subset of six CSOs with pure galaxy optical/NIR spectra and for which accurate spectroscopic redshifts place the redshifted 21cm line in a RFI-free spectral window. It is likely that the presence of ubiquitous RFI and the absence of accurate spectroscopic redshifts preclude HI detections in similar sources (only one detection out of the remaining 22 sources observed, 14 of which have only photometric redshifts). Future searches for highly-redshifted HI and molecular absorption can easily find more distant CSOs among bright, blank field' radio sources but will be severely hampered by an inability to determine accurate spectroscopic redshifts for them due to their lack of rest-frame UV continuum.

preprint2014arXiv

Approximating the inverse of a balanced symmetric matrix with positive elements

For an $n\times n$ balanced symmetric matrix $T=(t_{i,j})$ with positive elements satisfying $t_{i,i}= \sum_{j\neq i} t_{i,j}$ and certain bounding conditions, we propose to use the matrix $S=(s_{i,j})$ to approximate its inverse, where $s_{i,j}=δ_{i,j}/t_{i,i}-1/t_{..}$, $δ_{i,j}$ is the Kronecker delta function, and $t_{..}=\sum_{i,j=1 }^{n}(1-δ_{i,j}) t_{i,j}$. An explicit bound on the approximation error is obtained, showing that the inverse is well approximated to order $1/(n-1)^2$ uniformly.

preprint2014arXiv

Asymptotic normality in the maximum entropy models on graphs with an increasing number of parameters

Maximum entropy models, motivated by applications in neuron science, are natural generalizations of the $β$-model to weighted graphs. Similar to the $β$-model, each vertex in maximum entropy models is assigned a potential parameter, and the degree sequence is the natural sufficient statistic. Hillar and Wibisono (2013) has proved the consistency of the maximum likelihood estimators. In this paper, we further establish the asymptotic normality for any finite number of the maximum likelihood estimators in the maximum entropy models with three types of edge weights, when the total number of parameters goes to infinity. Simulation studies are provided to illustrate the asymptotic results.

preprint2014arXiv

Ranking in the generalized Bradley-Terry models when the strong connection condition fails

For nonbalanced paired comparisons, a wide variety of ranking methods have been proposed. One of the best popular methods is the Bradley-Terry model in which the ranking of a set of objects is decided by the maximum likelihood estimates (MLEs) of merits parameters. However, the existence of MLE for the Bradley-Terry model and its generalized models to allow for tied observation or home-field advantage or both to occur, crucially depends on the strong connection condition on the directed graph constructed by a win-loss matrix. When this condition fails, the MLE does not exist and hence there is no solution of ranking. In this paper, we propose an improved version of the $\varepsilon$ singular perturbation proposed by Conner and Grant (2000), to address this problem and extend it to the generalized Bradley-Terry models. Some necessary and sufficient conditions for the existence and uniqueness of the penalized MLEs for these generalized Bradley-Terry-$\varepsilon$ models are derived. Numerical studies show that the ranking is robust to the different $\varepsilon$. We apply the proposed methods to the data of the 2008 NFL regular season.

preprint2013arXiv

A central limit theorem in the $β$-model for undirected random graphs with a diverging number of vertices

Chatterjee, Diaconis and Sly (2011) recently established the consistency of the maximum likelihood estimate in the $β$-model when the number of vertices goes to infinity. By approximating the inverse of the Fisher information matrix, we obtain its asymptotic normality under mild conditions. Simulation studies and a data example illustrate the theoretical results.

preprint2012arXiv

Grouped sparse paired comparisons in the Bradley-Terry model

In a wide class of paired comparisons, especially in the sports games, in which all subjects are divided into several groups, the intragroup comparisons are dense and the intergroup comparisons are sparse. Typical examples include the NFL regular season. Motivated by these situations, we propose group sparsity for paired comparisons and show the consistency and asymptotical normality of the maximum likelihood estimate in the Bradley-Terry model when the number of parameters goes to infinity in this paper. Simulations are carried out to illustrate the group sparsity and asymptotical results.

preprint2012arXiv

The influence of fallback discs on the spectral and timing properties of neutron stars

Fallback discs around neutron stars (NSs) are believed to be an expected outcome of supernova explosions. Here we investigate the consequences of such a common outcome for the timing and spectral properties of the associated NS population, using Monte Carlo population synthesis models. We find that the long-term torque exerted by the fallback disc can substantially influence the late-time period distribution, but with quantitative differences which depend on whether the initial spin distribution is dominated by slow or fast pulsars. For the latter, a single-peaked initial spin distribution becomes bimodal at later times. Timing ages tend to underestimate the real age of older pulsars, and overestimate the age of younger ones. Braking indices cluster in the range 1.5 <~ n <~ 3 for slow-born pulsars, and -0.5 <~ n <~ 5 for fast-born pulsars, with the younger objects found predominantly below n <~ 3. Large values of n, while not common, are possible, and associated with torque transitions in the NS+disc system. The 0.1-10 keV thermal luminosity of the NS+disc system is found to be generally dominated by the disc emission at early times, t <~ 10^3 yr, but this declines faster than the thermal surface emission of the NS. Depending on the initial parameters, there can be occasional periods in which some NSs switch from the propeller to the accretion phase, increasing their luminosity up to the Eddington limit for ~ 10^3-10^4 years.