Source author record

Prakash Balachandran

Prakash Balachandran appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

4works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

4 published item(s)

preprint2016arXiv

On the Propagation of Low-Rate Measurement Error to Subgraph Counts in Large Networks

Our work in this paper is inspired by a statistical observation that is both elementary and broadly relevant to network analysis in practice -- that the uncertainty in approximating some true network graph $G=(V,E)$ by some estimated graph $\hat{G}=(V,\hat{E})$ manifests as errors in the status of (non)edges that must necessarily propagate to any estimates of network summaries $η(G)$ we seek. Motivated by the common practice of using plug-in estimates $η(\hat{G})$ as proxies for $η(G)$, our focus is on the problem of characterizing the distribution of the discrepancy $D=η(\hat{G}) - η(G)$, in the case where $η(\cdot)$ is a subgraph count. Specifically, we study the fundamental case where the statistic of interest is $|E|$, the number of edges in $G$. Our primary contribution in this paper is to show that in the empirically relevant setting of large graphs with low-rate measurement errors, the distribution of $D_E=|\hat{E}| - |E|$ is well-characterized by a Skellam distribution, when the errors are independent or weakly dependent. Under an assumption of independent errors, we are able to further show conditions under which this characterization is strictly better than that of an appropriate normal distribution. These results derive from our formulation of a general result, quantifying the accuracy with which the difference of two sums of dependent Bernoulli random variables may be approximated by the difference of two independent Poisson random variables, i.e., by a Skellam distribution. This general result is developed through the use of Stein's method, and may be of some general interest. We finish with a discussion of possible extension of our work to subgraph counts $η(G)$ of higher order.

preprint2014arXiv

Think Locally, Act Locally: The Detection of Small, Medium-Sized, and Large Communities in Large Networks

It is common in the study of networks to investigate meso-scale features to try to gain an understanding of network structure and function. For example, numerous algorithms have been developed to try to identify "communities," which are typically construed as sets of nodes with denser connections internally than with the remainder of a network. In this paper, we adopt a complementary perspective that "communities" are associated with bottlenecks of locally-biased dynamical processes that begin at seed sets of nodes, and we employ several different community-identification procedures (using diffusion-based and geodesic-based dynamics) to investigate community quality as a function of community size. Using several empirical and synthetic networks, we identify several distinct scenarios for ``size-resolved community structure'' that can arise in real (and realistic) networks. Depending on which scenario holds, one may or may not be able to successfully identify ``good'' communities in a given network, the manner in which different small communities fit together to form meso-scale network structures can be very different, and processes such as viral propagation and information diffusion can exhibit very different dynamics.In addition, our results suggest that, for many large realistic networks, the output of locally-biased methods that focus on communities that are centered around a given seed node might have better conceptual grounding and greater practical utility than the output of global community-detection methods. They also illustrate subtler structural properties that are important to consider in the development of better benchmark networks to test methods for community detection. [Note: Because of space limitations in the arXiv's abstract field, this is an abridged version of the paper's abstract.]

preprint2013arXiv

Exponential-type Inequalities Involving Ratios of the Modified Bessel Function of the First Kind and their Applications

The modified Bessel function of the first kind, $I_ν(x)$, arises in numerous areas of study, such as physics, signal processing, probability, statistics, etc. As such, there has been much interest in recent years in deducing properties of functionals involving $I_ν(x)$, in particular, of the ratio ${I_{ν+1}(x)}/{I_ν(x)}$, when $ν,x\geq 0$. In this paper we establish sharp upper and lower bounds on $H(ν,x)=\sum_{k=1}^{\infty} {I_{ν+k}(x)}/{I_ν(x)}$ for $ν,x\geq 0$ that appears as the complementary cumulative hazard function for a Skellam$(λ,λ)$ probability distribution in the statistical analysis of networks. Our technique relies on bounding existing estimates of ${I_{ν+1}(x)}/{I_ν(x)}$ from above and below by quantities with nicer algebraic properties, namely exponentials, to better evaluate the sum, while optimizing their rates in the regime when $ν+1\leq x$ in order to maintain their precision. We demonstrate the relevance of our results through applications, providing an improvement for the well-known asymptotic $\exp(-x)I_ν(x)\sim {1}/{\sqrt{2πx}}$ as $x\rightarrow \infty$, upper and lower bounding $\mathbb{P}\left[W=ν\right]$ for $W\sim Skellam(λ_1,λ_2)$, and deriving a novel concentration inequality on the $Skellam(λ,λ)$ probability distribution from above and below.

preprint2013arXiv

Inference of Network Summary Statistics Through Network Denoising

Consider observing an undirected network that is `noisy' in the sense that there are Type I and Type II errors in the observation of edges. Such errors can arise, for example, in the context of inferring gene regulatory networks in genomics or functional connectivity networks in neuroscience. Given a single observed network then, to what extent are summary statistics for that network representative of their analogues for the true underlying network? Can we infer such statistics more accurately by taking into account the noise in the observed network edges? In this paper, we answer both of these questions. In particular, we develop a spectral-based methodology using the adjacency matrix to `denoise' the observed network data and produce more accurate inference of the summary statistics of the true network. We characterize performance of our methodology through bounds on appropriate notions of risk in the $L^2$ sense, and conclude by illustrating the practical impact of this work on synthetic and real-world data.