Source author record

Linyuan Lu

Linyuan Lu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

38works
11topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

38 published item(s)

preprint2022arXiv

Mean field information Hessian matrices on graphs

We derive mean-field information Hessian matrices on finite graphs. The "information" refers to entropy functions on the probability simplex. And the "mean-field" means nonlinear weight functions of probabilities supported on graphs. These two concepts define a mean-field optimal transport type metric. In this metric space, we first derive Hessian matrices of energies on graphs, including linear, interaction energies, entropies. We name their smallest eigenvalues as mean-field Ricci curvature bounds on graphs. We next provide examples on two-point spaces and graph products. We last present several applications of the proposed matrices. E.g., we prove discrete Costa's entropy power inequalities on a two-point space.

preprint2020arXiv

On Hamiltonian Berge cycles in $3$-uniform hypergraphs

Given a set $R$, a hypergraph is $R$-uniform if the size of every hyperedge belongs to $R$. A hypergraph $\mathcal{H}$ is called \textit{covering} if every vertex pair is contained in some hyperedge in $\mathcal{H}$. In this note, we show that every covering $[3]$-uniform hypergraph on $n\geq 6$ vertices contains a Berge cycle $C_s$ for any $3\leq s\leq n$. As an application, we determine the maximum Lagrangian of $k$-uniform Berge-$C_{t}$-free hypergraphs and Berge-$P_{t}$-free hypergraphs.

preprint2016arXiv

An Upper Bound on Burning Number of Graphs

The burning number $b(G)$ of a graph $G$ was introduced by Bonato, Janssen, and Roshanbin [Lecture Notes in Computer Science 8882 (2014)] for measuring the speed of the spread of contagion in a graph. They proved for any connected graph $G$ of order $n$, $b(G)\leq 2\lceil \sqrt{n} \rceil-1$, and conjectured that $b(G)\leq \lceil \sqrt{n} \rceil$. In this paper, we proved $b(G)\leq \lceil\frac{-3+\sqrt{24n+33}}{4}\rceil$, which is roughly $\frac{\sqrt{6}}{2}\sqrt{n}$. We also settled the following conjecture of Bonato-Janssen-Roshanbin: $b(G)b(\bar G)\leq n+4$ provided both $G$ and $\bar G$ are connected.

preprint2015arXiv

A combinatorial identity on Galton-Watson process

Let $f(m,c)=\sum_{k=0}^{\infty} (km+1)^{k-1} c^k e^{-c(km+1)/m} / (m^kk!)$. For any positive integer $m$ and positive real $c$, the identity $f(m,c)=f(1,c)^{1/m}$ arises in the random graph theory. In this paper, we present two elementary proofs of this identity: a pure combinatorial proof and a power-serial proof. We also proved that this identity holds for any positive reals $m$ and $c$.

preprint2015arXiv

Effective spreading from multiple leaders identified by percolation in social networks

Social networks constitute a new platform for information propagation, but its success is crucially dependent on the choice of spreaders who initiate the spreading of information. In this paper, we remove edges in a network at random and the network segments into isolated clusters. The most important nodes in each cluster then form a group of influential spreaders, such that news propagating from them would lead to an extensive coverage and minimal redundancy. The method well utilizes the similarities between the pre-percolated state and the coverage of information propagation in each social cluster to obtain a set of distributed and coordinated spreaders. Our tests on the Facebook networks show that this method outperforms conventional methods based on centrality. The suggested way of identifying influential spreaders thus sheds light on a new paradigm of information propagation on social networks.

preprint2014arXiv

A new asymptotic enumeration technique: the Lovasz Local Lemma

Our previous paper applied a lopsided version of the Lovász Local Lemma that allows negative dependency graphs to the space of random injections from an $m$-element set to an $n$-element set. Equivalently, the same story can be told about the space of random matchings in $K_{n,m}$. Now we show how the cited version of the Lovász Local Lemma applies to the space of random matchings in $K_{2n}$. We also prove tight upper bounds that asymptotically match the lower bound given by the Lovász Local Lemma. As a consequence, we give new proofs to results on the enumeration of $d$-regular graphs. The tight upper bounds can be modified to the space of matchings in $K_{n,m}$, where they yield as application asymptotic formulas for permutation and Latin rectangle enumeration problems. The strength of the method is shown by a new result: enumeration of graphs by degree sequence or bipartite degree sequence and girth. As another application, we provide a new proof to the classical probabilistic result of Erd\H os that showed the existence of graphs with arbitrary large girth and chromatic number. If the degree sequence satisfies some mild conditions, almost all graphs with this degree sequence and prescribed girth have high chromatic number.

preprint2014arXiv

Computing Diffusion State Distance using Green's Function and Heat Kernel on Graphs

The diffusion state distance (DSD) was introduced by Cao-Zhang-Park-Daniels-Crovella-Cowen-Hescott [{\em PLoS ONE, 2013}] to capture functional similarity in protein-protein interaction networks. They proved the convergence of DSD for non-bipartite graphs. In this paper, we extend the DSD to bipartite graphs using lazy-random walks and consider the general $L_q$-version of DSD. We discovered the connection between the DSD $L_q$-distance and Green's function, which was studied by Chung and Yau [{\em J. Combinatorial Theory (A), 2000}]. Based on that, we computed the DSD $L_q$-distance for Paths, Cycles, Hypercubes, as well as random graphs $G(n,p)$ and $G(w_1,..., w_n)$. We also examined the DSD distances of two biological networks.

preprint2014arXiv

Connected Hypergraphs with Small Spectral Radius

In 1970 Smith classified all connected graphs with the spectral radius at most $2$. Here the spectral radius of a graph is the largest eigenvalue of its adjacency matrix. Recently, the definition of spectral radius has been extended to $r$-uniform hypergraphs. In this paper, we generalize the Smith's theorem to $r$-uniform hypergraphs. We show that the smallest limit point of the spectral radii of connected $r$-uniform hypergraphs is $ρ_r=(r-1)!\sqrt[r]{4}$. We discovered a novel method for computing the spectral radius of hypergraphs, and classified all connected $r$-uniform hypergraphs with spectral radius at most $ρ_r$.

preprint2014arXiv

Hypergraphs with Spectral Radius at most $(r-1)!\sqrt[r]{2+\sqrt{5}}$

In our previous paper, we classified all $r$-uniform hypergraphs with spectral radius at most $(r-1)!\sqrt[r]{4}$, which directly generalizes Smith's theorem for the graph case $r=2$. It is nature to ask the structures of the hypergraphs with spectral radius slightly beyond $(r-1)!\sqrt[r]{4}$. For $r=2$, the graphs with spectral radius at most $\sqrt{2+\sqrt{5}}$ are classified by [{\em Brouwer-Neumaier, Linear Algebra Appl., 1989}]. Here we consider the $r$-uniform hypergraphs $H$ with spectral radius at most $(r-1)!\sqrt[r]{2+\sqrt{5}}$. We show that $H$ must have a quipus-structure, which is similar to the graphs with spectral radius at most $\frac{3}{2}\sqrt{2}$ [{\em Woo-Neumaier, Graphs Combin., 2007}].

preprint2014arXiv

Set families with forbidden subposets

Let $F$ be a family of subsets of $\{1,\ldots,n\}$. We say that $F$ is $P$-free if the inclusion order on $F$ does not contain $P$ as an induced subposet. The \emph{Turán function} of $P$, denoted $π^*(n,P)$, is the maximum size of a $P$-free family of subsets of $\{1,\ldots,n\}$. We show that $π^*(n,P) \le (4r + O(\sqrt{r}))\binom{n}{n/2}$ if $P$ is an $r$-element poset of height at most $2$. We also show that $π^*(n,S_r) = (r+O(\sqrt{r}))\binom{n}{n/2}$ where $S_r$ is the standard example on $2r$ elements, and that $π^*(n,B_2) \le (2.583+o(1))\binom{n}{n/2}$, where $B_2$ is the $2$-dimensional Boolean lattice.

preprint2014arXiv

Strong Jumps and Lagrangians of Non-Uniform Hypergraphs

The hypergraph jump problem and the study of Lagrangians of uniform hypergraphs are two classical areas of study in the extremal graph theory. In this paper, we refine the concept of jumps to strong jumps and consider the analogous problems over non-uniform hypergraphs. Strong jumps have rich topological and algebraic structures. The non-strong-jump values are precisely the densities of the hereditary properties, which include the Turán densities of families of hypergraphs as special cases. Our method uses a generalized Lagrangian for non-uniform hypergraphs. We also classify all strong jump values for $\{1,2\}$-hypergraphs.

preprint2014arXiv

Unavoidable Multicoloured Families of Configurations

Balogh and Bollobás [{\em Combinatorica 25, 2005}] prove that for any $k$ there is a constant $f(k)$ such that any set system with at least $f(k)$ sets reduces to a $k$-star, an $k$-costar or an $k$-chain. They proved $f(k)<(2k)^{2^k}$. Here we improve it to $f(k)<2^{ck^2}$ for some constant $c>0$. This is a special case of the following result on the multi-coloured forbidden configurations at 2 colours. Let $r$ be given. Then there exists a constant $c_r$ so that a matrix with entries drawn from $\{0,1,...,r-1\}$ with at least $2^{c_rk^2}$ different columns will have a $k\times k$ submatrix that can have its rows and columns permuted so that in the resulting matrix will be either $I_k(a,b)$ or $T_k(a,b)$ (for some $a\ne b\in \{0,1,..., r-1\}$), where $I_k(a,b)$ is the $k\times k$ matrix with $a$'s on the diagonal and $b$'s else where, $T_k(a,b)$ the $k\times k$ matrix with $a$'s below the diagonal and $b$'s elsewhere. We also extend to considering the bound on the number of distinct columns, given that the number of rows is $m$, when avoiding a $t k\times k$ matrix obtained by taking any one of the $k \times k$ matrices above and repeating each column $t$ times. We use Ramsey Theory.

preprint2013arXiv

Boolean algebras and Lubell functions

Let $2^{[n]}$ denote the power set of $[n]:=\{1,2,..., n\}$. A collection $\B\subset 2^{[n]}$ forms a $d$-dimensional {\em Boolean algebra} if there exist pairwise disjoint sets $X_0, X_1,..., X_d \subseteq [n]$, all non-empty with perhaps the exception of $X_0$, so that $\B={X_0\cup \bigcup_{i\in I} X_i\colon I\subseteq [d]}$. Let $b(n,d)$ be the maximum cardinality of a family $\F\subset 2^X$ that does not contain a $d$-dimensional Boolean algebra. Gunderson, Rödl, and Sidorenko proved that $b(n,d) \leq c_d n^{-1/2^d} \cdot 2^n$ where $c_d= 10^d 2^{-2^{1-d}}d^{d-2^{-d}}$. In this paper, we use the Lubell function as a new measurement for large families instead of cardinality. The Lubell value of a family of sets $\F$ with $\F\subseteq \tsupn$ is defined by $h_n(\F):=\sum_{F\in \F}1/{n\choose |F|}$. We prove the following Turán type theorem. If $\F\subseteq 2^{[n]}$ contains no $d$-dimensional Boolean algebra, then $h_n(\F)\leq 2(n+1)^{1-2^{1-d}}$ for sufficiently large $n$. This results implies $b(n,d) \leq C n^{-1/2^d} \cdot 2^n$, where $C$ is an absolute constant independent of $n$ and $d$. As a consequence, we improve several Ramsey-type bounds on Boolean algebras. We also prove a canonical Ramsey theorem for Boolean algebras.

preprint2013arXiv

Repeated columns and an old chestnut

Let $t\ge 1$ be a given integer. Let ${\cal F}$ be a family of subsets of $[m]=\{1,2,\ldots,m\}$. Assume that for every pair of disjoint sets $S,T\subset [m]$ with $|S|=|T|=k$, there do not exist $2t$ sets in ${\cal F}$ where $t$ subsets of ${\cal F}$ contain $S$ and are disjoint from $T$ and $t$ subsets of ${\cal F}$ contain $T$ and are disjoint from $S$. We show that $|{\cal F}|$ is $O(m^{k})$. Our main new ingredient is allowing, during the inductive proof, multisets of subsets of $[m]$ where the multiplicity of a given set is bounded by $t-1$. We use a strong stability result of Anstee and Keevash. This is further evidence for a conjecture of Anstee and Sali. These problems can be stated in the language of matrices Let $t\cdot M$ denote $t$ copies of the matrix $M$ concatenated together. We have established the conjecture for those configurations $t\cdot F$ for any $k\times 2$ (0,1)-matrix $F$.

preprint2013arXiv

Ricci-flat graphs with girth at least five

A graph is called Ricci-flat if its Ricci-curvatures vanish on all edges. Here we use the definition of Ricci-cruvature on graphs given in [Lin-Lu-Yau, Tohoku Math., 2011], which is a variation of [Ollivier, J. Funct. Math., 2009]. In this paper, we classified all Ricci-flat connected graphs with girth at least five: they are the infinite path, cycle $C_n$ ($n\geq 6$), the dodecahedral graph, the Petersen graph, and the half-dodecahedral graph. We also construct many Ricci-flat graphs with girth 3 or 4 by using the root systems of simple Lie algebras.

preprint2013arXiv

Turan Problems on Non-uniform Hypergraphs

A non-uniform hypergraph $H=(V,E)$ consists of a vertex set $V$ and an edge set $E\subseteq 2^V$; the edges in $E$ are not required to all have the same cardinality. The set of all cardinalities of edges in $H$ is denoted by $R(H)$, the set of edge types. For a fixed hypergraph $H$, the Turán density $π(H)$ is defined to be $\lim_{n\to\infty}\max_{G_n}h_n(G_n)$, where the maximum is taken over all $H$-free hypergraphs $G_n$ on $n$ vertices satisfying $R(G_n)\subseteq R(H)$, and $h_n(G_n)$, the so called Lubell function, is the expected number of edges in $G_n$ hit by a random full chain. This concept, which generalizes the Turán density of $k$-uniform hypergraphs, is motivated by recent work on extremal poset problems. The details connecting these two areas will be revealed in the end of this paper. Several properties of Turán density, such as supersaturation, blow-up, and suspension, are generalized from uniform hypergraphs to non-uniform hypergraphs. Other questions such as "Which hypergraphs are degenerate?" are more complicated and don't appear to generalize well. In addition, we completely determine the Turán densities of ${1,2}$-hypergraphs.

preprint2012arXiv

Enhancing topology adaptation in information-sharing social networks

The advent of Internet and World Wide Web has led to unprecedent growth of the information available. People usually face the information overload by following a limited number of sources which best fit their interests. It has thus become important to address issues like who gets followed and how to allow people to discover new and better information sources. In this paper we conduct an empirical analysis on different on-line social networking sites, and draw inspiration from its results to present different source selection strategies in an adaptive model for social recommendation. We show that local search rules which enhance the typical topological features of real social communities give rise to network configurations that are globally optimal. These rules create networks which are effective in information diffusion and resemble structures resulting from real social systems.

preprint2012arXiv

On crown-free families of subsets

The crown $\Oh_{2t}$ is a height-2 poset whose Hasse diagram is a cycle of length $2t$. A family $\F$ of subsets of $[n]:=\{1,2..., n\}$ is {\em $\Oh_{2t}$-free} if $\Oh_{2t}$ is not a weak subposet of $(\F,\subseteq)$. Let $\La(n,\Oh_{2t})$ be the largest size of $\Oh_{2t}$-free families of subsets of $[n]$. De Bonis-Katona-Swanepoel proved $\La(n,\Oh_{4})= {n\choose \lfloor \frac{n}{2} \rfloor} + {n\choose \lceil \frac{n}{2} \rceil}$. Griggs and Lu proved that $\La(n,\Oh_{2t})=(1+o(1))\nchn$ for all even $t\ge 4$. In this paper, we prove $\La(n,\Oh_{2t})=(1+o(1))\nchn$ for all odd $t\geq 7$.

preprint2012arXiv

Scaling Laws in Human Language

Zipf's law on word frequency is observed in English, French, Spanish, Italian, and so on, yet it does not hold for Chinese, Japanese or Korean characters. A model for writing process is proposed to explain the above difference, which takes into account the effects of finite vocabulary size. Experiments, simulations and analytical solution agree well with each other. The results show that the frequency distribution follows a power law with exponent being equal to 1, at which the corresponding Zipf's exponent diverges. Actually, the distribution obeys exponential form in the Zipf's plot. Deviating from the Heaps' law, the number of distinct words grows with the text length in three stages: It grows linearly in the beginning, then turns to a logarithmical form, and eventually saturates. This work refines previous understanding about Zipf's law and Heaps' law in language systems.

preprint2012arXiv

Spectra of edge-independent random graphs

Let $G$ be a random graph on the vertex set $\{1,2,..., n\}$ such that edges in $G$ are determined by independent random indicator variables, while the probability $p_{ij}$ for $\{i,j\}$ being an edge in $G$ is not assumed to be equal. Spectra of the adjacency matrix and the normalized Laplacian matrix of $G$ are recently studied by Oliveira and Chung-Radcliffe. Let $A$ be the adjacency matrix of $G$, $\bar A=\E(A)$, and $Δ$ be the maximum expected degree of $G$. Oliveira first proved that almost surely $\|A-\bar A\|=O(\sqrt{Δ\ln n})$ provided $Δ\geq C \ln n$ for some constant $C$. Chung-Radcliffe improved the hidden constant in the error term using a new Chernoff-type inequality for random matrices. Here we prove that almost surely $\|A-\bar A\|\leq (2+o(1))\sqrtΔ$ with a slightly stronger condition $Δ\gg \ln^4 n$. For the Laplacian $L$ of $G$, Oliveira and Chung-Radcliffe proved similar results $\|L-\bar L|=O(\sqrt{\ln n}/\sqrtδ)$ provided the minimum expected degree $δ\gg \ln n$; we also improve their results by removing the $\sqrt{\ln n}$ multiplicative factor from the error term under some mild conditions. Our results naturally apply to the classic Erdős-Rényi random graphs, random graphs with given expected degree sequences, and bond percolation of general graphs.

preprint2012arXiv

The Fractional Chromatic Number of Triangle-free Graphs with $Δ\leq 3$

Let $G$ be any triangle-free graph with maximum degree $Δ\leq 3$. Staton proved that the independence number of $G$ is at least 5/14n. Heckman and Thomas conjectured that Staton's result can be strengthened into a bound on the fractional chromatic number of $G$, namely $χ_f(G)\leq 14/5. Recently, Hatami and Zhu proved $χ_f(G) \leq 3 -{3/64}$. In this paper, we prove $χ_f(G) \leq 3- 3/43$.

preprint2012arXiv

Unrolling residues to avoid progressions

We consider the problem of coloring $[n]={1,2,...,n}$ with $r$ colors to minimize the number of monochromatic $k$ term arithmetic progressions (or $k$-APs for short). We show how to extend colorings of $\mathbb{Z}_m$ which avoid nontrivial $k$-APs to colorings of $[n]$ by an unrolling process. In particular, by using residues to color $\mathbb{Z}_m$ we produce the best known colorings for minimizing the number of monochromatic $k$-APs for coloring with $r$ colors for several small values of $r$ and $k$.

preprint2011arXiv

A Fractional Analogue of Brooks' Theorem

Let $Δ(G)$ be the maximum degree of a graph $G$. Brooks' theorem states that the only connected graphs with chromatic number $χ(G)=Δ(G)+1$ are complete graphs and odd cycles. We prove a fractional analogue of Brooks' theorem in this paper. Namely, we classify all connected graphs $G$ such that the fractional chromatic number $χ_f(G)$ is at least $Δ(G)$. These graphs are complete graphs, odd cycles, $C^2_8$, $C_5\boxtimes K_2$, and graphs whose clique number $ω(G)$ equals the maximum degree $Δ(G)$. Among the two sporadic graphs, the graph $C^2_8$ is the square graph of cycle $C_8$ while the other graph $C_5\boxtimes K_2$ is the strong product of $C_5$ and $K_2$. In fact, we prove a stronger result; if a connected graph $G$ with $Δ(G)\geq 4$ is not one of the graphs listed above, then we have $χ_f(G)\leq Δ(G)- 2/67$.

preprint2011arXiv

Coarse Graining for Synchronization in Directed Networks

Coarse graining model is a promising way to analyze and visualize large-scale networks. The coarse-grained networks are required to preserve the same statistical properties as well as the dynamic behaviors as the initial networks. Some methods have been proposed and found effective in undirected networks, while the study on coarse graining in directed networks lacks of consideration. In this paper, we proposed a Topology-aware Coarse Graining (TCG) method to coarse grain the directed networks. Performing the linear stability analysis of synchronization and numerical simulation of the Kuramoto model on four kinds of directed networks, including tree-like networks and variants of Barabási-Albert networks, Watts-Strogatz networks and Erdös-Rényi networks, we find our method can effectively preserve the network synchronizability.

preprint2011arXiv

Diameters of Graphs with Spectral Radius at most $3/2\sqrt{2}$

The spectral radius $ρ(G)$ of a graph $G$ is the largest eigenvalue of its adjacency matrix. Woo and Neumaier discovered that a connected graph $G$ with $ρ(G)\leq 3/2{\sqrt{2}}$ is either a dagger, an open quipu, or a closed quipu. The reverse statement is not true. Many open quipus and closed quipus have spectral radius greater than $3/2{\sqrt{2}}$. In this paper we proved the following results. For any open quipu $G$ on $n$ vertices ($n\geq 6$) with spectral radius less than $3/2{\sqrt{2}}$, its diameter $D(G)$ satisfies $D(G)\geq (2n-4)/3$. This bound is tight. For any closed quipu $G$ on $n$ vertices ($n\geq 13$) with spectral radius less than $3/2{\sqrt{2}}$, its diameter $D(G)$ satisfies $\frac{n}{3}< D(G)\leq \frac{2n-2}{3}$. The upper bound is tight while the lower bound is asymptotically tight. Let $G^{min}_{n,D}$ be a graph with minimal spectral radius among all connected graphs on $n$ vertices with diameter $D$. We applied the results and found $G^{min}_{n,D}$ for some range of $D$. For $n\geq 13$ and $D\in [\frac{n}{2}, \frac{2n-7}{3}]$, we proved that $G^{min}_{n,D}$ is the graph obtained by attaching two paths of length $D-\lfloor\frac{n}{2}\rfloor$ and $D-\lceil\frac{n}{2}\rceil$ to a pair of antipodal vertices of the even cycle $C_{2(n-D)}$. Thus we settled a conjecture of Cioab-van Dam-Koolen-Lee, who previously proved a special case $D=\frac{n+e}{2}$ for $e=1,2,3,4$.

preprint2011arXiv

Diamond-free Families

Given a finite poset P, we consider the largest size La(n,P) of a family of subsets of $[n]:=\{1,...,n\}$ that contains no subposet P. This problem has been studied intensively in recent years, and it is conjectured that $π(P):= \lim_{n\rightarrow\infty} La(n,P)/{n choose n/2}$ exists for general posets P, and, moreover, it is an integer. For $k\ge2$ let $\D_k$ denote the $k$-diamond poset $\{A< B_1,...,B_k < C\}$. We study the average number of times a random full chain meets a $P$-free family, called the Lubell function, and use it for $P=\D_k$ to determine $π(\D_k)$ for infinitely many values $k$. A stubborn open problem is to show that $π(\D_2)=2$; here we make progress by proving $π(\D_2)\le 2 3/11$ (if it exists).

preprint2011arXiv

Empirical analysis of web-based user-object bipartite networks

Understanding the structure and evolution of web-based user-object networks is a significant task since they play a crucial role in e-commerce nowadays. This Letter reports the empirical analysis on two large-scale web sites, audioscrobbler.com and del.icio.us, where users are connected with music groups and bookmarks, respectively. The degree distributions and degree-degree correlations for both users and objects are reported. We propose a new index, named collaborative clustering coefficient, to quantify the clustering behavior based on the collaborative selection. Accordingly, the clustering properties and clustering-degree correlations are investigated. We report some novel phenomena well characterizing the selection mechanism of web users and outline the relevance of these phenomena to the information recommendation problem.

preprint2011arXiv

Graphs with Diameter $n-e$ Minimizing the Spectral Radius

The spectral radius $ρ(G)$ of a graph $G$ is the largest eigenvalue of its adjacency matrix $A(G)$. For a fixed integer $e\ge 1$, let $G^{min}_{n,n-e}$ be a graph with minimal spectral radius among all connected graphs on $n$ vertices with diameter $n-e$. Let $P_{n_1,n_2,...,n_t,p}^{m_1,m_2,...,m_t}$ be a tree obtained from a path of $p$ vertices ($0 \sim 1 \sim 2 \sim ... \sim (p-1)$) by linking one pendant path $P_{n_i}$ at $m_i$ for each $i\in\{1,2,...,t\}$. For $e=1,2,3,4,5$, $G^{min}_{n,n-e}$ were determined in the literature. Cioabǎ-van Dam-Koolen-Lee \cite{CDK} conjectured for fixed $e\geq 6$, $G^{min}_{n,n-e}$ is in the family ${\cal P}_{n,e}=\{P_{2,1,...1,2,n-e+1}^{2,m_2,...,m_{e-4},n-e-2}\mid 2<m_2<...<m_{e-4}<n-e-2\}$. For $e=6,7$, they conjectured $G^{min}_{n,n-6}=P^{2,\lceil\frac{D-1}{2}\rceil,D-2}_{2,1,2,n-5}$ and $G^{min}_{n,n-7}=P^{2,\lfloor\frac{D+2}{3}\rfloor,D- \lfloor\frac{D+2}{3}\rfloor, D-2}_{2,1,1,2,n-6}$. In this paper, we settle their three conjectures positively. We also determine $G^{min}_{n,n-8}$ in this paper.

preprint2011arXiv

High-ordered Random Walks and Generalized Laplacians on Hypergraphs

Despite of the extreme success of the spectral graph theory, there are relatively few papers applying spectral analysis to hypergraphs. Chung first introduced Laplacians for regular hypergraphs and showed some useful applications. Other researchers treated hypergraphs as weighted graphs and then studied the Laplacians of the corresponding weighted graphs. In this paper, we aim to unify these very different versions of Laplacians for hypergraphs. We introduce a set of Laplacians for hypergraphs through studying high-ordered random walks on hypergraphs. We prove the eigenvalues of these Laplacians can effectively control the mixing rate of high-ordered random walks, the generalized distances/diameters, and the edge expansions.

preprint2011arXiv

Information filtering via preferential diffusion

Recommender systems have shown great potential to address information overload problem, namely to help users in finding interesting and relevant objects within a huge information space. Some physical dynamics, including heat conduction process and mass or energy diffusion on networks, have recently found applications in personalized recommendation. Most of the previous studies focus overwhelmingly on recommendation accuracy as the only important factor, while overlook the significance of diversity and novelty which indeed provide the vitality of the system. In this paper, we propose a recommendation algorithm based on the preferential diffusion process on user-object bipartite network. Numerical analyses on two benchmark datasets, MovieLens and Netflix, indicate that our method outperforms the state-of-the-art methods. Specifically, it can not only provide more accurate recommendations, but also generate more diverse and novel recommendations by accurately recommending unpopular objects.

preprint2011arXiv

Leaders in Social Networks, the Delicious Case

Finding pertinent information is not limited to search engines. Online communities can amplify the influence of a small number of power users for the benefit of all other users. Users' information foraging in depth and breadth can be greatly enhanced by choosing suitable leaders. For instance in delicious.com, users subscribe to leaders' collection which lead to a deeper and wider reach not achievable with search engines. To consolidate such collective search, it is essential to utilize the leadership topology and identify influential users. Google's PageRank, as a successful search algorithm in the World Wide Web, turns out to be less effective in networks of people. We thus devise an adaptive and parameter-free algorithm, the LeaderRank, to quantify user influence. We show that LeaderRank outperforms PageRank in terms of ranking effectiveness, as well as robustness against manipulations and noisy data. These results suggest that leaders who are aware of their clout may reinforce the development of social networks, and thus the power of collective search.

preprint2011arXiv

Loose Laplacian spectra of random hypergraphs

Let $H=(V,E)$ be an $r$-uniform hypergraph with the vertex set $V$ and the edge set $E$. For $1\leq s \leq r/2$, we define a weighted graph $G^{(s)}$ on the vertex set ${V\choose s}$ as follows. Every pair of $s$-sets $I$ and $J$ is associated with a weight $w(I,J)$, which is the number of edges in $H$ passing through $I$ and $J$ if $I\cap J=\emptyset$, and 0 if $I\cap J\not=\emptyset$. The $s$-th Laplacian $Ł^{(s)}$ of $H$ is defined to be the normalized Laplacian of $G^{(s)}$. The eigenvalues of $\mathcal L^{(s)}$ are listed as $λ^{(s)}_0, λ^{(s)}_1,..., λ^{(s)}_{{n\choose s}-1}$ in non-decreasing order. Let $\barλ^{(s)}(H)=\max_{i\not=0}\{|1-λ^{(s)}_i|\}$. The parameters $\barλ^{(s)}(H)$ and $λ^{(s)}_1(H)$, which were introduced in our previous paper, have a number of connections to the mixing rate of high-ordered random walks, the generalized distances/diameters, and the edge expansions. For $0< p<1$, let $H^r(n,p)$ be a random $r$-uniform hypergraph over $[n]:={1,2,..., n}$, where each $r$-set of $[n]$ has probability $p$ to be an edge independently. For $1 \leq s \leq r/2$, $p(1-p)\gg \frac{\log^4 n}{n^{r-s}}$, and $1-p\gg \frac{\log n}{n^2}$, we prove that almost surely $$\barλ^{(s)}(H^r(n,p))\leq \frac{s}{n-s}+ (3+o(1))\sqrt{\frac{1-p}{{n-s\choose r-s}p}}.$$ We also prove that the empirical distribution of the eigenvalues of $Ł^{(s)}$ for $H^r(n,p)$ follows the Semicircle Law if $p(1-p)\gg \frac{\log^{1/3} n}{n^{r-s}}$ and $1-p\gg \frac{\log n}{n^{2+2r-2s}}$.

preprint2011arXiv

Monochromatic 4-term arithmetic progressions in 2-colorings of $\mathbb Z_n$

This paper is motivated by a recent result of Wolf \cite{wolf} on the minimum number of monochromatic 4-term arithmetic progressions(4-APs, for short) in $\Z_p$, where $p$ is a prime number. Wolf proved that there is a 2-coloring of $\Z_p$ with 0.000386% fewer monochromatic 4-APs than random 2-colorings; the proof is probabilistic and non-constructive. In this paper, we present an explicit and simple construction of a 2-coloring with 9.3% fewer monochromatic 4-APs than random 2-colorings. This problem leads us to consider the minimum number of monochromatic 4-APs in $\Z_n$ for general $n$. We obtain both lower bound and upper bound on the minimum number of monochromatic 4-APs in all 2-colorings of $\Z_n$. Wolf proved that any 2-coloring of $\Z_p$ has at least $(1/16+o(1))p^2$ monochromatic 4-APs. We improve this lower bound into $(7/96+o(1))p^2$. Our results on $\Z_n$ naturally apply to the similar problem on $[n]$ (i.e., $\{1,2,..., n\}$). In 2008, Parillo, Robertson, and Saracino \cite{prs} constructed a 2-coloring of $[n]$ with 14.6% fewer monochromatic 3-APs than random 2-colorings. In 2010, Butler, Costello, and Graham \cite{BCG} extended their methods and used an extensive computer search to construct a 2-coloring of $[n]$ with 17.35% fewer monochromatic 4-APs (and 26.8% fewer monochromatic 5-APs) than random 2-colorings. Our construction gives a 2-coloring of $[n]$ with 33.33% fewer monochromatic 4-APs (and 57.89% fewer monochromatic 5-APs) than random 2-colorings.

preprint2011arXiv

The Randic index and the diameter of graphs

The {\it Randić index} $R(G)$ of a graph $G$ is defined as the sum of 1/\sqrt{d_ud_v} over all edges $uv$ of $G$, where $d_u$ and $d_v$ are the degrees of vertices $u$ and $v,$ respectively. Let $D(G)$ be the diameter of $G$ when $G$ is connected. Aouchiche-Hansen-Zheng conjectured that among all connected graphs $G$ on $n$ vertices the path $P_n$ achieves the minimum values for both $R(G)/D(G)$ and $R(G)- D(G)$. We prove this conjecture completely. In fact, we prove a stronger theorem: If $G$ is a connected graph, then $R(G)-(1/2)D(G)\geq \sqrt{2}-1$, with equality if and only if $G$ is a path with at least three vertices.

preprint2010arXiv

Link Prediction Based on Local Random Walk

The problem of missing link prediction in complex networks has attracted much attention recently. Two difficulties in link prediction are the sparsity and huge size of the target networks. Therefore, the design of an efficient and effective method is of both theoretical interests and practical significance. In this Letter, we proposed a method based on local random walk, which can give competitively good prediction or even better prediction than other random-walk-based methods while has a lower computational complexity.

preprint2010arXiv

Link Prediction in Complex Networks: A Survey

Link prediction in complex networks has attracted increasing attention from both physical and computer science communities. The algorithms can be used to extract missing information, identify spurious interactions, evaluate network evolving mechanisms, and so on. This article summaries recent progress about link prediction algorithms, emphasizing on the contributions from physical perspectives and approaches, such as the random-walk-based methods and the maximum likelihood methods. We also introduce three typical applications: reconstruction of networks, evaluation of network evolving mechanism and classification of partially labelled networks. Finally, we introduce some applications and outline future challenges of link prediction algorithms.

preprint2010arXiv

Similarity-Based Classification in Partially Labeled Networks

We propose a similarity-based method, using the similarity between nodes, to address the problem of classification in partially labeled networks. The basic assumption is that two nodes are more likely to be categorized into the same class if they are more similar. In this paper, we introduce ten similarity indices, including five local ones and five global ones. Empirical results on the co-purchase network of political books show that the similarity-based method can give high accurate classification even when the labeled nodes are sparse which is one of the difficulties in classification. Furthermore, we find that when the target network has many labeled nodes, the local indices can perform as good as those global indices do, while when the data is sparce the global indices perform better. Besides, the similarity-based method can to some extent overcome the unconsistency problem which is another difficulty in classification.

preprint2010arXiv

Zipf's Law Leads to Heaps' Law: Analyzing Their Relation in Finite-Size Systems

Background: Zipf's law and Heaps' law are observed in disparate complex systems. Of particular interests, these two laws often appear together. Many theoretical models and analyses are performed to understand their co-occurrence in real systems, but it still lacks a clear picture about their relation. Methodology/Principal Findings: We show that the Heaps' law can be considered as a derivative phenomenon if the system obeys the Zipf's law. Furthermore, we refine the known approximate solution of the Heaps' exponent provided the Zipf's exponent. We show that the approximate solution is indeed an asymptotic solution for infinite systems, while in the finite-size system the Heaps' exponent is sensitive to the system size. Extensive empirical analysis on tens of disparate systems demonstrates that our refined results can better capture the relation between the Zipf's and Heaps' exponents. Conclusions/Significance: The present analysis provides a clear picture about the relation between the Zipf's law and Heaps' law without the help of any specific stochastic model, namely the Heaps' law is indeed a derivative phenomenon from Zipf's law. The presented numerical method gives considerably better estimation of the Heaps' exponent given the Zipf's exponent and the system size. Our analysis provides some insights and implications of real complex systems, for example, one can naturally obtained a better explanation of the accelerated growth of scale-free networks.