Source author record

Zi-Ke Zhang

Zi-Ke Zhang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

35works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

35 published item(s)

preprint2022arXiv

Influence Maximization in Hypergraphs

Influence maximization in complex networks, i.e., maximizing the size of influenced nodes via selecting K seed nodes for a given spreading process, has attracted great attention in recent years. However, the influence maximization problem in hypergraphs, in which the hyperedges are leveraged to represent the interactions among more than two nodes, is still an open question. In this paper, we propose an adaptive degree-based heuristic algorithm, i.e., Heuristic Degree Discount (HDD), which iteratively selects nodes with low influence overlap as seeds, to solve the influence maximization problem in hypergraphs. We further extend algorithms from ordinary networks as baselines and compare the performance of the proposed algorithm and baselines on both real data and synthetic hypergraphs. Results show that HDD outperforms the baselines in terms of both effectiveness and efficiency. Moreover, the experiments on synthetic hypergraphs indicate that HDD shows high performance, especially in hypergraphs with heterogeneous degree distribution.

preprint2022arXiv

Toward Structural Controllability and Predictability in Directed Networks

The lack of studying the complex organization of directed network usually limits to the understanding of underlying relationship between network structures and functions. Structural controllability and structural predictability, two seemingly unrelated subjects, are revealed in this paper to be both highly dependent on the critical links previously thought to only be able to influence the number of driver nodes in controllable directed networks. Here, we show that critical links can not only contribute to structural controllability, but they can also have a significant impact on the structural predictability of networks, suggesting the universal pattern of structural reciprocity in directed networks. In addition, it is shown that the fraction and location of critical links have a strong influence on the performance of prediction algorithms. Moreover, these empirical results are interpreted by introducing the link centrality based on corresponding line graphs. This work bridges the gap between the two independent research fields, and it provides indications of developing advanced control strategies and prediction algorithms from a microscopic perspective.

preprint2022arXiv

Vital node identification in hypergraphs via gravity model

Hypergraphs that can depict interactions beyond pairwise edges have emerged as an appropriate representation for modeling polyadic relations in complex systems. With the recent surge of interest in researching hypergraphs, the centrality problem has attracted abundant attention due to the challenge of how to utilize the higher-order structure for the definition of centrality metrics. In this paper, we propose a new centrality method (HGC) on the basis of the gravity model as well as a semi-local HGC (LHGC) which can achieve a balance between accuracy and computational complexity. Meanwhile, two comprehensive evaluation metrics, i.e., a complex contagion model in hypergraphs that mimics the group influence during the spreading process and network s-efficiency based on the higher-order distance between nodes, are first proposed to evaluate the effectiveness of our methods. The results show that our methods can filter out nodes that have fast spreading ability and are vital in terms of hypergraph connectivity.

preprint2020arXiv

Information Spreading Dynamics on Adaptive Social Networks

There is currently growing interest in modeling the information diffusion on social networks across multi-disciplines. The majority of the corresponding research has focused on information diffusion independently, ignoring the network evolution in the diffusion process. Therefore, it is more reasonable to describe the real diffusion systems by the co-evolution between network topologies and information states. In this work, we propose a mechanism considering the coevolution between information states and network topology simultaneously, in which the information diffusion was executed as an SIS process and network topology evolved based on the adaptive assumption. The theoretical analyses based on the Markov approach were very consistent with simulation. Both simulation results and theoretical analyses indicated that the adaptive process, in which informed individuals would rewire the links between the informed neighbors to a random non-neighbor node, can enhance information diffusion (leading to much broader spreading). In addition, we obtained that two threshold values exist for the information diffusion on adaptive networks, i.e., if the information propagation probability is less than the first threshold, information cannot diffuse and dies out immediately; if the propagation probability is between the first and second threshold, information will spread to a finite range and die out gradually; and if the propagation probability is larger than the second threshold, information will diffuse to a certain size of population in the network. These results may shed some light on understanding the co-evolution between information diffusion and network topology.

preprint2016arXiv

Identifying the Academic Rising Stars

Predicting the fast-rising young researchers (Academic Rising Stars) in the future provides useful guidance to the research community, e.g., offering competitive candidates to university for young faculty hiring as they are expected to have success academic careers. In this work, given a set of young researchers who have published the first first-author paper recently, we solve the problem of how to effectively predict the top k% researchers who achieve the highest citation increment in Δt years. We explore a series of factors that can drive an author to be fast-rising and design a novel impact increment ranking learning (IIRL) algorithm that leverages those factors to predict the academic rising stars. Experimental results on the large ArnetMiner dataset with over 1.7 million authors demonstrate the effectiveness of IIRL. Specifically, it outperforms all given benchmark methods, with over 8% average improvement. Further analysis demonstrates that the prediction models for different research topics follow the similar pattern. We also find that temporal features are the best indicators for rising stars prediction, while venue features are less relevant.

preprint2015arXiv

Epidemic Dynamics On Information-Driven Adaptive Networks

can evolve simultaneously. For the information-driven adaptive process, susceptible (infected) individuals who have abilities to recognize the disease would break the links of their infected (susceptible) neighbors to prevent the epidemic from further spreading. Simulation results and numerical analyses based on the pairwise approach indicate that the information-driven adaptive process can not only slow down the speed of epidemic spreading, but can also diminish the epidemic prevalence at the final state significantly. In addition, the disease spreading and information diffusion pattern on the lattice give a visual representation about how the disease is trapped into an isolated field with the information-driven adaptive process. Furthermore, we perform the local bifurcation analysis on four types of dynamical regions, including healthy, oscillatory, bistable and endemic, to understand the evolution of the observed dynamical behaviors. This work may shed some lights on understanding how information affects human activities on responding to epidemic spreading.

preprint2015arXiv

Events Determine Spreading Patterns: Information Transmission via Internal and External Influences on Social Networks

Recently, information transmission models motivated by the classical epidemic propagation, have been applied to a wide-range of social systems, generally assume that information mainly transmits among individuals via peer-to-peer interactions on social networks. In this paper, we consider one more approach for users to get information: the out-of-social-network influence. Empirical analyses of eight typical events' diffusion on a very large micro-blogging system, \emph{Sina Weibo}, show that the external influence has significant impact on information spreading along with social activities. In addition, we propose a theoretical model to interpret the spreading process via both internal and external channels, considering three essential properties: (i) memory effect; (ii) role of spreaders; and (iii) non-redundancy of contacts. Experimental and mathematical results indicate that the information indeed spreads much quicker and broader with mutual effects of the internal and external influences. More importantly, the present model reveals that the event characteristic would highly determine the essential spreading patterns once the network structure is established. The results may shed some light on the in-depth understanding of the underlying dynamics of information transmission on real social networks.

preprint2015arXiv

Multi-Linear Interactive Matrix Factorization

Recommender systems, which can significantly help users find their interested items from the information era, has attracted an increasing attention from both the scientific and application society. One of the widest applied recommendation methods is the Matrix Factorization (MF). However, most of MF based approaches focus on the user-item rating matrix, but ignoring the ingredients which may have significant influence on users' preferences on items. In this paper, we propose a multi-linear interactive MF algorithm (MLIMF) to model the interactions between the users and each event associated with their final decisions. Our model considers not only the user-item rating information but also the pairwise interactions based on some empirically supported factors. In addition, we compared the proposed model with three typical other methods: user-based collaborative filtering (UCF), item-based collaborative filtering (ICF) and regularized MF (RMF). Experimental results on two real-world datasets, \emph{MovieLens} 1M and \emph{MovieLens} 100k, show that our method performs much better than other three methods in the accuracy of recommendation. This work may shed some light on the in-depth understanding of modeling user online behaviors and the consequent decisions.

preprint2015arXiv

Mutual Feedback Between Epidemic Spreading and Information Diffusion

The impact that information diffusion has on epidemic spreading has recently attracted much attention. As a disease begins to spread in the population, information about the disease is transmitted to others, which in turn has an effect on the spread of disease. In this paper, using empirical results of the propagation of H7N9 and information about the disease, we clearly show that the spreading dynamics of the two-types of processes influence each other. We build a mathematical model in which both types of spreading dynamics are described using the SIS process in order to illustrate the influence of information diffusion on epidemic spreading. Both the simulation results and the pairwise analysis reveal that information diffusion can increase the threshold of an epidemic outbreak, decrease the final fraction of infected individuals and significantly decrease the rate at which the epidemic propagates. Additionally, we find that the multi-outbreak phenomena of epidemic spreading, along with the impact of information diffusion, is consistent with the empirical results. These findings highlight the requirement to maintain social awareness of diseases even when the epidemics seem to be under control in order to prevent a subsequent outbreak. These results may shed light on the in-depth understanding of the interplay between the dynamics of epidemic spreading and information diffusion.

preprint2014arXiv

Epidemic Spreading on Weighted Complex Networks

Nowadays, the emergence of online services provides various multi-relation information to support the comprehensive understanding of the epidemic spreading process. In this Letter, we consider the edge weights to represent such multi-role relations. In addition, we perform detailed analysis of two representative metrics, outbreak threshold and epidemic prevalence, on SIS and SIR models. Both theoretical and simulation results find good agreements with each other. Furthermore, experiments show that, on fully mixed networks, the weight distribution on edges would not affect the epidemic results once the average weight of whole network is fixed. This work may shed some light on the in-depth understanding of epidemic spreading on multi-relation and weighted networks.

preprint2014arXiv

Evolution of citation networks with the hypergraph formalism

In this paper, we proposed an evolving model via the hypergraph to illustrate the evolution of the citation network. In the evolving model, we consider the mechanism combined with preferential attachment and the aging influence. Simulation results show that the proposed model can characterize the citation distribution of the real system very well. In addition, we give the analytical result of the citation distribution using the master equation. Detailed analysis showed that the time decay factor should be the origin of the same citation distribution between the proposed model and the empirical result. The proposed model might shed some lights in understanding the underlying laws governing the structure of real citation networks.

preprint2014arXiv

Gravity Effects on Information Filtering and Network Evolving

In this paper, based on the gravity principle of classical physics, we propose a tunable gravity-based model, which considers tag usage pattern to weigh both the mass and distance of network nodes. We then apply this model in solving the problems of information filtering and network evolving. Experimental results on two real-world data sets, \emph{Del.icio.us} and \emph{MovieLens}, show that it can not only enhance the algorithmic performance, but can also better characterize the properties of real networks. This work may shed some light on the in-depth understanding of the effect of gravity model.

preprint2014arXiv

Information Filtering on Coupled Social Networks

In this paper, based on the coupled social networks (CSN), we propose a hybrid algorithm to nonlinearly integrate both social and behavior information of online users. Filtering algorithm based on the coupled social networks, which considers the effects of both social influence and personalized preference. Experimental results on two real datasets, \emph{Epinions} and \emph{Friendfeed}, show that hybrid pattern can not only provide more accurate recommendations, but also can enlarge the recommendation coverage while adopting global metric. Further empirical analyses demonstrate that the mutual reinforcement and rich-club phenomenon can also be found in coupled social networks where the identical individuals occupy the core position of the online system. This work may shed some light on the in-depth understanding structure and function of coupled social networks.

preprint2014arXiv

Information Filtering via Collaborative User Clustering Modeling

The past few years have witnessed the great success of recommender systems, which can significantly help users find out personalized items for them from the information era. One of the most widely applied recommendation methods is the Matrix Factorization (MF). However, most of researches on this topic have focused on mining the direct relationships between users and items. In this paper, we optimize the standard MF by integrating the user clustering regularization term. Our model considers not only the user-item rating information, but also takes into account the user interest. We compared the proposed model with three typical other methods: User-Mean (UM), Item-Mean (IM) and standard MF. Experimental results on a real-world dataset, \emph{MovieLens}, show that our method performs much better than other three methods in the accuracy of recommendation.

preprint2014arXiv

Promoting cold-start items in recommender systems

As one of major challenges, cold-start problem plagues nearly all recommender systems. In particular, new items will be overlooked, impeding the development of new products online. Given limited resources, how to utilize the knowledge of recommender systems and design efficient marketing strategy for new items is extremely important. In this paper, we convert this ticklish issue into a clear mathematical problem based on a bipartite network representation. Under the most widely used algorithm in real e-commerce recommender systems, so-called the item-based collaborative filtering, we show that to simply push new items to active users is not a good strategy. To our surprise, experiments on real recommender systems indicate that to connect new items with some less active users will statistically yield better performance, namely these new items will have more chance to appear in other users' recommendation lists. Further analysis suggests that the disassortative nature of recommender systems contributes to such observation. In a word, getting in-depth understanding on recommender systems could pave the way for the owners to popularize their cold-start products with low costs.

preprint2013arXiv

Emergence of Blind Areas in Information Spreading

Recently, contagion-based (disease, information, etc.) spreading on social networks has been extensively studied. In this paper, other than traditional full interaction, we propose a partial interaction based spreading model, considering that the informed individuals would transmit information to only a certain fraction of their neighbors due to the transmission ability in real-world social networks. Simulation results on three representative networks (BA, ER, WS) indicate that the spreading efficiency is highly correlated with the network heterogeneity. In addition, a special phenomenon, namely \emph{Information Blind Areas} where the network is separated by several information-unreachable clusters, will emerge from the spreading process. Furthermore, we also find that the size distribution of such information blind areas obeys power-law-like distribution, which has very similar exponent with that of site percolation. Detailed analyses show that the critical value is decreasing along with the network heterogeneity for the spreading process, which is complete the contrary to that of random selection. Moreover, the critical value in the latter process is also larger that of the former for the same network. Those findings might shed some lights in in-depth understanding the effect of network properties on information spreading.

preprint2013arXiv

Geography and similarity of regional cuisines in China

Food occupies a central position in every culture and it is therefore of great interest to understand the evolution of food culture. The advent of the World Wide Web and online recipe repositories has begun to provide unprecedented opportunities for data-driven, quantitative study of food culture. Here we harness an online database documenting recipes from various Chinese regional cuisines and investigate the similarity of regional cuisines in terms of geography and climate. We found that the geographical proximity, rather than climate proximity is a crucial factor that determines the similarity of regional cuisines. We develop a model of regional cuisine evolution that provides helpful clues to understand the evolution of cuisines and cultures.

preprint2013arXiv

Heterogeneity Involved Network-based Algorithm Leads to Accurate and Personalized Recommendations

Heterogeneity of both the source and target objects is taken into account in a network-based algorithm for the directional resource transformation between objects. Based on a biased heat conduction recommendation method (BHC) which considers the heterogeneity of the target object, we propose a heterogeneous heat conduction algorithm (HHC), by further taking the source object degree as the weight of diffusion. Tested on three real datasets, the Netflix, RYM and MovieLens, the HHC algorithm is found to present a better recommendation in both the accuracy and personalization than two excellent algorithms, i.e., the original BHC and a hybrid algorithm of heat conduction and mass diffusion (HHM), while not requiring any other accessorial information or parameter. Moreover, the HHC even elevates the recommendation accuracy on cold objects, referring to the so-called cold start problem, for effectively relieving the recommendation bias on objects with different level of popularity.

preprint2013arXiv

Influence of Reciprocal links in Social Networks

In this Letter, we empirically study the influence of reciprocal links, in order to understand its role in affecting the structure and function of directed social networks. Experimental results on two representative datesets, Sina Weibo and Douban, demonstrate that the reciprocal links indeed play a more important role than non-reciprocal ones in both spreading information and maintaining the network robustness. In particular, the information spreading process can be significantly enhanced by considering the reciprocal effect. In addition, reciprocal links are largely responsible for the connectivity and efficiency of directed networks. This work may shed some light on the in-depth understanding and application of the reciprocal effect in directed online social networks.

preprint2013arXiv

Information spreading on dynamic social networks

Nowadays, information spreading on social networks has triggered an explosive attention in various disciplines. Most of previous works in this area mainly focus on discussing the effects of spreading probability or immunization strategy on static networks. However, in real systems, the peer-to-peer network structure changes constantly according to frequently social activities of users. In order to capture this dynamical property and study its impact on information spreading, in this paper, a link rewiring strategy based on the Fermi function is introduced. In the present model, the informed individuals tend to break old links and reconnect to their second-order friends with more uninformed neighbors. Simulation results on the susceptible-infected-recovered (\textit{SIR}) model with fixed recovery time $T=1$ indicate that the information would spread more faster and broader with the proposed rewiring strategy. Extensive analyses of the information cascade size distribution show that the spreading process of the initial steps plays a very important role, that is to say, the information will spread out if it is still survival at the beginning time. The proposed model may shed some light on the in-depth understanding of information spreading on dynamical social networks.

preprint2012arXiv

A two-step Recommendation Algorithm via Iterative Local Least Squares

Recommender systems can change our life a lot and help us select suitable and favorite items much more conveniently and easily. As a consequence, various kinds of algorithms have been proposed in last few years to improve the performance. However, all of them face one critical problem: data sparsity. In this paper, we proposed a two-step recommendation algorithm via iterative local least squares (ILLS). Firstly, we obtain the ratings matrix which is constructed via users' behavioral records, and it is normally very sparse. Secondly, we preprocess the "ratings" matrix through ProbS which can convert the sparse data to a dense one. Then we use ILLS to estimate those missing values. Finally, the recommendation list is generated. Experimental results on the three datasets: MovieLens, Netflix, RYM, suggest that the proposed method can enhance the algorithmic accuracy of AUC. Especially, it performs much better in dense datasets. Furthermore, since this methods can improve those missing value more accurately via iteration which might show light in discovering those inactive users' purchasing intention and eventually solving cold-start problem.

preprint2012arXiv

Anchoring Bias in Online Voting

Voting online with explicit ratings could largely reflect people's preferences and objects' qualities, but ratings are always irrational, because they may be affected by many unpredictable factors like mood, weather, as well as other people's votes. By analyzing two real systems, this paper reveals a systematic bias embedding in the individual decision-making processes, namely people tend to give a low rating after a low rating, as well as a high rating following a high rating. This so-called \emph{anchoring bias} is validated via extensive comparisons with null models, and numerically speaking, the extent of bias decays with interval voting number in a logarithmic form. Our findings could be applied in the design of recommender systems and considered as important complementary materials to previous knowledge about anchoring effects on financial trades, performance judgements, auctions, and so on.

preprint2012arXiv

Cultural evolution and personalization

In social sciences, there is currently no consensus on the mechanism for cultural evolution. The evolution of first names of newborn babies offers a remarkable example for the researches in the field. Here we perform statistical analyses on over 100 years of data in the United States. We focus in particular on how the frequency-rank distribution and inequality of baby names change over time. We propose a stochastic model where name choice is determined by personalized preference and social influence. Remarkably, variations on the strength of personalized preference can account satisfactorily for the observed empirical features. Therefore, we claim that personalization drives cultural evolution, at least in the example of baby names.

preprint2012arXiv

Emergence of scale-free close-knit friendship structure in online social networks

Despite the structural properties of online social networks have attracted much attention, the properties of the close-knit friendship structures remain an important question. Here, we mainly focus on how these mesoscale structures are affected by the local and global structural properties. Analyzing the data of four large-scale online social networks reveals several common structural properties. It is found that not only the local structures given by the indegree, outdegree, and reciprocal degree distributions follow a similar scaling behavior, the mesoscale structures represented by the distributions of close-knit friendship structures also exhibit a similar scaling law. The degree correlation is very weak over a wide range of the degrees. We propose a simple directed network model that captures the observed properties. The model incorporates two mechanisms: reciprocation and preferential attachment. Through rate equation analysis of our model, the local-scale and mesoscale structural properties are derived. In the local-scale, the same scaling behavior of indegree and outdegree distributions stems from indegree and outdegree of nodes both growing as the same function of the introduction time, and the reciprocal degree distribution also shows the same power-law due to the linear relationship between the reciprocal degree and in/outdegree of nodes. In the mesoscale, the distributions of four closed triples representing close-knit friendship structures are found to exhibit identical power-laws, a behavior attributed to the negligible degree correlations. Intriguingly, all the power-law exponents of the distributions in the local-scale and mesoscale depend only on one global parameter -- the mean in/outdegree, while both the mean in/outdegree and the reciprocity together determine the ratio of the reciprocal degree of a node to its in/outdegree.

preprint2012arXiv

Promotional effect on cold start problem and diversity in a data characteristic based recommendation method

Pure methods generally perform excellently in either recommendation accuracy or diversity, whereas hybrid methods generally outperform pure cases in both recommendation accuracy and diversity, but encounter the dilemma of optimal hybridization parameter selection for different recommendation focuses. In this article, based on a user-item bipartite network, we propose a data characteristic based algorithm, by relating the hybridization parameter to the data characteristic. Different from previous hybrid methods, the present algorithm adaptively assign the optimal parameter specifically for each individual items according to the correlation between the algorithm and the item degrees. Compared with a highly accurate pure method, and a hybrid method which is outstanding in both the recommendation accuracy and the diversity, our method shows a remarkably promotional effect on the long-standing challenging problem of the cold start, as well as the recommendation diversity, while simultaneously keeps a high overall recommendation accuracy. Even compared with an improved hybrid method which is highly efficient on the cold start problem, the proposed method not only further improves the recommendation accuracy of the cold items, but also enhances the recommendation diversity. Our work might provide a promising way to better solving the personal recommendation from the perspective of relating algorithms with dataset properties.

preprint2012arXiv

Recommender Systems

The ongoing rapid expansion of the Internet greatly increases the necessity of effective recommender systems for filtering the abundant information. Extensive research for recommender systems is conducted by a broad range of communities including social and computer scientists, physicists, and interdisciplinary researchers. Despite substantial theoretical and practical achievements, unification and comparison of different approaches are lacking, which impedes further advances. In this article, we review recent developments in recommender systems and discuss the major challenges. We compare and evaluate available algorithms and examine their roles in the future developments. In addition to algorithms, physical aspects are described to illustrate macroscopic behavior of recommender systems. Potential impacts and future directions are discussed. We emphasize that recommendation has a great scientific depth and combines diverse research fields which makes it of interests for physicists as well as interdisciplinary researchers.

preprint2012arXiv

Scaling Laws in Human Language

Zipf's law on word frequency is observed in English, French, Spanish, Italian, and so on, yet it does not hold for Chinese, Japanese or Korean characters. A model for writing process is proposed to explain the above difference, which takes into account the effects of finite vocabulary size. Experiments, simulations and analytical solution agree well with each other. The results show that the frequency distribution follows a power law with exponent being equal to 1, at which the corresponding Zipf's exponent diverges. Actually, the distribution obeys exponential form in the Zipf's plot. Deviating from the Heaps' law, the number of distinct words grows with the text length in three stages: It grows linearly in the beginning, then turns to a logarithmical form, and eventually saturates. This work refines previous understanding about Zipf's law and Heaps' law in language systems.

preprint2012arXiv

Social Recommender Systems Based on Coupling Network Structure Analysis

The past few years has witnessed the great success of recommender systems, which can significantly help users find relevant and interesting items for them in the information era. However, a vast class of researches in this area mainly focus on predicting missing links in bipartite user-item networks (represented as behavioral networks). Comparatively, the social impact, especially the network structure based properties, is relatively lack of study. In this paper, we firstly obtain five corresponding network-based features, including user activity, average neighbors' degree, clustering coefficient, assortative coefficient and discrimination, from social and behavioral networks, respectively. A hybrid algorithm is proposed to integrate those features from two respective networks. Subsequently, we employ a machine learning process to use those features to provide recommendation results in a binary classifier method. Experimental results on a real dataset, Flixster, suggest that the proposed method can significantly enhance the algorithmic accuracy. In addition, as network-based properties consider not only the social activities, but also take into account user preferences in the behavioral networks, therefore, it performs much better than that from either social or behavioral networks. Furthermore, since the features based on the behavioral network contain more diverse and meaningfully structural information, they play a vital role in uncovering users' potential preference, which, might show light in deeply understanding the structure and function of the social and behavioral networks.

preprint2012arXiv

Tag-Aware Recommender Systems: A State-of-the-art Survey

In the past decade, Social Tagging Systems have attracted increasing attention from both physical and computer science communities. Besides the underlying structure and dynamics of tagging systems, many efforts have been addressed to unify tagging information to reveal user behaviors and preferences, extract the latent semantic relations among items, make recommendations, and so on. Specifically, this article summarizes recent progress about tag-aware recommender systems, emphasizing on the contributions from three mainstream perspectives and approaches: network-based methods, tensor-based methods, and the topic-based methods. Finally, we outline some other tag-related works and future challenges of tag-aware recommendation algorithms.

preprint2011arXiv

Emergence of scale-free leadership structure in social recommender systems

The study of the organization of social networks is important for understanding of opinion formation, rumor spreading, and the emergence of trends and fashion. This paper reports empirical analysis of networks extracted from four leading sites with social functionality (Delicious, Flickr, Twitter and YouTube) and shows that they all display a scale-free leadership structure. To reproduce this feature, we propose an adaptive network model driven by social recommending. Artificial agent-based simulations of this model highlight a "good get richer" mechanism where users with broad interests and good judgments are likely to become popular leaders for the others. Simulations also indicate that the studied social recommendation mechanism can gradually improve the user experience by adapting to tastes of its users. Finally we outline implications for real online resource-sharing systems.

preprint2011arXiv

Self-organization in social tagging systems

Individuals often imitate each other to fall into the typical group, leading to a self-organized state of typical behaviors in a community. In this paper, we model self-organization in social tagging systems and illustrate the underlying interaction and dynamics. Specifically, we introduce a model in which individuals adjust their own tagging tendency to imitate the average tagging tendency. We found that when users are of low confidence, they tend to imitate others and lead to a self-organized state with active tagging. On the other hand, when users are of high confidence and are stubborn for changes, tagging becomes inactive. We observe a phase transition at a critical level of user confidence when the system changes from one regime to the other. The distributions of post length obtained from the model are compared to real data which show good agreements.

preprint2010arXiv

Bridgeness: A Local Index on Edge Significance in Maintaining Global Connectivity

Edges in a network can be divided into two kinds according to their different roles: some enhance the locality like the ones inside a cluster while others contribute to the global connectivity like the ones connecting two clusters. A recent study by Onnela et al uncovered the weak ties effects in mobile communication. In this article, we provide complementary results on document networks, that is, the edges connecting less similar nodes in content are more significant in maintaining the global connectivity. We propose an index named bridgeness to quantify the edge significance in maintaining connectivity, which only depends on local information of network topology. We compare the bridgeness with content similarity and some other structural indices according to an edge percolation process. Experimental results on document networks show that the bridgeness outperforms content similarity in characterizing the edge significance. Furthermore, extensive numerical results on disparate networks indicate that the bridgeness is also better than some well-known indices on edge significance, including the Jaccard coefficient, degree product and betweenness centrality.

preprint2010arXiv

Hypergraph model of social tagging networks

The past few years have witnessed the great success of a new family of paradigms, so-called folksonomy, which allows users to freely associate tags to resources and efficiently manage them. In order to uncover the underlying structures and user behaviors in folksonomy, in this paper, we propose an evolutionary hypergrah model to explain the emerging statistical properties. The present model introduces a novel mechanism that one can not only assign tags to resources, but also retrieve resources via collaborative tags. We then compare the model with a real-world dataset: \emph{Del.icio.us}. Indeed, the present model shows considerable agreement with the empirical data in following aspects: power-law hyperdegree distributions, negtive correlation between clustering coefficients and hyperdegrees, and small average distances. Furthermore, the model indicates that most tagging behaviors are motivated by labeling tags to resources, and tags play a significant role in effectively retrieving interesting resources and making acquaintance with congenial friends. The proposed model may shed some light on the in-depth understanding of the structure and function of folksonomy.

preprint2010arXiv

Solving the Cold-Start Problem in Recommender Systems with Social Tags

In this paper, based on the user-tag-object tripartite graphs, we propose a recommendation algorithm, which considers social tags as an important role for information retrieval. Besides its low cost of computational time, the experiment results of two real-world data sets, \emph{Del.icio.us} and \emph{MovieLens}, show it can enhance the algorithmic accuracy and diversity. Especially, it can obtain more personalized recommendation results when users have diverse topics of tags. In addition, the numerical results on the dependence of algorithmic accuracy indicates that the proposed algorithm is particularly effective for small degree objects, which reminds us of the well-known \emph{cold-start} problem in recommender systems. Further empirical study shows that the proposed algorithm can significantly solve this problem in social tagging systems with heterogeneous object degree distributions.

preprint2010arXiv

Zipf's Law Leads to Heaps' Law: Analyzing Their Relation in Finite-Size Systems

Background: Zipf's law and Heaps' law are observed in disparate complex systems. Of particular interests, these two laws often appear together. Many theoretical models and analyses are performed to understand their co-occurrence in real systems, but it still lacks a clear picture about their relation. Methodology/Principal Findings: We show that the Heaps' law can be considered as a derivative phenomenon if the system obeys the Zipf's law. Furthermore, we refine the known approximate solution of the Heaps' exponent provided the Zipf's exponent. We show that the approximate solution is indeed an asymptotic solution for infinite systems, while in the finite-size system the Heaps' exponent is sensitive to the system size. Extensive empirical analysis on tens of disparate systems demonstrates that our refined results can better capture the relation between the Zipf's and Heaps' exponents. Conclusions/Significance: The present analysis provides a clear picture about the relation between the Zipf's law and Heaps' law without the help of any specific stochastic model, namely the Heaps' law is indeed a derivative phenomenon from Zipf's law. The presented numerical method gives considerably better estimation of the Heaps' exponent given the Zipf's exponent and the system size. Our analysis provides some insights and implications of real complex systems, for example, one can naturally obtained a better explanation of the accelerated growth of scale-free networks.