Source author record

Loet Leydesdorff

Loet Leydesdorff appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

159works
17topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

159 published item(s)

preprint2022arXiv

A discussion of measuring the top-1 percent most-highly cited publications: Quality and impact of Chinese papers

The top 1 percent most highly cited articles are watched closely as the vanguards of the sciences. Using Web of Science data, one can find that China had overtaken the USA in the relative participation in the top 1 percent in 2019, after outcompeting the EU on this indicator in 2015. However, this finding contrasts with repeated reports of Western agencies that the quality of Chinese output in science is lagging other advanced nations, even as it has caught up in numbers of articles. The difference between the results presented here and the previous results depends mainly upon field normalizations, which classify source journals by discipline. Average citation rates of these subsets are commonly used as a baseline so that one can compare among disciplines. However, the expected value of the top 1 percent of a sample of N papers is N 100, ceteris paribus. Using the average citation rates as expected values, errors are introduced by using the mean of highly skewed distributions and a specious precision in the delineations of the subsets. Classifications can be used for the decomposition, but not for the normalization. When the data is thus decomposed, the USA ranks ahead of China in biomedical fields such as virology. Although the number of papers is smaller, China outperforms the US in the field of Business and Finance in the Social Sciences Citation Index when p is less than .05. Using percentile ranks, subsets other than indexing based classifications can be tested for the statistical significance of differences among them.

preprint2016arXiv

Citation algorithms for identifying research milestones driving biomedical innovation

Scientific activity plays a major role in innovation for biomedicine and healthcare. For instance, fundamental research on disease pathologies and mechanisms can generate potential targets for drug therapy. This co-evolution is punctuated by papers which provide new perspectives and open new domains. Despite the relationship between scientific discovery and biomedical advancement, identifying these research milestones that truly impact biomedical innovation can be difficult and is largely based solely on the opinions of subject matter experts. Here, we consider whether a new class of citation algorithms that identify seminal scientific works in a field, Reference Publication Year Spectroscopy (RPYS) and multi-RPYS, can identify the connections between innovation (e.g. therapeutic treatments) and the foundational research underlying them. Specifically, we assess whether the results of these analytic techniques converge with expert opinions on research milestones driving biomedical innovation in the treatment of Basal Cell Carcinoma. Our results show that these algorithms successfully identify the majority of milestone papers detailed by experts (Wong and Dlugosz 2014) thereby validating the power of these algorithms to converge on independent opinions of seminal scientific works derived by subject matter experts. These advances offer an opportunity to identify scientific activities enabling innovation in biomedicine.

preprint2016arXiv

Citations: Indicators of Quality? The Impact Fallacy

We argue that citation is a composed indicator: short-term citations can be considered as currency at the research front, whereas long-term citations can contribute to the codification of knowledge claims into concept symbols. Knowledge claims at the research front are more likely to be transitory and are therefore problematic as indicators of quality. Citation impact studies focus on short-term citation, and therefore tend to measure not epistemic quality, but involvement in current discourses in which contributions are positioned by referencing. We explore this argument using three case studies: (1) citations of the journal Soziale Welt as an example of a venue that tends not to publish papers at a research front, unlike, for example, JACS; (2) Robert Merton as a concept symbol across theories of citation; and (3) the Multi-RPYS ("Multi-Referenced Publication Year Spectroscopy") of the journals Scientometrics, Gene, and Soziale Welt. We show empirically that the measurement of "quality" in terms of citations can further be qualified: short-term citation currency at the research front can be distinguished from longer-term processes of incorporation and codification of knowledge claims into bodies of knowledge. The recently introduced Multi-RPYS can be used to distinguish between short-term and long-term impacts.

preprint2016arXiv

Cited References and Medical Subject Headings (MeSH) as Two Different Knowledge Representations: Clustering and Mappings at the Paper Level

For the biomedical sciences, the Medical Subject Headings (MeSH) make available a rich feature which cannot currently be merged properly with widely used citing/cited data. Here, we provide methods and routines that make MeSH terms amenable to broader usage in the study of science indicators: using Web-of-Science (WoS) data, one can generate the matrix of citing versus cited documents; using PubMed/MEDLINE data, a matrix of the citing documents versus MeSH terms can be generated analogously. The two matrices can also be reorganized into a 2-mode matrix of MeSH terms versus cited references. Using the abbreviated journal names in the references, one can, for example, address the question whether MeSH terms can be used as an alternative to WoS Subject Categories for the purpose of normalizing citation data. We explore the applicability of the routines in the case of a research program about the amyloid cascade hypothesis in Alzheimer's disease (AD). One conclusion is that referenced journals provide archival structures, whereas MeSH terms indicate mainly variation (including novelty) at the research front. Furthermore, we explore the option of using the citing/cited matrix for main-path analysis as a by-product of the software.

preprint2016arXiv

Co-word Maps and Topic Modeling: A Comparison Using Small and Medium-Sized Corpora (n < 1000)

Induced by "big data," "topic modeling" has become an attractive alternative to mapping co-words in terms of co-occurrences and co-absences using network techniques. Does topic modeling provide an alternative for co-word mapping in research practices using moderately sized document collections? We return to the word/document matrix using first a single text with a strong argument ("The Leiden Manifesto") and then upscale to a sample of moderate size (n = 687) to study the pros and cons of the two approaches in terms of the resulting possibilities for making semantic maps that can serve an argument. The results from co-word mapping (using two different routines) versus topic modeling are significantly uncorrelated. Whereas components in the co-word maps can easily be designated, the topic models provide sets of words that are very differently organized. In these samples, the topic models seem to reveal similarities other than semantic ones (e.g., linguistic ones). In other words, topic modeling does not replace co-word mapping in small and medium-sized sets; but the paper leaves open the possibility that topic modeling would work well for the semantic mapping of large sets.

preprint2016arXiv

Construction of a Pragmatic Base Line for Journal Classifications and Maps Based on Aggregated Journal-Journal Citation Relations

A number of journal classification systems have been developed in bibliometrics since the launch of the Citation Indices by the Institute of Scientific Information (ISI) in the 1960s. These systems are used to normalize citation counts with respect to field-specific citation patterns. The best known system is the so-called "Web-of-Science Subject Categories" (WCs). In other systems papers are classified by algorithmic solutions. Using the Journal Citation Reports 2014 of the Science Citation Index and the Social Science Citation Index (n of journals = 11,149), we examine options for developing a new system based on journal classifications into subject categories using aggregated journal-journal citation data. Combining routines in VOSviewer and Pajek, a tree-like classification is developed. At each level one can generate a map of science for all the journals subsumed under a category. Nine major fields are distinguished at the top level. Further decomposition of the social sciences is pursued for the sake of example with a focus on journals in information science (LIS) and science studies (STS). The new classification system improves on alternative options by avoiding the problem of randomness in each run that has made algorithmic solutions hitherto irreproducible. Limitations of the new system are discussed (e.g. the classification of multi-disciplinary journals). The system's usefulness for field-normalization in bibliometrics should be explored in future studies.

preprint2016arXiv

Full and Fractional Counting in Bibliometric Networks

In their study entitled "Constructing bibliometric networks: A comparison between full and fractional counting," Perianes-Rodriguez, Waltman, & van Eck (2016; henceforth abbreviated as PWvE) provide arguments for the use of fractional counting at the network level as different from the level of publications. Whereas fractional counting in the latter case divides the credit among co-authors (countries, institutions, etc.), fractional counting at the network level can normalize the relative weights of links and thereby clarify the structures in the network. PWvE, however, propose a counting scheme for fractional counting that is one among other possible ones. Alternative schemes proposed by Batagelj and Cerinšek (2013) and Park, Yoon, & Leydesdorff (2016; henceforth abbreviated as PYL) are discussed in an appendix. However, our approach is not correctly identified as identical to their Equation A3. Here below, we distinguish three approaches analytically; routines for applying these approaches to bibliometric data are also provided.

preprint2016arXiv

Growth of International Cooperation in Science: Revisiting Six Case Studies

International collaboration in science continues to grow at a remarkable rate, but little agreement exists about dynamics of growth and organization at the discipline level. Some suggest that disciplines differ in their collaborative tendencies, reflecting their epistemic culture. This study examines collaborative patterns in six previously studied specialties to add new data and conduct analyses over time. Our findings show that the global network of collaboration continues to add new nations and new participants; each specialty has added many new nations to its lists of collaborating partners since 1990. We also find that the scope of international collaboration is positively related to impact. Network characteristics for the six specialties are notable in that instead of reflecting underlying culture, they tend towards convergence. This observation suggests that the global level may represent next-order dynamics that feed back to the national and local levels (as subsystems) in a complex, networked hierarchy.

preprint2016arXiv

Identification of long-term concept-symbols among citations: Can documents be clustered in terms of common intellectual histories?

"Citation classics" are not only highly cited, but also cited during several decades. We test whether the peaks in the spectrograms generated by Reference Publication Years Spectroscopy (RPYS) indicate such long-term impact by comparing across RPYS for subsequent time intervals. Multi-RPYS enables us to distinguish between short-term citation peaks at the research front that decay within ten years versus historically constitutive (long-term) citations that function as concept symbols (Small, 1978). Using these constitutive citations, one is able to cluster document sets (e.g., journals) in terms of intellectually shared histories. We test this premise by clustering 40 journals in the Web of Science Category of Information and Library Science using multi-RPYS. It follows that RPYS can not only be used for retrieving roots of sets under study (cited), but also for algorithmic historiography of the citing sets. Significant references are historically rooted symbols among other citations that function as currency.

preprint2016arXiv

Introducing CitedReferencesExplorer (CRExplorer): A program for Reference Publication Year Spectroscopy with Cited References Standardization

We introduce a new tool - the CitedReferencesExplorer (CRExplorer, www.crexplorer.net) - which can be used to disambiguate and analyze the cited references (CRs) of a publication set downloaded from the Web of Science (WoS). The tool is especially suitable to identify those publications which have been frequently cited by the researchers in a field and thereby to study for example the historical roots of a research field or topic. CRExplorer simplifies the identification of key publications by enabling the user to work with both a graph for identifying most frequently cited reference publication years (RPYs) and the list of references for the RPYs which have been most frequently cited. A further focus of the program is on the standardization of CRs. It is a serious problem in bibliometrics that there are several variants of the same CR in the WoS. In this study, CRExplorer is used to study the CRs of all papers published in the Journal of Informetrics. The analyses focus on the most important papers published between 1980 and 1990.

preprint2016arXiv

Measuring the match between evaluators and evaluees: Cognitive distances between panel members and research groups at the journal level

When research groups are evaluated by an expert panel, it is an open question how one can determine the match between panel and research groups. In this paper, we outline two quantitative approaches that determine the cognitive distance between evaluators and evaluees, based on the journals they have published in. We use example data from four research evaluations carried out between 2009 and 2014 at the University of Antwerp. While the barycenter approach is based on a journal map, the similarity-adapted publication vector (SAPV) approach is based on the full journal similarity matrix. Both approaches determine an entity's profile based on the journals in which it has published. Subsequently, we determine the Euclidean distance between the barycenter or SAPV profiles of two entities as an indicator of the cognitive distance between them. Using a bootstrapping approach, we determine confidence intervals for these distances. As such, the present article constitutes a refinement of a previous proposal that operates on the level of Web of Science subject categories.

preprint2016arXiv

New features of CitedReferencesExplorer (CRExplorer)

CRExplorer version 1.6.7 was released on July 5, 2016. This version includes the following new features and improvements: Scopus: Using "File" - "Import" - "Scopus", CRExplorer reads files from Scopus. The file format "CSV" (including citations, abstracts and references) should be chosen in Scopus for downloading records. Export facilities: Using "File" - "Export" - "Scopus", CRExplorer exports files in the Scopus format. Using "File" - "Export" - "Web of Science", CRExplorer exports files in the Web of Science format. These files can be imported in other bibliometric programs (e.g. VOSviewer). Space bar: Select a specific cited reference in the cited references table, press the space bar, and all bibliographic details of the CR are shown. Internal file format: Using "File" - "Save", working files are saved in the internal file format "*.cre". The files include all data including matching results and manual matching corrections. The files can be opened by using "File" - "Open".

preprint2016arXiv

Patent Portfolio Analysis of Cities: Statistics and Maps of Technological Inventiveness

Cities are engines of the knowledge-based economy, because they are the primary sites of knowledge production activities that subsequently shape the rate and direction of technological change and economic growth. Patents provide a wealth of information to analyse the knowledge specialization at specific places, such as technological details and information on inventors and entities involved, including address information. The technology codes on each patent document indicate the specialization and scope of the underlying technological knowledge of a given invention. In this paper we introduce tools for portfolio analysis in terms of patents that provide insights into the technological specialization of cities. The mapping and analysis of patent portfolios of cities using data of the Unites States Patent and Trademark Office (USPTO) website (at http://www.uspto.gov) and dedicated tools (at http://www.leydesdorff.net/portfolio) can be used to analyse the specialisation patterns of inventive activities among cities. The results allow policy makers and other stakeholders to identify promising areas of further knowledge development and 'smart specialisation' strategies.

preprint2016arXiv

Professional and Citizen Bibliometrics: Complementarities and ambivalences in the development and use of indicators

Bibliometric indicators such as journal impact factors, h-indices, and total citation counts are algorithmic artifacts that can be used in research evaluation and management. These artifacts have no meaning by themselves, but receive their meaning from attributions in institutional practices. We distinguish four main stakeholders in these practices: (1) producers of bibliometric data and indicators; (2) bibliometricians who develop and test indicators; (3) research managers who apply the indicators; and (4) the scientists being evaluated with potentially competing career interests. These different positions may lead to different and sometimes conflicting perspectives on the meaning and value of the indicators. The indicators can thus be considered as boundary objects which are socially constructed in translations among these perspectives. This paper proposes an analytical clarification by listing an informed set of (sometimes unsolved) problems in bibliometrics which can also shed light on the tension between simple but invalid indicators that are widely used (e.g., the h-index) and more sophisticated indicators that are not used or cannot be used in evaluation practices because they are not transparent for users, cannot be calculated, or are difficult to interpret.

preprint2016arXiv

Referenced Publication Year Spectroscopy (RPYS) and Algorithmic Historiography: The Bibliometric Reconstruction of András Schubert's Œuvre

Referenced Publication Year Spectroscopy (RPYS) was recently introduced as a method to analyze the historical roots of research fields and groups or institutions. RPYS maps the distribution of the publication years of the cited references in a document set. In this study, we apply this methodology to the œuvre of an individual researcher on the occasion of a Festschrift for András Schubert's 70th birthday. We discuss the different options of RPYS in relation to one another (e.g. Multi-RPYS), and in relation to the longer-term research program of algorithmic historiography (e.g., HistCite) based on Schubert's publications (n=172) and cited references therein as a bibliographic domain in scientometrics. Main path analysis and Multi-RPYS of the citation network are used to show the changes and continuities in Schubert's intellectual career. Diachronic and static decomposition of a document set can lead to different results, while the analytically distinguishable lines of research may overlap and interact over time, and intermittent.

preprint2016arXiv

RPYS i/o: A web-based tool for the historiography and visualization of citation classics, sleeping beauties, and research fronts

Reference Publication Year Spectroscopy (RPYS) and Multi-RPYS provide algorithmic approaches to reconstructing the intellectual histories of scientific fields. With this brief communication, we describe a technical advancement for developing research historiographies by introducing RPYS i/o, an online tool for performing standard RPYS and Multi-RPYS analyses interactively (at http://comins.leydesdorff.net/). The tool enables users to explore seminal works underlying a research field and to plot the influence of these seminal works over time. This suite of visualizations offers the potential to analyze and visualize the myriad of temporal dynamics of scientific influence, such as citation classics, sleeping beauties, and the dynamics of research fronts. We demonstrate the features of the tool by analyzing--as an example--the references in documents published in the journal Philosophy of Science.

preprint2016arXiv

Science Visualization and Discursive Knowledge

Positional and relational perspectives on network data have led to two different research traditions in textual analysis and social network analysis, respectively. Latent Semantic Analysis (LSA) focuses on the latent dimensions in textual data; social network analysis (SNA) on the observable networks. The two coupled topographies of information-processing in the network space and meaning-processing in the vector space operate with different (nonlinear) dynamics. The historical dynamics of information processing in observable networks organizes the system into instantiations; the systems dynamics, however, can be considered as self-organizing in terms of fluxes of communication along the various dimensions that operate with different codes. The development over time adds evolutionary differentiation to the historical integration; a richer structure can process more complexity.

preprint2016arXiv

Skewness of citation impact data and covariates of citation distributions: A large-scale empirical analysis based on Web of Science data

Using percentile shares, one can visualize and analyze the skewness in bibliometric data across disciplines and over time. The resulting figures can be intuitively interpreted and are more suitable for detailed analysis of the effects of independent and control variables on distributions than regression analysis. We show this by using percentile shares to analyze so-called "factors influencing citation impact" (FICs; e.g., the impact factor of the publishing journal) across year and disciplines. All articles (n= 2,961,789) covered by WoS in 1990 (n= 637,301), 2000 (n= 919,485), and 2010 (n= 1,405,003) are used. In 2010, nearly half of the citation impact is accounted for by the 10% most-frequently cited papers; the skewness is largest in the humanities (68.5% in the top-10% layer) and lowest in agricultural sciences (40.6%). The comparison of the effects of the different FICs (the number of cited references, number of authors, number of pages, and JIF) on citation impact shows that JIF has indeed the strongest correlations with the citation scores. However, the correlation between FICs and citation impact is lower, if citations are normalized instead of using raw citation counts.

preprint2016arXiv

The "Tournaments" Metaphor in Citation Impact Studies: Power-Weakness Ratios (PWR) as a Journal Indicator

Ramanujacharyulu's (1964) Power-Weakness Ratio (PWR) measures impact by recursively multiplying the citation matrix by itself until convergence is reached in both the cited and citing dimensions; the quotient of these values is defined as PWR, whereby "cited" is considered as power and "citing" as weakness. Analytically, PWR is an attractive candidate for measuring journal impact because of its symmetrical handling of the rows and columns in the asymmetrical citation matrix, its recursive algorithm, and its mathematical elegance. In this study, PWR is discussed and critically assessed in relation to other size-independent recursive metrics. A test using the set of 83 journals in "information and library science" (according to the Web-of-Science categorization) converged, but did not provide interpretable results. Further decomposition of this set into homogeneous sub-graphs shows that--like most other journal indicators--PWR can perhaps be used within homogeneous sets, but not across citation communities.

preprint2016arXiv

The Normalization of Co-authorship Networks in the Bibliometric Evaluation: The Government Stimulation Programs of China and Korea

Using co-authored publications between China and Korea in Web of Science (WoS) during the one-year period of 2014, we evaluate the government stimulation program for collaboration between China and Korea. In particular, we apply dual approaches, full integer vs. fractional counting, to collaborative publications in order to better examine both the patterns and contents of Sino-Korean collaboration networks in terms of individual countries and institutions. We first conduct a semi-automatic network analysis of Sino-Korean publications based on the full-integer counting method, and then compare our categorization with contextual rankings using the fractional technique; routines for fractional counting of WoS data are made available at http://www.leydesdorff.net/software/fraction . Increasing international collaboration leads paradoxically to lower numbers of publications and citations using fractional counting for performance measurement. However, integer counting is not an appropriate measure for the evaluation of the stimulation of collaborations. Both integer and fractional analytics can be used to identify important countries and institutions, but with other research questions.

preprint2016arXiv

The Self-Organization of Meaning and the Reflexive Communication of Information

Following a suggestion of Warren Weaver, we extend the Shannon model of communication piecemeal into a complex systems model in which communication is differentiated both vertically and horizontally. This model enables us to bridge the divide between Niklas Luhmann's theory of the self-organization of meaning in communications and empirical research using information theory. First, we distinguish between communication relations and correlations among patterns of relations. The correlations span a vector space in which relations are positioned and can be provided with meaning. Second, positions provide reflexive perspectives. Whereas the different meanings are integrated locally, each instantiation opens global perspectives--"horizons of meaning"--along eigenvectors of the communication matrix. These next-order codifications of meaning can be expected to generate redundancies when interacting in instantiations. Increases in redundancy indicate new options and can be measured as local reduction of prevailing uncertainty (in bits). The systemic generation of new options can be considered as a hallmark of the knowledge-based economy.

preprint2015arXiv

A Review of Theory and Practice in Scientometrics

Scientometrics is the study of the quantitative aspects of the process of science as a communication system. It is centrally, but not only, concerned with the analysis of citations in the academic literature. In recent years it has come to play a major role in the measurement and evaluation of research performance. In this review we consider: the historical development of scientometrics, sources of citation data, citation metrics and the "laws" of scientometrics, normalisation, journal impact factors and other journal metrics, visualising and mapping science, evaluation and policy, and future developments.

preprint2015arXiv

Bibliometrics/Citation networks

In addition to shaping social networks, for example, in terms of co-authorship relations, scientific communications induce and reproduce cognitive structures. Scientific literature is intellectually organized in terms of disciplines and specialties; these structures are reproduced and networked reflexively by making references to the authors, concepts and texts embedded in these literatures. The concept of a cognitive structure was introduced in social network analysis (SNA) in 1987 by David Krackhardt, but the focus in SNA has hitherto been on cognition as a psychological attribute of human agency. In bibliometrics, and in science and technology studies (STS) more generally, socio-cognitive structures refer to intellectual organization at the supra-individual level. This intellectual organization emerges and is reproduced by the collectives of authors who are organized not only in terms of inter-personal relations, but also more abstractly in terms of codes of communication that are field-specific. Citations can serve as indicators of this codification process.

preprint2015arXiv

Can "Hot Spots" in the Sciences Be Mapped Using the Dynamics of Aggregated Journal-Journal Citation Relations?

Using three years of the Journal Citation Reports (2011, 2012, and 2013), indicators of transitions in 2012 (between 2011 and 2013) are studied using methodologies based on entropy statistics. Changes can be indicated at the level of journals using the margin totals of entropy production along the row or column vectors, but also at the level of links among journals by importing the transition matrices into network analysis and visualization programs (and using community-finding algorithms). Seventy-four journals are flagged in terms of discontinuous changes in their citations; but 3,114 journals are involved in "hot" links. Most of these links are embedded in a main component; 78 clusters (containing 172 journals) are flagged as potential "hot spots" emerging at the network level. An additional finding is that PLoS ONE introduced a new communication dynamics into the database. The limitations of the methodology are elaborated using an example. The results of the study indicate where developments in the citation dynamics can be considered as significantly unexpected. This can be used as heuristic information; but what a "hot spot" in terms of the entropy statistics of aggregated citation relations means substantively can be expected to vary from case to case.

preprint2015arXiv

Can Intellectual Processes in the Sciences Also Be Simulated? The Anticipation and Visualization of Possible Future States

Socio-cognitive action reproduces and changes both social and cognitive structures. The analytical distinction between these dimensions of structure provides us with richer models of scientific development. In this study, I assume that (i) social structures organize expectations into belief structures that can be attributed to individuals and communities; (ii) expectations are specified in scholarly literature; and (iii) intellectually the sciences (disciplines, specialties) tend to self-organize as systems of rationalized expectations. Whereas social organizations remain localized, academic writings can circulate, and expectations can be stabilized and globalized using symbolically generalized codes of communication. The intellectual restructuring, however, remains latent as a second-order dynamics that can be accessed by participants only reflexively. Yet, the emerging "horizons of meaning" provide feedback to the historically developing organizations by constraining the possible future states as boundary conditions. I propose to model these possible future states using incursive and hyper-incursive equations from the computation of anticipatory systems. Simulations of these equations enable us to visualize the couplings among the historical--i.e., recursive--progression of social structures along trajectories, the evolutionary--i.e., hyper-incursive--development of systems of expectations at the regime level, and the incursive instantiations of expectations in actions, organizations, and texts.

preprint2015arXiv

Can Technology Life-Cycles Be Indicated by Diversity in Patent Classifications? The crucial role of variety

In a previous study of patent classifications in nine material technologies for photovoltaic cells, Leydesdorff et al. (2015) reported cyclical patterns in the longitudinal development of Rao-Stirling diversity. We suggested that these cyclical patterns can be used to indicate technological life-cycles. Upon decomposition, however, the cycles are exclusively due to increases and decreases in the variety of the classifications, and not to disparity or technological distance, measured as (1 - cosine). A single frequency component can accordingly be shown in the periodogram. Furthermore, the cyclical patterns are associated with the numbers of inventors in the respective technologies. Sometimes increased variety leads to a boost in the number of inventors, but in early phases--when the technology is still under constructio--it can also be the other way round. Since the development of the cycles thus seems independent of technological distances among the patents, the visualization in terms of patent maps can be considered as addressing an analytically different set of research questions.

preprint2015arXiv

Highly-cited papers in Library and Information Science (LIS): Authors, institutions, and network structures

As a follow-up to the highly-cited authors list published by Thomson Reuters in June 2014, we analyze the top-1% most frequently cited papers published between 2002 and 2012 included in the Web of Science (WoS) subject category "Information Science & Library Science." 798 authors contributed to 305 top-1% publications; these authors were employed at 275 institutions. The authors at Harvard University contributed the largest number of papers, when the addresses are whole-number counted. However, Leiden University leads the ranking, if fractional counting is used. Twenty-three of the 798 authors were also listed as most highly-cited authors by Thomson Reuters in June 2014 (http://highlycited.com/). Twelve of these 23 authors were involved in publishing four or more of the 305 papers under study. Analysis of co-authorship relations among the 798 highly-cited scientists shows that co-authorships are based on common interests in a specific topic. Three topics were important between 2002 and 2012: (1) collection and exploitation of information in clinical practices, (2) the use of internet in public communication and commerce, and (3) scientometrics.

preprint2015arXiv

Hyperincursive Cogitata and Incursive Cogitantes: Scholarly Discourse as a Strongly Anticipatory System

Strongly anticipatory systems-that is, systems which use models of themselves for their further development-and which additionally may be able to run hyperincursive routines-that is, develop only with reference to their future states-cannot exist in res extensa, but can only be envisaged in res cogitans. One needs incursive routines in cogitantes to instantiate these systems. Unlike historical systems (with recursion), these hyper-incursive routines generate redundancies by opening horizons of other possible states. Thus, intentional systems can enrich our perceptions of the cases that have happened to occur. The perspective of hindsight codified at the above-individual level enables us furthermore to intervene technologically. The theory and computation of anticipatory systems have made these loops between supra-individual hyper-incursion, individual incursion (in instantiation), and historical recursion accessible for modeling and empirical investigation.

preprint2015arXiv

Journal Portfolio Analysis for Countries, Cities, and Organizations: Maps and Comparisons

Using Web-of-Science data, portfolio analysis in terms of journal coverage can be projected on a base map for units of analysis such as countries, cities, universities, and firms. The units of analysis under study can be compared statistically across the 10,000+ journals. The interdisciplinarity of the portfolios is measured using Rao-Stirling diversity or Zhang et al.'s (in press) improved measure 2D3. At the country level we find regional differentiation (e.g., Latin-American or Asian countries), but also a major divide between advanced and less-developed countries. Israel and Israeli cities outperform other nations and cities in terms of diversity. Universities appear to be specifically related to firms when a number of these units are exploratively compared. The instrument is relatively simple and straightforward, and one can generalize the application to any document set retrieved from WoS. Further instruction is provided online at http://www.leydesdorff.net/portfolio .

preprint2015arXiv

Networks of reader and country status: An analysis of Mendeley reader statistics

The number of papers published in journals indexed by the Web of Science core collection is steadily increasing. In recent years, nearly two million new papers were published each year; somewhat more than one million papers when primary research papers are considered only (articles and reviews are the document types where primary research is usually reported or reviewed). However, who reads these papers? More precisely, which groups of researchers from which (self-assigned) scientific disciplines and countries are reading these papers? Is it possible to visualize readership patterns for certain countries, scientific disciplines, or academic status groups? One popular method to answer these questions is a network analysis. In this study, we analyze Mendeley readership data of a set of 1,133,224 articles and 64,960 reviews with publication year 2012 to generate three different kinds of networks: (1) The network based on disciplinary affiliations of Mendeley readers contains four groups: (i) biology, (ii) social science and humanities (including relevant computer science), (iii) bio-medical sciences, and (iv) natural science and engineering. In all four groups, the category with the addition "miscellaneous" prevails. (2) The network of co-readers in terms of professional status shows that a common interest in papers is mainly shared among PhD students, Master's students, and postdocs. (3) The country network focusses on global readership patterns: a group of 53 nations is identified as core to the scientific enterprise, including Russia and China as well as two thirds of the OECD (Organisation for Economic Co-operation and Development) countries.

preprint2015arXiv

Recent Developments in China-U.S. Cooperation in Science

China's remarkable gains in science over the past 25 years have been well documented (e.g., Jin and Rousseau, 2005a; Zhou and Leydesdorff, 2006; Shelton & Foland, 2009) but it is less well known that China and the United States have become each other's top collaborating country. Science and technology has been a primary vehicle for growing the bilateral relationship between China and the United States since the opening of relations between the two countries in the late 1970s. During the 2000s, the scientific relationship between China and the United States--as measured in coauthored papers--showed significant growth. Chinese scientists claim first authorship much more frequently than U.S. counterparts by the end of the decade. The sustained rate of increase of collaboration with one other country is unprecedented on the U.S. side. Even growth in relations with eastern European nations does not match the growth in the relationship between China and the United States. Both countries can benefit from the relationship, but for the U.S., greater benefit would come from a more targeted strategy.

preprint2015arXiv

Regional and Global Science: Latin American and Caribbean publications in the SciELO Citation Index and the Web of Science

We compare the visibility of Latin American and Caribbean (LAC) publications in the Core Collection indexes of the Web of Science (WoS)--Science Citation Index Expanded, Social Sciences Citation Index, and Arts & Humanities Citation Index--and the SciELO Citation Index (SciELO CI) which was integrated into the larger WoS platform in 2014. The purpose of this comparison is to contribute to our understanding of the communication of scientific knowledge produced in Latin America and the Caribbean, and to provide some reflections on the potential benefits of the articulation of regional indexing exercises into WoS for a better understanding of geographic and disciplinary contributions. How is the regional level of SciELO CI related to the global range of WoS? In WoS, LAC authors are integrated at the global level in international networks, while SciELO has provided a platform for interactions among LAC researchers. The articulation of SciELO into WoS may improve the international visibility of the regional journals, but at the cost of independent journal inclusion criteria.

preprint2015arXiv

Replicability and the public/private divide

In a recent letter, Carlos Vilchez-Roman criticizes Bornmann et al. (2015) for using data which cannot be reproduced without access to an in-house version of the Web-of-Science (WoS) at the Max Planck Digital Libraries (MPDL, Munich). We agree with the norm of replicability and therefore returned to our data. Is the problem only a practical one of automation or does the in-house processing add analytical value to the data? Is the newly emerging situation in any sense different from a further professionalization of the field? In our opinion, a political economy of science indicators has in the meantime emerged with a competitive dynamic that affects the intellectual organization of the field.

preprint2015arXiv

Scientometrics and Science Studies: From Words and Co-Words to Information and Probabilistic Entropy

The tension between qualitative theorizing and quantitative methods is pervasive in the social sciences, and poses a constant challenge to empirical research. But in science studies as an interdisciplinary specialty, there are additional reasons why a more reflexive consciousness of the differences among the relevant disciplines is necessary. How can qualitative insights from the history of ideas and the sociology of science be combined with the quantitative perspective? By using the example of the lexical and semantic value of word occurrences, the issue of qualitatively different meanings of the same phenomena is discussed as a methodological problem. Nine criteria for methods which are needed for the development of science studies as an integrated enterprise can then be specified. Information calculus is suggested as a method which can comply with these criteria.

preprint2015arXiv

Simple arithmetic versus intuitive understanding: The case of the impact factor

We show that as a consequence of basic properties of elementary arithmetic journal impact factors show a counterintuitive behaviour with respect to adding non-cited articles. Synchronous as well as diachronous journal impact factors are affected. Our findings provide a rationale for not taking uncitable publications into account in impact factor calculations, at least if these items are truly uncitable.

preprint2015arXiv

Strategic Intelligence on Emerging Technologies: Scientometric Overlay Mapping

This paper examines the use of scientometric overlay mapping as a tool of 'strategic intelligence' to aid the governance of emerging technologies. We develop an integrative synthesis of different overlay mapping techniques and associated perspectives on technological emergence across the geographical, social, and cognitive spaces. To do so, we longitudinally analyse (with publication and patent data) three case-studies of emerging technologies in the medical domain. These are: RNA interference (RNAi), Human Papilloma Virus (HPV) testing technologies for cervical cancer, and Thiopurine Methyltransferase (TPMT) genetic testing. Given the flexibility (i.e. adaptability to different sources of data) and granularity (i.e. applicability across multiple levels of data aggregation) of overlay mapping techniques, we argue that these techniques can favour the integration and comparison of results from different contexts and cases, thus potentially functioning as platform for a 'distributed' strategic intelligence for analysts and decision-makers.

preprint2015arXiv

The Dynamics of Triads in Aggregated Journal-Journal Citation Relations: Specialty Developments at the Above-Journal Level

Dyads of journals related by citations can agglomerate into specialties through the mechanism of triadic closure. Using the Journal Citation Reports 2011, 2012, and 2013, we analyze triad formation as indicators of integration (specialty growth) and disintegration (restructuring). The strongest integration is found among the large journals that report on studies in different scientific specialties, such as PLoS ONE, Nature Communications, Nature, and Science. This tendency towards large-scale integration has not yet stabilized. Using the Islands algorithm, we also distinguish 51 local maxima of integration. We zoom into the cited articles that carry the integration for: (i) a new development within high-energy physics and (ii) an emerging interface between the journals Applied Mathematical Modeling and the International Journal of Advanced Manufacturing Technology. In the first case, integration is brought about by a specific communication reaching across specialty boundaries, whereas in the second, the dyad of journals indicates an emerging interface between specialties. These results suggest that integration picks up substantive developments at the specialty level. An advantage of the bottom-up method is that no ex ante classification of journals is assumed in the dynamic analysis.

preprint2015arXiv

The Globalization of Academic Entrepreneurship? The Recent Growth (2009-2014) in University Patenting Decomposed

The contribution of academia to US patents has become increasingly global. Following a pause, with a relatively flat rate, from 1998 to 2008, the long-term trend of university patenting rising as a share of all patenting has resumed, driven by the internationalization of academic entrepreneurship and the persistence of US university technology transfer. We disaggregate this recent growth in university patenting at the US Patent and Trademark Organization (USPTO) in terms of nations and patent classes. Foreign patenting in the US has almost doubled during the period 2009-2014, mainly due to patenting by universities in Taiwan, Korea, China, and Japan. These nations compete with the US in terms of patent portfolios, whereas most European countries--with the exception of the UK--have more specific portfolios, mainly in the bio-medical fields. In the case of China, Tsinghua University holds 63% of the university patents in USPTO, followed by King Fahd University with 55.2% of the national portfolio.

preprint2015arXiv

The Normalization of Occurrence and Co-occurrence Matrices in Bibliometrics using Cosine Similarities and Ochiai Coefficients

We prove that Ochiai similarity of the co-occurrence matrix is equal to cosine similarity in the underlying occurrence matrix. Neither the cosine nor the Pearson correlation should be used for the normalization of co-occurrence matrices because the similarity is then normalized twice, and therefore over-estimated; the Ochiai coefficient can be used instead. Results are shown using a small matrix (5 cases, 4 variables) for didactic reasons, and also Ahlgren et al.'s (2003) co-occurrence matrix of 24 authors in library and information sciences. The over-estimation is shown numerically and will be illustrated using multidimensional scaling and cluster dendograms. If the occurrence matrix is not available (such as in internet research or author co-citation analysis) using Ochiai for the normalization is preferable to using the cosine.

preprint2014arXiv

Aggregated journal-journal citation relations in Scopus and Web-of-Science matched and compared in terms of networks, maps, and interactive overlays

We compare the network of aggregated journal-journal citation relations provided by the Journal Citation Reports (JCR) 2012 of the Science and Social Science Citation Indexes (SCI and SSCI) with similar data based on Scopus 2012. First, global maps were developed for the two sets separately; sets of documents can then be compared using overlays to both maps. Using fuzzy-string matching and ISSN numbers, we were able to match 10,524 journal names between the two sets; that is, 96.4% of the 10,936 journals contained in JCR or 51.2% of the 20,554 journals covered by Scopus. Network analysis was then pursued on the set of journals shared between the two databases and the two sets of unique journals. Citations among the shared journals are more comprehensively covered in JCR than Scopus, so the network in JCR is denser and more connected than in Scopus. The ranking of shared journals in terms of indegree (that is, numbers of citing journals) or total citations is similar in both databases overall (Spearman's \r{ho} > 0.97), but some individual journals rank very differently. Journals that are unique to Scopus seem to be less important--they are citing shared journals rather than being cited by them--but the humanities are covered better in Scopus than in JCR.

preprint2014arXiv

BRICS countries and scientific excellence: A bibliometric analysis of most frequently-cited papers

The BRICS countries (Brazil, Russia, India, and China, and South Africa) are noted for their increasing participation in science and technology. The governments of these countries have been boosting their investments in research and development to become part of the group of nations doing research at a world-class level. This study investigates the development of the BRICS countries in the domain of top-cited papers (top 10% and 1% most frequently cited papers) between 1990 and 2010. To assess the extent to which these countries have become important players on the top level, we compare the BRICS countries with the top-performing countries worldwide. As the analyses of the (annual) growth rates show, with the exception of Russia, the BRICS countries have increased their output in terms of most frequently-cited papers at a higher rate than the top-cited countries worldwide. In a further step of analysis for this study, we generate co-authorship networks among authors of highly cited papers for four time points to view changes in BRICS participation (1995, 2000, 2005, and 2010). Here, the results show that all BRICS countries succeeded in becoming part of this network, whereby the Chinese collaboration activities focus on the USA.

preprint2014arXiv

Journal Maps, Interactive Overlays, and the Measurement of Interdisciplinarity on the Basis of Scopus Data (1996-2012)

Using Scopus data, we construct a global map of science based on aggregated journal-journal citations from 1996-2012 (N of journals = 20,554). This base map enables users to overlay downloads from Scopus interactively. Using a single year (e.g., 2012), results can be compared with mappings based on the Journal Citation Reports at the Web-of-Science (N = 10,936). The Scopus maps are more detailed at both the local and global levels because of their greater coverage, including, for example, the arts and humanities. The base maps can be interactively overlaid with journal distributions in sets downloaded from Scopus, for example, for the purpose of portfolio analysis. Rao-Stirling diversity can be used as a measure of interdisciplinarity in the sets under study. Maps at the global and the local level, however, can be very different because of the different levels of aggregation involved. Two journals, for example, can both belong to the humanities in the global map, but participate in different specialty structures locally. The base map and interactive tools are available online (with instructions) at http://www.leydesdorff.net/scopus_ovl.

preprint2014arXiv

Matching MEDLINE/PubMed Data with Web of Science (WoS): A Routine in R language

We present a novel routine, namely medlineR, based on R-language, that enables the user to match data from MEDLINE/PubMed with records indexed in the ISI Web of Science (WoS) database. The matching allows exploiting the rich and controlled vocabulary of Medical Subject Headings (MeSH) of MEDLINE/PubMed with additional fields of WoS. The integration provides data (e.g. citation data, list of cited reference, full list of the addresses of authors' host organisations, WoS subject categories) to perform a variety of scientometric analyses. This brief communication describes medlineR, the methodology on which it relies, and the steps the user should follow to perform the matching across the two databases. In order to specify the differences from Leydesdorff and Opthof (2013), we conclude the brief communication by testing the routine on the case of the "Burgada Syndrome".

preprint2014arXiv

Measuring Triple-Helix Synergy in the Russian Innovation Systems at Regional, Provincial, and National Levels

We measure synergy for the Russian national, provincial, and regional innovation systems as reduction of uncertainty using mutual information among the three distributions of firm sizes, technological knowledge-bases of firms, and geographical locations. Half a million data at firm level in 2011 were obtained from the Orbis database of Bureau Van Dijk. The firm level data were aggregated at the levels of eight Federal Districts, the regional level of 83 Federal Subjects, and the single level of the Russian Federation. Not surprisingly, the knowledge base of the economy is concentrated in the Moscow region (22.8%); St. Petersburg follows with 4.0%. Only 0.4% of the firms are classified as high-tech, and 2.7% as medium-tech manufacturing (NACE, Rev. 2). Except in Moscow itself, high-tech manufacturing does not add synergy to any other unit at any of the various levels of geographical granularity; instead it disturbs regional coordination even in the region surrounding Moscow ("Moscow Region"). In the case of medium-tech manufacturing, there is also synergy in St. Petersburg. Knowledge-intensive services (KIS; including laboratories) contribute 12.8% to the economy in terms of establishments and contribute to the synergy in all Federal Districts (except the North-Caucasian Federal District), but only in 30 of the 83 Federal Subjects. The synergy in KIS is concentrated in centers of administration. Unlike Western European countries, the knowledge-intensive services (which are often state-affiliated) thus provide backbone to an emerging knowledge-based economy at the level of Federal Districts, but the economy is otherwise not knowledge-based (except for the Moscow region).

preprint2014arXiv

Patents as Instruments for Exploring Innovation Dynamics: Geographic and Technological Perspectives on "Photovoltaic Cells"

The dynamics of innovation are nonlinear and complex: geographical, technological, and economic selection environments can be expected to interact. Can patents provide an analytical lens to this process in terms of different attributes such as inventor addresses, classification codes, backward and forward citations, etc.? Two recently developed patent maps with interactive overlay techniques--Google Maps and maps based on citation relations among International Patent Classifications (IPC)--are elaborated into dynamic versions that allow for online animations and comparisons by using split screens. Various forms of animation are explored. The recently developed Cooperative Patent Classifications (CPC) of the U.S. Patent and Trade Office (USPTO) and the European Patent Office (EPO) provide new options for a precise delineation of samples in both USPTO data and the Worldwide Patent Statistics Database (PatStat) of EPO. Among the "technologies for the mitigation of climate change" (class Y02), we zoom in on nine material technologies for photovoltaic cells; and focus on one of them (CuInSe2) as a lead case. The longitudinal development of Rao-Stirling diversity in the IPC-based maps provides a heuristics for studying technological generations during the period under study (1975-2012). The sequencing of generations prevails in USPTO data more than in PatStat data because PatStat aggregates patent information from countries in different stages of technological development, whereas one can expect USPTO patents to be competitive at the technological edge.

preprint2014arXiv

Redundancy Generation in University-Industry-Government Relations: The Triple Helix Modeled, Measured, and Simulated

A Triple Helix (TH) of bi- and trilateral relations among universities, industries, and governments can be considered as an ecosystem in which uncertainty can be reduced auto-catalytically. The correlations among the distributions of relations span a vector space in which two vectors (P and Q) represent "sending" and "receiving," respectively. These vectors can also be understood in terms of the generation versus reduction of uncertainty in the communication field that results from interactions among the three (bi-lateral) communication channels. We specify a set of Lotka-Volterra equations between the vectors that can be solved. Redundancy generation can then be simulated and the results can be decomposed in terms of the TH components. Among other things, we show that the strength and frequency of the relations are independent parameters. Different components in terms of frequencies in triple-helix systems can also be distinguished and interpreted using Fourier analysis of the empirical time-series. The case of co-authorship relations in Japan is analyzed as an empirical example; but "triple contingencies" in an ecosystem of relations can also be considered more generally as a model for redundancy generation by providing meaning to the (Shannon-type) information in inter-human communications.

preprint2014arXiv

The European Union, China, and the United States in the Top-1% and Top-10% Layers of Most-Frequently-Cited Publications: Competition and Collaborations

The percentages of shares of world publications of the European Union and its member states, China, and the United States have been represented differently as a result of using different databases. An analytical variant of the Web-of-Science (of Thomson Reuters) enables us to study the dynamics in the world publication system in terms of the field-normalized top-1% and top-10% most-frequently-cited publications. Comparing the EU28, USA, and China at the global level shows a top-level dynamics that is different from the analysis in terms of shares of publications: the United States remains far more productive in the top-1% of all papers; China drops out of the competition for elite status; and the EU28 increased its share among the top-cited papers from 2000-2010. Some of the EU28 member states overtook the U.S. during this decade, but a clear divide remains between EU15 (Western Europe) and the Accession Countries. Network analysis shows that internationally co-authored top-1% publications perform far above expectation and also above top-10% ones. In 2005, China was embedded in this top-layer of internationally co-authored publications. These publications often involve more than a single European nation.

preprint2014arXiv

The Generation of Large Networks from Web-of-Science Data

During the 1990s, one of us developed a series of freeware routines (http://www.leydesdorff.net/indicators) that enable the user to organize downloads from the Web-of-Science (Thomson Reuters) into a relational database, and then to export matrices for further analysis in various formats (for example, for co-author analysis). The basic format of the matrices displays each document as a case in a row that can be attributed different variables in the columns. One limitation to this approach was hitherto that relational databases typically have an upper limit for the number of variables, such as 256 or 1024. In this brief communication, we report on a way to circumvent this limitation by using txt2Pajek.exe, available as freeware from http://www.pfeffer.at/txt2pajek/.

preprint2014arXiv

The Operationalization of "Fields" as WoS Subject Categories (WCs) in Evaluative Bibliometrics: The cases of "Library and Information Science" and "Science & Technology Studies"

Normalization of citation scores using reference sets based on Web-of-Science Subject Categories (WCs) has become an established ("best") practice in evaluative bibliometrics. For example, the Times Higher Education World University Rankings are, among other things, based on this operationalization. However, WCs were developed decades ago for the purpose of information retrieval and evolved incrementally with the database; the classification is machine-based and partially manually corrected. Using the WC "information science & library science" and the WCs attributed to journals in the field of "science and technology studies," we show that WCs do not provide sufficient analytical clarity to carry bibliometric normalization in evaluation practices because of "indexer effects." Can the compliance with "best practices" be replaced with an ambition to develop "best possible practices"? New research questions can then be envisaged.

preprint2013arXiv

Challenges for regional innovation policies in CEE countries: Spatial concentration and foreign control of US patenting

Using techniques of data collection and mapping as overlays to Google Maps--on the basis of patent information available online at the U.S. Patent and Trademark Office (USPTO)--we point at two major and interconnected challenges that policy-makers face in Central and Eastern Europe (CEE) when combating the lagging innovation performance. First, we address the spatial concentration by using a distribution analysis at the city level. The results suggest that patenting is concentrated in post-socialist territories more than in western nations and regions. However, there is not a single outstanding hub in CEE when one compares USPTO patents normalized for the respective population sizes. Secondly, we argue that dominance of foreign control over USPTO patents is mostly embodied in international co-operations at the individual level, and only rarely spilled-over to MNE subsidiaries. In our opinion, catching-up of CEE in terms of patenting is unlikely, unless innovation policy measures focus on growing hubs and target both domestic inventors and international relations of companies.

preprint2013arXiv

Detecting the historical roots of research fields by reference publication year spectroscopy (RPYS)

We introduce the quantitative method named "reference publication year spectroscopy" (RPYS). With this method one can determine the historical roots of research fields and quantify their impact on current research. RPYS is based on the analysis of the frequency with which references are cited in the publications of a specific research field in terms of the publication years of these cited references. The origins show up in the form of more or less pronounced peaks mostly caused by individual publications which are cited particularly frequently. In this study, we use research on graphene and on solar cells to illustrate how RPYS functions, and what results it can deliver.

preprint2013arXiv

Field-normalized Impact Factors: A Comparison of Rescaling versus Fractionally Counted IFs

Two methods for comparing impact factors and citation rates across fields of science are tested against each other using citations to the 3,705 journals in the Science Citation Index 2010 (CD-Rom version of SCI) and the 13 field categories used for the Science and Engineering Indicators of the US National Science Board. We compare (i) normalization by counting citations in proportion to the length of the reference list (1/N of references) with (ii) rescaling by dividing citation scores by the arithmetic mean of the citation rate of the cluster. Rescaling is analytical and therefore independent of the quality of the attribution to the sets, whereas fractional counting provides an empirical strategy for normalization among sets (by evaluating the between-group variance). By the fairness test of Radicchi & Castellano (2012a), rescaling outperforms fractional counting of citations for reasons that we consider.

preprint2013arXiv

Group-Based Trajectory Modeling of Citations in Scholarly Literature: Dynamic Qualities of "Transient" and "Sticky Knowledge Claims"

Group-based Trajectory Modeling (GBTM) is applied to the citation curves of articles in six journals and to all citable items in a single field of science (Virology, 24 journals), in order to distinguish among the developmental trajectories in subpopulations. Can highly-cited citation patterns be distinguished in an early phase as "fast-breaking" papers? Can "late bloomers" or "sleeping beauties" be identified? Most interesting, we find differences between "sticky knowledge claims" that continue to be cited more than ten years after publication, and "transient knowledge claims" that show a decay pattern after reaching a peak within a few years. Only papers following the trajectory of a "sticky knowledge claim" can be expected to have a sustained impact. These findings raise questions about indicators of "excellence" that use aggregated citation rates after two or three years (e.g., impact factors). Because aggregated citation curves can also be composites of the two patterns, 5th-order polynomials (with four bending points) are needed to capture citation curves precisely. For the journals under study, the most frequently cited groups were furthermore much smaller than ten percent. Although GBTM has proved a useful method for investigating differences among citation trajectories, the methodology does not enable us to define a percentage of highly-cited papers inductively across different fields and journals. Using multinomial logistic regression, we conclude that predictor variables such as journal names, number of authors, etc., do not affect the stickiness of knowledge claims in terms of citations, but only the levels of aggregated citations (that are field-specific).

preprint2013arXiv

How have the Eastern European countries of the former Warsaw Pact developed since 1990? A bibliometric study

Did the demise of the Soviet Union in 1991 influence the scientific performance of the researchers in Eastern European countries? Did this historical event affect international collaboration by researchers from the Eastern European countries with those of Western countries? Did it also change international collaboration among researchers from the Eastern European countries? Trying to answer these questions, this study aims to shed light on international collaboration by researchers from the Eastern European countries (Russia, Ukraine, Belarus, Moldova, Bulgaria, the Czech Republic, Hungary, Poland, Romania and Slovakia). The number of publications and normalized citation impact values are compared for these countries based on InCites (Thomson Reuters), from 1981 up to 2011. The international collaboration by researchers affiliated to institutions in Eastern European countries at the time points of 1990, 2000 and 2011 was studied with the help of Pajek and VOSviewer software, based on data from the Science Citation Index (Thomson Reuters). Our results show that the breakdown of the communist regime did not lead, on average, to a huge improvement in the publication performance of the Eastern European countries and that the increase in international co-authorship relations by the researchers affiliated to institutions in these countries was smaller than expected. Most of the Eastern European countries are still subject to changes and are still awaiting their boost in scientific development.

preprint2013arXiv

How to improve the prediction based on citation impact percentiles for years shortly after the publication date?

The findings of Bornmann, Leydesdorff, and Wang (in press) revealed that the consideration of journal impact improves the prediction of long-term citation impact. This paper further explores the possibility of improving citation impact measurements on the base of a short citation window by the consideration of journal impact and other variables, such as the number of authors, the number of cited references, and the number of pages. The dataset contains 475,391 journal papers published in 1980 and indexed in Web of Science (WoS, Thomson Reuters), and all annual citation counts (from 1980 to 2010) for these papers. As an indicator of citation impact, we used percentiles of citations calculated using the approach of Hazen (1914). Our results show that citation impact measurement can really be improved: If factors generally influencing citation impact are considered in the statistical analysis, the explained variance in the long-term citation impact can be much increased. However, this increase is only visible when using the years shortly after publication but not when using later years.

preprint2013arXiv

Innovation as a Nonlinear Process, the Scientometric Perspective, and the Specification of an "Innovation Opportunities Explorer"

The process of innovation follows non-linear patterns across the domains of science, technology, and the economy. Novel bibliometric mapping techniques can be used to investigate and represent distinctive, but complementary perspectives on the innovation process (e.g., "demand" and "supply") as well as the interactions among these perspectives. The perspectives can be represented as "continents" of data related to varying extents over time. For example, the different branches of Medical Subject Headings (MeSH) in the Medline database provide sources of such perspectives (e.g., "Diseases" versus "Drugs and Chemicals"). The multiple-perspective approach enables us to reconstruct facets of the dynamics of innovation, in terms of selection mechanisms shaping localizable trajectories and/or resulting in more globalized regimes. By expanding the data with patents and scholarly publications, we demonstrate the use of this multi-perspective approach in the case of RNA Interference (RNAi). The possibility to develop an "Innovation Opportunities Explorer" is specified.

preprint2013arXiv

Interactive Overlays of Journals and the Measurement of Interdisciplinarity on the basis of Aggregated Journal-Journal Citations

Using "Analyze Results" at the Web of Science, one can directly generate overlays onto global journal maps of science. The maps are based on the 10,000+ journals contained in the Journal Citation Reports (JCR) of the Science and Social Science Citation Indices (2011). The disciplinary diversity of the retrieval is measured in terms of Rao-Stirling's "quadratic entropy." Since this indicator of interdisciplinarity is normalized between zero and one, the interdisciplinarity can be compared among document sets and across years, cited or citing. The colors used for the overlays are based on Blondel et al.'s (2008) community-finding algorithms operating on the relations journals included in JCRs. The results can be exported from VOSViewer with different options such as proportional labels, heat maps, or cluster density maps. The maps can also be web-started and/or animated (e.g., using PowerPoint). The "citing" dimension of the aggregated journal-journal citation matrix was found to provide a more comprehensive description than the matrix based on the cited archive. The relations between local and global maps and their different functions in studying the sciences in terms of journal litteratures are further discussed: local and global maps are based on different assumptions and can be expected to serve different purposes for the explanation.

preprint2013arXiv

Interdisciplinarity at the Journal and Specialty Level: The changing knowledge bases of the journal Cognitive Science

Using the referencing patterns in articles in Cognitive Science over three decades, we analyze the knowledge base of this literature in terms of its changing disciplinary composition. Three periods are distinguished: (1) construction of the interdisciplinary space in the 1980s; (2) development of an interdisciplinary orientation in the 1990s; (3) reintegration into "cognitive psychology" in the 2000s. The fluidity and fuzziness of the interdisciplinary delineations in the different visualizations can be reduced and clarified using factor analysis. We also explore newly available routines ("CorText") to analyze this development in terms of "tubes" using an alluvial map, and compare the results with an animation (using "visone"). The historical specificity of this development can be compared with the development of "artificial intelligence" into an integrated specialty during this same period. "Interdisciplinarity" should be defined differently at the level of journals and of specialties.

preprint2013arXiv

International Co-authorship Relations in the Social Science Citation Index: Is Internationalization Leading the Network?

We analyze international co-authorship relations in the Social Science Citation Index 2011 using all citable items in the DVD-version of this index. Network statistics indicate four groups of nations: (i) an Asian-Pacific one to which all Anglo-Saxon nations (including the UK and Ireland) are attributed; (ii) a continental European one including also the Latin-American countries; (iii) the Scandinavian nations; and (iv) a community of African nations. Within the EU-28 (including Croatia), eleven of the EU-15 states have dominant positions. Collapsing the EU-28 into a single node leads to a bi-polar structure between the US and EU-28; China is part of the US-pole. We develop an information-theoretical test to distinguish whether international collaborations or domestic collaborations prevail; the results are mixed, but the international dimension is more important than the national one in the aggregated sets (this was found in both SSCI and SCI). In France, however, the national distribution is more important than the international one, while the reverse is true for most European nations in the core group (UK, Germany, the Netherlands, etc.). Decomposition of the USA in terms of states shows a similarly mixed result; more US states are domestically oriented in SSCI, whereas more internationally in SCI. The international networks have grown during the last decades in addition to the national ones, but not by replacing them.

preprint2013arXiv

International collaboration clusters in Africa

Recent discussion about the increase in international research collaboration suggests a comprehensive global network centred around a group of core countries and driven by generic socio-economic factors where the global system influences all national and institutional outcomes. In counterpoint, we demonstrate that the collaboration pattern for countries in Africa is far from universal. Instead, it exhibits layers of internal clusters and external links that are explained not by monotypic global influences but by regional geography and, perhaps even more strongly, by history, culture and language. Analysis of these bottom-up, subjective, human factors is required in order to provide the fuller explanation useful for policy and management purposes.

preprint2013arXiv

International Collaboration in Science: The Global Map and the Network

The network of international co-authorship relations has been dominated by certain European nations and the USA, but this network is rapidly expanding at the global level. Between 40 and 50 countries appear in the center of the international network in 2011, and almost all (201) nations are nowadays involved in international collaboration. In this brief communication, we present both a global map with the functionality of a Google Map (zooming, etc.) and network maps based on normalized relations. These maps reveal complementary aspects of the network. International collaboration in the generation of knowledge claims (that is, the context of discovery) changes the structural layering of the sciences. Previously, validation was at the global level and discovery more dependent on local contexts. This changing relationship between the geographical and intellectual dimensions of the sciences also has implications for national science policies.

preprint2013arXiv

Measuring the Knowledge-Based Economy of China in terms of Synergy among Technological, Organizational, and Geographic Attributes of Firms

Using the possible synergy among geographic, size, and technological distributions of firms in the Orbis database, we find the greatest reduction of uncertainty at the level of the 31 provinces of China, and an additional 18.0% at the national level. Some of the coastal provinces stand out as expected, but the metropolitan areas of Beijing and Shanghai are (with Tianjan and Chonqing) most pronounced at the next-lower administrative level of (339) prefectures, since these four metropoles are administratively defined at both levels. Focusing on high- and medium-tech manufacturing, a shift toward Beijing and Shanghai is indicated, and the synergy is on average enhanced (as expected; but not for all provinces). Unfortunately, the Orbis data is incomplete since it was collected for commercial and not for administrative or governmental purposes. However, we show a methodology that can be used by others who may have access to higher-quality statistical data for the measurement.

preprint2013arXiv

Mutual Redundancies in Inter-human Communication Systems: Steps Towards a Calculus of Processing Meaning

The study of inter-human communication requires a more complex framework than Shannon's (1948) mathematical theory of communication because "information" is defined in the latter case as meaningless uncertainty. Assuming that meaning cannot be communicated, we extend Shannon's theory by defining mutual redundancy as a positional counterpart of the relational communication of information. Mutual redundancy indicates the surplus of meanings that can be provided to the exchanges in reflexive communications. The information is redundant because based on "pure sets," that is, without subtraction of mutual information in the overlaps. We show that in the three-dimensional case (e.g., of a Triple Helix of university-industry-government relations), mutual redundancy is equal to mutual information (Rxyz = Txyz); but when the dimensionality is even, the sign is different. We generalize to the measurement in N dimensions and proceed to the interpretation. Using Luhmann's social-systems theory and/or Giddens' structuration theory, mutual redundancy can be provided with an interpretation in the sociological case: different meaning-processing structures code and decode with other algorithms. A surplus of ("absent") options can then be generated that add to the redundancy. Luhmann's "functional (sub)systems" of expectations or Giddens' "rule-resource sets" are positioned mutually, but coupled operationally in events or "instantiated" in actions. Shannon-type information is generated by the mediation, but the "structures" are (re-)positioned towards one another as sets of (potentially counterfactual) expectations. The structural differences among the coding and decoding algorithms provide a source of additional options in reflexive and anticipatory communications.

preprint2013arXiv

Referenced Publication Years Spectroscopy applied to iMetrics: Scientometrics, Journal of Informetrics, and a relevant subset of JASIST

We have developed a (freeware) routine for "referenced publication years spectroscopy" (RPYS) and apply this method to the historiography of "iMetrics," that is, the junction of the journals Scientometrics, Informetrics, and the relevant subset of JASIST (approx. 20%) that shapes the intellectual space for the development of information metrics (bibliometrics, scientometrics, informetrics, and webometrics). The application to information metrics (our own field of research) provides us with the opportunity to validate this methodology, and to add a reflection about using citations for the historical reconstruction. The results show that the field is rooted in individual contributions of the 1920s-1950s (e.g., Alfred J. Lotka), and was then shaped intellectually in the early 1960s by a confluence of the history of science (Derek de Solla Price), documentation (e.g., Michael M. Kessler's "bibliographic coupling"), and "citation indexing" (Eugene Garfield). Institutional development at the interfaces between science studies and information science has been reinforced by the new journal Informetrics since 2007. In a concluding reflection, we return to the question of how the historiography of science using algorithmic means--in terms of citation practices--can be different from an intellectual history of the field based, for example, on reading source materials.

preprint2013arXiv

Rotational Symmetry and the Transformation of Innovation Systems in a Triple Helix of University-Industry-Government Relations

Using a mathematical model, we show that a Triple Helix (TH) system contains self-interaction, and therefore self-organization of innovations can be expected in waves, whereas a Double Helix (DH) remains determined by its linear constituents. (The mathematical model is fully elaborated in the Appendices.) The ensuing innovation systems can be expected to have a fractal structure: innovation systems at different scales can be considered as spanned in a Cartesian space with the dimensions of (S)cience, (B)usiness, and (G)overnment. A national system, for example, contains sectorial and regional systems, and is a constituent part in technological and supra-national systems of innovation. The mathematical modeling enables us to clarify the mechanisms, and provides new possibilities for the prediction. Emerging technologies can be expected to be more diversified and their life cycles will become shorter than before. In terms of policy implications, the model suggests a shift from the production of material objects to the production of innovative technologies.

preprint2013arXiv

Scientometrics

The paper provides an overview of the field of scientometrics, that is: the study of science, technology, and innovation from a quantitative perspective. We cover major historical milestones in the development of this specialism from the 1960s to today and discuss its relationship with the sociology of scientific knowledge, the library and information sciences, and science policy issues such as indicator development. The disciplinary organization of scientometrics is analyzed both conceptually and empirically, using a map of journals cited in the core journal of the field, entitled Scientometrics. A state-of-the-art review of five major research threads is provided: (1) the measurement of impact; (2) the delineation of reference sets; (3) theories of citation; (4) mapping science; and (5) the policy and management contexts of indicator developments.

preprint2013arXiv

The "Academic Trace" of the Performance Matrix: A Mathematical Synthesis of the h-Index and the Integrated Impact Indicator (I3)

The h-index provides us with nine natural classes which can be written as a matrix of three vectors. The three vectors are: X=(X1, X2, X3) indicate publication distribution in the h-core, the h-tail, and the uncited ones, respectively; Y=(Y1, Y2, Y3) denote the citation distribution of the h-core, the h-tail and the so-called "excess" citations (above the h-threshold), respectively; and Z=(Z1, Z2, Z3)= (Y1-X1, Y2-X2, Y3-X3). The matrix V=(X,Y,Z)T constructs a measure of academic performance, in which the nine numbers can all be provided with meanings in different dimensions. The "academic trace" tr(V) of this matrix follows naturally, and contributes a unique indicator for total academic achievements by summarizing and weighting the accumulation of publications and citations. This measure can also be used to combine the advantages of the h-index and the Integrated Impact Indicator (I3) into a single number with a meaningful interpretation of the values. We illustrate the use of tr(V) for the cases of two journal sets, two universities, and ourselves as two individual authors.

preprint2013arXiv

The Disclosure of University Research for Third Parties: A Non-Market Perspective on an Italian University

Nations, universities, and regional governments commit resources to promote the dissemination of scientific and technical knowledge. One focuses on knowledge-based innovations and the economic function of the university in terms of technology transfer, intellectual property, university-industry-government relations, etc. Faculties other than engineering or applied sciences, however, may not be able to recognize opportunities in this "linear model" of technology transfer. We elaborate a non-market perspective on the third mission in terms of disclosure of the knowledge and areas of expertise available for disclosure to other audiences at a provincial university. The use of ICT can enhance communication between actors on the supply and demand sides. Using an idea originally developed in the context of the Dutch science shops, the university staff was questionnaired about keywords and areas of expertise with the specific purpose of disclosing this information to audiences other than academic colleagues. The results were brought online in a thesaurus-like structure that enables users to access the university at the level of individual email address. This model stimulates variation on both the supply and demand side of the innovation process, and strengthens the accessibility and embeddedness of the knowledge base in a regional economy.

preprint2013arXiv

The revised SNIP indicator of Elsevier's Scopus

The modified SNIP indicator of Elsevier, as recently explained by Waltman et al. (2013) in this journal, solves some of the problems which Leydesdorff & Opthof (2010 and 2011) indicated in relation to the original SNIP indicator (Moed, 2010 and 2011). The use of an arithmetic average, however, remains unfortunate in the case of scientometric distributions because these can be extremely skewed (Seglen, 1992 and 1997). The new indicator cannot (or hardly) be reproduced independently when used for evaluation purposes, and remains in this sense opaque from the perspective of evaluated units and scholars.

preprint2013arXiv

The Triple Helix of University-Industry-Government Relations at the Country Level, and Its Dynamic Evolution under the Pressures of Globalization

Using data from the Web of Science (WoS), we analyze the mutual information among university, industrial, and governmental addresses (U-I-G) at the country level for a number of countries. The dynamic evolution of the Triple Helix can thus be compared among developed and developing nations in terms of cross-sectorial co-authorship relations. The results show that the Triple-Helix interactions among the three subsystems U-I-G become less intensive over time, but unequally for different countries. We suggest that globalization erodes local Triple-Helix relations and thus can be expected to increase differentiation in national systems since the mid-1990s. This effect of globalization is more pronounced in developed countries than in developing ones. In the dynamic analysis, we focus on a more detailed comparison between China and the USA. The Chinese Academy of the (Social) Sciences changes increasingly from a public research institute to an academic one, and this has a measurable effect on China's position in the globalization.

preprint2013arXiv

Which percentile-based approach should be preferred for calculating normalized citation impact values? An empirical comparison of five approaches including a newly developed citation-rank approach (P100)

Percentile-based approaches have been proposed as a non-parametric alternative to parametric central-tendency statistics to normalize observed citation counts. Percentiles are based on an ordered set of citation counts in a reference set, whereby the fraction of papers at or below the citation counts of a focal paper is used as an indicator for its relative citation impact in the set. In this study, we pursue two related objectives: (1) although different percentile-based approaches have been developed, an approach is hitherto missing that satisfies a number of criteria such as scaling of the percentile ranks from zero (all other papers perform better) to 100 (all other papers perform worse), and solving the problem with tied citation ranks unambiguously. We introduce a new citation-rank approach having these properties, namely P100. (2) We compare the reliability of P100 empirically with other percentile-based approaches, such as the approaches developed by the SCImago group, the Centre for Science and Technology Studies (CWTS), and Thomson Reuters (InCites), using all papers published in 1980 in Thomson Reuters Web of Science (WoS). How accurately can the different approaches predict the long-term citation impact in 2010 (in year 31) using citation impact measured in previous time windows (years 1 to 30)? The comparison of the approaches shows that the method used by InCites overestimates citation impact (because of using the highest percentile rank when papers are assigned to more than a single subject category) whereas the SCImago indicator shows higher power in predicting the long-term citation impact on the basis of citation rates in early years. Since the results show a disadvantage in this predictive ability for P100 against the other approaches, there is still room for further improvements.

preprint2012arXiv

A bird's-eye view of scientific trading: Dependency relations among fields of science

We use a trading metaphor to study knowledge transfer in the sciences as well as the social sciences. The metaphor comprises four dimensions: (a) Discipline Self-dependence, (b) Knowledge Exports/Imports, (c) Scientific Trading Dynamics, and (d) Scientific Trading Impact. This framework is applied to a dataset of 221 Web of Science subject categories. We find that: (i) the Scientific Trading Impact and Dynamics of Materials Science And Transportation Science have increased; (ii) Biomedical Disciplines, Physics, And Mathematics are significant knowledge exporters, as is Statistics & Probability; (iii) in the social sciences, Economics, Business, Psychology, Management, And Sociology are important knowledge exporters; (iv) Discipline Self-dependence is associated with specialized domains which have ties to professional practice (e.g., Law, Ophthalmology, Dentistry, Oral Surgery & Medicine, Psychology, Psychoanalysis, Veterinary Sciences, And Nursing).

preprint2012arXiv

A Routine for Measuring Synergy in University-Industry-Government Relations: Mutual Information as a Triple-Helix and Quadruple-Helix Indicator

Mutual information in three (or more) dimensions can be considered as a Triple-Helix indicator of synergy in university-industry-government relations. An open-source routine th4.exe makes the computation of this indicator interactively available at the Internet, and thus applicable to large sets of data. Th4.exe computes all probabilistic entropies and mutual information in two, three, and, if available in the data, four dimensions among, for example, classes such as geographical addresses (cities, regions), technological codes (e.g., OECD's NACE codes), and size categories; or, alternatively, among institutional addresses (academic, industrial, public sector) in document sets. The relations between the Triple-Helix indicator -- as an indicator of synergy -- and the Triple-Helix model that specifies the possibility of feedback by an overlay of communications, are also discussed.

preprint2012arXiv

Accounting for the Uncertainty in the Evaluation of Percentile Ranks

In a recent paper entitled "Inconsistencies of Recently Proposed Citation Impact Indicators and how to Avoid Them," Schreiber (2012, at arXiv:1202.3861) proposed (i) a method to assess tied ranks consistently and (ii) fractional attribution to percentile ranks in the case of relatively small samples (e.g., for n < 100). Schreiber's solution to the problem of how to handle tied ranks is convincing, in my opinion (cf. Pudovkin & Garfield, 2009). The fractional attribution, however, is computationally intensive and cannot be done manually for even moderately large batches of documents. Schreiber attributed scores fractionally to the six percentile rank classes used in the Science and Engineering Indicators of the U.S. National Science Board, and thus missed, in my opinion, the point that fractional attribution at the level of hundred percentiles-or equivalently quantiles as the continuous random variable-is only a linear, and therefore much less complex problem. Given the quantile-values, the non-linear attribution to the six classes or any other evaluation scheme is then a question of aggregation. A new routine based on these principles (including Schreiber's solution for tied ranks) is made available as software for the assessment of documents retrieved from the Web of Science (at http://www.leydesdorff.net/software/i3).

preprint2012arXiv

Alternatives to the Journal Impact Factor: I3 and the Top-10% (or Top-25%?) of the Most-Highly Cited Papers

Journal Impact Factors (IFs) can be considered historically as the first attempt to normalize citation distributions by using averages over two years. However, it has been recognized that citation distributions vary among fields of science and that one needs to normalize for this. Furthermore, the mean-or any central-tendency statistics-is not a good representation of the citation distribution because these distributions are skewed. Important steps have been taken to solve these two problems during the last few years. First, one can normalize at the article level using the citing audience as the reference set. Second, one can use non-parametric statistics for testing the significance of differences among ratings. A proportion of most-highly cited papers (the top-10% or top-quartile) on the basis of fractional counting of the citations may provide an alternative to the current IF. This indicator is intuitively simple, allows for statistical testing, and accords with the state of the art.

preprint2012arXiv

An Evaluation of Impacts in "Nanoscience & nanotechnology:" Steps towards standards for citation analysis

One is inclined to conceptualize impact in terms of citations per publication, and thus as an average. However, citation distributions are skewed, and the average has the disadvantage that the number of publications is used in the denominator. Using hundred percentiles, one can integrate the normalized citation curve and develop an indicator that can be compared across document sets because percentile ranks are defined at the article level. I apply this indicator to the set of 58 journals in the ISI Subject Category of "Nanoscience & nanotechnology," and rank journals, countries, cities, and institutes using non-parametric statistics. The significance levels of results can thus be indicated. The results are first compared with the ISI-Impact Factors, but this Integrated Impact Indicator (I3) can be used with any set downloaded from the (Social) Science Citation Index. The software is made publicly available at the Internet. Visualization techniques are also specified for evaluation by positioning institutes on Google Map overlays.

preprint2012arXiv

An Integrated Impact Indicator (I3): A New Definition of "Impact" with Policy Relevance

Allocation of research funding, as well as promotion and tenure decisions, are increasingly made using indicators and impact factors drawn from citations to published work. A debate among scientometricians about proper normalization of citation counts has resolved with the creation of an Integrated Impact Indicator (I3) that solves a number of problems found among previously used indicators. The I3 applies non-parametric statistics using percentiles, allowing highly-cited papers to be weighted more than less-cited ones. It further allows unbundling of venues (i.e., journals or databases) at the article level. Measures at the article level can be re-aggregated in terms of units of evaluation. At the venue level, the I3 creates a properly weighted alternative to the journal impact factor. I3 has the added advantage of enabling and quantifying classifications such as the six percentile rank classes used by the National Science Board's Science & Engineering Indicators.

preprint2012arXiv

Betweenness Centrality as a Driver of Preferential Attachment in the Evolution of Research Collaboration Networks

We analyze whether preferential attachment in scientific coauthorship networks is different for authors with different forms of centrality. Using a complete database for the scientific specialty of research about "steel structures," we show that betweenness centrality of an existing node is a significantly better predictor of preferential attachment by new entrants than degree or closeness centrality. During the growth of a network, preferential attachment shifts from (local) degree centrality to betweenness centrality as a global measure. An interpretation is that supervisors of PhD projects and postdocs broker between new entrants and the already existing network, and thus become focal to preferential attachment. Because of this mediation, scholarly networks can be expected to develop differently from networks which are predicated on preferential attachment to nodes with high degree centrality.

preprint2012arXiv

Bibliometric Perspectives on Medical Innovation using the Medical Subject Headings (MeSH) of PubMed

Multiple perspectives on the nonlinear processes of medical innovations can be distinguished and combined using the Medical Subject Headings (MeSH) of the Medline database. Focusing on three main branches-"diseases," "drugs and chemicals," and "techniques and equipment"-we use base maps and overlay techniques to investigate the translations and interactions and thus to gain a bibliometric perspective on the dynamics of medical innovations. To this end, we first analyze the Medline database, the MeSH index tree, and the various options for a static mapping from different perspectives and at different levels of aggregation. Following a specific innovation (RNA interference) over time, the notion of a trajectory which leaves a signature in the database is elaborated. Can the detailed index terms describing the dynamics of research be used to predict the diffusion dynamics of research results? Possibilities are specified for further integration between the Medline database, on the one hand, and the Science Citation Index and Scopus (containing citation information), on the other.

preprint2012arXiv

Citation Analysis with Medical Subject Headings (MeSH) using the Web of Knowledge: A new routine

Citation analysis of documents retrieved from the Medline database (at the Web of Knowledge) has been possible only on a case-by-case basis. A technique is here developed for citation analysis in batch mode using both Medical Subject Headings (MeSH) at the Web of Knowledge and the Science Citation Index at the Web of Science. This freeware routine is applied to the case of "Brugada Syndrome," a specific disease and field of research (since 1992). The journals containing these publications, for example, are attributed to Web-of-Science Categories other than "Cardiac and Cardiovascular Systems"), perhaps because of the possibility of genetic testing for this syndrome in the clinic. With this routine, all the instruments available for citation analysis can now be used on the basis of MeSH terms. Other options for crossing between Medline, WoS, and Scopus are also reviewed.

preprint2012arXiv

Citation impact of papers published from six prolific countries: A national comparison based on InCites data

Using the InCites tool of Thomson Reuters, this study compares normalized citation impact values calculated for China, Japan, France, Germany, United States, and the UK throughout the time period from 1981 to 2010. The citation impact values are normalized to four subject areas: natural sciences; engineering and technology; medical and health sciences; and agricultural sciences. The results show an increasing trend in citation impact values for France, the UK and especially for Germany across the last thirty years in all subject areas. The citation impact of papers from China is still at a relatively low level (mostly below the world average), but the country follows an increasing trend line. The USA exhibits a relatively stable pattern of high citation impact values across the years. With small impact differences between the publication years, the US trend is increasing in engineering and technology but decreasing in medical and health sciences as well as in agricultural sciences. Similar to the USA, Japan follows increasing as well as decreasing trends in different subject areas, but the variability across the years is small. In most of the years, papers from Japan perform below or approximately at the world average in each subject area.

preprint2012arXiv

Does the specification of uncertainty hurt the progress of scientometrics?

In "Caveats for using statistical significance tests in research assessments,"--Journal of Informetrics 7(1)(2013) 50-62, available at arXiv:1112.2516 -- Schneider (2013) focuses on Opthof & Leydesdorff (2010) as an example of the misuse of statistics in the social sciences. However, our conclusions are theoretical since they are not dependent on the use of one statistics or another. We agree with Schneider insofar as he proposes to develop further statistical instruments (such as effect sizes). Schneider (2013), however, argues on meta-theoretical grounds against the specification of uncertainty because, in his opinion, the presence of statistics would legitimate decision-making. We disagree: uncertainty can also be used for opening a debate. Scientometric results in which error bars are suppressed for meta-theoretical reasons should not be trusted.

preprint2012arXiv

Edited Volumes, Monographs, and Book Chapters in the Book Citation Index (BKCI) and Science Citation Index (SCI, SoSCI, A&HCI)

In 2011, Thomson-Reuters introduced the Book Citation Index (BKCI) as part of the Science Citation Index (SCI). The interface of the Web of Science version 5 enables users to search for both "Books" and "Book Chapters" as new categories. Books and book chapters, however, were always among the cited references, and book chapters have been included in the database since 2005. We explore the two categories with both BKCI and SCI, and in the sister social sciences (SoSCI) and the arts & humanities (A&HCI) databases. Book chapters in edited volumes can be highly cited. Books contain many citing references but are relatively less cited. This may find its origin in the slower circulation of books than of journal articles. It is possible to distinguish between monographs and edited volumes among the "Books" scientometrically. Monographs may be underrated in terms of citation impact or overrated using publication performance indicators because individual chapters are counted as contributions separately in terms of articles, reviews, and/or book chapters.

preprint2012arXiv

Global Maps of Science based on the new Web-of-Science Categories

In August 2011, Thomson Reuters launched version 5 of the Science and Social Science Citation Index in the Web of Science (WoS). Among other things, the 222 ISI Subject Categories (SCs) for these two databases in version 4 of WoS were renamed and extended to 225 WoS Categories (WCs). A new set of 151 Subject Categories (SCs) was added, but at a higher level of aggregation. Since we previously used the ISI SCs as the baseline for a global map in Pajek (Rafols et al., 2010) and brought this facility online (at http://www.leydesdorff.net/overlaytoolkit), we recalibrated this map for the new WC categories using the Journal Citation Reports 2010. In the new installation, the base maps can also be made using VOSviewer (Van Eck & Waltman, 2010).

preprint2012arXiv

How Can Journal Impact Factors be Normalized across Fields of Science? An Assessment in terms of Percentile Ranks and Fractional Counts

Using the CD-ROM version of the Science Citation Index 2010 (N = 3,705 journals), we study the (combined) effects of (i) fractional counting on the impact factor (IF) and (ii) transformation of the skewed citation distributions into a distribution of 100 percentiles and six percentile rank classes (top-1%, top-5%, etc.). Do these approaches lead to field-normalized impact measures for journals? In addition to the two-year IF (IF2), we consider the five-year IF (IF5), the respective numerators of these IFs, and the number of Total Cites, counted both as integers and fractionally. These various indicators are tested against the hypothesis that the classification of journals into 11 broad fields by PatentBoard/National Science Foundation provides statistically significant between-field effects. Using fractional counting the between-field variance is reduced by 91.7% in the case of IF5, and by 79.2% in the case of IF2. However, the differences in citation counts are not significantly affected by fractional counting. These results accord with previous studies, but the longer citation window of a fractionally counted IF5 can lead to significant improvement in the normalization across fields.

preprint2012arXiv

How journal rankings can suppress interdisciplinary research. A comparison between Innovation Studies and Business & Management

This study provides quantitative evidence on how the use of journal rankings can disadvantage interdisciplinary research in research evaluations. Using publication and citation data, it compares the degree of interdisciplinarity and the research performance of a number of Innovation Studies units with that of leading Business & Management schools in the UK. On the basis of various mappings and metrics, this study shows that: (i) Innovation Studies units are consistently more interdisciplinary in their research than Business & Management schools; (ii) the top journals in the Association of Business Schools' rankings span a less diverse set of disciplines than lower-ranked journals; (iii) this results in a more favourable assessment of the performance of Business & Management schools, which are more disciplinary-focused. This citation-based analysis challenges the journal ranking-based assessment. In short, the investigation illustrates how ostensibly 'excellence-based' journal rankings exhibit a systematic bias in favour of mono-disciplinary research. The paper concludes with a discussion of implications of these phenomena, in particular how the bias is likely to affect negatively the evaluation and associated financial resourcing of interdisciplinary research organisations, and may result in researchers becoming more compliant with disciplinary authority over time.

preprint2012arXiv

Identifying Research Fields within Business and Management: A Journal Cross-Citation Analysis

A discipline such as business and management (B&M) is very broad and has many fields within it, ranging from fairly scientific ones such as management science or economics to softer ones such as information systems. There are at least two reasons why it is important to identify these sub-fields accurately. Firstly, for the purpose of normalizing citation data as it is well known that citation rates vary significantly between different disciplines. Secondly, because journal rankings and lists tend to split their classifications into different subjects, for example the the Association of Business Schools (ABS) list, which is a standard in the UK, has 22 different fields. Unfortunately, at the moment these are created in an ad hoc manner with no underlying rigour. The purpose of this paper is to identify possible sub-fields in B&M rigorously based on actual citation patterns. We have examined 450 journals in B&M which are included in the ISI Web of Science (WoS) and analysed the cross-citation rates between them enabling us to generate sets of coherent and consistent sub-fields that minimise the extent to which journals appear in several categories. Implications and limitations of the analysis are discussed

preprint2012arXiv

Information Metrics (iMetrics): A Research Specialty with a Socio-Cognitive Identity?

"Bibliometrics", "scientometrics", "informetrics", and "webometrics" can all be considered as manifestations of a single research area with similar objectives and methods, which we call "information metrics" or iMetrics. This study explores the cognitive and social distinctness of iMetrics with respect to the general information science (IS), focusing on a core of researchers, shared vocabulary and literature/knowledge base. Our analysis investigates the similarities and differences between four document sets. The document sets are drawn from three core journals for iMetrics research (Scientometrics, Journal of the American Society for Information Science and Technology, and Journal of Informetrics). We split JASIST into document sets containing iMetrics and general IS articles. The volume of publications in this representation of the specialty has increased rapidly during the last decade. A core of researchers that predominantly focus on iMetrics topics can thus be identified. This core group has developed a shared vocabulary as exhibited in high similarity of title words and one that shares a knowledge base. The research front of this field moves faster than the research front of information science in general, bringing it closer to Price's dream.

preprint2012arXiv

Interactive Overlay Maps for US Patent (USPTO) Data Based on International Patent Classifications (IPC)

We report on the development of an interface to the US Patent and Trademark Office (USPTO) that allows for the mapping of patent portfolios as overlays to basemaps constructed from citation relations among all patents contained in this database during the period 1976-2011. Both the interface and the data are in the public domain; the freeware programs VOSViewer and/or Pajek can be used for the visualization. These basemaps and overlays can be generated at both the 3-digit and 4-digit levels of the International Patent Classifications (IPC) of the World Intellectual Property Organization (WIPO). The basemaps can provide a stable mental framework for analysts to follow developments over searches for different years, which can be animated. The full flexibility of the advanced search engines of USPTO are available for generating sets of patents and/or patent applications which can thus be visualized and compared. This instrument allows for addressing questions about technological distance, diversity in portfolios, and animating the developments of both technologies and technological capacities of organizations over time.

preprint2012arXiv

Mapping (USPTO) Patent Data using Overlays to Google Maps

A technique is developed using patent information available online (at the US Patent and Trademark Office) for the generation of Google Maps. The overlays indicate both the quantity and quality of patents at the city level. This information is relevant for research questions in technology analysis, innovation studies and evolutionary economics, as well as economic geography. The resulting maps can also be relevant for technological innovation policies and R&D management, because the US market can be considered the leading market for patenting and patent competition. In addition to the maps, the routines provide quantitative data about the patents for statistical analysis. The cities on the map are colored according to the results of significance tests. The overlays are explored for the Netherlands as a "national system of innovations," and further elaborated in two cases of emerging technologies: "RNA interference" and "nanotechnology."

preprint2012arXiv

McCall's Area Transformation versus the Integrated Impact Indicator (I3)

In a study entitled "Skewed Citation Distributions and Bias Factors: Solutions to two core problems with the journal impact factor," Mutz & Daniel (2012) propose (i) McCall's (1922) Area Transformation of the skewed citation distribution so that this data can be considered as normally distributed (Krus & Kennedy, 1977), and (ii) to control for different document types as a co-variate (Rubin, 1977). This approach provides an alternative to Leydesdorff & Bornmann's (2011) Integrated Impact Indicator (I3). As the authors note, the two approaches are akin. Can something be said about the relative quality of the two approaches? To that end, I replicated the study of Mutz & Daniel for the 11 journals in the Subject Category "mathematical psychology," but using additionally I3 on the basis of continuous quantiles (Leydesdorff & Bornmann, in press) and its variant PR6 based on the six percentile rank classes distinguished by Bornmann & Mutz (2011) as follows: the top-1%, 95-99%, 90-95%, 75-90%, 50-75%, and bottom-50%.

preprint2012arXiv

Sociological and Communication-Theoretical Perspectives on the Commercialization of the Sciences

Both self-organization and organization are important for the further development of the sciences: the two dynamics condition and enable each other. Commercial and public considerations can interact and "interpenetrate" in historical organization; different codes of communication are then "recombined." However, self-organization in the symbolically generalized codes of communication can be expected to operate at the global level. The Triple Helix model allows for both a neo-institutional appreciation in terms of historical networks of university-industry-government relations and a neo-evolutionary interpretation in terms of three functions: (i) novelty production, (i) wealth generation, and (iii) political control. Using this model, one can appreciate both subdynamics. The mutual information in three dimensions enables us to measure the trade-off between organization and self-organization as a possible synergy. The question of optimization between commercial and public interests in the different sciences can thus be made empirical.

preprint2012arXiv

Statistical Tests and Research Assessments: A comment on Schneider (2012)

In a recent presentation at the 17th International Conference on Science and Technology Indicators, Schneider (2012) criticised the proposal of Bornmann, de Moya Anegon, and Leydesdorff (2012) and Leydesdorff and Bornmann (2012) to use statistical tests in order to evaluate research assessments and university rankings. We agree with Schneider's proposal to add statistical power analysis and effect size measures to research evaluations, but disagree that these procedures would replace significance testing. Accordingly, effect size measures were added to the Excel sheets that we bring online for testing performance differences between institutions in the Leiden Ranking and the SCImago Institutions Ranking.

preprint2012arXiv

Statistics for the Dynamic Analysis of Scientometric Data: The evolution of the sciences in terms of trajectories and regimes

The gap in statistics between multi-variate and time-series analysis can be bridged by using entropy statistics and recent developments in multi-dimensional scaling. For explaining the evolution of the sciences as non-linear dynamics, the configurations among variables can be important in addition to the statistics of individual variables and trend lines. Animations enable us to combine multiple perspectives (based on configurations of variables) and to visualize path-dependencies in terms of trajectories and regimes. Path-dependent transitions and systems formation can be tested using entropy statistics.

preprint2012arXiv

The Knowledge-Based Economy and the Triple Helix Model

1. Introduction - the metaphor of a "knowledge-based economy"; 2. The Triple Helix as a model of the knowledge-based economy; 3. Knowledge as a social coordination mechanism; 4. Neo-evolutionary dynamics in a Triple Helix of coordination mechanism; 5. The operation of the knowledge base; 6. The restructuring of knowledge production in a KBE; 7. The KBE and the systems-of-innovation approach; 8. The KBE and neo-evolutionary theories of innovation; 8.1 The construction of the evolving unit; 8.2 User-producer relations in systems of innovation; 8.3 'Mode-2' and the production of scientific knowledge; 8.4 A Triple Helix model of innovations; 9. Empirical studies and simulations using the TH model; 10. The KBE and the measurement; 10.1 The communication of meaning and information; 10.2 The expectation of social structure; 10.3 Configurations in a knowledge-based economy

preprint2012arXiv

The Swedish System of Innovation: Regional Synergies in a Knowledge-Based Economy

Based on the complete set of firm data for Sweden (N = 1,187,421; November 2011), we analyze the mutual information among the geographical, technological, and organizational distributions in terms of synergies at regional and national levels. Mutual information in three dimensions can become negative and thus indicate a net export of uncertainty by a system or, in other words, synergy in how knowledge functions are distributed over the carriers. Aggregation at the regional level (NUTS3) of the data organized at the municipal level (NUTS5) shows that 48.5% of the regional synergy is provided by the three metropolitan regions of Stockholm, Gothenburg, and Malmö/Lund. Sweden can be considered as a centralized and hierarchically organized system. Our results accord with other statistics, but this Triple Helix indicator measures synergy more specifically and quantitatively. The analysis also provides us with validation for using this measure in previous studies of more regionalized systems of innovation (such as Hungary and Norway).

preprint2012arXiv

The use of percentiles and percentile rank classes in the analysis of bibliometric data: Opportunities and limits

Percentiles have been established in bibliometrics as an important alternative to mean-based indicators for obtaining a normalized citation impact of publications. Percentiles have a number of advantages over standard bibliometric indicators used frequently: for example, their calculation is not based on the arithmetic mean which should not be used for skewed bibliometric data. This study describes the opportunities and limits and the advantages and disadvantages of using percentiles in bibliometrics. We also address problems in the calculation of percentiles and percentile rank classes for which there is not (yet) a satisfactory solution. It will be hard to compare the results of different percentile-based studies with each other unless it is clear that the studies were done with the same choices for percentile calculation and rank assignment.

preprint2012arXiv

The validation of (advanced) bibliometric indicators through peer assessments: A comparative study using data from InCites and F1000

The data of F1000 provide us with the unique opportunity to investigate the relationship between peers' ratings and bibliometric metrics on a broad and comprehensive data set with high-quality ratings. F1000 is a post-publication peer review system of the biomedical literature. The comparison of metrics with peer evaluation has been widely acknowledged as a way of validating metrics. Based on the seven indicators offered by InCites, we analyzed the validity of raw citation counts (Times Cited, 2nd Generation Citations, and 2nd Generation Citations per Citing Document), normalized indicators (Journal Actual/Expected Citations, Category Actual/Expected Citations, and Percentile in Subject Area), and a journal based indicator (Journal Impact Factor). The data set consists of 125 papers published in 2008 and belonging to the subject category cell biology or immunology. As the results show, Percentile in Subject Area achieves the highest correlation with F1000 ratings; we can assert that for further three other indicators (Times Cited, 2nd Generation Citations, and Category Actual/Expected Citations) the 'true' correlation with the ratings reaches at least a medium effect size.

preprint2012arXiv

Where is Synergy Indicated in the Norwegian Innovation System? Triple-Helix Relations among Technology, Organization, and Geography

Using information theory and data for all (0.5 million) Norwegian firms, the national and regional innovation systems are decomposed into three subdynamics: (i) economic wealth generation, (ii) technological novelty production, and (iii) government interventions and administrative control. The mutual information in three dimensions can then be used as an indicator of potential synergy, that is, reduction of uncertainty. We aggregate the data at the NUTS3 level for 19 counties, the NUTS2 level for seven regions, and the single NUTS1 level for the nation. Measured as in-between group reduction of uncertainty, 11.7 % of the synergy was found at the regional level, whereas only another 2.7% was added by aggregation at the national level. Using this triple-helix indicator, the counties along the west coast are indicated as more knowledge-based than the metropolitan area of Oslo or the geographical environment of the Technical University in Trondheim. Foreign direct investment seems to have larger knowledge spill-overs in Norway (oil, gas, offshore, chemistry, and marine) than the institutional knowledge infrastructure in established universities. The northern part of the country, which receives large government subsidies, shows a deviant pattern.

preprint2012arXiv

World Shares of Publications of the USA, EU-27, and China Compared and Predicted using the New Interface of the Web-of-Science versus Scopus

The new interface of the Web of Science (of Thomson Reuters) enables users to retrieve sets larger than 100,000 documents in a single search. This makes it possible to compare publication trends for China, the USA, EU-27, and a number of smaller countries. China no longer grew exponentially during the 2000s, but linearly. Contrary to previous predictions on the basis of exponential growth or Scopus data, the cross-over of the lines for China and the USA is postponed to the next decade (after 2020) according to this data. These long extrapolations, however, should be used only as indicators and not as predictions. Along with the dynamics in the publication trends, one also has to take into account the dynamics of the databases used for the measurement.

preprint2011arXiv

"Meaning" as a sociological concept: A review of the modeling, mapping, and simulation of the communication of knowledge and meaning

The development of discursive knowledge presumes the communication of meaning as analytically different from the communication of information. Knowledge can then be considered as a meaning which makes a difference. Whereas the communication of information is studied in the information sciences and scientometrics, the communication of meaning has been central to Luhmann's attempts to make the theory of autopoiesis relevant for sociology. Analytical techniques such as semantic maps and the simulation of anticipatory systems enable us to operationalize the distinctions which Luhmann proposed as relevant to the elaboration of Husserl's "horizons of meaning" in empirical research: interactions among communications, the organization of meaning in instantiations, and the self-organization of interhuman communication in terms of symbolically generalized media such as truth, love, and power. Horizons of meaning, however, remain uncertain orders of expectations, and one should caution against reification from the meta-biological perspective of systems theory.

preprint2011arXiv

"Structuration" by Intellectual Organization: The Configuration of Knowledge in Relations among Structural Components in Networks of Science

Using aggregated journal-journal citation networks, the measurement of the knowledge base in empirical systems is factor-analyzed in two cases of interdisciplinary developments during the period 1995-2005: (i) the development of nanotechnology in the natural sciences and (ii) the development of communication studies as an interdiscipline between social psychology and political science. The results are compared with a case of stable development: the citation networks of core journals in chemistry. These citation networks are intellectually organized by networks of expectations in the knowledge base at the specialty (that is, above-journal) level. This "structuration" of structural components (over time) can be measured as configurational information. The latter is compared with the Shannon-type information generated in the interactions among structural components: the difference between these two measures provides us with a measure for the redundancy generated by the specification of a model in the knowledge base of the system. This knowledge base incurs (against the entropy law) to variable extents on the knowledge infrastructures provided by the observable networks of relations.

preprint2011arXiv

A Rejoinder on Energy versus Impact Indicators

Citation distributions are so skewed that using the mean or any other central tendency measure is ill-advised. Unlike G. Prathap's scalar measures (Energy, Exergy, and Entropy or EEE), the Integrated Impact Indicator (I3) is based on non-parametric statistics using the (100) percentiles of the distribution. Observed values can be tested against expected ones; impact can be qualified at the article level and then aggregated.

preprint2011arXiv

Book review: Katy Börner, Atlas of Science: Visualizing What We Know. Cambridge, MA/ London UK: The MIT Press, 2010

Katy Börner has written a wonderful book about visualization that makes our field of scientometrics accessible to much larger audiences. The book is to be read in relation to the ongoing series of exhibitions entitled "Places & Spaces: Mapping Science" currently touring the world. The book also provides the scholarly background to the exhibitions. It celebrates scientometrics as the discipline in the background that enables us to visualize the evolution of knowledge as the acumen of human civilization.

preprint2011arXiv

Emerging Search Regimes: Measuring Co-evolutions among Research, Science, and Society

Scientometric data is used to investigate empirically the emergence of search regimes in Biotechnology, Genomics, and Nanotechnology. Complex regimes can emerge when three independent sources of variance interact. In our model, researchers can be considered as the nodes that carry the science system. Research is geographically situated with site-specific skills, tacit knowledge and infrastructures. Second, the emergent science level refers to the formal communication of codified knowledge published in journals. Third, the socio-economic dynamics indicate the ways in which knowledge production relates to society. Although Biotechnology, Genomics, and Nanotechnology can all be characterised by rapid growth and divergent dynamics, the regimes differ in terms of self-organization among these three sources of variance. The scope of opportunities for researchers to contribute within the constraints of the existing body of knowledge are different in each field. Furthermore, the relevance of the context of application contributes to the knowledge dynamics to various degrees.

preprint2011arXiv

Fractional counting of citations in research evaluation: A cross- and interdisciplinary assessment of the Tsinghua University in Beijing

In the case of the scientometric evaluation of multi- or interdisciplinary units one risks to compare apples with oranges: each paper has to be assessed in comparison to an appropriate reference set. We suggest that the set of citing papers can be considered as the relevant representation of the field of impact. In order to normalize for differences in citation behavior among fields, citations can be fractionally counted proportionately to the length of the reference lists in the citing papers. This new method enables us to compare among units with different disciplinary affiliations at the paper level and also to assess the statistical significance of differences among sets. Twenty-seven departments of the Tsinghua University in Beijing are thus compared. Among them, the Department of Chinese Language and Linguistics is upgraded from the 19th to the second position in the ranking. The overall impact of 19 of the 27 departments is not significantly different at the 5% level when thus normalized for different citation potentials.

preprint2011arXiv

Innovation Systems as Patent Networks: The Netherlands, India and Nanotech

Research in the domain of 'Innovation Studies' has been claimed to allow for the study of how technology will develop in the future. Some suggest that the National and Sectoral Innovation Systems literature has become bogged down, however, into case studies of how specific institutions affect innovation in a specific country. A useful notion for policy makers in particular, Balzat & Hanusch (2004) argued that there is a need for NIS studies to develop complementary and also quantitative methods in order to generate new insights that are comparable across national borders. We use data for patents granted by the World Intellectual Property Organization (WIPO) to map innovation systems. Groupings of patents into primary and secondary classes (co-classification) can be used as relational indicators. Knowledge from one class may be more easily used in another class when a co-classification relation exists. Using social network analysis, we map the co-classification of patents among classes and thus indicate what characterizes an innovation system. A main contribution of this paper is methodological, adding to the repertoire of methods NIS studies use and using information from patents in a different way. Policy makers may also find benefits in the social network analysis of the complete set of patents granted by the WIPO to firms and individuals in a country. Social network analysis indicates what innovation activity occurs in a countries and which fields of technology are likely to give rise to innovative products in the near future. We offer such analysis for the Dutch and Indian Innovation Systems. This social network analysis could also be done for a Sector Innovation System, and we do so for Nanotech to determine empirically the knowledge field relevant for this emerging scientific domain.

preprint2011arXiv

Integrated Impact Indicators (I3) compared with Impact Factors (IFs): An alternative research design with policy implications

In bibliometrics, the association of "impact" with central-tendency statistics is mistaken. Impacts add up, and citation curves should therefore be integrated instead of averaged. For example, the journals MIS Quarterly and JASIST differ by a factor of two in terms of their respective impact factors (IF), but the journal with the lower IF has the higher impact. Using percentile ranks (e.g., top-1%, top-10%, etc.), an integrated impact indicator (I3) can be based on integration of the citation curves, but after normalization of the citation curves to the same scale. The results across document sets can be compared as percentages of the total impact of a reference set. Total number of citations, however, should not be used instead because the shape of the citation curves is then not appreciated. I3 can be applied to any document set and any citation window. The results of the integration (summation) are fully decomposable in terms of journals or instititutional units such as nations, universities, etc., because percentile ranks are determined at the paper level. In this study, we first compare I3 with IFs for the journals in two ISI Subject Categories ("Information Science & Library Science" and "Multidisciplinary Sciences"). The LIS set is additionally decomposed in terms of nations. Policy implications of this possible paradigm shift in citation impact analysis are specified.

preprint2011arXiv

Interactive Overlays: A New Method for Generating Global Journal Maps from Web-of-Science Data

Recent advances in methods and techniques enable us to develop an interactive overlay to the global map of science based on aggregated citation relations among the 9,162 journals contained in the Science Citation Index and Social Science Citation Index 2009 combined. The resulting mapping is provided by VOSViewer. We first discuss the pros and cons of the various options: cited versus citing, multidimensional scaling versus spring-embedded algorithms, VOSViewer versus Gephi, and the various clustering algorithms and similarity criteria. Our approach focuses on the positions of journals in the multidimensional space spanned by the aggregated journal-journal citations. A number of choices can be left to the user, but we provide default options reflecting our preferences. Some examples are also provided; for example, the potential of using this technique to assess the interdisciplinarity of organizations and/or document sets.

preprint2011arXiv

Mapping excellence in the geography of science: An approach based on Scopus data

As research becomes an ever more globalized activity, there is growing interest in national and international comparisons of standards and quality in different countries and regions. A sign for this trend is the increasing interest in rankings of universities according to their research performance, both inside but also outside the scientific environment. New methods presented in this paper, enable us to map centers of excellence around the world using programs that are freely available. Based on Scopus data, field-specific excellence can be identified and agglomerated in regions and cities where recently highly-cited papers were published. Differences in performance rates can be visualized on the map using colors and sizes of the marks.

preprint2011arXiv

Percentile Ranks and the Integrated Impact Indicator (I3)

We tested Rousseau's (in press) recent proposal to define percentile classes in the case of the Integrated Impact Indicator (I3) so that the largest number in a set always belongs to the highest (100th) percentile rank class. In the case a set of nine uncited papers and one with citation, however, the uncited papers would all be placed in the 90th percentile rank. A lowly-cited document set would thus be advantaged when compared with a highly-cited one. Notwithstanding our reservations, we extended the program for computing I3 in Web-of-Science data (at http://www.leydesdorff.net/software/i3) with this option; the quantiles without a correction are now the default. As Rousseau mentions, excellence indicators (e.g., the top-10%) can be considered as special cases of I3: only two percentile rank classes are distinguished for the evaluation. Both excellence and impact indicators can be tested statistically using the z-test for independent proportions.

preprint2011arXiv

The Communication of Meaning in Anticipatory Systems: A Simulation Study of the Dynamics of Intentionality in Social Interactions

Psychological and social systems provide us with a natural domain for the study of anticipations because these systems are based on and operate in terms of intentionality. Psychological systems can be expected to contain a model of themselves and their environments social systems can be strongly anticipatory and therefore co-construct their environments, for example, in techno-economic (co-)evolutions. Using Duboi's hyper-incursive and incursive formulations of the logistic equation, these two types of systems and their couplings can be simulated. In addition to their structural coupling, psychological and social systems are also coupled by providing meaning reflexively to each other's meaning-processing. Luhmann's distinctions among (1) interactions between intentions at the micro-level, (2) organization at the meso-level, and (3) self-organization of the fluxes of meaningful communication at the global level can be modeled and simulated using three hyper-incursive equations. The global level of self-organizing interactions among fluxes of communication is retained at the meso-level of organization. In a knowledge-based economy, these two levels of anticipatory structuration can be expected to propel each other at the supra-individual level.

preprint2011arXiv

The Local Emergence and Global Diffusion of Research Technologies: An Exploration of Patterns of Network Formation

Grasping the fruits of "emerging technologies" is an objective of many government priority programs in a knowledge-based and globalizing economy. We use the publication records (in the Science Citation Index) of two emerging technologies to study the mechanisms of diffusion in the case of two innovation trajectories: small interference RNA (siRNA) and nano-crystalline solar cells (NCSC). Methods for analyzing and visualizing geographical and cognitive diffusion are specified as indicators of different dynamics. Geographical diffusion is illustrated with overlays to Google Maps; cognitive diffusion is mapped using an overlay to a map based on the ISI Subject Categories. The evolving geographical networks show both preferential attachment and small-world characteristics. The strength of preferential attachment decreases over time, while the network evolves into an oligopolistic control structure with small-world characteristics. The transition from disciplinary-oriented ("mode-1") to transfer-oriented ("mode-2") research is suggested as the crucial difference in explaining the different rates of diffusion between siRNA and NCSC.

preprint2011arXiv

The new Excellence Indicator in the World Report of the SCImago Institutions Rankings 2011

The new excellence indicator in the World Report of the SCImago Institutions Rankings (SIR) makes it possible to test differences in the ranking in terms of statistical significance. For example, at the 17th position of these rankings, UCLA has an output of 37,994 papers with an excellence indicator of 28.9. Stanford University follows at the 19th position with 37,885 papers and 29.1 excellence, and z = - 0.607. The difference between these two institution thus is not statistically significant. We provide a calculator at http://www.leydesdorff.net/scimago11/scimago11.xls in which one can fill out this test for any two institutions and also for each institution on whether its score is significantly above or below expectation (assuming that 10% of the papers are for stochastic reasons in the top-10% set).

preprint2011arXiv

The semantic mapping of words and co-words in contexts

Meaning can be generated when information is related at a systemic level. Such a system can be an observer, but also a discourse, for example, operationalized as a set of documents. The measurement of semantics as similarity in patterns (correlations) and latent variables (factor analysis) has been enhanced by computer techniques and the use of statistics; for example, in "Latent Semantic Analysis". This communication provides an introduction, an example, pointers to relevant software, and summarizes the choices that can be made by the analyst. Visualization ("semantic mapping") is thus made more accessible.

preprint2011arXiv

The structure of the Arts & Humanities Citation Index: A mapping on the basis of aggregated citations among 1,157 journals

Using the Arts & Humanities Citation Index (A&HCI) 2008, we apply mapping techniques previously developed for mapping journal structures in the Science and Social Science Citation Indices. Citation relations among the 110,718 records were aggregated at the level of 1,157 journals specific to the A&HCI, and the journal structures are questioned on whether a cognitive structure can be reconstructed and visualized. Both cosine-normalization (bottom up) and factor analysis (top down) suggest a division into approximately twelve subsets. The relations among these subsets are explored using various visualization techniques. However, we were not able to retrieve this structure using the ISI Subject Categories, including the 25 categories which are specific to the A&HCI. We discuss options for validation such as against the categories of the Humanities Indicators of the American Academy of Arts and Sciences, the panel structure of the European Reference Index for the Humanities (ERIH), and compare our results with the curriculum organization of the Humanities Section of the College of Letters and Sciences of UCLA as an example of institutional organization.

preprint2011arXiv

The Triple Helix, Quadruple Helix, . . ., and an N-tuple of Helices: Explanatory Models for Analyzing the Knowledge-based Economy?

Using the Triple Helix model of university-industry-government relations, one can measure the extent to which innovation has become systemic instead of assuming the existence of national (or regional) systems of innovations on a priori grounds. Systemness of innovation patterns, however, can be expected to remain in transition because of integrating and differentiating forces. Integration among the functions of wealth creation, knowledge production, and normative control takes place at the interfaces in organizations, while exchanges on the market, scholarly communication in knowledge production, and political discourse tend to differentiate globally. The neo-institutional and the neo-evolutionary versions of the Triple Helix model enable us to capture this tension reflexively. Empirical studies inform us whether more than three helices are needed for the explanation. The Triple Helix indicator can be extended algorithmically, for example, with local-global as a fourth dimension or, more generally, to an N-tuple of helices.

preprint2011arXiv

Turning the tables in citation analysis one more time: Principles for comparing sets of documents

We submit newly developed citation impact indicators based not on arithmetic averages of citations but on percentile ranks. Citation distributions are-as a rule-highly skewed and should not be arithmetically averaged. With percentile ranks, the citation of each paper is rated in terms of its percentile in the citation distribution. The percentile ranks approach allows for the formulation of a more abstract indicator scheme that can be used to organize and/or schematize different impact indicators according to three degrees of freedom: the selection of the reference sets, the evaluation criteria, and the choice of whether or not to define the publication sets as independent. Bibliometric data of seven principal investigators (PIs) of the Academic Medical Center of the University of Amsterdam is used as an exemplary data set. We demonstrate that the proposed indicators [R(6), R(100), R(6,k), R(100,k)] are an improvement of averages-based indicators because one can account for the shape of the distributions of citations over papers.

preprint2011arXiv

Visualization and Analysis of Frames in Collections of Messages: Content Analysis and the Measurement of Meaning

A step-to-step introduction is provided on how to generate a semantic map from a collection of messages (full texts, paragraphs or statements) using freely available software and/or SPSS for the relevant statistics and the visualization. The techniques are discussed in the various theoretical contexts of (i) linguistics (e.g., Latent Semantic Analysis), (ii) sociocybernetics and social systems theory (e.g., the communication of meaning), and (iii) communication studies (e.g., framing and agenda-setting). We distinguish between the communication of information in the network space (social network analysis) and the communication of meaning in the vector space. The vector space can be considered a generated as an architecture by the network of relations in the network space; words are then not only related, but also positioned. These positions are expected rather than observed and therefore one can communicate meaning. Knowledge can be generated when these meanings can recursively be communicated and therefore also further codified.

preprint2011arXiv

Which are the best cities for psychology research worldwide? A map visualizing city ratios of observed and expected numbers of highly-cited papers

We present scientometric results about world-wide centers of excellence in psychology. Based on Web of Science data, domain-specific excellence can be identified for cities where highly cited papers are published. Data refer to all psychology articles published in 2007 which are documented in the Social Science Citation Index and to their citation frequencies from 2007 to May 2011. Visualized are 214 cities with an article output of at least 50 in 2007. Statistical z tests are used for the evaluation of the degree to which an observed number of top-cited papers (top-10%) for a city differs from the number expected on the basis of randomness in the selection of papers. Map visualizing city ratios on significant differences between observed and expected numbers of highly-cited papers point at excellence centers in cities at the East and West Coast of the United States as well as in Great Britain, Germany, the Netherlands, Ireland, Belgium, Sweden, Finland, Australia, and Taiwan. Furthermore, positive but non-significant differences in favor of high citation rates are documented for some cities in the United States, Great Britain, the Netherlands, the Scandinavian and the German-speaking countries, Belgium, France, Spain, Israel, South Korea, and China. Scientometric results show convincingly that highly-cited psychological research articles come from the Anglo-American countries and some of the non-English European countries in which the number of English-language publications has increased during the last decades.

preprint2011arXiv

Which cities produce excellent papers worldwide more than can be expected? A new mapping approach--using Google Maps--based on statistical significance testing

The methods presented in this paper allow for a statistical analysis revealing centers of excellence around the world using programs that are freely available. Based on Web of Science data, field-specific excellence can be identified in cities where highly-cited papers were published significantly. Compared to the mapping approaches published hitherto, our approach is more analytically oriented by allowing the assessment of an observed number of excellent papers for a city (in the sample) against the expected number. Using this test, the approach cannot only identify the top performers in output but the "true jewels." These are cities locating authors who publish significantly more top cited papers than can be expected. As the examples in this paper show for physics, chemistry, and psychology, these cities do not necessarily have a high output of excellent papers.

preprint2011arXiv

Which cities' paper output and citation impact are above expectation in information science? Some improvements of our previous mapping approaches

Bornmann and Leydesdorff (in press) proposed methods based on Web-of-Science data to identify field-specific excellence in cities where highly-cited papers were published more frequently than can be expected. Top performers in output are cities in which authors are located who publish a number of highly-cited papers that is statistically significantly higher than can be expected for these cities. Using papers published between 1989 and 2009 in information science improvements to the methods of Bornmann and Leydesdorff (in press) are presented and an alternative mapping approach based on the indicator I3 is introduced here. The I3 indicator was introduced by Leydesdorff and Bornmann (in press).

preprint2010arXiv

A comparative study on communication structures of Chinese journals in the social sciences

We argue that the communication structures in the Chinese social sciences have not yet been sufficiently reformed. Citation patterns among Chinese domestic journals in three subject areas -- political science and marxism, library and information science, and economics -- are compared with their counterparts internationally. Like their colleagues in the natural and life sciences, Chinese scholars in the social sciences provide fewer references to journal publications than their international counterparts; like their international colleagues, social scientists provide fewer references than natural sciences. The resulting citation networks, therefore, are sparse. Nevertheless, the citation structures clearly suggest that the Chinese social sciences are far less specialized in terms of disciplinary delineations than their international counterparts. Marxism studies are more established than political science in China. In terms of the impact of the Chinese political system on academic fields, disciplines closely related to the political system are less specialized than those weakly related. In the discussion section, we explore reasons that may cause the current stagnation and provide policy recommendations.

preprint2010arXiv

Caveats for the journal and field normalizations in the CWTS ("Leiden") evaluations of research performance

The Center for Science and Technology Studies at Leiden University advocates the use of specific normalizations for assessing research performance with reference to a world average. The Journal Citation Score (JCS) and Field Citation Score (FCS) are averaged for the research group or individual researcher under study, and then these values are used as denominators of the (mean) Citations per publication (CPP). Thus, this normalization is based on dividing two averages. This procedure only generates a legitimate indicator in the case of underlying normal distributions. Given the skewed distributions under study, one should average the observed versus expected values which are to be divided first for each publication. We show the effects of the Leiden normalization for a recent evaluation where we happened to have access to the underlying data.

preprint2010arXiv

Communicative Competencies and the Structuration of Expectations: The creative tension between Habermas' critical theory and Luhmann's social systems theory

I elaborate on the tension between Luhmann's social systems theory and Habermas' theory of communicative action, and argue that this tension can be resolved by focusing on language as the interhuman medium of the communication which enables us to develop symbolically generalized media of communication such as truth, love, power, etc. Following Luhmann, the layers of self-organization among the differently codified subsystems of communication versus organization of meaning at contingent interfaces can analytically be distinguished as compatible, yet empirically researchable alternatives to Habermas' distinction between "system" and "lifeworld." Mediation by a facilitator can then be considered as a special case of organizing historically contingent translations among the evolutionarily developing fluxes of intentions and expectations. Accordingly, I suggest modifying Giddens' terminology into "a theory of the structuration of expectations."

preprint2010arXiv

Competing Technologies: Disturbance, Selection, and the Possibilities of Lock-in

Arthur's (1988) model for competing technologies is discussed from the perspective of evolution theory. Using Arthur's own model for the simulation, we show that 'lock-ins' can be suppressed by adding reflexivity or uncertainty on the side of consumers. Competing technologies then tend to remain in competition. From an evolutionary perspective, lock-ins and prevailing equilibrium can be considered as different trajectories of the techno-economic systems under study. Our simulation results suggest that technological developments which affect the natural preferences of consumers do not induce changes in trajectory, while changes in network parameters of a technology sometimes induce ordered substitution processes. These substitution processes have been shown empirically (e.g., Fisher & Prey, 1971), but hitherto they have been insufficiently understood from the perspective of evolutionary modelling. Implications for technology policies are discussed.

preprint2010arXiv

Distributed scientific communication in the European information society: Some cases of "Mode 2" fields of research

Can self-organization of scientific communication be specified by using literature-based indicators? In this study, we explore this question by applying entropy measures to typical "Mode-2" fields of knowledge production. We hypothesized these scientific systems to be developing from a self-organization of the interaction between cognitive and institutional levels: European subsidized research programs aim at creating an institutional network, while a cognitive reorganization is continuously ongoing at the scientific field level. The results indicate that the European system develops towards a stable level of distribution of cited references and title-words among the European member states. We suggested that this distribution could be a property of the emerging European system. In order to measure to degree of specialization with respect to the respective distributions of countries, cited references and title words, the mutual information among the three frequency distributions was calculated. The so-called transmission values informed us that the European system shows increasing levels of differentiation.

preprint2010arXiv

Fractional counting of citations in research evaluation: An option for cross- and interdisciplinary assessments

In the case of the scientometric evaluation of multi- or interdisciplinary units one risks to compare apples with oranges: each paper has to assessed in comparison to an appropriate reference set. We suggest that the set of citing papers first can be considered as the relevant representation of the field of impact. In order to normalize for differences in citation behavior among fields, citations can be fractionally counted proportionately to the length of the reference lists in the citing papers. This new method enables us to compare among units with different disciplinary affiliations at the paper level and also to assess the statistical significance of differences among sets. Twenty-seven departments of the Tsinghua University in Beijing are thus compared. Among them, the Department of Chinese Language and Linguistics is upgraded from the 19th to the second position in the ranking. The overall impact of 19 of the 27 departments is not significantly different at the 5% level when thus normalized for different citation potentials.

preprint2010arXiv

How fractional counting affects the Impact Factor: Normalization in terms of differences in citation potentials among fields of science

The ISI-Impact Factors suffer from a number of drawbacks, among them the statistics-why should one use the mean and not the median?-and the incomparability among fields of science because of systematic differences in citation behavior among fields. Can these drawbacks be counteracted by counting citation weights fractionally instead of using whole numbers in the numerators? (i) Fractional citation counts are normalized in terms of the citing sources and thus would take into account differences in citation behavior among fields of science. (ii) Differences in the resulting distributions can be tested statistically for their significance at different levels of aggregation. (iii) Fractional counting can be generalized to any document set including journals or groups of journals, and thus the significance of differences among both small and large sets can be tested. A list of fractionally counted Impact Factors for 2008 is available online at http://www.leydesdorff.net/weighted_if/weighted_if.xls. The in-between group variance among the thirteen fields of science identified in the U.S. Science and Engineering Indicators is not statistically significant after this normalization. Although citation behavior differs largely between disciplines, the reflection of these differences in fractionally counted citation distributions could not be used as a reliable instrument for the classification.

preprint2010arXiv

How to evaluate universities in terms of their relative citation impacts: Fractional counting of citations and the normalization of differences among disciplines

Fractional counting of citations can improve on ranking of multi-disciplinary research units (such as universities) by normalizing the differences among fields of science in terms of differences in citation behavior. Furthermore, normalization in terms of citing papers abolishes the unsolved questions in scientometrics about the delineation of fields of science in terms of journals and normalization when comparing among different journals. Using publication and citation data of seven Korean research universities, we demonstrate the advantages and the differences in the rankings, explain the possible statistics, and suggest ways to visualize the differences in (citing) audiences in terms of a network.

preprint2010arXiv

Implicit media frames: Automated analysis of public debate on artificial sweeteners

The framing of issues in the mass media plays a crucial role in the public understanding of science and technology. This article contributes to research concerned with diachronic analysis of media frames by making an analytical distinction between implicit and explicit media frames, and by introducing an automated method for analysing diachronic changes of implicit frames. In particular, we apply a semantic maps method to a case study on the newspaper debate about artificial sweeteners, published in The New York Times (NYT) between 1980 and 2006. Our results show that the analysis of semantic changes enables us to filter out the dynamics of implicit frames, and to detect emerging metaphors in public debates. Theoretically, we discuss the relation between implicit frames in public debates and codification of information in scientific discourses, and suggest further avenues for research interested in the automated analysis of frame changes and trends in public debates.

preprint2010arXiv

Indicators of the Interdisciplinarity of Journals: Diversity, Centrality, and Citations

A citation-based indicator for interdisciplinarity has been missing hitherto among the set of available journal indicators. In this study, we investigate network indicators (betweenness centrality), journal indicators (Shannon entropy, the Gini coefficient), and more recently proposed Rao-Stirling measures for "interdisciplinarity." The latter index combines the statistics of both citation distributions of journals (vector-based) and distances in citation networks among journals (matrix-based). The effects of various normalizations are specified and measured using the matrix of 8,207 journals contained in the Journal Citation Reports of the (Social) Science Citation Index 2008. Betweenness centrality in symmetrical (1-mode) cosine-normalized networks provides an indicator outperforming betweenness in the asymmetrical (2-mode) citation network. Among the vector-based indicators, Shannon entropy performs better than the Gini coefficient, but is sensitive to size. Science and Nature, for example, are indicated at the top of the list. The new diversity measure provides reasonable results when (1 - cosine) is assumed as a measure for the distance, but results using Euclidean distances were difficult to interpret.

preprint2010arXiv

Is Inequality Among Universities Increasing? Gini Coefficients and the Elusive Rise of Elite Universities

One of the unintended consequences of the New Public Management (NPM) in universities is often feared to be a division between elite institutions focused on research and large institutions with teaching missions. However, institutional isomorphisms provide counter-incentives. For example, university rankings focus on certain output parameters such as publications, but not on others (e.g., patents). In this study, we apply Gini coefficients to university rankings in order to assess whether universities are becoming more unequal, at the level of both the world and individual nations. Our results do not support the thesis that universities are becoming more unequal. If anything, we predominantly find homogenization, both at the level of the global comparisons and nationally. In a more restricted dataset (using only publications in the natural and life sciences), we find increasing inequality for those countries, which used NPM during the 1990s, but not during the 2000s. Our findings suggest that increased output steering from the policy side leads to a global conformation to performance standards.

preprint2010arXiv

Is the European Monetary System converging to integration?

The emerging system at the European level can be conceptualized as a pattern of relations among member states that tends to be reproduced despite disturbances in individual trajectories. The Markov property is used as an indicator of systemness in the distribution. The individual trajectories of nations participating in the European Monetary System is assessed using an information theoretical model that is consistent with the Markov property in the multivariate case. Economic and monetary integration are analyzed using independent data sets. Increasing integration can be retrieved in both of these dimensions, notably since the currency crises of 1992 and 1993. However, the dynamics for countries which have strongly coupled their currency to the German Mark are different from those which did not. Additionally, developments in inflation and exchange rates at the European level are assessed in relation to global developments.

preprint2010arXiv

Knowledge-Based Innovation Systems and the Model of a Triple Helix of University-Industry-Government Relations

The (neo-)evolutionary model of a Triple Helix of University-Industry-Government Relations focuses on the overlay of expectations, communications, and interactions that potentially feed back on the institutional arrangements among the carrying agencies. From this perspective, the evolutionary perspective in economics can be complemented with the reflexive turn from sociology. The combination provides a richer understanding of how knowledge-based systems of innovation are shaped and reconstructed. The communicative capacities of the carrying agents become crucial to the system's further development, whereas the institutional arrangements (e.g., national systems) can be expected to remain under reconstruction. The tension of the differentiation no longer needs to be resolved, since the network configurations are reproduced by means of translations among historically changing codes. Some methodological and epistemological implications for studying innovation systems are explicated.

preprint2010arXiv

Mapping the Geography of Science: Distribution Patterns and Networks of Relations among Cities and Institutes

Using Google Earth, Google Maps and/or network visualization programs such as Pajek, one can overlay the network of relations among addresses in scientific publications on the geographic map. We discuss the pros en cons of the various options, and provide software (freeware) for bridging existing gaps between the Science Citation Indices and Scopus, on the one side, and these various visualization tools, on the other. At the level of city names, the global map can be drawn reliably on the basis of the available address information. At the level of the names of organizations and institutes, there are problems of unification both in the ISI-databases and Scopus. Pajek enables us to combine the visualization with statistical analysis, whereas the Google Maps and its derivates provide superior tools at the Internet.

preprint2010arXiv

Normalization at the field level: fractional counting of citations

Van Raan et al. (2010; arXiv:1003.2113) have proposed a new indicator (MNCS) for field normalization. Since field normalization is also used in the Leiden Rankings of universities, we elaborate our critique of journal normalization in Opthof & Leydesdorff (2010; arXiv:1002.2769) in this rejoinder concerning field normalization. Fractional citation counting thoroughly solves the issue of normalization for differences in citation behavior among fields. This indicator can also be used to obtain a normalized impact factor.

preprint2010arXiv

Normalization, CWTS indicators, and the Leiden Rankings: Differences in citation behavior at the level of fields

Van Raan et al. (2010; arXiv:1003.2113) have proposed a new indicator (MNCS) for field normalization. Since field normalization is also used in the Leiden Rankings of universities, we elaborate our critique of journal normalization in Opthof & Leydesdorff (2010; arXiv:1002.2769) in this rejoinder concerning field normalization. Fractional citation counting thoroughly solves the issue of normalization for differences in citation behavior among fields. This indicator can also be used to obtain a normalized impact factor.

preprint2010arXiv

Quality Control and Validation Boundaries in a Triple Helix of University-Industry-Government: 'Mode 2' and the Future of University Research

How is quality control organized in the new "Mode 2" of the production of scientific knowledge? When institutional boundaries are increasingly blurred in a Triple Helix of University-Industry-Government relations, criteria for quality control in the production of scientific knowledge can be expected to change at the interfaces. The categorization in terms of two modes of knowledge production was introduced by Gibbons et al. (1994) in order to describe changes in the networks of scientific communications (funding patterns, research configurations, styles of knowledge management, etc.). These changes were mainly specified as institutional parameters in order to deal with the subjects of R&D management and S&T policies, that is, ex ante (Spiegel-Ring 1973 Van den Daele et al. 1979). We focus on the 'validation boundaries' emerging from the differences between Mode 1 and Mode 2 that is, on the criteria for quality control that can analytically and reflexively be brought to the fore ex post. The shift from an institutional frame of reference to a focus on the dynamics of communications enables us to clarify several problems in the discussion of the future of university research.

preprint2010arXiv

Redundancy in Systems which Entertain a Model of Themselves: Interaction Information and the Self-organization of Anticipation

Mutual information among three or more dimensions (mu-star = - Q) has been considered as interaction information. However, Krippendorff (2009a, 2009b) has shown that this measure cannot be interpreted as a unique property of the interactions and has proposed an alternative measure of interaction information based on iterative approximation of maximum entropies. Q can then be considered as a measure of the difference between interaction information and redundancy generated in a model entertained by an observer. I argue that this provides us with a measure of the imprint of a second-order observing system -- a model entertained by the system itself -- on the underlying information processing. The second-order system communicates meaning hyper-incursively; an observation instantiates this meaning-processing within the information processing. The net results may add to or reduce the prevailing uncertainty. The model is tested empirically for the case where textual organization can be expected to contain intellectual organization in terms of distributions of title words, author names, and cited references.

preprint2010arXiv

Remaining problems with the "New Crown Indicator" (MNCS) of the CWTS

In their article, entitled "Towards a new crown indicator: some theoretical considerations," Waltman et al. (2010; at arXiv:1003.2167) show that the "old crown indicator" of CWTS in Leiden was mathematically inconsistent and that one should move to the normalization as applied in the "new crown indicator." Although we now agree about the statistical normalization, the "new crown indicator" inherits the scientometric problems of the "old" one in treating subject categories of journals as a standard for normalizing differences in citation behavior among fields of science. We further note that the "mean" is not a proper statistics for measuring differences among skewed distributions. Without changing the acronym of "MNCS," one could define the "Median Normalized Citation Score." This would relate the new crown indicator directly to the percentile approach that is, for example, used in the Science and Engineering Indicators of US National Science Board (2010). The median is by definition equal to the 50th percentile. The indicator can thus easily be extended with the 1% (= 99th percentile) most highly-cited papers (Bornmann et al., in press). The seeming disadvantage of having to use non-parametric statistics is more than compensated by possible gains in the precision.

preprint2010arXiv

Scaling Trajectories in Civil Aircraft (1913-1997)

Using entropy statistics we analyse scaling patterns in terms of changes in the ratios among product characteristics of 143 designs in civil aircraft. Two allegedly dominant designs, the piston propeller DC3 and the turbofan Boeing 707, are shown to have triggered a scaling trajectory at the level of the respective firms. Along these trajectories different variables have been scaled at different moments in time: this points to the versatility of a dominant design which allows a firm to react to a variety of user needs. Scaling at the level of the industry took off only after subsequently reengineered models were introduced, like the piston propeller Douglas DC4 and the turbofan Boeing 767. The two scaling trajectories in civil aircraft corresponding to the piston propeller and the turbofan paradigm can be compared with a single, less pronounced scaling trajectory in helicopter technology for which we have data during the period 1940-1996. Management and policy implications can be specified in terms of the phases of codification at the firm and the industry level.

preprint2010arXiv

Scientometrics and Communication Theory: Towards Theoretically Informed Indicators

A theory of citations should not consider cited and/or citing agents as its sole subject of study. One is able to study also the dynamics in the networks of communications. While communicating agents (e.g., authors, laboratories, journals) can be made comparable in terms of their publication and citation counts, one would expect the communication networks not to be homogeneous. The latent structures of the network indicate different codifications that span a space of possible 'translations'. The various subdynamics can be hypothesized from an evolutionary perspective. Using the network of aggregated journal-journal citations in Science & Technology Studies as an empirical case, the operation of such subdynamics can be demonstrated. Policy implications and the consequences for a theory-driven type of scientometrics will be elaborated.

preprint2010arXiv

Scopus' SNIP Indicator

Rejoinder to Moed [arXiv:1005.4906]: Our main objection is against developing new indicators which, like some of the older ones (for example, the "crown indicator" of CWTS), do not allow for indicating error because they do not provide a statistics, but are based, in our opinion, on a violation of the order of operations. The claim of validity for the SNIP indicator is hollow because the normalizations are based on field classifications which are not valid. Both problems can perhaps be solved by using fractional counting.

preprint2010arXiv

Scopus's Source Normalized Impact per Paper (SNIP) versus a Journal Impact Factor based on Fractional Counting of Citations

Impact factors (and similar measures such as the Scimago Journal Rankings) suffer from two problems: (i) citation behavior varies among fields of science and therefore leads to systematic differences, and (ii) there are no statistics to inform us whether differences are significant. The recently introduced SNIP indicator of Scopus tries to remedy the first of these two problems, but a number of normalization decisions are involved which makes it impossible to test for significance. Using fractional counting of citations-based on the assumption that impact is proportionate to the number of references in the citing documents-citations can be contextualized at the paper level and aggregated impacts of sets can be tested for their significance. It can be shown that the weighted impact of Annals of Mathematics (0.247) is not so much lower than that of Molecular Cell (0.386) despite a five-fold difference between their impact factors (2.793 and 13.156, respectively).

preprint2010arXiv

Technological Developments and Factor Substitution in a Complex and Dynamic System

Schumpeter's (1939) distinction between changes in the form of the production function corresponding to innovation, and shifts along the production function corresponding to factor substitution, does not preclude that the underlying dynamics interact. In an evolutionarily complex system, such interactions are expected: they lead to non-linear terms in the model, and therefore to stabilization and self-organization in addition to selection and variation. Relatively simple simulations enable us to specify various concepts used in 'evolutionary economics' in terms of non-linear dynamics. While a technological trajectory can be considered as a stabilization in a (distributed) environment, a technological regime can be defined as a next-higher-order globalization in a hyper-space. A regime is able to restore its order despite local disturbances, for example by the political system. Technology policies may be effective at the level of the (sub-)systems if they provide the relevant agents with room for 'creative destruction' of the globalized hyper-systems. Implications for firm behaviour and innovation policies are elaborated.

preprint2010arXiv

The Citation Field of Evolutionary Economics

Evolutionary economics has developed into an academic field of its own, institutionalized around, amongst others, the Journal of Evolutionary Economics (JEE). This paper analyzes the way and extent to which evolutionary economics has become an interdisciplinary journal, as its aim was: a journal that is indispensable in the exchange of expert knowledge on topics and using approaches that relate naturally with it. Analyzing citation data for the relevant academic field for the Journal of Evolutionary Economics, we use insights from scientometrics and social network analysis to find that, indeed, the JEE is a central player in this interdisciplinary field aiming mostly at understanding technological and regional dynamics. It does not, however, link firmly with the natural sciences (including biology) nor to management sciences, entrepreneurship, and organization studies. Another journal that could be perceived to have evolutionary acumen, the Journal of Economic Issues, does relate to heterodox economics journals and is relatively more involved in discussing issues of firm and industry organization. The JEE seems most keen to develop theoretical insights.

preprint2010arXiv

The Decline of University Patenting and the End of the Bayh-Dole Effect

University patenting has been heralded as a symbol of changing relations between universities and their social environments. The Bayh-Dole Act of 1980 in the USA was eagerly promoted by the OECD as a recipe for the commercialization of university research, and the law was imitated by a number of national governments. However, since the 2000s university patenting in the most advanced economies has been on the decline both as a percentage and in absolute terms. We suggest that the institutional incentives for university patenting have disappeared with the new regime of university ranking. Patents and spin-offs are not counted in university rankings. In the new arrangements of university-industry-government relations, universities have become very responsive to changes in their relevant environments.

preprint2010arXiv

The Development of the Journal Environment of Leonardo

We present animations based on the aggregated journal-journal citations of Leonardo during the period 1974-2008. Leonardo is mainly cited by journals outside the arts domain for cultural reasons, for example, in neuropsychology and physics. Articles in Leonardo itself cite a large number of journals, but with a focus on the arts. Animations at this level of aggregation enable us to show the history of the journal from a network perspective.

preprint2010arXiv

The Evolution of Communication Systems

One can study communications by using Shannon's (1948) mathematical theory of communication. In social communications, however, the channels are not "fixed", but themselves subject to change. Communication systems change by communicating information to related communication systems; co-variation among systems if repeated over time, can lead to co-evolution. Conditions for stabilization of higher-order systems are specifiable: segmentation, stratification, differentiation, reflection, and self-organization can be distinguished in terms of developmental stages of increasingly complex networks. In addition to natural and cultural evolution, a condition for the artificial evolution of communication systems can be specified.

preprint2010arXiv

The Non-linear Dynamics of Sociological Reflections

Actors are embedded in networks of communication: the relations of the actors can be represented as the rows of a matrix, while the column vectors represent their communications. The two systems are structurally coupled in the co-variation: each action can be considered as a communication with reference to the network. Co-variation among systems if repeated over time, may lead to co-evolution. Conditions for stabilization of higher-order systems are specifiable: segmentation, stratification, reflection, differentiation, and self-organization can be distinguished in terms of developmental stages of increasingly complex networks. The sociological theory of communication occupies a central position for the clarification of the possibility of a general theory of communication, since it confronts us with the limits of reflexivity in human understanding and reflexive discourse. The implications for modelling the relations among incommensurable discourses (e.g., paradigms) are elaborated.

preprint2010arXiv

The Production of Probabilistic Entropy in Structure/Action Contingency Relations

Luhmann (1984) defined society as a communication system which is structurally coupled to, but not an aggregate of, human action systems. The communication system is then considered as self-organizing ("autopoietic"), as are human actors. Communication systems can be studied by using Shannon's (1948) mathematical theory of communication. The update of a network by action at one of the local nodes is then a well-known problem in artificial intelligence (Pearl 1988). By combining these various theories, a general algorithm for probabilistic structure/action contingency can be derived. The consequences of this contingency for each system, its consequences for their further histories, and the stabilization on each side by counterbalancing mechanisms are discussed, in both mathematical and theoretical terms. An empirical example is elaborated.

preprint2010arXiv

The Triple Helix Model and the Meta-Stabilization of Urban Technologies in Smart Cities

The Triple Helix model of university-industry-government relations can be generalized from a neo-institutional model of networks to a neo-evolutionary model of how three selection environments operate upon one another. The neo-evolutionary model enables us to appreciate both organizational integration in university-industry-government relations and differentiation among functions like the generation of intellectual capital, creation of wealth, and their attending legislation. The specification of innovation systems in terms of nations, sectors, cities, and regions can then be formulated as empirical questions: is synergy generated among functions in networks of relations? This Triple Helix model enables us to study the knowledge base of an urban economy in terms of a trade-off between locally stabilized and (potentially locked-in) trajectories versus the techno-economic and cultural development regimes which work with one more degree of freedom at the global level. The meta-stabilizing potentials of urban technologies between these two levels can be used reflexively as the intelligence of a creative reconstruction making cities smart(er).

preprint2010arXiv

The Triple Helix Perspective of Innovation Systems

Alongside the neo-institutional model of networked relations among universities, industries, and governments, the Triple Helix can be provided with a neo-evolutionary interpretation as three selection environments operating upon one another: markets, organizations, and technological opportunities. How are technological innovation systems different from national ones? The three selection environments fulfill social functions: wealth creation, organization control, and organized knowledge production. The main carriers of this system-industry, government, and academia-provide the variation both recursively and by interacting among them under the pressure of competition. Empirical case studies enable us to understand how these evolutionary mechanisms can be expected to operate in historical instance. The model is needed for distinguishing, for example, between trajectories and regimes.

preprint2010arXiv

Uncertainty and the Communication of Time

Prigogine and Stengers (1988) have pointed to the centrality of the concepts of "time and eternity" for the cosmology contained in Newtonian physics, but they have not addressed this issue beyond the domain of physics. The construction of "time" in the cosmology dates back to debates among Huygens, Newton, and Leibniz. The deconstruction of this cosmology in terms of the philosophical questions of the 17th century suggests an uncertainty in the time dimension. While order has been conceived as an "harmonie préétablie", it is considered as emergent from an evolutionary perspective. In a "chaology", one should fully appreciate that different systems may use different clocks. Communication systems can be considered as contingent in space and time: substances contain force or action, and they communicate not only in (observable) extension, but also over time. While each communication system can be considered as a system of reference for a special theory of communication, the addition of an evolutionary perspective to the mathematical theory of communication opens up the possibility of a general theory of communication.

preprint2010arXiv

What Can Heterogeneity Add to the Scientometric Map? Steps towards algorithmic historiography

The Actor Network represents heterogeneous entities as actants (Callon et al., 1983; 1986). Although computer programs for the visualization of social networks increasingly allow us to represent heterogeneity in a network using different shapes and colors for the visualization, hitherto this possibility has scarcely been exploited (Mogoutov et al., 2008). In this contribution to the Festschrift, I study the question of what heterogeneity can add specifically to the visualization of a network. How does an integrated network improve on the one-dimensional ones (such as co-word and co-author maps)? The oeuvre of Michel Callon is used as the case materials, that is, his 65 papers which can be retrieved from the (Social) Science Citation Index since 1975.

preprint2010arXiv

What the Cited and Citing Environments Reveal of "Advances in Atmospheric Sciences"?

The networking ability of journals reflects their academic influence among peer journals. This paper analyzes the cited and citing environments of the journal--Advances in Atmospheric Sciences--using methods from social network analysis. The journal has been actively participating in the international journal environment, but one has a tendency to cite papers published in international journals. Advances in Atmospheric Sciences is intensely interrelated with international peer journals in terms of similar citing pattern. However, there is still room for an increase in its academic visibility given the comparatively smaller reception in terms of cited references.

preprint2009arXiv

The relation between Pearson's correlation coefficient r and Salton's cosine measure

The relation between Pearson's correlation coefficient and Salton's cosine measure is revealed based on the different possible values of the division of the L1-norm and the L2-norm of a vector. These different values yield a sheaf of increasingly straight lines which form together a cloud of points, being the investigated relation. The theoretical results are tested against the author co-citation relations among 24 informetricians for whom two matrices can be constructed, based on co-citations: the asymmetric occurrence matrix and the symmetric co-citation matrix. Both examples completely confirm the theoretical results. The results enable us to specify an algorithm which provides a threshold value for the cosine above which none of the corresponding Pearson correlations would be negative. Using this threshold value can be expected to optimize the visualization of the vector space.