Topic overview

Digital Libraries

719 works2190 researchers

Map preview

Start with the graph, then narrow the list

719works
2190researchers

Next steps

Use the topic as a working map

Open the full map for clusters, then return here to scan ranked papers and people.

Topic graph

See the topic as a live network

Open full explorer

Inspect nearby papers, researchers, institutions and communities without opening a separate graph page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Papers in this area

24 paper(s) to start with

preprint2016arXiv

From Ontology to Structured Applied Epistemology

Developing and organizing new knowledge is a core activity for scholars. Recently, ontologies have been introduced as an approach for organizing knowledge. However, most ontologies do not readily support the development and organization of new knowledge. By comparison, to ontology, epistemology is the study of what can be known. Aspects of epistemology include the acquisition of and justification for new knowledge. Thus, we need to coordinate ontology with epistemology. Because we are developing frameworks for capturing knowledge across several scholarly domains, we describe the work in this paper as exploring structured applied epistemology. Unlike other recent proposals for new approaches to scholarly publishing, we propose an integrated and comprehensive approach. We have explored direct representation based on the rigorous Basic Formal Ontology and in this paper, we consider how epistemology can be incorporated with that. In addition to highly-structured scientific research reports, we also consider how to develop highly-structured descriptions of historical events on which historical analyses can be based.

preprint2017arXiv

Citation indices and dimensional homogeneity

The importance of dimensional analysis and dimensional homogeneity in bibliometric studies is always overlooked. In this paper, we look at this issue systematically and show that most h-type indices have the dimensions of [P], where [P] is the basic dimensional unit in bibliometrics which is the unit publication or paper. The newly introduced Euclidean index, based on the Euclidean length of the citation vector has the dimensions [P3/2]. An empirical example is used to illustrate the concepts.

preprint2017arXiv

Computer-Assisted Processing of Intertextuality in Ancient Languages

The production of digital critical editions of texts using TEI is now a widely-adopted procedure within digital humanities. The work described in this paper extends this approach to the publication of gnomologia (anthologies of wise sayings), which formed a widespread literary genre in many cultures of the medieval Mediterranean. These texts are challenging because they were rarely copied straightforwardly; rather, sayings were selected, reorganised, modified or re-attributed between manuscripts, resulting in a highly interconnected corpus for which a standard approach to digital publication is insufficient. Focusing on Greek and Arabic collections, we address this challenge using semantic web techniques to create an ecosystem of texts, relationships and annotations, and consider a new model - organic, collaborative, interconnected, and open-ended - of what constitutes an edition. This semantic web-based approach allows scholars to add their own materials and annotations to the network of information and to explore the conceptual networks that arise from these interconnected sayings.

preprint2016arXiv

A two-sided academic landscape: portrait of highly-cited documents in Google Scholar (1950-2013)

The main objective of this paper is to identify the set of highly-cited documents in Google Scholar and to define their core characteristics (document types, language, free availability, source providers, and number of versions), under the hypothesis that the wide coverage of this search engine may provide a different portrait about this document set respect to that offered by the traditional bibliographic databases. To do this, a query per year was carried out from 1950 to 2013 identifying the top 1,000 documents retrieved from Google Scholar and obtaining a final sample of 64,000 documents, of which 40% provided a free full-text link. The results obtained show that the average highly-cited document is a journal article or a book (62% of the top 1% most cited documents of the sample), written in English (92.5% of all documents) and available online in PDF format (86.0% of all documents). Yet, the existence of errors especially when detecting duplicates and linking cites properly must be pointed out. The fact of managing with highly cited papers, however, minimizes the effects of these limitations. Given the high presence of books, and to a lesser extend of other document types (su

preprint2016arXiv

Growth of International Cooperation in Science: Revisiting Six Case Studies

International collaboration in science continues to grow at a remarkable rate, but little agreement exists about dynamics of growth and organization at the discipline level. Some suggest that disciplines differ in their collaborative tendencies, reflecting their epistemic culture. This study examines collaborative patterns in six previously studied specialties to add new data and conduct analyses over time. Our findings show that the global network of collaboration continues to add new nations and new participants; each specialty has added many new nations to its lists of collaborating partners since 1990. We also find that the scope of international collaboration is positively related to impact. Network characteristics for the six specialties are notable in that instead of reflecting underlying culture, they tend towards convergence. This observation suggests that the global level may represent next-order dynamics that feed back to the national and local levels (as subsystems) in a complex, networked hierarchy.

preprint2016arXiv

Citation success index - An intuitive pair-wise journal comparison metric

In this paper we present "citation success index", a metric for comparing the citation capacity of pairs of journals. Citation success index is the probability that a random paper in one journal has more citations than a random paper in another journal (50% means the two journals do equally well). Unlike the journal impact factor (IF), the citation success index depends on the broadness and the shape of citation distributions. Also, it is insensitive to sporadic highly-cited papers that skew the IF. Nevertheless, we show, based on 16,000 journals containing ~2.4 million articles, that the citation success index is a relatively tight function of the ratio of IFs of journals being compared, due to the fact that journals with same IF have quite similar citation distributions. The citation success index grows slowly as a function of IF ratio. It is substantial (>90%) only when the ratio of IFs exceeds ~6, whereas a factor of two difference in IF values translates into a modest advantage for the journal with higher IF (index of ~70%). We facilitate the wider adoption of this metric by providing an online calculator that takes as input parameters only the IFs of the pair of journ

preprint2016arXiv

Could freely available articles reduce faculty reliance on library for access? An analysis of items cited by faculty from Singapore Management University

Various studies have attempted to assess the amount of free full text available on the web and recent work have suggested that we are close to the 50% mark for freely available articles (Archambault et al. 2013; Bjork et al. 2010; Jamali and Nabavi 2015). It is natural to wonder if this might reduce researchers' reliance on library subscriptions for access. To do so, we need to determine not just what papers researchers are citing to that are free today, but to estimate if the papers they were citing were freely available at the time they were citing it. We attempt to do so for a sample of citations made by researchers in the Singapore Management University in the field of Economics.

preprint2016arXiv

ScienceWISE: Topic Modeling over Scientific Literature Networks

We provide an up-to-date view on the knowledge management system ScienceWISE (SW) and address issues related to the automatic assignment of articles to research topics. So far, SW has been proven to be an effective platform for managing large volumes of technical articles by means of ontological concept-based browsing. However, as the publication of research articles accelerates, the expressivity and the richness of the SW ontology turns into a double-edged sword: a more fine-grained characterization of articles is possible, but at the cost of introducing more spurious relations among them. In this context, the challenge of continuously recommending relevant articles to users lies in tackling a network partitioning problem, where nodes represent articles and co-occurring concepts create edges between them. In this paper, we discuss the three research directions we have taken for solving this issue: i) the identification of generic concepts to reinforce inter-article similarities; ii) the adoption of a bipartite network representation to improve scalability; iii) the design of a clustering algorithm to identify concepts for cross-disciplinary articles and obtain fine-grained topics

preprint2016arXiv

Three practical field normalised alternative indicator formulae for research evaluation

Although altmetrics and other web-based alternative indicators are now commonplace in publishers' websites, they can be difficult for research evaluators to use because of the time or expense of the data, the need to benchmark in order to assess their values, the high proportion of zeros in some alternative indicators, and the time taken to calculate multiple complex indicators. These problems are addressed here by (a) a field normalisation formula, the Mean Normalised Log-transformed Citation Score (MNLCS) that allows simple confidence limits to be calculated and is similar to a proposal of Lundberg, (b) field normalisation formulae for the proportion of cited articles in a set, the Equalised Mean-based Normalised Proportion Cited (EMNPC) and the Mean-based Normalised Proportion Cited (MNPC), to deal with mostly uncited data sets, (c) a sampling strategy to minimise data collection costs, and (d) free unified software to gather the raw data, implement the sampling strategy, and calculate the indicator formulae and confidence limits. The approach is demonstrated (but not fully tested) by comparing the Scopus citations, Mendeley readers and Wikipedia mentions of research funded

preprint2016arXiv

Prerequisites for International Exchanges of Health Information: Comparison of Australian, Austrian, Finnish, Swiss, and US Privacy Policies

Capabilities to exchange health information are critical to accelerate discovery and its diffusion to healthcare practice. However, the same ethical and legal policies that protect privacy hinder these data exchanges, and the issues accumulate if moving data across geographical or organizational borders. This can be seen as one of the reasons why many health technologies and research findings are limited to very narrow domains. In this paper, we compare how using and disclosing personal data for research purposes is addressed in Australian, Austrian, Finnish, Swiss, and US policies with a focus on text data analytics. Our goal is to identify approaches and issues that enable or hinder international health information exchanges. As expected, the policies within each country are not as diverse as across countries. Most policies apply the principles of accountability and/or adequacy and are thereby fundamentally similar. Their following requirements create complications with re-using and re-disclosing data and even secondary data: 1) informing data subjects about the purposes of data collection and use, before the dataset is collected; 2) assurance that the subjects are no longer iden

preprint2016arXiv

Patent Portfolio Analysis of Cities: Statistics and Maps of Technological Inventiveness

Cities are engines of the knowledge-based economy, because they are the primary sites of knowledge production activities that subsequently shape the rate and direction of technological change and economic growth. Patents provide a wealth of information to analyse the knowledge specialization at specific places, such as technological details and information on inventors and entities involved, including address information. The technology codes on each patent document indicate the specialization and scope of the underlying technological knowledge of a given invention. In this paper we introduce tools for portfolio analysis in terms of patents that provide insights into the technological specialization of cities. The mapping and analysis of patent portfolios of cities using data of the Unites States Patent and Trademark Office (USPTO) website (at http://www.uspto.gov) and dedicated tools (at http://www.leydesdorff.net/portfolio) can be used to analyse the specialisation patterns of inventive activities among cities. The results allow policy makers and other stakeholders to identify promising areas of further knowledge development and 'smart specialisation' strategies.

preprint2016arXiv

Analyzing Web Archives Through Topic and Event Focused Sub-collections

Web archives capture the history of the Web and are therefore an important source to study how societal developments have been reflected on the Web. However, the large size of Web archives and their temporal nature pose many challenges to researchers interested in working with these collections. In this work, we describe the challenges of working with Web archives and propose the research methodology of extracting and studying sub-collections of the archive focused on specific topics and events. We discuss the opportunities and challenges of this approach and suggest a framework for creating sub-collections.

preprint2016arXiv

The iCrawl Wizard -- Supporting Interactive Focused Crawl Specification

Collections of Web documents about specific topics are needed for many areas of current research. Focused crawling enables the creation of such collections on demand. Current focused crawlers require the user to manually specify starting points for the crawl (seed URLs). These are also used to describe the expected topic of the collection. The choice of seed URLs influences the quality of the resulting collection and requires a lot of expertise. In this demonstration we present the iCrawl Wizard, a tool that assists users in defining focused crawls efficiently and semi-automatically. Our tool uses major search engines and Social Media APIs as well as information extraction techniques to find seed URLs and a semantic description of the crawl intent. Using the iCrawl Wizard even non-expert users can create semantic specifications for focused crawlers interactively and efficiently.

preprint2016arXiv

Conceptual difficulties in the use of statistical inference in citation analysis

In this comment, I discuss the use of statistical inference in citation analysis. In a recent paper, Williams and Bornmann argue in favor of the use of statistical inference in citation analysis. I present a critical analysis of their arguments and of similar arguments provided elsewhere in the literature. My conclusion is that the use of statistical inference in citation analysis involves major conceptual difficulties and, consequently, that the usefulness of statistical inference in citation analysis is highly questionable.

preprint2016arXiv

Quantifying perceived impact of scientific publications

Citations are commonly held to represent scientific impact. To date, however, there is no empirical evidence in support of this postulate that is central to research assessment exercises and Science of Science studies. Here, we report on the first empirical verification of the degree to which citation numbers represent scientific impact as it is actually perceived by experts in their respective field. We run a large-scale survey of about 2000 corresponding authors who performed a pairwise impact assessment task across more than 20000 scientific articles. Results of the survey show that citation data and perceived impact do not align well, unless one properly accounts for strong psychological biases that affect the opinions of experts with respect to their own papers vs. those of others. First, researchers tend to largely prefer their own publications to the most cited papers in their field of research. Second, there is only a mild positive correlation between the number of citations of top-cited papers in given research areas and expert preference in pairwise comparisons. This also applies to pairs of papers with several orders of magnitude differences in their total number of accu

preprint2016arXiv

Web of Science: showing a bug today that can mislead scientific research output's prediction

As it happened in all domains of human activities, economic issues and the increase of people working in scientific research have altered the way scientific production is evaluated so as the objectives of performing the evaluation. Introduced in 2005 by J. E. Hirsch as an indicator able to measure individual scientific output not only in terms of quantity, but also in terms of quality, h index has spread throughout the world. In 2007, Hirsch proposed its adoption also as the best to predict future scientific achievement and, consequently, a useful guide for investments in research and for institutions when hiring members for their scientific staff. Since then, several authors have also been using the Thomson ISI Web of Science database to develop their proposals for evaluating research output. Here, using a software we have developed, we analyse more than 100 thousand articles and show that a subtle flaw in Web of Science can inflate the results of info collected, therefore compromising the exactness and, consequently, the effectiveness of Hirsch's proposal and its variations.

preprint2016arXiv

A critical comparative analysis of five world university rankings

To provide users insight into the value and limits of world university rankings, a comparative analysis is conducted of 5 ranking systems: ARWU, Leiden, THE, QS and U-Multirank. It links these systems with one another at the level of individual institutions, and analyses the overlap in institutional coverage, geographical coverage, how indicators are calculated from raw data, the skewness of indicator distributions, and statistical correlations between indicators. Four secondary analyses are presented investigating national academic systems and selected pairs of indicators. It is argued that current systems are still one-dimensional in the sense that they provide finalized, seemingly unrelated indicator values rather than offering a data set and tools to observe patterns in multi-faceted data. By systematically comparing different systems, more insight is provided into how their institutional coverage, rating methods, the selection of indicators and their normalizations influence the ranking positions of given institutions.

preprint2016arXiv

The Effect of Gender in the Publication Patterns in Mathematics

Despite the increasing number of women graduating in mathematics, a systemic gender imbalance persists and is signified by a pronounced gender gap in the distribution of active researchers and professors. Especially at the level of university faculty, women mathematicians continue being drastically underrepresented, decades after the first affirmative action measures have been put into place. A solid publication record is of paramount importance for securing permanent positions. Thus, the question arises whether the publication patterns of men and women mathematicians differ in a significant way. Making use of the zbMATH database, one of the most comprehensive metadata sources on mathematical publications, we analyze the scholarly output of ~150,000 mathematicians from the past four decades whose gender we algorithmically inferred. We focus on development over time, collaboration through coautorships, presumed journal quality and distribution of research topics -- factors known to have a strong impact on job perspectives. We report significant differences between genders which may put women at a disadvantage when pursuing an academic career in mathematics.

preprint2016arXiv

Skewness of citation impact data and covariates of citation distributions: A large-scale empirical analysis based on Web of Science data

Using percentile shares, one can visualize and analyze the skewness in bibliometric data across disciplines and over time. The resulting figures can be intuitively interpreted and are more suitable for detailed analysis of the effects of independent and control variables on distributions than regression analysis. We show this by using percentile shares to analyze so-called "factors influencing citation impact" (FICs; e.g., the impact factor of the publishing journal) across year and disciplines. All articles (n= 2,961,789) covered by WoS in 1990 (n= 637,301), 2000 (n= 919,485), and 2010 (n= 1,405,003) are used. In 2010, nearly half of the citation impact is accounted for by the 10% most-frequently cited papers; the skewness is largest in the humanities (68.5% in the top-10% layer) and lowest in agricultural sciences (40.6%). The comparison of the effects of the different FICs (the number of cited references, number of authors, number of pages, and JIF) on citation impact shows that JIF has indeed the strongest correlations with the citation scores. However, the correlation between FICs and citation impact is lower, if citations are normalized instead of using raw citation c

preprint2016arXiv

Full and Fractional Counting in Bibliometric Networks

In their study entitled "Constructing bibliometric networks: A comparison between full and fractional counting," Perianes-Rodriguez, Waltman, & van Eck (2016; henceforth abbreviated as PWvE) provide arguments for the use of fractional counting at the network level as different from the level of publications. Whereas fractional counting in the latter case divides the credit among co-authors (countries, institutions, etc.), fractional counting at the network level can normalize the relative weights of links and thereby clarify the structures in the network. PWvE, however, propose a counting scheme for fractional counting that is one among other possible ones. Alternative schemes proposed by Batagelj and Cerinšek (2013) and Park, Yoon, & Leydesdorff (2016; henceforth abbreviated as PYL) are discussed in an appendix. However, our approach is not correctly identified as identical to their Equation A3. Here below, we distinguish three approaches analytically; routines for applying these approaches to bibliometric data are also provided.

preprint2016arXiv

A conservation rule for constructing bibliometric network matrices

The social network analysis of bibliometric data needs matrices to be recast in a network framework. In this paper we argue that a simple conservation rule requires that this should be done only using fractional counting so that conservation at the paper level will be faithfully reproduced at higher levels ofaggregation (i.e. author, institute, country, journal etc.) of the complex network.

preprint2011arXiv

Correction of Noisy Sentences using a Monolingual Corpus

Correction of Noisy Natural Language Text is an important and well studied problem in Natural Language Processing. It has a number of applications in domains like Statistical Machine Translation, Second Language Learning and Natural Language Generation. In this work, we consider some statistical techniques for Text Correction. We define the classes of errors commonly found in text and describe algorithms to correct them. The data has been taken from a poorly trained Machine Translation system. The algorithms use only a language model in the target language in order to correct the sentences. We use phrase based correction methods in both the algorithms. The phrases are replaced and combined to give us the final corrected sentence. We also present the methods to model different kinds of errors, in addition to results of the working of the algorithms on the test set. We show that one of the approaches fail to achieve the desired goal, whereas the other succeeds well. In the end, we analyze the possible reasons for such a trend in performance.

People in this topic

12 visible researcher(s)