Source author record

Andrea Scharnhorst

Andrea Scharnhorst appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

30works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

30 published item(s)

preprint2022arXiv

Classifications as Linked Open Data. Challenges and Opportunities

Linked Data (LD) as a web--based technology enables in principle the seamless, machine--supported integration, interplay and augmentation of all kinds of knowledge, into what has been labeled a huge knowledge graph. Despite decades of web technology and, more recently, the LD approach, the task to fully exploit these new technologies in the public domain is only commencing. One specific challenge is to transfer techniques developed preweb to order our knowledge into the realm of Linked Open Data (LOD) This paper illustrates two different models in which a general analytico--synthetic classification can be published and made available as LD. In both cases, an LD solution deals with the intricacies of a pre--coordinated indexing language.

preprint2022arXiv

The Need for Knowledge Organization. Introduction to the book Linking Knowledge: Linked Open Data for Knowledge Organization

This book is not restricted to semantic web (SW) technologies. An aspiration was to contribute to the awakening of a dialogue between information and documentation concerned with knowledge organization systems (KOSs), and branches in computer science with an emphasis on machines, algorithms and ontologies. The technological evolution of the last decades has not only fostered the emergence of ever more KOSs but also semantic web technologies. Both the actions of 'making a KOS' and 'applying existing KOSs' represent research. The design of an information layer for a knowledge domain and the design of a domain specific research process are intrinsically interwoven. We extended our intervention to KOS practices into education, by presenting a translation of existing standards and recommendations about linked open data (LOD) publishing for non-experts. The chapters describe the state of the art in providing KOSs as semantic artefacts; how the state of the art is applied in new fields; how the state of the art is pushed towards new technological solutions by being confronted with new applications; how best practices need to be tailored towards specific solutions; and what challenges occur when merging new and old ways of expressing KOSs. The linked data (LD) ecosystem represents a source of knowledge generation, acquisition, production and dissemination. The underlying discourse shows historical vision alongside the promise of linking knowledge for interaction. The already maturing ecosystems of the SW are interlocking information institutions clearly devoted to the expansion of human experience through the growth of knowledge interaction.

preprint2020arXiv

Lost or found? Discovering data needed for research

Finding data is a necessary precursor to being able to reuse data, although relatively little large-scale empirical evidence exists about how researchers discover, make sense of and (re)use data for research. This study presents evidence from the largest known survey investigating how researchers discover and use data that they do not create themselves. We examine the data needs and discovery strategies of respondents, propose a typology for data reuse and probe the role of social interactions and literature search in data discovery. We consider how data communities can be conceptualized according to data uses and propose practical applications of our findings for designers of data discovery systems and repositories. Specifically, we consider how to design for a diversity of practices, how communities of use can serve as an entry point for design and the role of metadata in supporting both sensemaking and social interactions.

preprint2020arXiv

Searching Data: A Review of Observational Data Retrieval Practices in Selected Disciplines

A cross-disciplinary examination of the user behaviours involved in seeking and evaluating data is surprisingly absent from the research data discussion. This review explores the data retrieval literature to identify commonalities in how users search for and evaluate observational research data. Two analytical frameworks rooted in information retrieval and science technology studies are used to identify key similarities in practices as a first step toward developing a model describing data retrieval.

preprint2019arXiv

Understanding Data Search as a Socio-technical Practice

Open research data are heralded as having the potential to increase effectiveness, productivity, and reproducibility in science, but little is known about the actual practices involved in data search. The socio-technical problem of locating data for reuse is often reduced to the technological dimension of designing data search systems. We combine a bibliometric study of the current academic discourse around data search with interviews with data seekers. In this article, we explore how adopting a contextual, socio-technical perspective can help to understand user practices and behavior and ultimately help to improve the design of data discovery systems.

preprint2016arXiv

Bibliometrics and Information Retrieval: Creating Knowledge through Research Synergies

This panel brings together experts in bibliometrics and information retrieval to discuss how each of these two important areas of information science can help to inform the research of the other. There is a growing body of literature that capitalizes on the synergies created by combining methodological approaches of each to solve research problems and practical issues related to how information is created, stored, organized, retrieved and used. The session will begin with an overview of the common threads that exist between IR and metrics, followed by a summary of findings from the BIR workshops and examples of research projects that combine aspects of each area to benefit IR or metrics research areas, including search results ranking, semantic indexing and visualization. The panel will conclude with an engaging discussion with the audience to identify future areas of research and collaboration.

preprint2015arXiv

Ariadne's Thread - Interactive Navigation in a World of Networked Information

This work-in-progress paper introduces an interface for the interactive visual exploration of the context of queries using the ArticleFirst database, a product of OCLC. We describe a workflow which allows the user to browse live entities associated with 65 million articles. In the on-line interface, each query leads to a specific network representation of the most prevailing entities: topics (words), authors, journals and Dewey decimal classes linked to the set of terms in the query. This network represents the context of a query. Each of the network nodes is clickable: by clicking through, a user traverses a large space of articles along dimensions of authors, journals, Dewey classes and words simultaneously. We present different use cases of such an interface. This paper provides a link between the quest for maps of science and on-going debates in HCI about the use of interactive information visualisation to empower users in their search.

preprint2015arXiv

Bibliometric-enhanced Information Retrieval: 2nd International BIR Workshop

This workshop brings together experts of communities which often have been perceived as different once: bibliometrics / scientometrics / informetrics on the one side and information retrieval on the other. Our motivation as organizers of the workshop started from the observation that main discourses in both fields are different, that communities are only partly overlapping and from the belief that a knowledge transfer would be profitable for both sides. Bibliometric techniques are not yet widely used to enhance retrieval processes in digital libraries, although they offer value-added effects for users. On the other side, more and more information professionals, working in libraries and archives are confronted with applying bibliometric techniques in their services. This way knowledge exchange becomes more urgent. The first workshop set the research agenda, by introducing in each other methods, reporting about current research problems and brainstorming about common interests. This follow-up workshop continues the overall communication, but also puts one problem into the focus. In particular, we will explore how statistical modelling of scholarship can improve retrieval services for specific communities, as well as for large, cross-domain collections like Mendeley or ResearchGate. This second BIR workshop continues to raise awareness of the missing link between Information Retrieval (IR) and bibliometrics and contributes to create a common ground for the incorporation of bibliometric-enhanced services into retrieval at the scholarly search engine interface.

preprint2015arXiv

Contextualization of topics - browsing through terms, authors, journals and cluster allocations

This paper builds on an innovative Information Retrieval tool, Ariadne. The tool has been developed as an interactive network visualization and browsing tool for large-scale bibliographic databases. It basically allows to gain insights into a topic by contextualizing a search query (Koopman et al., 2015). In this paper, we apply the Ariadne tool to a far smaller dataset of 111,616 documents in astronomy and astrophysics. Labeled as the Berlin dataset, this data have been used by several research teams to apply and later compare different clustering algorithms. The quest for this team effort is how to delineate topics. This paper contributes to this challenge in two different ways. First, we produce one of the different cluster solution and second, we use Ariadne (the method behind it, and the interface - called LittleAriadne) to display cluster solutions of the different group members. By providing a tool that allows the visual inspection of the similarity of article clusters produced by different algorithms, we present a complementary approach to other possible means of comparison. More particular, we discuss how we can - with LittleAriadne - browse through the network of topical terms, authors, journals and cluster solutions in the Berlin dataset and compare cluster solutions as well as see their context.

preprint2015arXiv

Editorial for the Proceedings of the Workshop Knowledge Maps and Information Retrieval (KMIR2014) at Digital Libraries 2014

Knowledge maps are promising tools for visualizing the structure of large-scale information spaces, but still far away from being applicable for searching. The first international workshop on "Knowledge Maps and Information Retrieval (KMIR)", held as part of the International Conference on Digital Libraries 2014 in London, aimed at bringing together experts in Information Retrieval (IR) and knowledge mapping in order to discuss the potential of interactive knowledge maps for information seeking purposes.

preprint2015arXiv

Modelling the Structure and Dynamics of Science Using Books

Scientific research is a major driving force in a knowledge based economy. Income, health and wellbeing depend on scientific progress. The better we understand the inner workings of the scientific enterprise, the better we can prompt, manage, steer, and utilize scientific progress. Diverse indicators and approaches exist to evaluate and monitor research activities, from calculating the reputation of a researcher, institution, or country to analyzing and visualizing global brain circulation. However, there are very few predictive models of science that are used by key decision makers in academia, industry, or government interested to improve the quality and impact of scholarly efforts. We present a novel 'bibliographic bibliometric' analysis which we apply to a large collection of books relevant for the modelling of science. We explain the data collection together with the results of the data analyses and visualizations. In the final section we discuss how the analysis of books that describe different modelling approaches can inform the design of new models of science.

preprint2015arXiv

Walking through a library remotely - Why we need maps for collections and how KnoweScape can help us to make them?

There is no escape from the expansion of information, so that structuring and locating meaningful knowledge becomes ever more difficult. The question of how to order our knowledge is as old as the systematic acquisition, circulation, and storage of knowledge. Classification systems have been known since ancient times. On the Internet, one finds both classifications and taxonomies designed by information professionals and folksonomies based on social tagging. Nevertheless, a user navigating through large information spaces is still confronted with a text based search interface and a list of hits as outcome. There is still an obvious gap between a physical encounter with, for example, a librarys collection and browsing its content through an on-line catalogue. This paper starts from the need of digital scholarship for effective knowledge inquiry, revisits traditional ways to support knowledge ordering and information retrieval, and introduces into a newly funded research network where five different communities from all corners of the scientific landscape join forces in a quest for knowledge maps. It can be read as a manifesto for a newly funded specific research network KnoweScape. At the same time it is a general reflection about what one has to take into account when representing structure and evolution of data, information and knowledge and designing instruments to help scholars and others to navigate across the lands and oceans of knowledge.

preprint2014arXiv

Editorial for the Bibliometric-enhanced Information Retrieval Workshop at ECIR 2014

This first "Bibliometric-enhanced Information Retrieval" (BIR 2014) workshop aims to engage with the IR community about possible links to bibliometrics and scholarly communication. Bibliometric techniques are not yet widely used to enhance retrieval processes in digital libraries, although they offer value-added effects for users. In this workshop we will explore how statistical modelling of scholarship, such as Bradfordizing or network analysis of co-authorship network, can improve retrieval services for specific communities, as well as for large, cross-domain collections. This workshop aims to raise awareness of the missing link between information retrieval (IR) and bibliometrics / scientometrics and to create a common ground for the incorporation of bibliometric-enhanced services into retrieval at the digital library interface. Our interests include information retrieval, information seeking, science modelling, network analysis, and digital libraries. The goal is to apply insights from bibliometrics, scientometrics, and informetrics to concrete practical problems of information retrieval and browsing.

preprint2014arXiv

Knowledge Maps and Information Retrieval (KMIR)

Information systems usually show as a particular point of failure the vagueness between user search terms and the knowledge orders of the information space in question. Some kind of guided searching therefore becomes more and more important in order to precisely discover information without knowing the right search terms. Knowledge maps of digital library collections are promising navigation tools through knowledge spaces but still far away from being applicable for searching digital libraries. However, there is no continuous knowledge exchange between the "map makers" on the one hand and the Information Retrieval (IR) specialists on the other hand. Thus, there is also a lack of models that properly combine insights of the two strands. The proposed workshop aims at bringing together these two communities: experts in IR reflecting on visual enhanced search interfaces and experts in knowledge mapping reflecting on visualizations of the content of a collection that might also present a context for a search term in a visual manner. The intention of the workshop is to raise awareness of the potential of interactive knowledge maps for information seeking purposes and to create a common ground for experiments aiming at the incorporation of knowledge maps into IR models at the level of the user interface.

preprint2014arXiv

Modellierungskonzepte der Synergetik und der Theorie der Selbstorganisation

Mnay models situated in the current research landscape of modelling and simulating social processes have roots in physics. This is visible in the name of specialties as Econophysics or Sociophysics. This chapter describes the history of knowledge transfer from physics, in particular physics of self-organization and evolution, to the social sciences. We discuss why physicists felt called to describe social processes. Across models and simulations the question how to explain the emergence of something new is the most intriguing one. We present one model approach to this problem and introduce a game -- Evolino -- inviting a larger audience to get acquainted with abstract evolution-theory approaches to describe the quest for new ideas.

preprint2013arXiv

"Seed+Expand": A validated methodology for creating high quality publication oeuvres of individual researchers

The study of science at the individual micro-level frequently requires the disambiguation of author names. The creation of author's publication oeuvres involves matching the list of unique author names to names used in publication databases. Despite recent progress in the development of unique author identifiers, e.g., ORCID, VIVO, or DAI, author disambiguation remains a key problem when it comes to large-scale bibliometric analysis using data from multiple databases. This study introduces and validates a new methodology called seed+expand for semi-automatic bibliographic data collection for a given set of individual authors. Specifically, we identify the oeuvre of a set of Dutch full professors during the period 1980-2011. In particular, we combine author records from the National Research Information System (NARCIS) with publication records from the Web of Science. Starting with an initial list of 8,378 names, we identify "seed publications" for each author using five different approaches. Subsequently, we "expand" the set of publication in three different approaches. The different approaches are compared and resulting oeuvres are evaluated on precision and recall using a "gold standard" dataset of authors for which verified publications in the period 2001-2010 are available.

preprint2013arXiv

Bibliometric-enhanced Information Retrieval

Bibliometric techniques are not yet widely used to enhance retrieval processes in digital libraries, although they offer value-added effects for users. In this workshop we will explore how statistical modelling of scholarship, such as Bradfordizing or network analysis of coauthorship network, can improve retrieval services for specific communities, as well as for large, cross-domain collections. This workshop aims to raise awareness of the missing link between information retrieval (IR) and bibliometrics/scientometrics and to create a common ground for the incorporation of bibliometric-enhanced services into retrieval at the digital library interface.

preprint2013arXiv

Genericity versus expressivity - an exercise in semantic interoperable research information systems for Web Science

The web does not only enable new forms of science, it also creates new possibilities to study science and new digital scholarship. This paper brings together multiple perspectives: from individual researchers seeking the best options to display their activities and market their skills on the academic job market; to academic institutions, national funding agencies, and countries needing to monitor the science system and account for public money spending. We also address the research interests aimed at better understanding the self-organising and complex nature of the science system through researcher tracing, the identification of the emergence of new fields, and knowledge discovery using large-data mining and non-linear dynamics. In particular this paper draws attention to the need for standardisation and data interoperability in the area of research information as an indispensable pre-condition for any science modelling. We discuss which levels of complexity are needed to provide a globally, interoperable, and expressive data infrastructure for research information. With possible dynamic science model applications in mind, we introduce the need for a "middle-range" level of complexity for data representation and propose a conceptual model for research data based on a core international ontology with national and local extensions.

preprint2013arXiv

Mapping EINS -- An exercise in mapping the Network of Excellence in Internet Science

This paper demonstrates the application of bibliometric mapping techniques in the area of funded research networks. We discuss how science maps can be used to facilitate communication inside newly formed communities, but also to account for their activities to funding agencies. We present the mapping of EINS as case -- an FP7 funded Network of Excellence. Finally, we discuss how these techniques can be used to serve as knowledge maps for interdisciplinary working experts.

preprint2013arXiv

Training in Data Curation as Service in a Federated Data Infrastructure - the FrontOffice-BackOffice Model

The increasing volume and importance of research data leads to the emergence of research data infrastructures in which data management plays an important role. As a consequence, practices at digital archives and libraries change. In this paper, we focus on a possible alliance between archives and libraries around training activities in data curation. We introduce a so-called \emph{FrontOffice--BackOffice model} and discuss experiences of its implementation in the Netherlands. In this model, an efficient division of tasks relies on a distributed infrastructure in which research institutions (i.e., universities) use centralized storage and data curation services provided by national research data archives. The training activities are aimed at information professionals working at those research institutions, for instance as digital librarians. We describe our experiences with the course \emph{DataIntelligence4Librarians}. Eventually, we reflect about the international dimension of education and training around data curation and stewardship.

preprint2013arXiv

UDC in Action

The UDC (Universal Decimal Classification) is not only a classification language with a long history; it also presents a complex cognitive system worthy of the attention of complexity theory. The elements of the UDC: classes, auxiliaries, and operations are combined into symbolic strings, which in essence represent a complex networks of concepts. This network forms a backbone of ordering of knowledge and at the same time allows expression of different perspectives on various products of human knowledge production. In this paper we look at UDC strings derived from the holdings of libraries. In particular we analyse the subject headings of holdings of the university library in Leuven, and an extraction of UDC numbers from the OCLC WorldCat. Comparing those sets with the Master Reference File, we look into the length of strings, the occurrence of different auxiliary signs, and the resulting connections between UDC classes. We apply methods and representations from complexity theory. Mapping out basic statistics on UDC classes as used in libraries from a point of view of complexity theory bears different benefits. Deploying its structure could serve as an overview and basic information for users among the nature and focus of specific collections. A closer view into combined UDC numbers reveals the complex nature of the UDC as an example for a knowledge ordering system, which deserves future exploration from a complexity theoretical perspective.

preprint2012arXiv

Evolution of Wikipedia's Category Structure

Wikipedia, as a social phenomenon of collaborative knowledge creating, has been studied extensively from various points of views. The category system of Wikipedia, introduced in 2004, has attracted relatively little attention. In this study, we focus on the documentation of knowledge, and the transformation of this documentation with time. We take Wikipedia as a proxy for knowledge in general and its category system as an aspect of the structure of this knowledge. We investigate the evolution of the category structure of the English Wikipedia from its birth in 2004 to 2008. We treat the category system as if it is a hierarchical Knowledge Organization System, capturing the changes in the distributions of the top categories. We investigate how the clustering of articles, defined by the category system, matches the direct link network between the articles and show how it changes over time. We find the Wikipedia category network mostly stable, but with occasional reorganization. We show that the clustering matches the link structure quite well, except short periods preceding the reorganizations.

preprint2012arXiv

Looking at a digital research data archive - Visual interfaces to EASY

In this paper we explore visually the structure of the collection of a digital research data archive in terms of metadata for deposited datasets. We look into the distribution of datasets over different scientific fields; the role of main depositors (persons and institutions) in different fields, and main access choices for the deposited datasets. We argue that visual analytics of metadata of collections can be used in multiple ways: to inform the archive about structure and growth of its collection; to foster collections strategies; and to check metadata consistency. We combine visual analytics and visual enhanced browsing introducing a set of web-based, interactive visual interfaces to the archive's collection. We discuss how text based search combined with visual enhanced browsing enhances data access, navigation, and reuse.

preprint2012arXiv

The evolution of classification systems: Ontogeny of the UDC

To classify is to put things in meaningful groups, but the criteria for doing so can be problematic. Study of evolution of classification includes ontogenetic analysis of change in classification over time. We present an empirical analysis of the UDC over the entire period of its development. We demonstrate stability in main classes, with major change driven by 20th century scientific developments. But we also demonstrate a vast increase in the complexity of auxiliaries. This study illustrates an alternative to Tennis' "scheme-versioning" method.

preprint2011arXiv

Learning in a Landscape: Simulation-building as Reflexive Intervention

This article makes a dual contribution to scholarship in science and technology studies (STS) on simulation-building. It both documents a specific simulation-building project, and demonstrates a concrete contribution to interdisciplinary work of STS insights. The article analyses the struggles that arise in the course of determining what counts as theory, as model and even as a simulation. Such debates are especially decisive when working across disciplinary boundaries, and their resolution is an important part of the work involved in building simulations. In particular, we show how ontological arguments about the value of simulations tend to determine the direction of simulation-building. This dynamic makes it difficult to maintain an interest in the heterogeneity of simulations and a view of simulations as unfolding scientific objects. As an outcome of our analysis of the process and reflections about interdisciplinary work around simulations, we propose a chart, as a tool to facilitate discussions about simulations. This chart can be a means to create common ground among actors in a simulation-building project, and a support for discussions that address other features of simulations besides their ontological status. Rather than foregrounding the chart's classificatory potential, we stress its (past and potential) role in discussing and reflecting on simulation-building as interdisciplinary endeavor. This chart is a concrete instance of the kinds of contributions that STS can make to better, more reflexive practice of simulation-building.

preprint2011arXiv

Need to categorize: A comparative look at the categories of the Universal Decimal Classification system (UDC) and Wikipedia

This study analyzes the differences between the category structure of the Universal Decimal Classification (UDC) system (which is one of the widely used library classification systems in Europe) and Wikipedia. In particular, we compare the emerging structure of category-links to the structure of classes in the UDC. With this comparison we would like to scrutinize the question of how do knowledge maps of the same domain differ when they are created socially (i.e. Wikipedia) as opposed to when they are created formally (UDC) using classificatio theory. As a case study, we focus on the category of "Arts".

preprint2010arXiv

Highly connected - a recipe for success

In this paper, we tackle the problem of innovation spreading from a modeling point of view. We consider a networked system of individuals, with a competition between two groups. We show its relation to the innovation spreading issues. We introduce an abstract model and show how it can be interpreted in this framework, as well as what conclusions we can draw form it. We further explain how model-derived conclusions can help to investigate the original problem, as well as other, similar problems. The model is an agent-based model assuming simple binary attributes of those agents. It uses a majority dynamics (Ising model to be exact), meaning that individuals attempt to be similar to the majority of their peers, barring the occasional purely individual decisions that are modeled as random. We show that this simplistic model can be related to the decision-making during innovation adoption processes. The majority dynamics for the model mean that when a dominant attribute, representing an existing practice or solution, is already established, it will persists in the system. We show however, that in a two group competition, a smaller group that represents innovation users can still convince the larger group, if it has high self-support. We argue that this conclusion, while drawn from a simple model, can be applied to real cases of innovation spreading. We also show that the model could be interpreted in different ways, allowing different problems to profit from our conclusions.

preprint2010arXiv

Tracing scientific influence

Scientometrics is the field of quantitative studies of scholarly activity. It has been used for systematic studies of the fundamentals of scholarly practice as well as for evaluation purposes. Although advocated from the very beginning the use of scientometrics as an additional method for science history is still under explored. In this paper we show how a scientometric analysis can be used to shed light on the reception history of certain outstanding scholars. As a case, we look into citation patterns of a specific paper by the American sociologist Robert K. Merton.