Researcher profile

Göran Kauermann

Göran Kauermann contributes to research discovery and scholarly infrastructure.

ResearcherAffiliation not importedOpen to collaborate

Trust snapshot

Quick read

Trust 21 - EmergingVerification L1Unclaimed author
11works
0followers
7topics
4close collaborators

Actions

Decide how to stay connected

Follow researcher0

Identity and collaboration

How to connect with this researcher

Claiming links this public author record to a researcher profile and unlocks direct collaboration workflows.

Log in to claim

Direct collaboration

Open a focused conversation when the fit is right

Claim this author entity first to unlock direct invitations.

Research graph

See the researcher in context

Open full explorer

Inspect adjacent work, topics, institutions and collaborators without jumping out to a separate graph page.

Building this graph slice

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

11 published item(s)

preprint2022arXiv

All that Glitters is not Gold: Relational Events Models with Spurious Events

As relational event models are an increasingly popular model for studying relational structures, the reliability of large-scale event data collection becomes more and more important. Automated or human-coded events often suffer from non-negligible false-discovery rates in event identification. And most sensor data is primarily based on actors' spatial proximity for predefined time windows; hence, the observed events could relate either to a social relationship or random co-location. Both examples imply spurious events that may bias estimates and inference. We propose the Relational Event Model for Spurious Events (REMSE), an extension to existing approaches for interaction data. The model provides a flexible solution for modeling data while controlling for spurious events. Estimation of our model is carried out in an empirical Bayesian approach via data augmentation. Based on a simulation study, we investigate the properties of the estimation procedure. To demonstrate its usefulness in two distinct applications, we employ this model to combat events from the Syrian civil war and student co-location data. Results from the simulation and the applications identify the REMSE as a suitable approach to modeling relational event data in the presence of spurious events.

preprint2022arXiv

Going Beyond One-Hot Encoding in Classification: Can Human Uncertainty Improve Model Performance?

Technological and computational advances continuously drive forward the broad field of deep learning. In recent years, the derivation of quantities describing theuncertainty in the prediction - which naturally accompanies the modeling process - has sparked general interest in the deep learning community. Often neglected in the machine learning setting is the human uncertainty that influences numerous labeling processes. As the core of this work, label uncertainty is explicitly embedded into the training process via distributional labels. We demonstrate the effectiveness of our approach on image classification with a remote sensing data set that contains multiple label votes by domain experts for each image: The incorporation of label uncertainty helps the model to generalize better to unseen data and increases model performance. Similar to existing calibration methods, the distributional labels lead to better-calibrated probabilities, which in turn yield more certain and trustworthy predictions.

preprint2022arXiv

Modelling the large and dynamically growing bipartite network of German patents and inventors

We analyse the bipartite dynamic network of inventors and patents registered within the main area of electrical engineering in Germany to explore the driving forces behind innovation. The data at hand leads to a bipartite network, where an edge between an inventor and a patent is present if the inventor is a co-owner of the respective patent. Since more than a hundred thousand patents were filed by similarly as many inventors during the observational period, this amounts to a massive bipartite network, too large to be analysed as a whole. Therefore, we decompose the bipartite network by utilising an essential characteristic of the network: most inventors tend to stay active only for a relatively short period, while new ones become active at each point in time. Consequently, the adjacency matrix carries several structural zeros. To accommodate for these, we propose a bipartite variant of the Temporal Exponential Random Graph Model (TERGM) in which we let the actor set vary over time, differentiate between inventors that already submitted patents and those that did not, and account for pairwise statistics of inventors. Our results corroborate the hypotheses that inventor characteristics and knowledge flows play a crucial role in the dynamics of invention.

preprint2022arXiv

Stochastic Block Smooth Graphon Model

The paper proposes the combination of stochastic blockmodels with smooth graphon models. The first allow for partitioning the set of individuals in a network into blocks which represent groups of nodes that presumably connect stochastically equivalently, therefore often also called communities. Smooth graphon models instead assume that the network's nodes can be arranged on a one-dimensional scale such that closeness implies a similar connectivity behavior. Both models belong to the model class of node-specific latent variables, entailing a natural relationship. While these model strands have developed more or less completely independently, this paper proposes their generalization towards stochastic block smooth graphon models. This approach enables to exploit the advantages of both worlds. We pursue a general EM-type algorithm for estimation and demonstrate the usability by applying the model to both simulated and real-world examples.

preprint2021arXiv

Matrix-free Penalized Spline Smoothing with Multiple Covariates

The paper motivates high dimensional smoothing with penalized splines and its numerical calculation in an efficient way. If smoothing is carried out over three or more covariates the classical tensor product spline bases explode in their dimension bringing the estimation to its numerical limits. A recent approach by Siebenborn and Wagner(2019) circumvents storage expensive implementations by proposing matrix-free calculations which allows to smooth over several covariates. We extend their approach here by linking penalized smoothing and its Bayesian formulation as mixed model which provides a matrix-free calculation of the smoothing parameter to avoid the use of high-computational cross validation. Further, we show how to extend the ideas towards generalized regression models. The extended approach is applied to remote sensing satellite data in combination with spatial smoothing.

preprint2021arXiv

Regional now- and forecasting for data reported with delay: Towards surveillance of COVID-19 infections

Governments around the world continue to act to contain and mitigate the spread of COVID-19. The rapidly evolving situation compels officials and executives to continuously adapt policies and social distancing measures depending on the current state of the spread of the disease. In this context, it is crucial for policymakers to have a firm grasp on what the current state of the pandemic is as well as to have an idea of how the infective situation is going to unfold in the next days. However, as in many other situations of compulsorily-notifiable diseases and beyond, cases are reported with delay to a central register, with this delay deferring an up-to-date view of the state of things. We provide a stable tool for monitoring current infection levels as well as predicting infection numbers in the immediate future at the regional level. We accomplish this through nowcasting of cases that have not yet been reported as well as through predictions of future infections. We apply our model to German data, for which our focus lies in predicting and explain infectious behavior by district.

preprint2020arXiv

A smooth dynamic network model for patent collaboration data

The development and application of models, which take the evolution of network dynamics into account are receiving increasing attention. We contribute to this field and focus on a profile likelihood approach to model time-stamped event data for a large-scale dynamic network. We investigate the collaboration of inventors using EU patent data. As event we consider the submission of a joint patent and we explore the driving forces for collaboration between inventors. We propose a flexible semiparametric model, which includes external and internal covariates, where the latter are built from the network history.

preprint2020arXiv

Estimation of Latent Network Flows in Bike-Sharing Systems

Estimation of latent network flows is a common problem in statistical network analysis. The typical setting is that we know the margins of the network, i.e. in- and outdegrees, but the flows are unobserved. In this paper, we develop a mixed regression model to estimate network flows in a bike-sharing network if only the hourly differences of in- and outdegrees at bike stations are known. We also include exogenous covariates such as weather conditions. Two different parameterizations of the model are considered to estimate 1) the whole network flow and 2) the network margins only. The estimation of the model parameters is proposed via an iterative penalized maximum likelihood approach. This is exemplified by modeling network flows in the Vienna Bike-Sharing Network. Furthermore, a simulation study is conducted to show the performance of the model. For practical purposes it is crucial to predict when and at which station there is a lack or an excess of bikes. For this application, our model shows to be well suited by providing quite accurate predictions.

preprint2020arXiv

Intensity Estimation on Geometric Networks with Penalized Splines

In the past decades, the growing amount of network data has lead to many novel statistical models. In this paper we consider so called geometric networks. Typical examples are road networks or other infrastructure networks. But also the neurons or the blood vessels in a human body can be interpreted as a geometric network embedded in a three-dimensional space. In all these applications a network specific metric rather than the Euclidean metric is usually used, which makes the analyses on network data challenging. We consider network based point processes and our task is to estimate the intensity (or density) of the process which allows to detect high- and low- intensity regions of the underlying stochastic processes. Available routines that tackle this problem are commonly based on kernel smoothing methods. However, kernel based estimation in general exhibits some drawbacks such as suffering from boundary effects and the locality of the smoother. In an Euclidean space, the disadvantages of kernel methods can be overcome by using penalized spline smoothing. We here extend penalized spline smoothing towards smooth intensity estimation on geometric networks and apply the approach to both, simulated and real world data. The results show that penalized spline based intensity estimation is numerically efficient and outperforms kernel based methods. Furthermore, our approach easily allows to incorporate covariates, which allows to respect the network geometry in a regression model framework.

preprint2020arXiv

Mixture Models and Networks -- Overview of Stochastic Blockmodelling

Mixture models are probabilistic models aimed at uncovering and representing latent subgroups within a population. In the realm of network data analysis, the latent subgroups of nodes are typically identified by their connectivity behaviour, with nodes behaving similarly belonging to the same community. In this context, mixture modelling is pursued through stochastic blockmodelling. We consider stochastic blockmodels and some of their variants and extensions from a mixture modelling perspective. We also survey some of the main classes of estimation methods available, and propose an alternative approach. In addition to the discussion of inferential properties and estimating procedures, we focus on the application of the models to several real-world network datasets, showcasing the advantages and pitfalls of different approaches.

preprint2019arXiv

Tempus Volat, Hora Fugit -- A Survey of Tie-Oriented Dynamic Network Models in Discrete and Continuous Time

Given the growing number of available tools for modeling dynamic networks, the choice of a suitable model becomes central. The goal of this survey is to provide an overview of tie-oriented dynamic network models. The survey is focused on introducing binary network models with their corresponding assumptions, advantages, and shortfalls. The models are divided according to generating processes, operating in discrete and continuous time. First, we introduce the Temporal Exponential Random Graph Model (TERGM) and the Separable TERGM (STERGM), both being time-discrete models. These models are then contrasted with continuous process models, focusing on the Relational Event Model (REM). We additionally show how the REM can handle time-clustered observations, i.e., continuous time data observed at discrete time points. Besides the discussion of theoretical properties and fitting procedures, we specifically focus on the application of the models on two networks that represent international arms transfers and email exchange. The data allow to demonstrate the applicability and interpretation of the network models.