Source author record

Johannes Rauh

Johannes Rauh appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

23works
16topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

23 published item(s)

preprint2022arXiv

Conditional independence ideals with hidden variables

We study a class of determinantal ideals that are related to conditional independence (CI) statements with hidden variables. Such CI statements correspond to determinantal conditions on a matrix whose entries are probabilities of events involving the observed random variables. We focus on an example that generalizes the CI ideals of the intersection axiom. In this example, the minimal primes are again determinantal ideals, which is not true in general.

preprint2016arXiv

Hierarchical Models as Marginals of Hierarchical Models

We investigate the representation of hierarchical models in terms of marginals of other hierarchical models with smaller interactions. We focus on binary variables and marginals of pairwise interaction models whose hidden variables are conditionally independent given the visible variables. In this case the problem is equivalent to the representation of linear subspaces of polynomials by feedforward neural networks with soft-plus computational units. We show that every hidden variable can freely model multiple interactions among the visible variables, which allows us to generalize and improve previous results. In particular, we show that a restricted Boltzmann machine with less than $[ 2(\log(v)+1) / (v+1) ] 2^v-1$ hidden binary variables can approximate every distribution of $v$ visible binary variables arbitrarily well, compared to $2^{v-1}-1$ from the best previously known result.

preprint2015arXiv

Lifting Markov Bases and Higher Codimension Toric Fiber Products

We study how to lift Markov bases and Gröbner bases along linear maps of lattices. We give a lifting algorithm that allows to compute such bases iteratively provided a certain associated semigroup is normal. Our main application is the toric fiber product of toric ideals, where lifting gives Markov bases of the factor ideals that satisfy the compatible projection property. We illustrate the technique by computing Markov bases of various infinite families of hierarchical models. The methodology also implies new finiteness results for iterated toric fiber products.

preprint2015arXiv

Mode Poset Probability Polytopes

A mode of a probability vector is a local maximum with respect to some vicinity structure on the set of elementary events. The mode inequalities cut out a polytope from the simplex of probability vectors. Related to this is the concept of strong modes. A strong mode of a distribution is an elementary event that has more probability mass than all its direct neighbors together. The set of probability distributions with a given set of strong modes is again a polytope. We study the vertices, the facets, and the volume of such polytopes depending on the sets of (strong) modes and the vicinity structures.

preprint2015arXiv

Quantifying Morphological Computation based on an Information Decomposition of the Sensorimotor Loop

The question how an agent is affected by its embodiment has attracted growing attention in recent years. A new field of artificial intelligence has emerged, which is based on the idea that intelligence cannot be understood without taking into account embodiment. We believe that a formal approach to quantifying the embodiment's effect on the agent's behaviour is beneficial to the fields of artificial life and artificial intelligence. The contribution of an agent's body and environment to its behaviour is also known as morphological computation. Therefore, in this work, we propose a quantification of morphological computation, which is based on an information decomposition of the sensorimotor loop into shared, unique and synergistic information. In numerical simulation based on a formal representation of the sensorimotor loop, we show that the unique information of the body and environment is a good measure for morphological computation. The results are compared to our previously derived quantification of morphological computation.

preprint2014arXiv

Expressive Power and Approximation Errors of Restricted Boltzmann Machines

We present explicit classes of probability distributions that can be learned by Restricted Boltzmann Machines (RBMs) depending on the number of units that they contain, and which are representative for the expressive power of the model. We use this to show that the maximal Kullback-Leibler divergence to the RBM model with $n$ visible and $m$ hidden units is bounded from above by $n - \left\lfloor \log(m+1) \right\rfloor - \frac{m+1}{2^{\left\lfloor\log(m+1)\right\rfloor}} \approx (n -1) - \log(m+1)$. In this way we can specify the number of hidden units that guarantees a sufficiently rich model containing different classes of distributions and respecting a given error tolerance.

preprint2014arXiv

On the Fisher Metric of Conditional Probability Polytopes

We consider three different approaches to define natural Riemannian metrics on polytopes of stochastic matrices. First, we define a natural class of stochastic maps between these polytopes and give a metric characterization of Chentsov type in terms of invariance with respect to these maps. Second, we consider the Fisher metric defined on arbitrary polytopes through their embeddings as exponential families in the probability simplex. We show that these metrics can also be characterized by an invariance principle with respect to morphisms of exponential families. Third, we consider the Fisher metric resulting from embedding the polytope of stochastic matrices in a simplex of joint distributions by specifying a marginal distribution. All three approaches result in slight variations of products of Fisher metrics. This is consistent with the nature of polytopes of stochastic matrices, which are Cartesian products of probability simplices. The first approach yields a scaled product of Fisher metrics; the second, a product of Fisher metrics; and the third, a product of Fisher metrics scaled by the marginal distribution.

preprint2014arXiv

Quantifying unique information

We propose new measures of shared information, unique information and synergistic information that can be used to decompose the multi-information of a pair of random variables $(Y,Z)$ with a third random variable $X$. Our measures are motivated by an operational idea of unique information which suggests that shared information and unique information should depend only on the pair marginal distributions of $(X,Y)$ and $(X,Z)$. Although this invariance property has not been studied before, it is satisfied by other proposed measures of shared information. The invariance property does not uniquely determine our new measures, but it implies that the functions that we define are bounds to any other measures satisfying the same invariance property. We study properties of our measures and compare them to other candidate measures.

preprint2014arXiv

Reconsidering unique information: Towards a multivariate information decomposition

The information that two random variables $Y$, $Z$ contain about a third random variable $X$ can have aspects of shared information (contained in both $Y$ and $Z$), of complementary information (only available from $(Y,Z)$ together) and of unique information (contained exclusively in either $Y$ or $Z$). Here, we study measures $\widetilde{SI}$ of shared, $\widetilde{UI}$ unique and $\widetilde{CI}$ complementary information introduced by Bertschinger et al., which are motivated from a decision theoretic perspective. We find that in most cases the intuitive rule that more variables contain more information applies, with the exception that $\widetilde{SI}$ and $\widetilde{CI}$ information are not monotone in the target variable $X$. Additionally, we show that it is not possible to extend the bivariate information decomposition into $\widetilde{SI}$, $\widetilde{UI}$ and $\widetilde{CI}$ to a non-negative decomposition on the partial information lattice of Williams and Beer. Nevertheless, the quantities $\widetilde{UI}$, $\widetilde{SI}$ and $\widetilde{CI}$ have a well-defined interpretation, even in the multivariate setting.

preprint2014arXiv

Toric fiber products versus Segre products

The toric fiber product is an operation that combines two ideals that are homogeneous with respect to a grading by an affine monoid. The Segre product is a related construction that combines two multigraded rings. The quotient ring by a toric fiber product of two ideals is a subring of the Segre product, but in general this inclusion is strict. We contrast the two constructions and show that any Segre product can be presented as a toric fiber product without changing the involved quotient rings. This allows to apply previous results about toric fiber products to the study of Segre products. We give criteria for the Segre product of two affine toric varieties to be dense in their toric fiber product, and for the map from the Segre product to the toric fiber product to be finite. We give an example that shows that the quotient ring of a toric fiber product of normal ideals need not be normal. In rings with Veronese type gradings, we find examples of toric fiber products that are always Segre products, and we show that iterated toric fiber products of Veronese ideals over Veronese rings are normal.

preprint2013arXiv

Maximal Information Divergence from Statistical Models defined by Neural Networks

We review recent results about the maximal values of the Kullback-Leibler information divergence from statistical models defined by neural networks, including naive Bayes models, restricted Boltzmann machines, deep belief networks, and various classes of exponential families. We illustrate approaches to compute the maximal divergence from a given model starting from simple sub- or super-models. We give a new result for deep and narrow belief networks with finite-valued units.

preprint2013arXiv

Scaling of Model Approximation Errors and Expected Entropy Distances

We compute the expected value of the Kullback-Leibler divergence to various fundamental statistical models with respect to canonical priors on the probability simplex. We obtain closed formulas for the expected model approximation errors, depending on the dimension of the models and the cardinalities of their sample spaces. For the uniform prior, the expected divergence from any model containing the uniform distribution is bounded by a constant $1-γ$, and for the models that we consider, this bound is approached if the state space is very large and the models' dimension does not grow too fast. For Dirichlet priors the expected divergence is bounded in a similar way, if the concentration parameters take reasonable values. These results serve as reference values for more complicated statistical models.

preprint2012arXiv

Generalized Binomial Edge Ideals

This paper studies a class of binomial ideals associated to graphs with finite vertex sets. They generalize the binomial edge ideals, and they arise in the study of conditional independence ideals. A Gröbner basis can be computed by studying paths in the graph. Since these Gröbner bases are square-free, generalized binomial edge ideals are radical. To find the primary decomposition a combinatorial problem involving the connected components of subgraphs has to be solved. The irreducible components of the solution variety are all rational.

preprint2012arXiv

Positive margins and primary decomposition

We study random walks on contingency tables with fixed marginals, corresponding to a (log-linear) hierarchical model. If the set of allowed moves is not a Markov basis, then there exist tables with the same marginals that are not connected. We study linear conditions on the values of the marginals that ensure that all tables in a given fiber are connected. We show that many graphical models have the positive margins property, which says that all fibers with strictly positive marginals are connected by the quadratic moves that correspond to conditional independence statements. The property persists under natural operations such as gluing along cliques, but we also construct examples of graphical models not enjoying this property. We also provide a negative answer to a question of Engström, Kahle, and Sullivant by demonstrating that the global Markov ideal of the complete bipartite graph K_(3,3) is not radical. Our analysis of the positive margins property depends on computing the primary decomposition of the associated conditional independence ideal. The main technical results of the paper are primary decompositions of the conditional independence ideals of graphical models of the $N$-cycle and the complete bipartite graph $K_(2,N-2)$, with various restrictions on the size of the nodes.

preprint2012arXiv

Robustness, Canalyzing Functions and Systems Design

We study a notion of robustness of a Markov kernel that describes a system of several input random variables and one output random variable. Robustness requires that the behaviour of the system does not change if one or several of the input variables are knocked out. If the system is required to be robust against too many knockouts, then the output variable cannot distinguish reliably between input states and must be independent of the input. We study how many input states the output variable can distinguish as a function of the required level of robustness. Gibbs potentials allow a mechanistic description of the behaviour of the system after knockouts. Robustness imposes structural constraints on these potentials. We show that interaction families of Gibbs potentials allow to describe robust systems. Given a distribution of the input random variables and the Markov kernel describing the system, we obtain a joint probability distribution. Robustness implies a number of conditional independence statements for this joint distribution. The set of all probability distributions corresponding to robust systems can be decomposed into a finite union of components, and we find parametrizations of the components. The decomposition corresponds to a primary decomposition of the conditional independence ideal and can be derived from more general results about generalized binomial edge ideals.

preprint2012arXiv

Shared Information -- New Insights and Problems in Decomposing Information in Complex Systems

How can the information that a set ${X_{1},...,X_{n}}$ of random variables contains about another random variable $S$ be decomposed? To what extent do different subgroups provide the same, i.e. shared or redundant, information, carry unique information or interact for the emergence of synergistic information? Recently Williams and Beer proposed such a decomposition based on natural properties for shared information. While these properties fix the structure of the decomposition, they do not uniquely specify the values of the different terms. Therefore, we investigate additional properties such as strong symmetry and left monotonicity. We find that strong symmetry is incompatible with the properties proposed by Williams and Beer. Although left monotonicity is a very natural property for an information measure it is not fulfilled by any of the proposed measures. We also study a geometric framework for information decompositions and ask whether it is possible to represent shared information by a family of posterior distributions. Finally, we draw connections to the notions of shared knowledge and common knowledge in game theory. While many people believe that independent variables cannot share information, we show that in game theory independent agents can have shared knowledge, but not common knowledge. We conclude that intuition and heuristic arguments do not suffice when arguing about information.

preprint2011arXiv

Optimally approximating exponential families

This article studies exponential families $\mathcal{E}$ on finite sets such that the information divergence $D(P\|\mathcal{E})$ of an arbitrary probability distribution from $\mathcal{E}$ is bounded by some constant $D>0$. A particular class of low-dimensional exponential families that have low values of $D$ can be obtained from partitions of the state space. The main results concern optimality properties of these partition exponential families. Exponential families where $D=\log(2)$ are studied in detail. This case is special, because if $D<\log(2)$, then $\mathcal{E}$ contains all probability measures with full support.

preprint2011arXiv

Robustness and Conditional Independence Ideals

We study notions of robustness of Markov kernels and probability distribution of a system that is described by $n$ input random variables and one output random variable. Markov kernels can be expanded in a series of potentials that allow to describe the system's behaviour after knockouts. Robustness imposes structural constraints on these potentials. Robustness of probability distributions is defined via conditional independence statements. These statements can be studied algebraically. The corresponding conditional independence ideals are related to binary edge ideals. The set of robust probability distributions lies on an algebraic variety. We compute a Gröbner basis of this ideal and study the irreducible decomposition of the variety. These algebraic results allow to parametrize the set of all robust probability distributions.

preprint2011arXiv

Support Sets in Exponential Families and Oriented Matroid Theory

The closure of a discrete exponential family is described by a finite set of equations corresponding to the circuits of an underlying oriented matroid. These equations are similar to the equations used in algebraic statistics, although they need not be polynomial in the general case. This description allows for a combinatorial study of the possible support sets in the closure of an exponential family. If two exponential families induce the same oriented matroid, then their closures have the same support sets. Furthermore, the positive cocircuits give a parameterization of the closure of the exponential family.

preprint2010arXiv

Knockouts, Robustness and Cell Cycles

The response to a knockout of a node is a characteristic feature of a networked dynamical system. Knockout resilience in the dynamics of the remaining nodes is a sign of robustness. Here we study the effect of knockouts for binary state sequences and their implementations in terms of Boolean threshold networks. Beside random sequences with biologically plausible constraints, we analyze the cell cycle sequence of the species Saccharomyces cerevisiae and the Boolean networks implementing it. Comparing with an appropriate null model we do not find evidence that the yeast wildtype network is optimized for high knockout resilience. Our notion of knockout resilience weakly correlates with the size of the basin of attraction, which has also been considered a measure of robustness.

preprint2009arXiv

Finding the Maximizers of the Information Divergence from an Exponential Family

This paper investigates maximizers of the information divergence from an exponential family $E$. It is shown that the $rI$-projection of a maximizer $P$ to $E$ is a convex combination of $P$ and a probability measure $P_-$ with disjoint support and the same value of the sufficient statistics $A$. This observation can be used to transform the original problem of maximizing $D(\cdot||E)$ over the set of all probability measures into the maximization of a function $\Dbar$ over a convex subset of $\ker A$. The global maximizers of both problems correspond to each other. Furthermore, finding all local maximizers of $\Dbar$ yields all local maximizers of $D(\cdot||E)$. This paper also proposes two algorithms to find the maximizers of $\Dbar$ and applies them to two examples, where the maximizers of $D(\cdot||E)$ were not known before.