Source author record

Jan-Christian Hütter

Jan-Christian Hütter appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

8works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

8 published item(s)

preprint2026arXiv

AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational model of cellular behavior that could accelerate biological discovery. One of the most compelling promises of this vision is the ability to perform in silico phenotypic screens, in which a model predicts the effects of cellular perturbations in unseen biological contexts. This task combines heterogeneous textual inputs with diverse phenotypic outputs, making it particularly well-suited to LLMs and agentic systems. Yet, no standard benchmark currently exists for this task, as existing efforts focus on narrower molecular readouts that are only indirectly aligned with the phenotypic endpoints driving many real-world drug discovery workflows. In this work, we present AssayBench, a benchmark for phenotypic screen prediction, built from 1,920 publicly available CRISPR screens spanning five broad classes of cellular phenotypes. We formulate the screen prediction task as a gene rank prediction for each screen and introduce the adjusted nDCG, a continuous metric for comparing performance across heterogeneous assays. Our extensive evaluation shows that existing methods remain far from empirically estimated performance ceilings and zero-shot generalist LLMs outperform biology-specific LLMs and trainable baselines. Optimization techniques such as fine-tuning, ensembling, and prompt optimization can further improve LLM performance on this task. Overall, AssayBench offers a practical testbed for measuring progress toward in silico phenotypic screening and, more broadly, virtual cell models.

preprint2023arXiv

NODAGS-Flow: Nonlinear Cyclic Causal Structure Learning

Learning causal relationships between variables is a well-studied problem in statistics, with many important applications in science. However, modeling real-world systems remain challenging, as most existing algorithms assume that the underlying causal graph is acyclic. While this is a convenient framework for developing theoretical developments about causal reasoning and inference, the underlying modeling assumption is likely to be violated in real systems, because feedback loops are common (e.g., in biological systems). Although a few methods search for cyclic causal models, they usually rely on some form of linearity, which is also limiting, or lack a clear underlying probabilistic model. In this work, we propose a novel framework for learning nonlinear cyclic causal graphical models from interventional data, called NODAGS-Flow. We perform inference via direct likelihood optimization, employing techniques from residual normalizing flows for likelihood estimation. Through synthetic experiments and an application to single-cell high-content perturbation screening data, we show significant performance improvements with our approach compared to state-of-the-art methods with respect to structure recovery and predictive performance.

preprint2020arXiv

Minimax estimation of smooth optimal transport maps

Brenier's theorem is a cornerstone of optimal transport that guarantees the existence of an optimal transport map $T$ between two probability distributions $P$ and $Q$ over $\mathbb{R}^d$ under certain regularity conditions. The main goal of this work is to establish the minimax estimation rates for such a transport map from data sampled from $P$ and $Q$ under additional smoothness assumptions on $T$. To achieve this goal, we develop an estimator based on the minimization of an empirical version of the semi-dual optimal transport problem, restricted to truncated wavelet expansions. This estimator is shown to achieve near minimax optimality using new stability arguments for the semi-dual and a complementary minimax lower bound. Furthermore, we provide numerical experiments on synthetic data supporting our theoretical findings and highlighting the practical benefits of smoothness regularization. These are the first minimax estimation rates for transport maps in general dimension.

preprint2020arXiv

Optimal Rates for Estimation of Two-Dimensional Totally Positive Distributions

We study minimax estimation of two-dimensional totally positive distributions. Such distributions pertain to pairs of strongly positively dependent random variables and appear frequently in statistics and probability. In particular, for distributions with $β$-Hölder smooth densities where $β\in (0, 2)$, we observe polynomially faster minimax rates of estimation when, additionally, the total positivity condition is imposed. Moreover, we demonstrate fast algorithms to compute the proposed estimators and corroborate the theoretical rates of estimation by simulation studies.

preprint2016arXiv

Optimal rates for total variation denoising

Motivated by its practical success, we show that the two-dimensional total variation denoiser satisfies a sharp oracle inequality that leads to near optimal rates of estimation for a large class of image models such as bi-isotonic, Hölder smooth and cartoons. Our analysis hinges on properties of the unnormalized Laplacian of the two-dimensional grid such as eigenvector delocalization and spectral decay. We also present extensions to more than two dimensions as well as several other graphs.

preprint2014arXiv

Asymptotic Behavior of Gradient Flows Driven by Nonlocal Power Repulsion and Attraction Potentials in One Dimension

We study the long time behavior of the Wasserstein gradient flow for an energy functional consisting of two components: particles are attracted to a fixed profile $ω$ by means of an interaction kernel $ψ_a(z)=|z|^{q_a}$,and they repel each other by means of another kernel $ψ_r(z)=|z|^{q_r}$. We focus on the case of one space dimension and assume that $1\le q_r\le q_a\le 2$. Our main result is that the flow converges to an equilibrium if either $q_r<q_a$ or $1\le q_r=q_a\le4/3$,and if the solution has the same (conserved) mass as the reference state $ω$. In the cases $q_r=1$ and $q_r=2$, we are able to discuss the behavior for different masses as well, and we explicitly identify the equilibrium state, which is independent of the initial condition. Our proofs heavily use the inverse distribution function of the solution.

preprint2013arXiv

Consistency of Probability Measure Quantization by Means of Power Repulsion-Attraction Potentials

This paper is concerned with the study of the consistency of a variational method for probability measure quantization, deterministically realized by means of a minimizing principle, balancing power repulsion and attraction potentials. The proof of consistency is based on the construction of a target energy functional whose unique minimizer is actually the given probability measure ωto be quantized. Then we show that the discrete functionals, defining the discrete quantizers as their minimizers, actually Γ-converge to the target energy with respect to the narrow topology on the space of probability measures. A key ingredient is the reformulation of the target functional by means of a Fourier representation, which extends the characterization of conditionally positive semi-definite functions from points in generic position to probability measures. As a byproduct of the Fourier representation, we also obtain compactness of sublevels of the target energy in terms of uniform moment bounds, which already found applications in the asymptotic analysis of corresponding gradient flows. To model situations where the given probability is affected by noise, we additionally consider a modified energy, with the addition of a regularizing total variation term and we investigate again its point mass approximations in terms of Γ-convergence. We show that such a discrete measure representation of the total variation can be interpreted as an additional nonlinear potential, repulsive at a short range, attractive at a medium range, and at a long range not having effect, promoting a uniform distribution of the point masses.

preprint2013arXiv

Minimizers and Gradient Flows of Attraction-Repulsion Functionals with Power Kernels and Their Total Variation Regularization

We study properties of an attractive-repulsive energy functional based on power-kernels, which can be used for halftoning of images. In the first part of this work, using a variational framework for probability measures, we examine existence and behavior of minimizers to the functional and to a regularization of it by a total variation term. Moreover, we introduce particle approximations to the functional and to its regularized version and prove their consistency in terms of Gamma-convergence, which we additionally illustrate by numerical examples. In the second part, we consider the gradient flow of the functional in the 2-Wasserstein space and prove statements about its asymptotic behavior for large times, for which we employ the pseudo-inverse technique for probability measures in 1D. Depending on the parameter range, this includes existence of a subsequence converging to a steady state or even convergence of the whole trajectory to a limit which we can specify explicitly. For both parts of the work, a key ingredient is the generalized Fourier transform, which allows us to verify the conditional positive definiteness of the interaction kernel for coinciding attractive and repulsive exponents.