Source author record

David Schwab

David Schwab appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Artificial Intelligence Machine Learning Cell Behavior cond-mat.stat-mech Genomics math-ph math.MP Neurons and Cognition nlin.PS Populations and Evolution

Catalog footprint

What is connected

6works

10topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

Data curation has shifted the quality-compute frontier for language-model and contrastive image-text pretraining, but its role for vision-language models (VLMs) is far less established. We ask how far data curation alone can take VLM performance, holding architecture, training recipe, and compute fixed and varying only the training data. Our pipeline, applied to the MAmmoTH-VL single-image subset, lifts performance by +11.7pp on average across 20 public VLM benchmarks (spanning grounding, VQA, OCR/documents, captioning, spatial/3D, counting, charts, math, brand-ID, and multi-image reasoning) and by +11.3pp on average across all nine capability axes of DatBench, our high-fidelity VLM eval suite. At 2B, our curated model surpasses InternVL3.5-2B by 9.9pp at ~17x less training compute and closes the gap to Qwen3-VL-2B to within 1.8pp at ~87x less compute, from pretraining alone. Beyond accuracy, curation delivers four further properties: (1) Reliability: per-capability std across training seeds drops by ~67% and the lift survives a 4k-to-16k context-length sweep; (2) OOD generalization: the 9-eval OOD average rises by +7.2pp, and multi-image BLINK rises by +3.09pp despite single-image-only training, with Visual Correspondence gaining +11.8pp; (3) Behavioral gains beyond benchmarks: across ~1,100 open-ended queries the curated 2B is more honest and more specific than the matched-compute baseline, and more concise and less refusal-prone than a frontier 2B reference; (4) Pareto-dominance on inference cost: at every scale (1B, 2B, 4B) the curated model raises accuracy while lowering response FLOPs vs. the matched-compute baseline, and the curated 4B matches near-frontier accuracy at 3.3x lower response FLOPs than Qwen3-VL-4B. Data curation is a high-leverage tool for building better VLMs, reaching near-frontier accuracy at up to ~150x less training compute.

preprint2026arXiv

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

Empirical evaluation serves as the primary compass guiding research progress in foundation models. Despite a large body of work focused on training frontier vision-language models (VLMs), approaches to their evaluation remain nascent. To guide their maturation, we propose three desiderata that evaluations should satisfy: (1) faithfulness to the modality and application, (2) discriminability between models of varying quality, and (3) efficiency in compute. Through this lens, we identify critical failure modes that violate faithfulness and discriminability, misrepresenting model capabilities: (i) multiple-choice formats reward guessing, poorly reflect downstream use cases, and saturate early as models improve; (ii) blindly solvable questions, which can be answered without images, constitute up to 70% of some evaluations; and (iii) mislabeled or ambiguous samples compromise up to 42% of examples in certain datasets. Regarding efficiency, the computational burden of evaluating frontier models has become prohibitive: by some accounts, nearly 20% of development compute is devoted to evaluation alone. Rather than discarding existing benchmarks, we curate them via transformation and filtering to maximize fidelity and discriminability. We find that converting multiple-choice questions to generative tasks reveals sharp capability drops of up to 35%. In addition, filtering blindly solvable and mislabeled samples improves discriminative power while simultaneously reducing computational cost. We release DatBench-Full, a cleaned evaluation suite of 33 datasets spanning nine VLM capabilities, and DatBench, a discriminative subset that achieves 13x average speedup (up to 50x) while closely matching the discriminative power of the original datasets. Our work outlines a path toward evaluation practices that are both rigorous and sustainable as VLMs continue to scale.

preprint2026arXiv

How Much is Brain Data Worth for Machine Learning?

If a person can solve a task, can measuring their brain make it easier to train a model to solve that task too? Recent NeuroAI work suggests that supplementing task training with neural recordings can modestly improve model performance and robustness. However, it is unclear when there should be a benefit from using neural data and how much benefit to expect. We formulate this question mathematically, and begin to address it theoretically using a simple, analytically tractable linear gaussian model of task targets and neural recordings. For a multimodal estimator trained on both brain data and task labels, we derive scaling laws for how performance scales with the numbers of brain and task samples. From these laws we derive relative value and exchange rates between brain samples and task samples, quantifying how much extra task samples neural data is worth as a function of task-brain alignment, neural and task noise, latent dimension, and brain data sample size. We also analyze test distribution shift, to identify conditions where brain-regularized learning can produce substantial robustness gains through learned invariances. Finally, under a fixed collection budget, we characterize the regimes in which brain data is worth collecting. Our results provide a foundation for understanding how valuable brain data could be for improving machine learning.

preprint2021arXiv

Understanding Species Abundance Distributions in Complex Ecosystems of Interacting Species

Niche and neutral theory are two prevailing, yet much debated, ideas in ecology proposed to explain the patterns of biodiversity. Whereas niche theory emphasizes selective differences between species and interspecific interactions in shaping the community, neutral theory supposes functional equivalence between species and points to stochasticity as the primary driver of ecological dynamics. In this work, we draw a bridge between these two opposing theories. Starting from a Lotka-Volterra (LV) model with demographic noise and random symmetric interactions, we analytically derive the stationary population statistics and species abundance distribution (SAD). Using these results, we demonstrate that the model can exhibit three classes of SADs commonly found in niche and neutral theories and found conditions that allow an ecosystem to transition between these various regimes. Thus, we reconcile how neutral-like statistics may arise from a diverse community with niche differentiation.

preprint2014arXiv

Multiscale modeling of oscillations and spiral waves in Dictyostelium populations

Unicellular organisms exhibit elaborate collective behaviors in response to environmental cues. These behaviors are controlled by complex biochemical networks within individual cells and coordinated through cell-to-cell communication. Describing these behaviors requires new mathematical models that can bridge scales -- from biochemical networks within individual cells to spatially structured cellular populations. Here, we present a family of multiscale models for the emergence of spiral waves in the social amoeba Dictyostelium discoideum. Our models exploit new experimental advances that allow for the direct measurement and manipulation of the small signaling molecule cAMP used by Dictyostelium cells to coordinate behavior in cellular populations. Inspired by recent experiments, we model the Dictyostelium signaling network as an excitable system coupled to various pre-processing modules. We use this family of models to study spatially unstructured populations by constructing phase diagrams that relate the properties of population-level oscillations to parameters in the underlying biochemical network. We then extend our models to include spatial structure and show how they naturally give rise to spiral waves. Our models exhibit a wide range of novel phenomena including a density dependent frequency change, bistability, and dynamic death due to slow cAMP dynamics. Our modeling approach provides a powerful tool for bridging scales in modeling of Dictyostelium populations.

preprint2010arXiv

Statistical mechanics of transcription-factor binding site discovery using Hidden Markov Models

Hidden Markov Models (HMMs) are a commonly used tool for inference of transcription factor (TF) binding sites from DNA sequence data. We exploit the mathematical equivalence between HMMs for TF binding and the "inverse" statistical mechanics of hard rods in a one-dimensional disordered potential to investigate learning in HMMs. We derive analytic expressions for the Fisher information, a commonly employed measure of confidence in learned parameters, in the biologically relevant limit where the density of binding sites is low. We then use techniques from statistical mechanics to derive a scaling principle relating the specificity (binding energy) of a TF to the minimum amount of training data necessary to learn it.

David Schwab

What is connected

Connect this record

See the researcher in context

Building this map preview

6 published item(s)

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone

DatBench: Discriminative, Faithful, and Efficient VLM Evaluations

How Much is Brain Data Worth for Machine Learning?

Understanding Species Abundance Distributions in Complex Ecosystems of Interacting Species

Multiscale modeling of oscillations and spiral waves in Dictyostelium populations

Statistical mechanics of transcription-factor binding site discovery using Hidden Markov Models