Source author record

James Withers

James Withers appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Computation and Language Cryptography and Security Machine Learning math.PR Populations and Evolution

Catalog footprint

What is connected

3works

5topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2021arXiv

N-grams Bayesian Differential Privacy

Differential privacy has gained popularity in machine learning as a strong privacy guarantee, in contrast to privacy mitigation techniques such as k-anonymity. However, applying differential privacy to n-gram counts significantly degrades the utility of derived language models due to their large vocabularies. We propose a differential privacy mechanism that uses public data as a prior in a Bayesian setup to provide tighter bounds on the privacy loss metric epsilon, and thus better privacy-utility trade-offs. It first transforms the counts to log space, approximating the distribution of the public and private data as Gaussian. The posterior distribution is then evaluated and softmax is applied to produce a probability distribution. This technique achieves up to 85% reduction in KL divergence compared to previously known mechanisms at epsilon equals 0.1. We compare our mechanism to k-anonymity in a n-gram language modelling task and show that it offers competitive performance at large vocabulary sizes, while also providing superior privacy protection.

preprint2021arXiv

Training Data Leakage Analysis in Language Models

Recent advances in neural network based language models lead to successful deployments of such models, improving user experience in various applications. It has been demonstrated that strong performance of language models comes along with the ability to memorize rare training samples, which poses serious privacy threats in case the model is trained on confidential user content. In this work, we introduce a methodology that investigates identifying the user content in the training data that could be leaked under a strong and realistic threat model. We propose two metrics to quantify user-level data leakage by measuring a model's ability to produce unique sentence fragments within training data. Our metrics further enable comparing different models trained on the same data in terms of privacy. We demonstrate our approach through extensive numerical studies on both RNN and Transformer based models. We further illustrate how the proposed metrics can be utilized to investigate the efficacy of mitigations like differentially private training or API hardening.

preprint2019arXiv

Transient amplifiers of selection and reducers of fixation for death-Birth updating on graphs

The spatial structure of an evolving population affects which mutations become fixed. Some structures amplify selection, increasing the likelihood that beneficial mutations become fixed while deleterious mutations do not. Other structures suppress selection, reducing the effect of fitness differences and increasing the role of random chance. This phenomenon can be modeled by representing spatial structure as a graph, with individuals occupying vertices. Births and deaths occur stochastically, according to a specified update rule. We study death-Birth updating: An individual is chosen to die and then its neighbors compete to reproduce into the vacant spot. Previous numerical experiments suggested that amplifiers of selection for this process are either rare or nonexistent. We introduce a perturbative method for this problem for weak selection regime, meaning that mutations have small fitness effects. We show that fixation probability under weak selection can be calculated in terms of the coalescence times of random walks. This result leads naturally to a new definition of effective population size. Using this and other methods, we uncover the first known examples of transient amplifiers of selection (graphs that amplify selection for a particular range of fitness values) for the death-Birth process. We also exhibit new families of "reducers of fixation", which decrease the fixation probability of all mutations, whether beneficial or deleterious.