Source author record

Thomas Gueudre

Thomas Gueudre appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

3works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

3 published item(s)

preprint2022arXiv

Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems

We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction.

preprint2012arXiv

Directed polymer near a hard wall and KPZ equation in the half-space

We study the directed polymer with fixed endpoints near an absorbing wall, in the continuum and in presence of disorder, equivalent to the KPZ equation on the half space with droplet initial conditions. From a Bethe Ansatz solution of the equivalent attractive boson model we obtain the exact expression for the free energy distribution at all times. It converges at large time to the Tracy Widom distribution $F_4$ of the Gaussian Symplectic Ensemble (GSE). We compare our results with numerical simulations of the lattice directed polymer, both at zero and high temperature.

preprint2012arXiv

Short time growth of a KPZ interface with flat initial conditions

The short time behavior of the 1+1 dimensional KPZ growth equation with a flat initial condition is obtained from the exact expressions of the moments of the partition function of a directed polymer with one endpoint free and the other fixed. From these expressions, the short time expansions of the lowest cumulants of the KPZ height field are exactly derived. The results for these two classes of cumulants are checked in high precision lattice numerical simulations. The short time limit considered here is relevant for the study of the interface growth in the large diffusivity/weak noise limit, and describes the universal crossover between the Edwards-Wilkinson and KPZ universality classes for an initially flat interface.