Source author record

Alessandro Manzotti

Alessandro Manzotti appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2022arXiv

Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems

We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction.

preprint2016arXiv

External priors for the next generation of CMB experiments

Planned cosmic microwave background (CMB) experiments can dramatically improve what we know about neutrino physics, inflation, and dark energy. The low level of noise, together with improved angular resolution, will increase the signal to noise of the CMB polarized signal as well as the reconstructed lensing potential of high redshift large scale structure. Projected constraints on cosmological parameters are extremely tight, but these can be improved even further with information from external experiments. Here, we examine quantitatively the extent to which external priors can lead to improvement in projected constraints from a CMB-Stage IV (S4) experiment on neutrino and dark energy properties. We find that CMB S4 constraints on neutrino mass could be strongly enhanced by external constraints on the cold dark matter density $Ω_{c}h^{2}$ and the Hubble constant $H_{0}$. If polarization on the largest scales ($\ell<50$) will not be measured, an external prior on the primordial amplitude $A_{s}$ or the optical depth $τ$ will also be important. A CMB constraint on the number of relativistic degrees of freedom, $N_{\rm eff}$, will benefit from an external prior on the spectral index $n_{s}$ and the baryon energy density $Ω_{b}h^{2}$. Finally, an external prior on $H_{0}$ will help constrain the dark energy equation of state ($w$).

preprint2014arXiv

A coarse grained perturbation theory for the Large Scale Structure, with cosmology and time independence in the UV

Standard cosmological perturbation theory (SPT) for the Large Scale Structure (LSS) of the Universe fails at small scales (UV) due to strong nonlinearities and to multistreaming effects. In Pietroni et al. 2011 a new framework was proposed in which the large scales (IR) are treated perturbatively while the information on the UV, mainly small scale velocity dispersion, is obtained by nonlinear methods like N-body simulations. Here we develop this approach, showing that it is possible to reproduce the fully nonlinear power spectrum (PS) by combining a simple (and fast) 1-loop computation for the IR scales and the measurement of a single, dominant, correlator from N-body simulations for the UV ones. We measure this correlator for a suite of seven different cosmologies, and we show that its inclusion in our perturbation scheme reproduces the fully non-linear PS with percent level accuracy, for wave numbers up to $k\sim 0.4\, h~{\rm Mpc^{-1}}$ down to $z=0$. We then show that, once this correlator has been measured in a given cosmology, there is no need to run a new simulation for a different cosmology in the suite. Indeed, by rescaling this correlator by a proper function computable in SPT, the reconstruction procedure works also for the other cosmologies and for all redshifts, with comparable accuracy. Finally, we clarify the relation of this approach to the Effective Field Theory methods recently proposed in the LSS context.

preprint2014arXiv

Super-Sample CMB Lensing

Lensing of the CMB by modes that are larger than the size of the survey dilates intrinsic scales in the temperature and polarization fields and coherently shifts their observed power spectra with respect to the ensemble or all-sky mean. The effect can be simply encapsulated as a contribution to the power spectrum covariance matrix in accordance with the lensing trispectrum or as an additional parameter, the mean convergence in the field, for parameter estimation. It should be included for upcoming surveys that precisely measure acoustic polarization features deep into the damping tail at multipoles of $\ell \gtrsim 1500$ with less than $10\%$ of sky. Its omission may lead to seemingly conflicting values for the angular scale of the sound horizon which may then provide erroneous cosmological parameters when compared to baryon acoustic oscillation measurements.

preprint2012arXiv

Prospects for early localization of gravitational-wave signals from compact binary coalescences with advanced detectors

A leading candidate source of detectable gravitational waves is the inspiral and merger of pairs of stellar-mass compact objects. The advanced LIGO and advanced Virgo detectors will allow scientists to detect inspiral signals from more massive systems and at earlier times in the detector band, than with first generation detectors. The signal from a coalescence of two neutron stars is expected to stay in the sensitive band of advanced detectors for several minutes, thus allowing detection before the final coalescence of the system. In this work, the prospects of detecting inspiral signals prior to coalescence, and the possibility to derive a suitable sky area for source locations are investigated. As a large fraction of the signal is accumulated in the last ~10 seconds prior to coalescence, bandwidth and timing accuracy are largely accrued in the very last moments prior to coalescence. We use Monte Carlo techniques to estimate the accuracy of sky localization through networks of ground-based interferometers such as aLIGO and aVirgo. With the addition of the Japanese KAGRA detector, it is shown that the detection and triangulation before coalescence may be feasible.

preprint2010arXiv

Effective target arrangement in a deterministic scale-free graph

We study the random walk problem on a deterministic scale-free network, in the presence of a set of static, identical targets; due to the strong inhomogeneity of the underlying structure the mean first-passage time (MFPT), meant as a measure of transport efficiency, is expected to depend sensitively on the position of targets. We consider several spatial arrangements for targets and we calculate, mainly rigorously, the related MFPT, where the average is taken over all possible starting points and over all possible paths. For all the cases studied, the MFPT asymptotically scales like N^{theta}, being N the volume of the substrate and theta ranging from (1 - log 2/log3), for central target(s), to 1, for a single peripheral target.