Source author record

David Sánchez

David Sánchez appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

19works
13topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

19 published item(s)

preprint2022arXiv

Defending against the Label-flipping Attack in Federated Learning

Federated learning (FL) provides autonomy and privacy by design to participating peers, who cooperatively build a machine learning (ML) model while keeping their private data in their devices. However, that same autonomy opens the door for malicious peers to poison the model by conducting either untargeted or targeted poisoning attacks. The label-flipping (LF) attack is a targeted poisoning attack where the attackers poison their training data by flipping the labels of some examples from one class (i.e., the source class) to another (i.e., the target class). Unfortunately, this attack is easy to perform and hard to detect and it negatively impacts on the performance of the global model. Existing defenses against LF are limited by assumptions on the distribution of the peers' data and/or do not perform well with high-dimensional models. In this paper, we deeply investigate the LF attack behavior and find that the contradicting objectives of attackers and honest peers on the source class examples are reflected in the parameter gradients corresponding to the neurons of the source and target classes in the output layer, making those gradients good discriminative features for the attack detection. Accordingly, we propose a novel defense that first dynamically extracts those gradients from the peers' local updates, and then clusters the extracted gradients, analyzes the resulting clusters and filters out potential bad updates before model aggregation. Extensive empirical analysis on three data sets shows the proposed defense's effectiveness against the LF attack regardless of the data distribution or model dimensionality. Also, the proposed defense outperforms several state-of-the-art defenses by offering lower test error, higher overall accuracy, higher source class accuracy, lower attack success rate, and higher stability of the source class accuracy.

preprint2022arXiv

The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization

We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods. Text anonymization, defined as the task of editing a text document to prevent the disclosure of personal information, currently suffers from a shortage of privacy-oriented annotated text resources, making it difficult to properly evaluate the level of privacy protection offered by various anonymization methods. This paper presents TAB (Text Anonymization Benchmark), a new, open-source annotated corpus developed to address this shortage. The corpus comprises 1,268 English-language court cases from the European Court of Human Rights (ECHR) enriched with comprehensive annotations about the personal information appearing in each document, including their semantic category, identifier type, confidential attributes, and co-reference relations. Compared to previous work, the TAB corpus is designed to go beyond traditional de-identification (which is limited to the detection of predefined semantic categories), and explicitly marks which text spans ought to be masked in order to conceal the identity of the person to be protected. Along with presenting the corpus and its annotation layers, we also propose a set of evaluation metrics that are specifically tailored towards measuring the performance of text anonymization, both in terms of privacy protection and utility preservation. We illustrate the use of the benchmark and the proposed metrics by assessing the empirical performance of several baseline text anonymization models. The full corpus along with its privacy-oriented annotation guidelines, evaluation scripts and baseline models are available on: https://github.com/NorskRegnesentral/text-anonymisation-benchmark

preprint2022arXiv

Trivial and topological bound states in bilayer graphene quantum dots and rings

We discuss and compare two different types of confinement in bilayer graphene by top and bottom gating with symmetrical microelectrodes. Trivial confinement corresponds to the same polarity of all top gates, which is opposed to that of all bottom ones. Topological confinement requires the polarity of part of the top-bottom pairs of gates to be reversed. We show that the main qualitative difference between trivial and topological bound states manifests itself in the magnetic field dependence. We illustrate our finding with an explicit calculation of the energy spectrum for quantum dots and rings. Trivial confinement shows bunching of levels into degenerate Landau bands, with a non-centered gap, while topological confinement shows no field-induced gap and a sequence of state branches always crossing zero-energy.

preprint2021arXiv

Geometry effects in topologically confined bilayer graphene loops

We investigate the electronic confinement in bilayer graphene by topological loops of different shapes. These loops are created by lateral gates acting via gap inversion on the two graphene sheets. For large-area loops the spectrum is well described by a quantization rule depending only on the loop perimeter. For small sizes, the spectrum depends on the loop shape. We find that zero-energy states exhibit a characteristic pattern that strongly depends on the spatial symmetry. We show this by considering loops of higher to lower symmetry (circle, square, rectangle and irregular polygon). Interestingly, magnetic field causes valley splittings of the states, an asymmetry between energy reversal states, flux periodicities and the emergence of persistent currents.

preprint2021arXiv

Scattering of topological kink-antikink states in bilayer graphene structures

Gapped bilayer graphene can support the presence of intragap states due to kink gate potentials applied to the graphene layers. Electrons in these states display valley-momentum locking, which makes them attractive for topological valleytronics. Here, we show that kink-antikink local potentials enable modulated scattering of topological currents. We find that the kink-antikink coupling leads to anomalous steps in the junction conductance. Further, when the constriction detaches from the propagating modes, forming a loop, the conductance reveals the system energy spectrum. Remarkably, these kink-antikink devices can also work as valley filters with tiny magnetic fields by tuning a central gate.

preprint2017arXiv

Toward sensitive document release with privacy guarantees

Privacy has become a serious concern for modern Information Societies. The sensitive nature of much of the data that are daily exchanged or released to untrusted parties requires that responsible organizations undertake appropriate privacy protection measures. Nowadays, much of these data are texts (e.g., emails, messages posted in social media, healthcare outcomes, etc.) that, because of their unstructured and semantic nature, constitute a challenge for automatic data protection methods. In fact, textual documents are usually protected manually, in a process known as document redaction or sanitization. To do so, human experts identify sensitive terms (i.e., terms that may reveal identities and/or confidential information) and protect them accordingly (e.g., via removal or, preferably, generalization). To relieve experts from this burdensome task, in a previous work we introduced the theoretical basis of C-sanitization, an inherently semantic privacy model that provides the basis to the development of automatic document redaction/sanitization algorithms and offers clear and a priori privacy guarantees on data protection; even though its potential benefits C-sanitization still presents some limitations when applied to practice (mainly regarding flexibility, efficiency and accuracy). In this paper, we propose a new more flexible model, named (C, g(C))-sanitization, which enables an intuitive configuration of the trade-off between the desired level of protection (i.e., controlled information disclosure) and the preservation of the utility of the protected data (i.e., amount of semantics to be preserved). Moreover, we also present a set of technical solutions and algorithms that provide an efficient and scalable implementation of the model and improve its practical accuracy, as we also illustrate through empirical experiments.

preprint2016arXiv

Cotunneling drag effect in Coulomb-coupled quantum dots

In Coulomb drag, a current flowing in one conductor can induce a voltage across an adjacent conductor via the Coulomb interaction. The mechanisms yielding drag effects are not always understood, even though drag effects are sufficiently general to be seen in many low-dimensional systems. In this Letter, we observe Coulomb drag in a Coulomb-coupled double quantum dot (CC-DQD) and, through both experimental and theoretical arguments, identify cotunneling as essential to obtaining a correct qualitative understanding of the drag behavior.

preprint2016arXiv

Coulomb-blockade effect in nonlinear mesoscopic capacitors

We consider an interacting quantum dot working as a coherent source of single electrons. The dot is tunnel coupled to a reservoir and capacitively coupled to a gate terminal with an applied ac potential. At low frequencies, this is the quantum analog of the RC circuit with a purely dynamical response. We investigate the quantized dynamics as a consequence of ac pulses with large amplitude. Within a Keldysh-Green function formalism we derive the time-dependent current in the Coulomb blockade regime. Our theory thus extends previous models that considered either noninteracting electrons in nonlinear response or interacting electrons in the linear regime. We prove that the electron emission and absorption resonances undergo a splitting when the charging energy is larger than the tunnel broadening. For very large charging energies, the additional peaks collapse and the original resonances are recovered, though with a reduced amplitude. Quantization of the charge emitted by the capacitor is reduced due to Coulomb repulsion and additional plateaus arise. Additionally, we discuss the differential capacitance and resistance as a function of time. We find that to leading order in driving frequency the current can be expressed as a weighted sum of noninteracting currents shifted by the charging energy.

preprint2016arXiv

Enforcing transparent access to private content in social networks by means of automatic sanitization

Social networks have become an essential meeting point for millions of individuals willing to publish and consume huge quantities of heterogeneous information. Some studies have shown that the data published in these platforms may contain sensitive personal information and that external entities can gather and exploit this knowledge for their own benefit. Even though some methods to preserve the privacy of social networks users have been proposed, they generally apply rigid access control measures to the protected content and, even worse, they do not enable the users to understand which contents are sensitive. Last but not least, most of them require the collaboration of social network operators or they fail to provide a practical solution capable of working with well-known and already deployed social platforms. In this paper, we propose a new scheme that addresses all these issues. The new system is envisaged as an independent piece of software that does not depend on the social network in use and that can be transparently applied to most existing ones. According to a set of privacy requirements intuitively defined by the users of a social network, the proposed scheme is able to: (i) automatically detect sensitive data in users' publications; (ii) construct sanitized versions of such data; and (iii) provide privacy-preserving transparent access to sensitive contents by disclosing more or less information to readers according to their credentials toward the owner of the publications. We also study the applicability of the proposed system in general and illustrate its behavior in two case studies.

preprint2016arXiv

Privacy-driven Access Control in Social Networks by Means of Automatic Semantic Annotation

In online social networks (OSN), users quite usually disclose sensitive information about themselves by publishing messages. At the same time, they are (in many cases) unable to properly manage the access to this sensitive information due to the following issues: i) the rigidness of the access control mechanism implemented by the OSN, and ii) many users lack of technical knowledge about data privacy and access control. To tackle these limitations, in this paper, we propose a dynamic, transparent and privacy-driven access control mechanism for textual messages published in OSNs. The notion of privacy-driven is achieved by analyzing the semantics of the messages to be published and, according to that, assessing the degree of sensitiveness of their contents. For this purpose, the proposed system relies on an automatic semantic annotation mechanism that, by using knowledge bases and linguistic tools, is able to associate a meaning to the information to be published. By means of this annotation, our mechanism automatically detects the information that is sensitive according to the privacy requirements of the publisher of data, with regard to the type of reader that may access such data. Finally, our access control mechanism automatically creates sanitized versions of the users' publications according to the type of reader that accesses them. As a result, our proposal, which can be integrated in already existing social networks, provides an automatic, seamless and content-driven protection of user publications, which are coherent with her privacy requirements and the type of readers that access them. Complementary to the system design, we also discuss the feasibility of the system by illustrating it through a real example and evaluate its accuracy and effectiveness over standard approaches.

preprint2016arXiv

Supplementary Materials for "How to Avoid Reidentification with Proper Anonymization"- Comment on "Unique in the shopping mall: on the reidentifiability of credit card metadata"

The study by De Montjoye et al. ("Science", 30 January 2015, p. 536) claimed that most individuals can be reidentified from a deidentified credit card transaction database and that anonymization mechanisms are not effective against reidentification. Such claims deserve detailed quantitative scrutiny, as they might seriously undermine the willingness of data owners and subjects to share data for research. In a recent Technical Comment published in "Science" (18 March 2016, p. 1274), we demonstrate that the reidentification risk reported by De Montjoye et al. was significantly overestimated (due to a misunderstanding of the reidentification attack) and that the alleged ineffectiveness of anonymization is due to the choice of poor and undocumented methods and to a general disregard of 40 years of anonymization literature. The technical comment also shows how to properly anonymize data, in order to reduce unequivocal reidentifications to zero while retaining even more analytical utility than with the poor anonymization mechanisms employed by De Montjoye et al. In conclusion, data owners, subjects and users can be reassured that sound privacy models and anonymization methods exist to produce safe and useful anonymized data. Supplementary materials detailing the data sets, algorithms and extended results of our study are available here. Moreover, unlike the De Montjoye et al.'s data set, which was never made available, our data, anonymized results, and anonymization algorithms can be freely downloaded from http://crises-deim.urv.cat/opendata/SPD_Science.zip

preprint2016arXiv

The Intrinsic Shape of Sagittarius A* at 3.5-mm Wavelength

The radio emission from Sgr A$^\ast$ is thought to be powered by accretion onto a supermassive black hole of $\sim\! 4\times10^6~ \rm{M}_\odot$ at the Galactic Center. At millimeter wavelengths, Very Long Baseline Interferometry (VLBI) observations can directly resolve the bright innermost accretion region of Sgr A$^\ast$. Motivated by the addition of many sensitive, long baselines in the north-south direction, we developed a full VLBI capability at the Large Millimeter Telescope Alfonso Serrano (LMT). We successfully detected Sgr A$^\ast$ at 3.5~mm with an array consisting of 6 Very Long Baseline Array telescopes and the LMT. We model the source as an elliptical Gaussian brightness distribution and estimate the scattered size and orientation of the source from closure amplitude and self-calibration analysis, obtaining consistent results between methods and epochs. We then use the known scattering kernel to determine the intrinsic two dimensional source size at 3.5 mm: $(147\pm7~μ\rm{as}) \times (120\pm12~μ\rm{as})$, at position angle $88^\circ\pm7^\circ$ east of north. Finally, we detect non-zero closure phases on some baseline triangles, but we show that these are consistent with being introduced by refractive scattering in the interstellar medium and do not require intrinsic source asymmetry to explain.

preprint2015arXiv

t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation

Microaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate $k$-anonymous data sets, where the identity of each subject is hidden within a group of $k$ subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers and avoiding discretization of numerical data. $k$-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of $k$ subjects is too small. To address this issue, several refinements of $k$-anonymity have been proposed, among which $t$-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate $t$-close data sets are based on generalization and suppression (they are extensions of $k$-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate $k$-anonymous $t$-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for $k$-anonymous $t$-closeness are presented and empirically evaluated.

preprint2015arXiv

Utility-Preserving Differentially Private Data Releases Via Individual Ranking Microaggregation

Being able to release and exploit open data gathered in information systems is crucial for researchers, enterprises and the overall society. Yet, these data must be anonymized before release to protect the privacy of the subjects to whom the records relate. Differential privacy is a privacy model for anonymization that offers more robust privacy guarantees than previous models, such as $k$-anonymity and its extensions. However, it is often disregarded that the utility of differentially private outputs is quite limited, either because of the amount of noise that needs to be added to obtain them or because utility is only preserved for a restricted type and/or a limited number of queries. On the contrary, $k$-anonymity-like data releases make no assumptions on the uses of the protected data and, thus, do not restrict the number and type of doable analyses. Recently, some authors have proposed mechanisms to offer general-purpose differentially private data releases. This paper extends such works with a specific focus on the preservation of the utility of the protected data. Our proposal builds on microaggregation-based anonymization, which is more flexible and utility-preserving than alternative anonymization methods used in the literature, in order to reduce the amount of noise needed to satisfy differential privacy. In this way, we improve the utility of differentially private data releases. Moreover, the noise reduction we achieve does not depend on the size of the data set, but just on the number of attributes to be protected, which is a more desirable behavior for large data sets. The utility benefits brought by our proposal are empirically evaluated and compared with related works for several data sets and metrics.

preprint2014arXiv

Crowdsourcing Dialect Characterization through Twitter

We perform a large-scale analysis of language diatopic variation using geotagged microblogging datasets. By collecting all Twitter messages written in Spanish over more than two years, we build a corpus from which a carefully selected list of concepts allows us to characterize Spanish varieties on a global scale. A cluster analysis proves the existence of well defined macroregions sharing common lexical properties. Remarkably enough, we find that Spanish language is split into two superdialects, namely, an urban speech used across major American and Spanish citites and a diverse form that encompasses rural areas and small towns. The latter can be further clustered into smaller varieties with a stronger regional character.

preprint2013arXiv

Magnetic-field asymmetry of nonlinear thermoelectric and heat transport

Nonlinear transport coefficients do not obey, in general, reciprocity relations. We here discuss the magnetic-field asymmetries that arise in thermoelectric and heat transport of mesoscopic systems. Based on a scattering theory of weakly nonlinear transport, we analyze the leading-order symmetry parameters in terms of the screening potential response to either voltage or temperature shifts. We apply our general results to a quantum Hall antidot system. Interestingly, we find that certain symmetry parameters show a dependence on the measurement configuration.

preprint2013arXiv

Nonlinear thermovoltage and thermocurrent in quantum dots

Quantum dots are model systems for quantum thermoelectric behavior because of the ability to control and measure the effects of electron-energy filtering and quantum confinement on thermoelectric properties. Interestingly, nonlinear thermoelectric properties of such small systems can modify the efficiency of thermoelectric power conversion. Using quantum dots embedded in semiconductor nanowires, we measure thermovoltage and thermocurrent that are strongly nonlinear in the applied thermal bias. We show that most of the observed nonlinear effects can be understood in terms of a renormalization of the quantum-dot energy levels as a function of applied thermal bias and provide a theoretical model of the nonlinear thermovoltage taking renormalization into account. Furthermore, we propose a theory that explains a possible source of the observed, pronounced renormalization effect by the melting of Kondo correlations in the mixed-valence regime. The ability to control nonlinear thermoelectric behavior expands the range in which quantum thermoelectric effects may be used for efficient energy conversion.

preprint2012arXiv

Thermally driven ballistic rectifer

The response of electric devices to an applied thermal gradient has, so far, been studied almost exclusively in two-terminal devices. Here we present measurements of the response to a thermal bias of a four-terminal, quasi-ballistic junction with a central scattering site. We find a novel transverse thermovoltage measured across isothermal contacts. Using a multi-terminal scattering model extended to the weakly non-linear voltage regime, we show that the device's response to a thermal bias can be predicted from its nonlinear response to an electric bias. Our approach forms a foundation for the discovery and understanding of advanced, nonlocal, thermoelectric phenomena that in the future may lead to novel thermoelectric device concepts.

preprint2010arXiv

Mesoscopic Coulomb drag, broken detailed balance and fluctuation relations

When a biased conductor is put in proximity with an unbiased conductor a drag current can be induced in the absence of detailed balance. This is known as the Coulomb drag effect. However, even in this situation far away from equilibrium where detailed balance is explicitly broken, theory predicts that fluctuation relations are satisfied. This surprising effect has, to date, not been confirmed experimentally. Here we propose a system consisting of a capacitively coupled double quantum dot where the nonlinear fluctuation relations are verified in the absence of detailed balance.