Catalog footprint

What is connected

25works
18topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

25 published item(s)

preprint2022arXiv

A BIC-based Mixture Model Defense against Data Poisoning Attacks on Classifiers

Data Poisoning (DP) is an effective attack that causes trained classifiers to misclassify their inputs. DP attacks significantly degrade a classifier's accuracy by covertly injecting attack samples into the training set. Broadly applicable to different classifier structures, without strong assumptions about the attacker, an {\it unsupervised} Bayesian Information Criterion (BIC)-based mixture model defense against "error generic" DP attacks is herein proposed that: 1) addresses the most challenging {\it embedded} DP scenario wherein, if DP is present, the poisoned samples are an {\it a priori} unknown subset of the training set, and with no clean validation set available; 2) applies a mixture model both to well-fit potentially multi-modal class distributions and to capture poisoned samples within a small subset of the mixture components; 3) jointly identifies poisoned components and samples by minimizing the BIC cost defined over the whole training set, with the identified poisoned data removed prior to classifier training. Our experimental results, for various classifier structures and benchmark datasets, demonstrate the effectiveness and universality of our defense under strong DP attacks, as well as its superiority over other works.

preprint2022arXiv

Anomaly Detection of Adversarial Examples using Class-conditional Generative Adversarial Networks

Deep Neural Networks (DNNs) have been shown vulnerable to Test-Time Evasion attacks (TTEs, or adversarial examples), which, by making small changes to the input, alter the DNN's decision. We propose an unsupervised attack detector on DNN classifiers based on class-conditional Generative Adversarial Networks (GANs). We model the distribution of clean data conditioned on the predicted class label by an Auxiliary Classifier GAN (AC-GAN). Given a test sample and its predicted class, three detection statistics are calculated based on the AC-GAN Generator and Discriminator. Experiments on image classification datasets under various TTE attacks show that our method outperforms previous detection methods. We also investigate the effectiveness of anomaly detection using different DNN layers (input features or internal-layer features) and demonstrate, as one might expect, that anomalies are harder to detect using features closer to the DNN's output layer.

preprint2022arXiv

Improved Constraints on Effective Top Quark Interactions using Edge Convolution Networks

We explore the potential of Graph Neural Networks (GNNs) to improve the performance of high-dimensional effective field theory parameter fits to collider data beyond traditional rectangular cut-based differential distribution analyses. In this study, we focus on a SMEFT analysis of $pp \to t\bar t$ production, including top decays, where the linear effective field deformation is parametrised by thirteen independent Wilson coefficients. The application of GNNs allows us to condense the multidimensional phase space information available for the discrimination of BSM effects from the SM expectation by considering all available final state correlations directly. The number of contributing new physics couplings very quickly leads to statistical limitations when the GNN output is directly employed as an EFT discrimination tool. However, a selection based on minimising the SM contribution enhances the fit's sensitivity when reflected as a (non-rectangular) selection on the inclusive data samples that are typically employed when looking for non-resonant deviations from the SM by means of differential distributions.

preprint2022arXiv

Post-Training Detection of Backdoor Attacks for Two-Class and Multi-Attack Scenarios

Backdoor attacks (BAs) are an emerging threat to deep neural network classifiers. A victim classifier will predict to an attacker-desired target class whenever a test sample is embedded with the same backdoor pattern (BP) that was used to poison the classifier's training set. Detecting whether a classifier is backdoor attacked is not easy in practice, especially when the defender is, e.g., a downstream user without access to the classifier's training set. This challenge is addressed here by a reverse-engineering defense (RED), which has been shown to yield state-of-the-art performance in several domains. However, existing REDs are not applicable when there are only {\it two classes} or when {\it multiple attacks} are present. These scenarios are first studied in the current paper, under the practical constraints that the defender neither has access to the classifier's training set nor to supervision from clean reference classifiers trained for the same domain. We propose a detection framework based on BP reverse-engineering and a novel {\it expected transferability} (ET) statistic. We show that our ET statistic is effective {\it using the same detection threshold}, irrespective of the classification domain, the attack configuration, and the BP reverse-engineering algorithm that is used. The excellent performance of our method is demonstrated on six benchmark datasets. Notably, our detection framework is also applicable to multi-class scenarios with multiple attacks. Code is available at https://github.com/zhenxianglance/2ClassBADetection.

preprint2020arXiv

Detection of Backdoors in Trained Classifiers Without Access to the Training Set

Recently, a special type of data poisoning (DP) attack targeting Deep Neural Network (DNN) classifiers, known as a backdoor, was proposed. These attacks do not seek to degrade classification accuracy, but rather to have the classifier learn to classify to a target class whenever the backdoor pattern is present in a test example. Launching backdoor attacks does not require knowledge of the classifier or its training process - it only needs the ability to poison the training set with (a sufficient number of) exemplars containing a sufficiently strong backdoor pattern (labeled with the target class). Here we address post-training detection of backdoor attacks in DNN image classifiers, seldom considered in existing works, wherein the defender does not have access to the poisoned training set, but only to the trained classifier itself, as well as to clean examples from the classification domain. This is an important scenario because a trained classifier may be the basis of e.g. a phone app that will be shared with many users. Detecting backdoors post-training may thus reveal a widespread attack. We propose a purely unsupervised anomaly detection (AD) defense against imperceptible backdoor attacks that: i) detects whether the trained DNN has been backdoor-attacked; ii) infers the source and target classes involved in a detected attack; iii) we even demonstrate it is possible to accurately estimate the backdoor pattern. We test our AD approach, in comparison with alternative defenses, for several backdoor patterns, data sets, and attack settings and demonstrate its favorability. Our defense essentially requires setting a single hyperparameter (the detection threshold), which can e.g. be chosen to fix the system's false positive rate.

preprint2020arXiv

L-RED: Efficient Post-Training Detection of Imperceptible Backdoor Attacks without Access to the Training Set

Backdoor attacks (BAs) are an emerging form of adversarial attack typically against deep neural network image classifiers. The attacker aims to have the classifier learn to classify to a target class when test images from one or more source classes contain a backdoor pattern, while maintaining high accuracy on all clean test images. Reverse-Engineering-based Defenses (REDs) against BAs do not require access to the training set but only to an independent clean dataset. Unfortunately, most existing REDs rely on an unrealistic assumption that all classes except the target class are source classes of the attack. REDs that do not rely on this assumption often require a large set of clean images and heavy computation. In this paper, we propose a Lagrangian-based RED (L-RED) that does not require knowledge of the number of source classes (or whether an attack is present). Our defense requires very few clean images to effectively detect BAs and is computationally efficient. Notably, we detect 56 out of 60 BAs using only two clean images per class in our experiments on CIFAR-10.

preprint2020arXiv

Phenomenology of GUT-inspired gauge-Higgs unification

We perform a detailed investigation of a Grand Unified Theory (GUT)-inspired theory of gauge-Higgs unification. Scanning the model's parameter space with adapted numerical techniques, we contrast the scenario's low energy limit with existing SM and collider search constraints. We discuss potential modifications of di-Higgs phenomenology at hadron colliders as sensitive probes of the gauge-like character of the Higgs self-interactions and find that for phenomenologically viable parameter choices modifications of the order of 20\% compared to the SM cross section can be expected. While these modifications are challenging to observe at the LHC, a future 100 TeV hadron collider might be able to constrain the scenario through more precise di-Higgs measurements. We point out alternative signatures that can be employed to constrain this model in the near future.

preprint2020arXiv

The Weinberg Angle and 5D RGE effects in a SO(11) GUT theory

The Weinberg angle is an important parameter in Grand Unified Theories (GUT) as its size is crucially influenced by the assumption of unification. In scenarios with different steps of symmetry breaking, in particular in models that involve gauge-Higgs unification, the connection of the ultraviolet theory and the TeV scale-relevant, effective Standard Model description is an important test of the models' validity. In this work, we consider a 6D gauge-Higgs unification GUT scenario and explore the TeV scale-GUT relation using a detailed RGE analysis in the 4D and 5D regimes of the theory, including constraints from LHC measurements. We show that such can be consistent with unification in the light of current constraints, while the Weinberg angle likely translates into concrete conditions on the fermion sector in the higher dimensional setup.

preprint2019arXiv

Adversarial Learning in Statistical Classification: A Comprehensive Review of Defenses Against Attacks

There is great potential for damage from adversarial learning (AL) attacks on machine-learning based systems. In this paper, we provide a contemporary survey of AL, focused particularly on defenses against attacks on statistical classifiers. After introducing relevant terminology and the goals and range of possible knowledge of both attackers and defenders, we survey recent work on test-time evasion (TTE), data poisoning (DP), and reverse engineering (RE) attacks and particularly defenses against same. In so doing, we distinguish robust classification from anomaly detection (AD), unsupervised from supervised, and statistical hypothesis-based defenses from ones that do not have an explicit null (no attack) hypothesis; we identify the hyperparameters a particular method requires, its computational complexity, as well as the performance measures on which it was evaluated and the obtained quality. We then dig deeper, providing novel insights that challenge conventional AL wisdom and that target unresolved issues, including: 1) robust classification versus AD as a defense strategy; 2) the belief that attack success increases with attack strength, which ignores susceptibility to AD; 3) small perturbations for test-time evasion attacks: a fallacy or a requirement?; 4) validity of the universal assumption that a TTE attacker knows the ground-truth class for the example to be attacked; 5) black, grey, or white box attacks as the standard for defense evaluation; 6) susceptibility of query-based RE to an AD defense. We also discuss attacks on the privacy of training data. We then present benchmark comparisons of several defenses against TTE, RE, and backdoor DP attacks on images. The paper concludes with a discussion of future work.

preprint2019arXiv

Non-volatile rewritable frequency tuning of a nanoelectromechanical resonator using photoinduced doping

Tuning the frequency of a resonant element is of vital importance in both the macroscopic world, such as when tuning a musical instrument, as well as at the nanoscale. In particular, precisely controlling the resonance frequency of isolated nanoelectromechanical resonators (NEMS) has enabled innovations such as tunable mechanical filtering and mixing as well as commercial technologies such as robust timing oscillators. Much like their electronic device counterparts, the potential of NEMS grows when they are built up into large-scale arrays. Such arrays have enabled neutral-particle mass spectroscopy and have been proposed for ultralow-power alternatives to traditional analog electronics as well as nanomechanical information technologies like memory, logic, and computing. A fundamental challenge to these applications is to precisely tune the vibrational frequency and coupling of all resonators in the array, since traditional tuning methods, like patterned electrostatic gating or dielectric tuning, become intractable when devices are densely packed. Here, we demonstrate a persistent, rewritable, scalable, and high-speed frequency tuning method for graphene-based NEMS. Our method uses a focused laser and two shared electrical contacts to photodope individual resonators by simultaneously applying optical and electrostatic fields. After the fields are removed, the trapped charge created by this process persists and applies a local electrostatic tension to the resonators, tuning their frequencies. By providing a facile means to locally address the strain of a NEMS resonator, this approach lays the groundwork for fully programmable large-scale NEMS lattices and networks.

preprint2016arXiv

A to Z of the Muon Anomalous Magnetic Moment in the MSSM with Pati-Salam at the GUT scale

We analyse the low energy predictions of the minimal supersymmetric standard model (MSSM) arising from a GUT scale Pati-Salam gauge group further constrained by an $A_4 \times Z_5$ family symmetry, resulting in four soft scalar masses at the GUT scale: one left-handed soft mass $m_0$ and three right-handed soft masses $m_1,m_2,m_3$, one for each generation. We demonstrate that this model, which was initially developed to describe the neutrino sector, can explain collider and non-collider measurements such as the dark matter relic density, the Higgs boson mass and, in particular, the anomalous magnetic moment of the muon $(g-2)_μ$. Since about two decades, $(g-2)_μ$ suffers a puzzling about 3$\,σ$ excess of the experimentally measured value over the theoretical prediction, which our model is able to fully resolve. As the consequence of this resolution, our model predicts specific regions of the parameter space with the specific properties including light smuons and neutralinos, which could also potentially explain di-lepton excesses observed by CMS and ATLAS.

preprint2016arXiv

ATD: Anomalous Topic Discovery in High Dimensional Discrete Data

We propose an algorithm for detecting patterns exhibited by anomalous clusters in high dimensional discrete data. Unlike most anomaly detection (AD) methods, which detect individual anomalies, our proposed method detects groups (clusters) of anomalies; i.e. sets of points which collectively exhibit abnormal patterns. In many applications this can lead to better understanding of the nature of the atypical behavior and to identifying the sources of the anomalies. Moreover, we consider the case where the atypical patterns exhibit on only a small (salient) subset of the very high dimensional feature space. Individual AD techniques and techniques that detect anomalies using all the features typically fail to detect such anomalies, but our method can detect such instances collectively, discover the shared anomalous patterns exhibited by them, and identify the subsets of salient features. In this paper, we focus on detecting anomalous topics in a batch of text documents, developing our algorithm based on topic models. Results of our experiments show that our method can accurately detect anomalous topics and salient features (words) under each such topic in a synthetic data set and two real-world text corpora and achieves better performance compared to both standard group AD and individual AD techniques. All required code to reproduce our experiments is available from https://github.com/hsoleimani/ATD

preprint2015arXiv

A global fit of top quark effective theory to data

In this paper we present a global fit of beyond the Standard Model (BSM) dimension six operators relevant to the top quark sector to currently available data. Experimental measurements include parton-level top-pair and single top production from the LHC and the Tevatron. Higher order QCD corrections are modelled using differential and global K-factors, and we use novel fast-fitting techniques developed in the context of Monte Carlo event generator tuning to perform the fit. This allows us to provide new, fully correlated and model-independent bounds on new physics effects in the top sector from the most current direct hadron-collider measurements in light of the involved theoretical and experimental systematics. As a by-product, our analysis constitutes a proof-of-principle that fast fitting of theory to data is possible in the top quark sector, and paves the way for a more detailed analysis including top quark decays, detector corrections and precision observables.

preprint2015arXiv

Asymmetric Independence Model for Detecting Interactions between Variables

Detecting complex interactions among risk factors in case-control studies is a fundamental task in clinical and population research. However, though hypothesis testing using logistic regression (LR) is a convenient solution, the LR framework is poorly powered and ill-suited under several common circumstances in practice including missing or unmeasured risk factors, imperfectly correlated "surrogates", and multiple disease sub-types. The weakness of LR in these settings is related to the way in which the null hypothesis is defined. Here we propose the Asymmetric Independence Model (AIM) as a biologically-inspired alternative to LR, based on the key observation that the mechanisms associated with acquiring a "disease" versus maintaining "health" are asymmetric. We prove mathematically that, unlike LR, AIM is a robust model under the abovementioned confounding scenarios. Further, we provide a mathematical definition of a "synergistic" interaction, and prove that theoretically AIM has better power than LR for such interactions. We then experimentally show the superior performance of AIM as compared to LR on both simulations and four real datasets. While the principal application here involves genetic or environmental variables in the life sciences, our methodology is readily applied to other types of measurements and inferences, e.g. in the social sciences.

preprint2015arXiv

Convex Analysis of Mixtures for Separating Non-negative Well-grounded Sources

Blind Source Separation (BSS) has proven to be a powerful tool for the analysis of composite patterns in engineering and science. We introduce Convex Analysis of Mixtures (CAM) for separating non-negative well-grounded sources, which learns the mixing matrix by identifying the lateral edges of the convex data scatter plot. We prove a sufficient and necessary condition for identifying the mixing matrix through edge detection, which also serves as the foundation for CAM to be applied not only to the exact-determined and over-determined cases, but also to the under-determined case. We show the optimality of the edge detection strategy, even for cases where source well-groundedness is not strictly satisfied. The CAM algorithm integrates plug-in noise filtering using sector-based clustering, an efficient geometric convex analysis scheme, and stability-based model order selection. We demonstrate the principle of CAM on simulated data and numerically mixed natural images. The superior performance of CAM against a panel of benchmark BSS techniques is demonstrated on numerically mixed gene expression data. We then apply CAM to dissect dynamic contrast-enhanced magnetic resonance imaging data taken from breast tumors and time-course microarray gene expression data derived from in-vivo muscle regeneration in mice, both producing biologically plausible decomposition results.

preprint2015arXiv

Detecting Clusters of Anomalies on Low-Dimensional Feature Subsets with Application to Network Traffic Flow Data

In a variety of applications, one desires to detect groups of anomalous data samples, with a group potentially manifesting its atypicality (relative to a reference model) on a low-dimensional subset of the full measured set of features. Samples may only be weakly atypical individually, whereas they may be strongly atypical when considered jointly. What makes this group anomaly detection problem quite challenging is that it is a priori unknown which subset of features jointly manifests a particular group of anomalies. Moreover, it is unknown how many anomalous groups are present in a given data batch. In this work, we develop a group anomaly detection (GAD) scheme to identify the subset of samples and subset of features that jointly specify an anomalous cluster. We apply our approach to network intrusion detection to detect BotNet and peer-to-peer flow clusters. Unlike previous studies, our approach captures and exploits statistical dependencies that may exist between the measured features. Experiments on real world network traffic data demonstrate the advantage of our proposed system, and highlight the importance of exploiting feature dependency structure, compared to the feature (or test) independence assumption made in previous studies.

preprint2015arXiv

Jet substructure and probes of CP violation in Vh production

We analyse the hVV (V = W, Z) vertex in a model independent way using Vh production. To that end, we consider possible corrections to the Standard Model Higgs Lagrangian, in the form of higher dimensional operators which parametrise the effects of new physics. In our analysis, we pay special attention to linear observables that can be used to probe CP violation in the same. By considering the associated production of a Higgs boson with a vector boson (W or Z), we use jet substructure methods to define angular observables which are sensitive to new physics effects, including an asymmetry which is linearly sensitive to the presence of CP odd effects. We demonstrate how to use these observables to place bounds on the presence of higher dimensional operators, and quantify these statements using a log likelihood analysis. Our approach allows one to probe separately the hZZ and hWW vertices, involving arbitrary combinations of BSM operators, at the Large Hadron Collider.

preprint2015arXiv

Next-to-leading order predictions for WW+jet production

In this work we report on a next-to-leading order calculation of WW + jet production at hadron colliders, with subsequent leptonic decays of the W-bosons included. The calculation of the one-loop contributions is performed using generalized unitarity methods in order to derive analytic expressions for the relevant amplitudes. These amplitudes have been implemented in the parton-level Monte Carlo generator MCFM, which we use to provide a complete next-to-leading order calculation. Predictions for total cross-sections, as well as differential distributions for several key observables, are computed both for the LHC operating at 14 TeV as well as for a possible future 100 TeV proton-proton collider.

preprint2014arXiv

Generation bidding game with flexible demand

For a simple model of price-responsive demand, we consider a deregulated electricity marketplace wherein the grid (ISO, retailer-distributor) accepts bids per-unit supply from generators (simplified herein neither to consider start-up/ramp-up expenses nor day-ahead or shorter-term load following) which are then averaged (by supply allocations via an economic dispatch) to a common "clearing" price borne by customers (irrespective of variations in transmission/distribution or generation prices), i.e., the ISO does not compensate generators based on their marginal costs. Rather, the ISO provides sufficient information for generators to sensibly adjust their bids. Notwithstanding our idealizations, the dispatch dynamics are complex. For a simple benchmark power system, we find a price-symmetric Nash equilibrium through numerical experiments.

preprint2014arXiv

Parsimonious Topic Models with Salient Word Discovery

We propose a parsimonious topic model for text corpora. In related models such as Latent Dirichlet Allocation (LDA), all words are modeled topic-specifically, even though many words occur with similar frequencies across different topics. Our modeling determines salient words for each topic, which have topic-specific probabilities, with the rest explained by a universal shared model. Further, in LDA all topics are in principle present in every document. By contrast our model gives sparse topic representation, determining the (small) subset of relevant topics for each document. We derive a Bayesian Information Criterion (BIC), balancing model complexity and goodness of fit. Here, interestingly, we identify an effective sample size and corresponding penalty specific to each parameter type in our model. We minimize BIC to jointly determine our entire model -- the topic-specific words, document-specific topics, all model parameter values, {\it and} the total number of topics -- in a wholly unsupervised fashion. Results on three text corpora and an image dataset show that our model achieves higher test set likelihood and better agreement with ground-truth class labels, compared to LDA and to a model designed to incorporate sparsity.

preprint2013arXiv

Boosting Higgs CP properties via VH Production at the Large Hadron Collider

We consider ZH and WH production at the Large Hadron Collider, where the Higgs decays to a bb pair. We use jet substructure techniques to reconstruct the Higgs boson and construct angular observables involving leptonic decay products of the vector bosons. These efficiently discriminate between the tensor structure of the HVV vertex expected in the Standard Model and that arising from possible new physics, as quantified by higher dimensional operators. This can then be used to examine the CP nature of the Higgs as well as CP mixing effects in the HZZ and HWW vertices separately.

preprint2012arXiv

Computation of the scaling factor of resistance forms of the pillow and fractalina fractals

Much is known in the analysis of a finitely ramified self-similar fractal when the fractal has a harmonic structure: a Dirichlet form which respects the self-similarity of a fractal. What is still an open question is when such structure exists in general. In this paper, we introduce two fractals, the fractalina and pillow, and compute their resistance scaling factor. This is the factor which dictates how the Dirichlet form scales with the self-similarity of the fractal. By knowing this factor one can compute the harmonic structure on the fractal. The fractalina has scaling factor $(3+\sqrt{41})/16$, and the pillow fractal has scaling factor $\sqrt[3]{2}$.

preprint2012arXiv

Game Theoretic Iterative Partitioning for Dynamic Load Balancing in Distributed Network Simulation

High fidelity simulation of large-sized complex networks can be realized on a distributed computing platform that leverages the combined resources of multiple processors or machines. In a discrete event driven simulation, the assignment of logical processes (LPs) to machines is a critical step that affects the computational and communication burden on the machines, which in turn affects the simulation execution time of the experiment. We study a network partitioning game wherein each node (LP) acts as a selfish player. We derive two local node-level cost frameworks which are feasible in the sense that the aggregate state information required to be exchanged between the machines is independent of the size of the simulated network model. For both cost frameworks, we prove the existence of stable Nash equilibria in pure strategies. Using iterative partition improvements, we propose game theoretic partitioning algorithms based on the two cost criteria and show that each descends in a global cost. To exploit the distributed nature of the system, the algorithm is distributed, with each node's decision based on its local information and on a few global quantities which can be communicated machine-to-machine. We demonstrate the performance of our partitioning algorithm on an optimistic discrete event driven simulation platform that models an actual parallel simulator.

preprint2011arXiv

An MRI-Derived Definition of MCI-to-AD Conversion for Long-Term, Automati c Prognosis of MCI Patients

Alzheimer's disease (AD) and mild cognitive impairment (MCI), continue to be widely studied. While there is no consensus on whether MCIs actually "convert" to AD, the more important question is not whether MCIs convert, but what is the best such definition. We focus on automatic prognostication, nominally using only a baseline image brain scan, of whether an MCI individual will convert to AD within a multi-year period following the initial clinical visit. This is in fact not a traditional supervised learning problem since, in ADNI, there are no definitive labeled examples of MCI conversion. Prior works have defined MCI subclasses based on whether or not clinical/cognitive scores such as CDR significantly change from baseline. There are concerns with these definitions, however, since e.g. most MCIs (and ADs) do not change from a baseline CDR=0.5, even while physiological changes may be occurring. These works ignore rich phenotypical information in an MCI patient's brain scan and labeled AD and Control examples, in defining conversion. We propose an innovative conversion definition, wherein an MCI patient is declared to be a converter if any of the patient's brain scans (at follow-up visits) are classified "AD" by an (accurately-designed) Control-AD classifier. This novel definition bootstraps the design of a second classifier, specifically trained to predict whether or not MCIs will convert. This second classifier thus predicts whether an AD-Control classifier will predict that a patient has AD. Our results demonstrate this new definition leads not only to much higher prognostic accuracy than by-CDR conversion, but also to subpopulations much more consistent with known AD brain region biomarkers. We also identify key prognostic region biomarkers, essential for accurately discriminating the converter and nonconverter groups.