Source author record

Feifei Wang

Feifei Wang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Artificial Intelligence Machine Learning Methodology astro-ph.HE Computation and Language Computer Vision Cryptography and Security physics.app-ph

Catalog footprint

What is connected

7works

8topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding

Automated international classification of diseases (ICD) coding aims to assign multiple disease codes to clinical documents and plays a critical role in healthcare informatics. However, its performance is hindered by the extreme long-tail distribution of the ICD ontology, where a few common codes dominate while thousands of rare codes have very few examples. To address this issue, we propose a Probability-Biased Directed Graph Attention model (ProBias) that partitions codes into common and rare sets and allows information to flow only from common to rare codes. Edge weights are determined by conditional co-occurrence probabilities, which guide the attention mechanism to enrich rare-code representations with clinically related signals. To provide higher-quality semantic representations as model inputs, we further employ large language models to generate enriched textual descriptions for ICD codes, offering external clinical context that complements statistical co-occurrence signals. Applied to automated ICD coding, our approach significantly improves the representation and prediction of rare codes, achieving state-of-the-art performance on three benchmark datasets. In particular, we observe substantial gains in macro-averaged F1 score, a key metric for long-tail classification.

preprint2026arXiv

Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks

Federated Learning (FL) has emerged as a promising privacy-preserving collaborative model training paradigm without sharing raw data. However, recent studies have revealed that private information can still be leaked through shared gradient information and attacked by Gradient Inversion Attacks (GIA). While many GIA methods have been proposed, a detailed analysis, evaluation, and summary of these methods are still lacking. Although various survey papers summarize existing privacy attacks in FL, few studies have conducted extensive experiments to unveil the effectiveness of GIA and their associated limiting factors in this context. To fill this gap, we first undertake a systematic review of GIA and categorize existing methods into three types, i.e., \textit{optimization-based} GIA (OP-GIA), \textit{generation-based} GIA (GEN-GIA), and \textit{analytics-based} GIA (ANA-GIA). Then, we comprehensively analyze and evaluate the three types of GIA in FL, providing insights into the factors that influence their performance, practicality, and potential threats. Our findings indicate that OP-GIA is the most practical attack setting despite its unsatisfactory performance, while GEN-GIA has many dependencies and ANA-GIA is easily detectable, making them both impractical. Finally, we offer a three-stage defense pipeline to users when designing FL frameworks and protocols for better privacy protection and share some future research directions from the perspectives of attackers and defenders that we believe should be pursued. We hope that our study can help researchers design more robust FL frameworks to defend against these attacks.

preprint2022arXiv

High-Capacity Rechargeable $Li/Cl_2$ Batteries with Graphite Positive Electrodes

Developing new types of high-capacity and high-energy density rechargeable battery is important to future generations of consumer electronics, electric vehicles, and mass energy storage applications. Recently we reported ~ 3.5 V sodium/chlorine $(Na/Cl_2)$ and lithium/chlorine $(Li/Cl_2)$ batteries with up to 1200 mAh $g^{-1}$ reversible capacity, using either a Na or Li metal as the negative electrode, an amorphous carbon nanosphere (aCNS) as the positive electrode, and aluminum chloride $(AlCl_3)$ dissolved in thionyl chloride $(SOCl_2)$ with fluoride-based additives as the electrolyte. The high surface area and large pore volume of aCNS in the positive electrode facilitated NaCl or LiCl deposition and trapping of $Cl_2$ for reversible $NaCl/Cl_2$ or $LiCl/Cl_2$ redox reactions and battery discharge/charge cycling. Here we report an initially low surface area/porosity graphite (DGr) material as the positive electrode in a $Li/Cl_2$ battery, attaining high battery performance after activation in carbon dioxide $(CO_2)$ at 1000 °C (DGr_ac) with the first discharge capacity ~ 1910 mAh $g^{-1}$ and a cycling capacity up to 1200 mAh $g^{-1}$. Ex situ Raman spectroscopy and X-ray diffraction (XRD) revealed the evolution of graphite over battery cycling, including intercalation/de-intercalation and exfoliation that generated sufficient pores for hosting $LiCl/Cl_2$ redox. This work opens up widely available, low-cost graphitic materials for high-capacity alkali metal/$Cl_2$ batteries. Lastly, we employed mass spectrometry to probe the $Cl_2$ trapped in the graphitic positive electrode, shedding light into the $Li/Cl_2$ battery operation.

preprint2022arXiv

Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning

Federated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models at https://github.com/Liangqiong/ViT-FL-main to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front.

preprint2020arXiv

Efficient Estimation for Generalized Linear Models on a Distributed System with Nonrandomly Distributed Data

Distributed systems have been widely used in practice to accomplish data analysis tasks of huge scales. In this work, we target on the estimation problem of generalized linear models on a distributed system with nonrandomly distributed data. We develop a Pseudo-Newton-Raphson algorithm for efficient estimation. In this algorithm, we first obtain a pilot estimator based on a small random sample collected from different Workers. Then conduct one-step updating based on the computed derivatives of log-likelihood functions in each Worker at the pilot estimator. The final one-step estimator is proved to be statistically efficient as the global estimator, even with nonrandomly distributed data. In addition, it is computationally efficient, in terms of both communication cost and storage usage. Based on the one-step estimator, we also develop a likelihood ratio test for hypothesis testing. The theoretical properties of the one-step estimator and the corresponding likelihood ratio test are investigated. The finite sample performances are assessed through simulations. Finally, an American Airline dataset is analyzed on a Spark cluster for illustration purpose.

preprint2019arXiv

A comprehensive statistical study on gamma-ray bursts

In order to obtain an overview of the gamma-ray bursts (GRBs), we need a full sample. In this paper, we collected 6289 GRBs (from GRB 910421 to GRB 160509A) from the literature, including prompt emission, afterglow and host galaxy properties. We hope to use this large sample to reveal the intrinsic properties of GRB. We have listed all the data in machine readable tables, including the properties of the GRBs, correlation coefficients and linear regression results of two arbitrary parameters, and linear regression results of any three parameters. These machine readable tables could be used as a data reservoir for further studies on the classifications or correlations. One may find some intrinsic properties from these statistical results. With this comprehensive table, it is possible to find relations between different parameters, and to classify the GRBs into different kinds of sub-groups. With the completion, it may reveal the nature of GRBs and may be used as tools like pseudo-redshift indicators, standard candles, etc. All the machine readable data and statistical results are available on the website of the journal.

preprint2016arXiv

Disease Mapping with Generative Models

Disease mapping focuses on learning about areal units presenting high relative risk. Disease mapping models for disease counts specify Poisson regressions in relative risks compared with the expected counts. These models typically incorporate spatial random effects to accomplish spatial smoothing. Fitting of these models customarily computes expected disease counts via internal standardization. This places the data on both sides of the model, i.e., the counts are on the left side but they are also used to obtain the expected counts on the right side. As a result, these internally standardized models are incoherent and not generative; probabilistically, they could not produce the observed data. Here, we argue for adopting the direct generative model for disease counts. We model disease incidence instead of relative risks, using a generalized logistic regression. We extract relative risks post model fitting. We also extend the generative model to dynamic settings. We compare the generative models with internally standardized models through simulated datasets and a well-examined lung cancer morbidity data in Ohio. Each model is a spatial smoother and they smooth the data similarly with regard to relative risks. However, the generative models tend to provide tighter credible intervals. Since the generative specification is no more difficult to fit, is coherent, and is at least as good inferentially, we suggest it should be the model of choice for spatial disease mapping.