Source author record

Chi Wang

Chi Wang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

24works
18topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

24 published item(s)

preprint2026arXiv

HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation

Generating controllable and physically plausible indoor scenes is a pivotal prerequisite for constructing high-fidelity simulation environments for embodied AI. However, existing deeplearning-based methods usually treat all objects as homogeneous instances within a unified generation process. While effective for sparse and simplistic layouts, they struggle to model realistic layouts with dense object arrangements and complex spatial dependencies, leadingto limited scalability and degraded physical plausibility. To deal with these challenges, we revisit indoor layout generation from the perspective of structural heterogeneity and decompose the objects into primary objects and secondary objects according to their distinct roles in shaping a scene. Based on this decomposition, we propose HetScene, a heterogeneous two-stage generation framework that decouples indoor layout synthesis into Structural Layout Generation (SLG) and Contextual Layout Generation (CLG). SLG first generates globally coherent structural layouts with only primary objects conditioned on text descriptions, top-down binary room masks, and spatial relation graphs, establishing a stable global macro-skeleton of large core furniture.

preprint2022arXiv

ACE: Adaptive Constraint-aware Early Stopping in Hyperparameter Optimization

Deploying machine learning models requires high model quality and needs to comply with application constraints. That motivates hyperparameter optimization (HPO) to tune model configurations under deployment constraints. The constraints often require additional computation cost to evaluate, and training ineligible configurations can waste a large amount of tuning cost. In this work, we propose an Adaptive Constraint-aware Early stopping (ACE) method to incorporate constraint evaluation into trial pruning during HPO. To minimize the overall optimization cost, ACE estimates the cost-effective constraint evaluation interval based on a theoretical analysis of the expected evaluation cost. Meanwhile, we propose a stratum early stopping criterion in ACE, which considers both optimization and constraint metrics in pruning and does not require regularization hyperparameters. Our experiments demonstrate superior performance of ACE in hyperparameter tuning of classification tasks under fairness or robustness constraints.

preprint2022arXiv

Active Boundary Loss for Semantic Segmentation

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted boundaries detected from the segmentation results using current network parameters, we formulate the boundary alignment problem as a differentiable direction vector prediction problem to guide the movement of predicted boundaries in each iteration. Our loss is model-agnostic and can be plugged in to the training of segmentation networks to improve the boundary details. Experimental results show that training with the active boundary loss can effectively improve the boundary F-score and mean Intersection-over-Union on challenging image and video object segmentation datasets.

preprint2022arXiv

Multivariate Sparse Group Lasso Joint Model for Radiogenomics Data

Radiogenomics is an emerging field in cancer research that combines medical imaging data with genomic data to predict patients clinical outcomes. In this paper, we propose a multivariate sparse group lasso joint model to integrate imaging and genomic data for building prediction models. Specifically, we jointly consider two models, one regresses imaging features on genomic features, and the other regresses patients clinical outcomes on genomic features. The regularization penalties through sparse group lasso allow incorporation of intrinsic group information, e.g. biological pathway and imaging category, to select both important intrinsic groups and important features within a group. To integrate information from the two models, in each model, we introduce a weight in the penalty term of each individual genomic feature, where the weight is inversely correlated with the model coefficient of that feature in the other model. This weight allows a feature to have a higher chance of selection by one model if it is selected by the other model. Our model is applicable to both continuous and time to event outcomes. It also allows the use of two separate datasets to fit the two models, addressing a practical challenge that many genomic datasets do not have imaging data available. Simulations and real data analyses demonstrate that our method outperforms existing methods in the literature.

preprint2022arXiv

Whistler Waves As a Signature of Converging Magnetic Holes in Space Plasmas

Magnetic holes are plasma structures that trap a large number of particles in a magnetic field that is weaker than the field in its surroundings. The unprecedented high time-resolution observations by NASA's Magnetospheric Multi-Scale (MMS) mission enable us to study the particle dynamics in magnetic holes in the Earth's magnetosheath in great detail. We reveal the local generation mechanism of whistler waves by a combination of Landau-resonant and cyclotron-resonant wave-particle interactions of electrons in response to the large-scale evolution of a magnetic hole. As the magnetic hole converges, a pair of counter-streaming electron beams form near the hole's center as a consequence of the combined action of betatron and Fermi effects. The beams trigger the generation of slightly-oblique whistler waves. Our conceptual prediction is supported by a remarkable agreement between our observations and numerical predictions from the Arbitrary Linear Plasma Solver (ALPS). Our study shows that wave-particle interactions are fundamental to the evolution of magnetic holes in space and astrophysical plasmas.

preprint2020arXiv

ALEX: An Updatable Adaptive Learned Index

Recent work on "learned indexes" has changed the way we look at the decades-old field of DBMS indexing. The key idea is that indexes can be thought of as "models" that predict the position of a key in a dataset. Indexes can, thus, be learned. The original work by Kraska et al. shows that a learned index beats a B+Tree by a factor of up to three in search time and by an order of magnitude in memory footprint. However, it is limited to static, read-only workloads. In this paper, we present a new learned index called ALEX which addresses practical issues that arise when implementing learned indexes for workloads that contain a mix of point lookups, short range queries, inserts, updates, and deletes. ALEX effectively combines the core insights from learned indexes with proven storage and indexing techniques to achieve high performance and low memory footprint. On read-only workloads, ALEX beats the learned index from Kraska et al. by up to 2.2X on performance with up to 15X smaller index size. Across the spectrum of read-write workloads, ALEX beats B+Trees by up to 4.1X while never performing worse, with up to 2000X smaller index size. We believe ALEX presents a key step towards making learned indexes practical for a broader class of database workloads with dynamic updates.

preprint2020arXiv

Evolution of the Earth's Magnetosheath Turbulence: A statistical study based on MMS observations

Composed of shocked solar wind, the Earth's magnetosheath serves as a natural laboratory to study the transition of turbulence from low Alfv{é}n Mach number, $M_\mathrm{A}$, to high $M_\mathrm{A}$. The simultaneous observations of magnetic field and plasma moments with unprecedented high temporal resolution provided by NASA's \textit{Magnetospheric Multiscale} Mission enable us to study the magnetosheath turbulence at both magnetohydrodynamics (MHD) and sub-ion scales. Based on 1841 burst-mode segments of MMS-1 from 2015/09 to 2019/06, comprehensive patterns of the spatial evolution of magnetosheath turbulences are obtained: (1) from the sub-solar region to the flanks, $M_\mathrm{A}$ increases from $<$ 1 to $>$ 5. At MHD scales, the spectral indices of the magnetic-field and velocity spectra present a positive and negative correlation with $M_\mathrm{A}$. However, no obvious correlations between the spectral indices and $M_\mathrm{A}$ are found at sub-ion scales. (2) from the bow shock to the magnetopause, the turbulent sonic Mach number, $M_{\mathrm{turb}}$, generally decreases from $>$ 0.4 to $<$ 0.1. All spectra steepen at MHD scales and flatten at sub-ion scales, representing a positive/negative correlations with $M_\mathrm{turb}$. The break frequency increases by 0.1 Hz when approaching the magnetopause for the magnetic-field and velocity spectra, while it remains at 0.3 Hz for the density spectra. (3) In spite of some differences, similar results are found for the quasi-parallel and quasi-perpendicular magnetosheath. In addition, the spatial evolution of magnetosheath turbulence is found to be independent of the upstream solar wind conditions, e.g., the Z-component of the interplanetary magnetic field and the solar wind speed.

preprint2020arXiv

GEE-TGDR: A longitudinal feature selection algorithm and its application to lncRNA expression profiles for psoriasis patients treated with immune therapies

With the fast evolution of high-throughput technology, longitudinal gene expression experiments have become affordable and increasingly common in biomedical fields. Generalized estimating equation (GEE) approach is a widely used statistical method for the analysis of longitudinal data. Feature selection is imperative in longitudinal omics data analysis. Among a variety of existing feature selection methods, an embedded method, namely, threshold gradient descent regularization (TGDR) stands out due to its excellent characteristics. An alignment of GEE with TGDR is a promising area for the purpose of identifying relevant markers that can explain the dynamic changes of outcomes across time. In this study, we proposed a new novel feature selection algorithm for longitudinal outcomes:GEE-TGDR. In the GEE-TGDR method, the corresponding quasi-likelihood function of a GEE model is the objective function to be optimized and the optimization and feature selection are accomplished by the TGDR method. We applied the GEE-TGDR method a longitudinal lncRNA gene expression dataset that examined the treatment response of psoriasis patients to immune therapy. Under different working correlation structures, a list including 10 relevant lncRNAs were identified with a predictive accuracy of 80 % and meaningful biological interpretation. To conclude, a widespread application of the proposed GEE-TGDR method in omics data analysis is anticipated.

preprint2020arXiv

Qd-tree: Learning Data Layouts for Big Data Analytics

Corporations today collect data at an unprecedented and accelerating scale, making the need to run queries on large datasets increasingly important. Technologies such as columnar block-based data organization and compression have become standard practice in most commercial database systems. However, the problem of best assigning records to data blocks on storage is still open. For example, today's systems usually partition data by arrival time into row groups, or range/hash partition the data based on selected fields. For a given workload, however, such techniques are unable to optimize for the important metric of the number of blocks accessed by a query. This metric directly relates to the I/O cost, and therefore performance, of most analytical queries. Further, they are unable to exploit additional available storage to drive this metric down further. In this paper, we propose a new framework called a query-data routing tree, or qd-tree, to address this problem, and propose two algorithms for their construction based on greedy and deep reinforcement learning techniques. Experiments over benchmark and real workloads show that a qd-tree can provide physical speedups of more than an order of magnitude compared to current blocking schemes, and can reach within 2X of the lower bound for data skipping based on selectivity, while providing complete semantic descriptions of created blocks.

preprint2020arXiv

TaxoExpan: Self-supervised Taxonomy Expansion with Position-Enhanced Graph Neural Network

Taxonomies consist of machine-interpretable semantics and provide valuable knowledge for many web applications. For example, online retailers (e.g., Amazon and eBay) use taxonomies for product recommendation, and web search engines (e.g., Google and Bing) leverage taxonomies to enhance query understanding. Enormous efforts have been made on constructing taxonomies either manually or semi-automatically. However, with the fast-growing volume of web content, existing taxonomies will become outdated and fail to capture emerging knowledge. Therefore, in many applications, dynamic expansions of an existing taxonomy are in great demand. In this paper, we study how to expand an existing taxonomy by adding a set of new concepts. We propose a novel self-supervised framework, named TaxoExpan, which automatically generates a set of <query concept, anchor concept> pairs from the existing taxonomy as training data. Using such self-supervision data, TaxoExpan learns a model to predict whether a query concept is the direct hyponym of an anchor concept. We develop two innovative techniques in TaxoExpan: (1) a position-enhanced graph neural network that encodes the local structure of an anchor concept in the existing taxonomy, and (2) a noise-robust training objective that enables the learned model to be insensitive to the label noise in the self-supervision data. Extensive experiments on three large-scale datasets from different domains demonstrate both the effectiveness and the efficiency of TaxoExpan for taxonomy expansion.

preprint2016arXiv

Plasma heating inside ICMEs by Alfvenic fluctuations dissipation

Nonlinear cascade of low-frequency Alfvenic fluctuations (AFs) is regarded as one candidate of the energy sources to heat plasma during the non-adiabatic expansion of interplanetary coronal mass ejections (ICMEs). However, AFs inside ICMEs were seldom reported in the literature. In this study, we investigate AFs inside ICMEs using observations from Voyager 2 between 1 and 6 au. It is found that AFs with high degree of Alfvenicity frequently occurred inside ICMEs, for almost all the identified ICMEs (30 out of 33 ICMEs), and 12.6% of ICME time interval. As ICMEs expand and move outward, the percentage of AF duration decays linearly in general. The occurrence rate of AFs inside ICMEs is much less than that in ambient solar wind, especially within 4 au. AFs inside ICMEs are more frequently presented in the center and at the boundaries of ICMEs. In addition, the proton temperature inside ICME has a similar distribution. These findings suggest significant contribution of AFs on local plasma heating inside ICMEs.

preprint2016arXiv

Properties of post-shock solar wind deduced from geomagnetic indices responses after sudden impulses

Interplanetary (IP) shock plays a key role in causing the global dynamic changes of the geospace environment. For the perspective of Solar-Terrestrial relationship, it will be of great importance to estimate the properties of post-shock solar wind simply and accurately. Motivated by this, we performed a statistical analysis of IP shocks during 1998-2008, focusing on the significantly different responses of two well-used geomagnetic indices (SYMH and AL) to the passive of two types of IP shocks. For the IP shocks with northward IMF (91 cases), the SYMH index keeps on the high level after the sudden impulses (SI) for a long time. Meanwhile, the change of AL index is relative small, with an mean value of only -29 nT. However, for the IP shocks with southward IMF (92 cases), the SYMH index suddenly decreases at a certain rate after SI, and the change of AL index is much significant, of -316 nT. Furthermore, the change rate of SYMH index after SI is found to be linearly correlated with the post-shock reconnection E-field (E$_{KL}$). Based on these facts, an inversion model of post-shock IMF orientation and E$_{KL}$ is developed. The model validity is also confirmed by studying 68 IP shocks in the period of 2009-2013. The inversion accuracy of IMF orientation is 88.24%, and the inversion efficiency of E$_{KL}$ is as high as 78%.

preprint2016arXiv

Temperature Dependence Calibration and Correction of the DAMPE BGO Electromagnetic Calorimeter

A BGO electromagnetic calorimeter (ECAL) is built for the DArk Matter Particle Explorer (DAMPE) mission. The effect of temperature on the BGO ECAL was investigated with a thermal vacuum experiment. The light output of a BGO crystal depends on temperature significantly. The temperature coefficient of each BGO crystal bar has been calibrated, and a correction method is also presented in this paper.

preprint2016arXiv

The calibration and electron energy reconstruction of the BGO ECAL of the DAMPE detector

The DArk Matter Particle Explorer (DAMPE) is a space experiment designed to search for dark matter indirectly by measuring the spectra of photons, electrons, and positrons up to 10 TeV. The BGO electromagnetic calorimeter (ECAL) is its main sub-detector for energy measurement. In this paper, the instrumentation and development of the BGO ECAL is briefly described. The calibration on the ground, including the pedestal, minimum ionizing particle (MIP) peak, dynode ratio, and attenuation length with the cosmic rays and beam particles is discussed in detail. Also, the energy reconstruction results of the electrons from the beam test are presented.

preprint2016arXiv

Weighted SAMGSR: combining significance analysis of microarray-gene set reduction algorithm with pathway topology-based weights to select relevant genes

Introduction It has been demonstrated that a pathway-based feature selection method which incorporates biological information within pathways into the process of feature selection usually outperform a gene-based feature selection algorithm in terms of predictive accuracy, stability, and biological interpretation. Significance analysis of microarray-gene set reduction algorithm (SAMGSR), an extension to a gene set analysis method with further reduction of the selected pathways to their respective core subsets, can be regarded as a pathway-based feature selection method. Results and Discussion In SAMGSR, whether a gene is selected is mainly determined by its expression difference between the phenotypes, and partially by the number of pathways to which this gene belongs, but ignoring the topology information among pathways. In this study, we propose a weighted version of the SAMGSR algorithm by constructing weights based on the connectivity among genes and then incorporating these weights in the test statistic. Conclusions Using both simulated and real-world data, we evaluate the performance of the proposed SAMGSR extension and demonstrate that gene connectivity is indeed informative for feature selection.

preprint2015arXiv

A study of energy correction for the electron beam data in the BGO ECAL of the DAMPE

The DArk Matter Particle Explorer (DAMPE) is an orbital experiment aiming at searching for dark matter indirectly by measuring the spectra of photons, electrons and positrons originating from deep space. The BGO electromagnetic calorimeter is one of the key sub-detectors of the DAMPE, which is designed for high energy measurement with a large dynamic range from 5 GeV to 10 TeV. In this paper, some methods for energy correction are discussed and tried, in order to reconstruct the primary energy of the incident electrons. Different methods are chosen for the appropriate energy ranges. The results of Geant4 simulation and beam test data (at CERN) are presented.

preprint2015arXiv

Feature selection for longitudinal microarray data by adapting a pathway analysis method

Introduction: Feature selection and gene set analysis are of increasing interest in bioinformatics. While these two approaches have been developed for different purposes, we describe how some gene set analysis methods can be used to conduct feature selection. Here we adapt the gene set analysis method, significance analysis of microarray gene set reduction (SAMGSR), for feature selection, and propose two extensions-simple SAMGSR and two-level SAMGSR to identify relevant features for longitudinal microarray data. Results and Discussion: When applied to a real-world application, both simple and two-level SAMGSR work comparably well. Using simulated data, we further demonstrate that both SAMGSR extensions have the ability to identify the true relevant genes. If the relevant genes are not highly correlated with the irrelevant ones, the final models given by the two SAMGSR extensions are parsimonious as well. Conclusions: By adapting SAMGSR for feature selection and applying the proposed algorithms on a longitudinal gene expression dataset, we demonstrate that a gene set analysis method can be used for the purpose of feature selection. We believe this work paves the way for more research to bridge feature selection and gene set analysis with the development of novel algorithms.

preprint2015arXiv

In situ Evidence of Breaking the Ion Frozen-in Condition via the Non-gyrotropic Pressure Effect in Magnetic Reconnection

For magnetic reconnection to proceed, the frozen-in condition for both ion fluid and electron fluid in a localized diffusion region must be violated by inertial effects, thermal pressure effects, or inter-species collisions. It has been unclear which underlying effects unfreeze ion fluid in the diffusion region. By analyzing in-situ THEMIS spacecraft measurements at the dayside magnetopause, we present clear evidence that the off-diagonal components of the ion pressure tensor is mainly responsible for breaking the ion frozen-in condition in reconnection. The off-diagonal pressure tensor, which corresponds to a nongyrotropic pressure effect, is a fluid manifestation of ion demagnetization in the diffusion region. From the perspective of the ion momentum equation, the reported non-gyrotropic ion pressure tensor is a fundamental aspect in specifying the reconnection electric field that controls how quickly reconnection proceeds.

preprint2015arXiv

No evidence of histology subtype-specific prognostic signatures among lung adenocarcinoma and squamous cell carcinoma patients at early stages

Background Non-small cell lung cancer (NSCLC) is the predominant histological type of lung cancer, accounting for up to 85% of cases. Disease stage is commonly used to determine adjuvant treatment eligibility of NSCLC patients, however, it is an imprecise predictor of the prognosis of an individual patient. Currently, many researchers resort to microarray technology for identifying relevant genetic prognostic markers, with particular attention on trimming or extending a Cox regression model. Among NSCLC, adenocarcinoma (AC) and squamous cell carcinoma (SCC) are two major histology subtypes. It has been demonstrated that there exist fundamental differences in the underlying mechanisms between them, which motivated us to postulate there might exist specific genes relevant to prognosis of each histology subtype. Results In this article, we propose a simple filterer feature selection algorithm with a Cox regression model as the base. Applying this method to a real-world microarray data, no evidence has been found to support the existence of histology-specific prognostic gene signature. Nevertheless, a 31-gene prognostic gene signature for the early-stage AC and SCC samples is obtained, which provides comparable performance when compared with other relevant signatures. Conclusions Our proposal is conceptually simple and straightforward to implement. Therefore, it is expected that other researchers, especially those with less statistical knowledge and experience, can adapt this method readily to test their own research hypotheses.

preprint2015arXiv

On Sun-to-Earth Propagation of Coronal Mass Ejections: 2. Slow Events and Comparison with Others

As a follow-up study on Sun-to-Earth propagation of fast coronal mass ejections (CMEs), we examine the Sun-to-Earth characteristics of slow CMEs combining heliospheric imaging and in situ observations. Three events of particular interest, the 2010 June 16, 2011 March 25 and 2012 September 25 CMEs, are selected for this study. We compare slow CMEs with fast and intermediate-speed events, and obtain key results complementing the attempt of \citet{liu13} to create a general picture of CME Sun-to-Earth propagation: (1) the Sun-to-Earth propagation of a typical slow CME can be approximately described by two phases, a gradual acceleration out to about 20-30 solar radii, followed by a nearly invariant speed around the average solar wind level, (2) comparison between different types of CMEs indicates that faster CMEs tend to accelerate and decelerate more rapidly and have shorter cessation distances for the acceleration and deceleration, (3) both intermediate-speed and slow CMEs would have a speed comparable to the average solar wind level before reaching 1 AU, (4) slow CMEs have a high potential to interact with other solar wind structures in the Sun-Earth space due to their slow motion, providing critical ingredients to enhance space weather, and (5) the slow CMEs studied here lack strong magnetic fields at the Earth but tend to preserve a flux-rope structure with axis generally perpendicular to the radial direction from the Sun. We also suggest a "best" strategy for the application of a triangulation concept in determining CME Sun-to-Earth kinematics, which helps to clarify confusions about CME geometry assumptions in the triangulation and to improve CME analysis and observations.

preprint2014arXiv

Propagation of the 2012 March Coronal Mass Ejections from the Sun to Heliopause

In 2012 March the Sun exhibited extraordinary activities. In particular, the active region NOAA AR 11429 emitted a series of large coronal mass ejections (CMEs) which were imaged by STEREO as it rotated with the Sun from the east to west. These sustained eruptions are expected to generate a global shell of disturbed material sweeping through the heliosphere. A cluster of shocks and interplanetary CMEs (ICMEs) were observed near the Earth, and are propagated outward from 1 AU using an MHD model. The transient streams interact with each other, which erases memory of the source and results in a large merged interaction region (MIR) with a preceding shock. The MHD model predicts that the shock and MIR would reach 120 AU around 2013 April 22, which agrees well with the period of radio emissions and the time of a transient disturbance in galactic cosmic rays detected by Voyager 1. These results are important for understanding the "fate" of CMEs in the outer heliosphere and provide confidence that the heliopause is located around 120 AU from the Sun.

preprint2014arXiv

Scalable and Robust Construction of Topical Hierarchies

Automated generation of high-quality topical hierarchies for a text collection is a dream problem in knowledge engineering with many valuable applications. In this paper a scalable and robust algorithm is proposed for constructing a hierarchy of topics from a text collection. We divide and conquer the problem using a top-down recursive framework, based on a tensor orthogonal decomposition technique. We solve a critical challenge to perform scalable inference for our newly designed hierarchical topic model. Experiments with various real-world datasets illustrate its ability to generate robust, high-quality hierarchies efficiently. Our method reduces the time of construction by several orders of magnitude, and its robust feature renders it possible for users to interactively revise the hierarchy.

preprint2014arXiv

Scalable Topical Phrase Mining from Text Corpora

While most topic modeling algorithms model text corpora with unigrams, human interpretation often relies on inherent grouping of terms into phrases. As such, we consider the problem of discovering topical phrases of mixed lengths. Existing work either performs post processing to the inference results of unigram-based topic models, or utilizes complex n-gram-discovery topic models. These methods generally produce low-quality topical phrases or suffer from poor scalability on even moderately-sized datasets. We propose a different approach that is both computationally efficient and effective. Our solution combines a novel phrase mining framework to segment a document into single and multi-word phrases, and a new topic model that operates on the induced document partition. Our approach discovers high quality topical phrases with negligible extra cost to the bag-of-words topic model in a variety of datasets including research publication titles, abstracts, reviews, and news articles.

preprint2013arXiv

KERT: Automatic Extraction and Ranking of Topical Keyphrases from Content-Representative Document Titles

We introduce KERT (Keyphrase Extraction and Ranking by Topic), a framework for topical keyphrase generation and ranking. By shifting from the unigram-centric traditional methods of unsupervised keyphrase extraction to a phrase-centric approach, we are able to directly compare and rank phrases of different lengths. We construct a topical keyphrase ranking function which implements the four criteria that represent high quality topical keyphrases (coverage, purity, phraseness, and completeness). The effectiveness of our approach is demonstrated on two collections of content-representative titles in the domains of Computer Science and Physics.