Source author record

Qingyang Li

Qingyang Li appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

9works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

9 published item(s)

preprint2022arXiv

\textsc{The Three Hundred} project: The \textsc{Gizmo-Simba} run

We introduce \textsc{Gizmo-Simba}, a new suite of galaxy cluster simulations within \textsc{The Three Hundred} project. \textsc{The Three Hundred} consists of zoom re-simulations of 324 clusters with $M_{200}\gtrsim 10^{14.8}M_\odot$ drawn from the MultiDark-Planck $N$-body simulation, run using several hydrodynamic and semi-analytic codes. The \textsc{Gizmo-Simba} suite adds a state-of-the-art galaxy formation model based on the highly successful {\sc Simba} simulation, mildly re-calibrated to match $z=0$ cluster stellar properties. Comparing to \textsc{The Three Hundred} zooms run with \textsc{Gadget-X}, we find intrinsic differences in the evolution of the stellar and gas mass fractions, BCG ages, and galaxy colour-magnitude diagrams, with \textsc{Gizmo-Simba} generally providing a good match to available data at $z \approx 0$. \textsc{Gizmo-Simba}'s unique black hole growth and feedback model yields agreement with the observed BH scaling relations at the intermediate-mass range and predicts a slightly different slope at high masses where few observations currently lie. \textsc{Gizmo-Simba} provides a new and novel platform to elucidate the co-evolution of galaxies, gas, and black holes within the densest cosmic environments.

preprint2022arXiv

A machine learning approach to infer the accreted stellar mass fractions of central galaxies in the TNG100 simulation

We propose a random forest (RF) machine learning approach to determine the accreted stellar mass fractions ($f_\mathrm{acc}$) of central galaxies, based on various dark matter halo and galaxy features. The RF is trained and tested using 2,710 galaxies with stellar mass $\log_{10}M_\ast/M_\odot>10.16$ from the TNG100 simulation. Galaxy size is the most important individual feature when calculated in 3-dimensions, which becomes less important after accounting for observational effects. For smaller galaxies, the rankings for features related to merger histories increase. When an entire set of halo and galaxy features are used, the prediction is almost unbiased, with root-mean-square error (RMSE) of $\sim$0.068. A combination of up to three features with different types (galaxy size, merger history and morphology) already saturates the power of prediction. If using observable features, the RMSE increases to $\sim$0.104, and a combined usage of stellar mass, galaxy size plus galaxy concentration achieves similar predictions. Lastly, when using galaxy density, velocity and velocity dispersion profiles as features, which approximately represent the maximum amount of information extracted from galaxy images and velocity maps, the prediction is not improved much. Hence the limiting precision of predicting $f_\mathrm{acc}$ is $\sim$0.1 with observables, and the multi-component decomposition of galaxy images should have similar or larger uncertainties. If the central black hole mass and the spin parameter of galaxies can be accurately measured in future observations, the RMSE is promising to be further decreased by $\sim$20%.

preprint2022arXiv

An Extended Halo-based Group/Cluster finder: application to the DESI legacy imaging surveys DR8

We extend the halo-based group finder developed by \citet[][]{Yang2005a} to use data {\it simultaneously} with either photometric or spectroscopic redshifts. A mock galaxy redshift survey constructed from a high-resolution N-body simulation is used to evaluate the performance of this extended group finder. For galaxies with magnitude ${\rm z\le 21}$ and redshift $0<z\le 1.0$ in the DESI legacy imaging surveys (the Legacy Surveys), our group finder successfully identifies more than 60\% of the members in about $90\%$ of halos with mass $\ga 10^{12.5}\msunh$. Detected groups with mass $\ga 10^{12.0}\msunh$ have a purity (the fraction of true groups) greater than 90\%. The halo mass assigned to each group has an uncertainty of about 0.2 dex at the high mass end $\ga 10^{13.5}\msunh$ and 0.40 dex at the low mass end. Groups with more than 10 members have a redshift accuracy of $\sim 0.008$. We apply this group finder to the Legacy Surveys DR8 and find 5.2 Million groups with at least 3 members. About 387,000 of these groups have at least 10 members. The resulting catalog containing 3D coordinates, richness, halo masses, and total group luminosities, is made publicly available.

preprint2022arXiv

Groups and protocluster candidates in the CLAUDS and HSC-SSP joint deep surveys

Using the extended halo-based group finder developed by Yang et al. (2021), which is able to deal with galaxies via spectroscopic and photometric redshifts simultaneously, we construct galaxy group and candidate protocluster catalogs in a wide redshift range ($0 < z < 6$) from the joint CFHT Large Area $U$-band Deep Survey (CLAUDS) and Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) deep data set. Based on a selection of 5,607,052 galaxies with $i$-band magnitude $m_{i} < 26$ and a sky coverage of $34.41\ {\rm deg}^2$, we identify a total of 2,232,134 groups, within which 402,947 groups have at least three member galaxies. We have visually checked and discussed the general properties of those richest groups at redshift $z>2.0$. By checking the galaxy number distributions within a $5-7\ h^{-1}\mathrm{Mpc}$ projected separation and a redshift difference $Δz \le 0.1$ around those richest groups at redshift $z>2$, we identified a list of 761, 343 and 43 protocluster candidates in the redshift bins $2\leq z<3$, $3\leq z<4$ and $z \geq 4$, respectively. In general, these catalogs of galaxy groups and protocluster candidates will provide useful environmental information in probing galaxy evolution along the cosmic time.

preprint2022arXiv

What to expect from dynamical modelling of cluster haloes II. Investigating dynamical state indicators with Random Forest

We investigate the importances of various dynamical features in predicting the dynamical state (DS) of galaxy clusters, based on the Random Forest (RF) machine learning approach. We use a large sample of galaxy clusters from the Three Hundred Project of hydrodynamical zoomed-in simulations, and construct dynamical features from the raw data as well as from the corresponding mock maps in the optical, X-ray, and Sunyaev-Zel'dovich (SZ) channels. Instead of relying on the impurity based feature importance of the RF algorithm, we directly use the out-of-bag (OOB) scores to evaluate the importances of individual features and different feature combinations. Among all the features studied, we find the virial ratio, $η$, to be the most important single feature. The features calculated directly from the simulations and in 3-dimensions carry more information on the DS than those constructed from the mock maps. Compared with the features based on X-ray or SZ maps, features related to the centroid positions are more important. Despite the large number of investigated features, a combination of up to three features of different types can already saturate the score of the prediction. Lastly, we show that the most sensitive feature $η$ is strongly correlated with the well-known half-mass bias in dynamical modelling. Without a selection in DS, cluster halos have an asymmetric distribution in $η$, corresponding to an overall positive half-mass bias. Our work provides a quantitative reference for selecting the best features to discriminate the DS of galaxy clusters in both simulations and observations.

preprint2020arXiv

Hierarchical Adaptive Contextual Bandits for Resource Constraint based Recommendation

Contextual multi-armed bandit (MAB) achieves cutting-edge performance on a variety of problems. When it comes to real-world scenarios such as recommendation system and online advertising, however, it is essential to consider the resource consumption of exploration. In practice, there is typically non-zero cost associated with executing a recommendation (arm) in the environment, and hence, the policy should be learned with a fixed exploration cost constraint. It is challenging to learn a global optimal policy directly, since it is a NP-hard problem and significantly complicates the exploration and exploitation trade-off of bandit algorithms. Existing approaches focus on solving the problems by adopting the greedy policy which estimates the expected rewards and costs and uses a greedy selection based on each arm's expected reward/cost ratio using historical observation until the exploration resource is exhausted. However, existing methods are hard to extend to infinite time horizon, since the learning process will be terminated when there is no more resource. In this paper, we propose a hierarchical adaptive contextual bandit method (HATCH) to conduct the policy learning of contextual bandits with a budget constraint. HATCH adopts an adaptive method to allocate the exploration resource based on the remaining resource/time and the estimation of reward distribution among different user contexts. In addition, we utilize full of contextual feature information to find the best personalized recommendation. Finally, in order to prove the theoretical guarantee, we present a regret bound analysis and prove that HATCH achieves a regret bound as low as $O(\sqrt{T})$. The experimental results demonstrate the effectiveness and efficiency of the proposed method on both synthetic data sets and the real-world applications.

preprint2020arXiv

The Three Hundred Project: the stellar and gas profiles

Using the catalogues of galaxy clusters from The Three Hundred project, modelled with both hydrodynamic simulations, (Gadget-X and Gadget-MUSIC), and semi-analytic models (SAMs), we study the scatter and self-similarity of the profiles and distributions of the baryonic components of the clusters: the stellar and gas mass, metallicity, the stellar age, gas temperature, and the (specific) star formation rate. Through comparisons with observational results, we find that the shape and the scatter of the gas density profiles matches well the observed trends including the reduced scatter at large radii which is a signature of self-similarity suggested in previous studies. One of our simulated sets, Gadget-X, reproduces well the shape of the observed temperature profile, while Gadget-MUSIC has a higher and flatter profile in the cluster centre and a lower and steeper profile at large radii. The gas metallicity profiles from both simulation sets, despite following the observed trend, have a relatively lower normalisation. The cumulative stellar density profiles from SAMs are in better agreement with the observed result than both hydrodynamic simulations which show relatively higher profiles. The scatter in these physical profiles, especially in the cluster centre region, shows a dependence on the cluster dynamical state and on the cool-core/non-cool-core dichotomy. The stellar age, metallicity and (s)SFR show very large scatter, which are then presented in 2D maps. We also do not find any clear radial dependence of these properties. However, the brightest central galaxies have distinguishable features compared to the properties of the satellite galaxies.

preprint2016arXiv

Large-scale Collaborative Imaging Genetics Studies of Risk Genetic Factors for Alzheimer's Disease Across Multiple Institutions

Genome-wide association studies (GWAS) offer new opportunities to identify genetic risk factors for Alzheimer's disease (AD). Recently, collaborative efforts across different institutions emerged that enhance the power of many existing techniques on individual institution data. However, a major barrier to collaborative studies of GWAS is that many institutions need to preserve individual data privacy. To address this challenge, we propose a novel distributed framework, termed Local Query Model (LQM) to detect risk SNPs for AD across multiple research institutions. To accelerate the learning process, we propose a Distributed Enhanced Dual Polytope Projection (D-EDPP) screening rule to identify irrelevant features and remove them from the optimization. To the best of our knowledge, this is the first successful run of the computationally intensive model selection procedure to learn a consistent model across different institutions without compromising their privacy while ranking the SNPs that may collectively affect AD. Empirical studies are conducted on 809 subjects with 5.9 million SNP features which are distributed across three individual institutions. D-EDPP achieved a 66-fold speed-up by effectively identifying irrelevant features.

preprint2014arXiv

Stochastic Coordinate Coding and Its Application for Drosophila Gene Expression Pattern Annotation

\textit{Drosophila melanogaster} has been established as a model organism for investigating the fundamental principles of developmental gene interactions. The gene expression patterns of \textit{Drosophila melanogaster} can be documented as digital images, which are annotated with anatomical ontology terms to facilitate pattern discovery and comparison. The automated annotation of gene expression pattern images has received increasing attention due to the recent expansion of the image database. The effectiveness of gene expression pattern annotation relies on the quality of feature representation. Previous studies have demonstrated that sparse coding is effective for extracting features from gene expression images. However, solving sparse coding remains a computationally challenging problem, especially when dealing with large-scale data sets and learning large size dictionaries. In this paper, we propose a novel algorithm to solve the sparse coding problem, called Stochastic Coordinate Coding (SCC). The proposed algorithm alternatively updates the sparse codes via just a few steps of coordinate descent and updates the dictionary via second order stochastic gradient descent. The computational cost is further reduced by focusing on the non-zero components of the sparse codes and the corresponding columns of the dictionary only in the updating procedure. Thus, the proposed algorithm significantly improves the efficiency and the scalability, making sparse coding applicable for large-scale data sets and large dictionary sizes. Our experiments on Drosophila gene expression data sets demonstrate the efficiency and the effectiveness of the proposed algorithm.