Source author record

Malgorzata Bogdan

Malgorzata Bogdan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

12works
13topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

12 published item(s)

preprint2024arXiv

Inferring the redshift of more than 150 GRBs with a Machine Learning Ensemble model

Gamma-Ray Bursts (GRBs), due to their high luminosities are detected up to redshift 10, and thus have the potential to be vital cosmological probes of early processes in the universe. Fulfilling this potential requires a large sample of GRBs with known redshifts, but due to observational limitations, only 11\% have known redshifts ($z$). There have been numerous attempts to estimate redshifts via correlation studies, most of which have led to inaccurate predictions. To overcome this, we estimated GRB redshift via an ensemble supervised machine learning model that uses X-ray afterglows of long-duration GRBs observed by the Neil Gehrels Swift Observatory. The estimated redshifts are strongly correlated (a Pearson coefficient of 0.93) and have a root mean square error, namely the square root of the average squared error $\langleΔz^2\rangle$, of 0.46 with the observed redshifts showing the reliability of this method. The addition of GRB afterglow parameters improves the predictions considerably by 63\% compared to previous results in peer-reviewed literature. Finally, we use our machine learning model to infer the redshifts of 154 GRBs, which increase the known redshifts of long GRBs with plateaus by 94\%, a significant milestone for enhancing GRB population studies that require large samples with redshift.

preprint2023arXiv

Shedding new light on the Hubble constant tension through Supernovae Ia

The standard cosmological model, the $Λ$CDM model, is the most suitable description for our universe. This framework can explain the accelerated expansion phase of the universe but still is not immune to open problems when it comes to the comparison with observations. One of the most critical issues is the so-called Hubble constant ($H_0$) tension, namely, the difference of about $5σ$ as an average between the value of $H_0$ estimated locally and the cosmological value measured from the Last Scattering Surface. The value of this tension changes from 4 to 6 $σ$ according to the data used. The current analysis explores the $H_0$ tension in the \textit{Pantheon} sample (PS) of SNe Ia. Through the division of the PS in 3 and 4 bins, the value of $H_0$ is estimated for each bin and all the values are fitted with a decreasing function of the redshift ($z$). Remarkably, $H_0$ undergoes a slow decreasing evolution with $z$, having an evolutionary coefficient compatible with zero up to $5.8σ$. If this trend is not caused by hidden astrophysical biases or $z$-selection effects, then the $f(R)$ modified theories of gravity represent a valid model for explaining such a trend.

preprint2022arXiv

On the evolution of the Hubble constant with the SNe Ia Pantheon Sample and Baryon Acoustic Oscillations: a feasibility study for GRB-cosmology in 2030

The difference from 4 to 6 $σ$ in the Hubble constant ($H_0$) between the values observed with the local (Cepheids and Supernovae Ia, SNe Ia) and the high-z probes (CMB obtained by the Planck data) still challenges the astrophysics and cosmology community. Previous analysis has shown that there is an evolution in the Hubble constant that scales as $f(z)=\mathcal{H}_{0}/(1+z)^η$, where $\mathcal{H}_0$ is $H_{0}(z=0)$ and $η$ is the evolutionary parameter. Here, we investigate if this evolution still holds by using the SNe Ia gathered in the Pantheon sample and the BAOs. We assume $H_{0}=70 \textrm{km s}^{-1}\textrm{Mpc}^{-1}$ as the local value and divide the Pantheon into 3 bins ordered in increasing values of redshift. Similar to our previous analysis but varying two cosmological parameters contemporaneously ($H_0$, $Ω_{0m}$ in the $Λ$CDM model and $H_0$, $w_a$ in the $w_{0}w_{a}$CDM model), for each bin we implement a MCMC analysis obtaining the value of $H_0$ $[...]$. Subsequently, the values of $H_0$ are fitted with the model $f(z)$. Our results show that a decreasing trend with $η\sim10^{-2}$ is still visible in this sample. The $η$ coefficient reaches zero in 2.0 $σ$ for the $Λ$CDM model up to 5.8 $σ$ for $w_{0}w_{a}$CDM model. This trend, if not due to statistical fluctuations, could be explained through a hidden astrophysical bias, such as the effect of stretch evolution, or it requires new theoretical models, a possible proposition is the modified gravity theories, $f(R)$ $[...]$. This work is also a preparatory to understand how the combined probes still show an evolution of the $H_0$ by redshift and what is the current status of simulations on GRB cosmology to obtain the uncertainties on the $Ω_{0m}$ comparable with the ones achieved through SNe Ia.

preprint2022arXiv

Predicting the redshift of gamma-ray loud AGNs using Supervised Machine Learning: Part 2

Measuring the redshift of active galactic nuclei (AGNs) requires the use of time-consuming and expensive spectroscopic analysis. However, obtaining redshift measurements of AGNs is crucial as it can enable AGN population studies, provide insight into the star formation rate, the luminosity function, and the density rate evolution. Hence, there is a requirement for alternative redshift measurement techniques. In this project, we aim to use the Fermi gamma-ray space telescope's 4LAC Data Release (DR2) catalog to train a machine learning model capable of predicting the redshift reliably. In addition, this project aims at improving and extending with the new 4LAC Catalog the predictive capabilities of the machine learning (ML) methodology published in Dainotti et al. (2021). Furthermore, we implement feature engineering to expand the parameter space and a bias correction technique to our final results. This study uses additional machine learning techniques inside the ensemble method, the SuperLearner, previously used in Dainotti et al.(2021). Additionally, we also test a novel ML model called Sorted L-One Penalized Estimation (SLOPE). Using these methods we provide a catalog of estimated redshift values for those AGNs that do not have a spectroscopic redshift measurement. These estimates can serve as a redshift reference for the community to verify as updated Fermi catalogs are released with more redshift measurements.

preprint2022arXiv

Using Multivariate Imputation by Chained Equations to Predict Redshifts of Active Galactic Nuclei

Redshift measurement of active galactic nuclei (AGNs) remains a time-consuming and challenging task, as it requires follow up spectroscopic observations and detailed analysis. Hence, there exists an urgent requirement for alternative redshift estimation techniques. The use of machine learning (ML) for this purpose has been growing over the last few years, primarily due to the availability of large-scale galactic surveys. However, due to observational errors, a significant fraction of these data sets often have missing entries, rendering that fraction unusable for ML regression applications. In this study, we demonstrate the performance of an imputation technique called Multivariate Imputation by Chained Equations (MICE), which rectifies the issue of missing data entries by imputing them using the available information in the catalog. We use the Fermi-LAT Fourth Data Release Catalog (4LAC) and impute 24% of the catalog. Subsequently, we follow the methodology described in Dainotti et al. (2021) and create an ML model for estimating the redshift of 4LAC AGNs. We present results which highlight positive impact of MICE imputation technique on the machine learning models performance and obtained redshift estimation accuracy.

preprint2016arXiv

Bayesian dimensionality reduction with PCA using penalized semi-integrated likelihood

We discuss the problem of estimating the number of principal components in Principal Com- ponents Analysis (PCA). Despite of the importance of the problem and the multitude of solutions proposed in the literature, it comes as a surprise that there does not exist a coherent asymptotic framework which would justify different approaches depending on the actual size of the data set. In this paper we address this issue by presenting an approximate Bayesian approach based on Laplace approximation and introducing a general method for building the model selection criteria, called PEnalized SEmi-integrated Likelihood (PESEL). Our general framework encompasses a variety of existing approaches based on probabilistic models, like e.g. Bayesian Information Criterion for the Probabilistic PCA (PPCA), and allows for construction of new criteria, depending on the size of the data set at hand. Specifically, we define PESEL when the number of variables substantially exceeds the number of observations. We also report results of extensive simulation studies and real data analysis, which illustrate good properties of our proposed criteria as compared to the state-of- the-art methods and very recent proposals. Specifially, these simulations show that PESEL based criteria can be quite robust against deviations from the probabilistic model assumptions. Selected PESEL based criteria for the estimation of the number of principal components are implemented in R package varclust, which is available on github (https://github.com/psobczyk/varclust).

preprint2016arXiv

False Discoveries Occur Early on the Lasso Path

In regression settings where explanatory variables have very low correlations and there are relatively few effects, each of large magnitude, we expect the Lasso to find the important variables with few errors, if any. This paper shows that in a regime of linear sparsity---meaning that the fraction of variables with a non-vanishing effect tends to a constant, however small---this cannot really be the case, even when the design variables are stochastically independent. We demonstrate that true features and null features are always interspersed on the Lasso path, and that this phenomenon occurs no matter how strong the effect sizes are. We derive a sharp asymptotic trade-off between false and true positive rates or, equivalently, between measures of type I and type II errors along the Lasso path. This trade-off states that if we ever want to achieve a type II error (false negative rate) under a critical value, then anywhere on the Lasso path the type I error (false positive rate) will need to exceed a given threshold so that we can never have both errors at a low level at the same time. Our analysis uses tools from approximate message passing (AMP) theory as well as novel elements to deal with a possibly adaptive selection of the Lasso regularizing parameter.

preprint2016arXiv

Fast Saddle-Point Algorithm for Generalized Dantzig Selector and FDR Control with the Ordered l1-Norm

In this paper we propose a primal-dual proximal extragradient algorithm to solve the generalized Dantzig selector (GDS) estimation problem, based on a new convex-concave saddle-point (SP) reformulation. Our new formulation makes it possible to adopt recent developments in saddle-point optimization, to achieve the optimal $O(1/k)$ rate of convergence. Compared to the optimal non-SP algorithms, ours do not require specification of sensitive parameters that affect algorithm performance or solution quality. We also provide a new analysis showing a possibility of local acceleration to achieve the rate of $O(1/k^2)$ in special cases even without strong convexity or strong smoothness. As an application, we propose a GDS equipped with the ordered $\ell_1$-norm, showing its false discovery rate control properties in variable selection. Algorithm performance is compared between ours and other alternatives, including the linearized ADMM, Nesterov's smoothing, Nemirovski's mirror-prox, and the accelerated hybrid proximal extragradient techniques.

preprint2016arXiv

Group SLOPE - adaptive selection of groups of predictors

Sorted L-One Penalized Estimation (SLOPE) is a relatively new convex optimization procedure which allows for adaptive selection of regressors under sparse high dimensional designs. Here we extend the idea of SLOPE to deal with the situation when one aims at selecting whole groups of explanatory variables instead of single regressors. Such groups can be formed by clustering strongly correlated predictors or groups of dummy variables corresponding to different levels of the same qualitative predictor. We formulate the respective convex optimization problem, gSLOPE (group SLOPE), and propose an efficient algorithm for its solution. We also define a notion of the group false discovery rate (gFDR) and provide a choice of the sequence of tuning parameters for gSLOPE so that gFDR is provably controlled at a prespecified level if the groups of variables are orthogonal to each other. Moreover, we prove that the resulting procedure adapts to unknown sparsity and is asymptotically minimax with respect to the estimation of the proportions of variance of the response variable explained by regressors from different groups. We also provide a method for the choice of the regularizing sequence when variables in different groups are not orthogonal but statistically independent and illustrate its good properties with computer simulations. Finally, we illustrate the advantages of gSLOPE in the context of Genome Wide Association Studies. R package grpSLOPE with implementation of our method is available on CRAN.

preprint2013arXiv

Statistical estimation and testing via the sorted L1 norm

We introduce a novel method for sparse regression and variable selection, which is inspired by modern ideas in multiple testing. Imagine we have observations from the linear model y = X beta + z, then we suggest estimating the regression coefficients by means of a new estimator called SLOPE, which is the solution to minimize 0.5 ||y - Xb\|_2^2 + lambda_1 |b|_(1) + lambda_2 |b|_(2) + ... + lambda_p |b|_(p); here, lambda_1 >= λ_2 >= ... >= λ_p >= 0 and |b|_(1) >= |b|_(2) >= ... >= |b|_(p) is the order statistic of the magnitudes of b. The regularizer is a sorted L1 norm which penalizes the regression coefficients according to their rank: the higher the rank, the larger the penalty. This is similar to the famous BHq procedure [Benjamini and Hochberg, 1995], which compares the value of a test statistic taken from a family to a critical threshold that depends on its rank in the family. SLOPE is a convex program and we demonstrate an efficient algorithm for computing the solution. We prove that for orthogonal designs with p variables, taking lambda_i = F^{-1}(1-q_i) (F is the cdf of the errors), q_i = iq/(2p), controls the false discovery rate (FDR) for variable selection. When the design matrix is nonorthogonal there are inherent limitations on the FDR level and the power which can be obtained with model selection methods based on L1-like penalties. However, whenever the columns of the design matrix are not strongly correlated, we demonstrate empirically that it is possible to select the parameters lambda_i as to obtain FDR control at a reasonable level as long as the number of nonzero coefficients is not too large. At the same time, the procedure exhibits increased power over the lasso, which treats all coefficients equally. The paper illustrates further estimation properties of the new selection rule through comprehensive simulation studies.

preprint2011arXiv

Asymptotic Bayes optimality under sparsity for generally distributed effect sizes under the alternative

Recent results concerning asymptotic Bayes-optimality under sparsity (ABOS) of multiple testing procedures are extended to fairly generally distributed effect sizes under the alternative. An asymptotic framework is considered where both the number of tests m and the sample size m go to infinity, while the fraction p of true alternatives converges to zero. It is shown that under mild restrictions on the loss function nontrivial asymptotic inference is possible only if n increases to infinity at least at the rate of log m. Based on this assumption precise conditions are given under which the Bonferroni correction with nominal Family Wise Error Rate (FWER) level alpha and the Benjamini- Hochberg procedure (BH) at FDR level alpha are asymptotically optimal. When n is proportional to log m then alpha can remain fixed, whereas when n increases to infinity at a quicker rate, then alpha has to converge to zero roughly like n^(-1/2). Under these conditions the Bonferroni correction is ABOS in case of extreme sparsity, while BH adapts well to the unknown level of sparsity. In the second part of this article these optimality results are carried over to model selection in the context of multiple regression with orthogonal regressors. Several modifications of Bayesian Information Criterion are considered, controlling either FWER or FDR, and conditions are provided under which these selection criteria are ABOS. Finally the performance of these criteria is examined in a brief simulation study.

preprint2010arXiv

A model selection approach to genome wide association studies

For the vast majority of genome wide association studies (GWAS) published so far, statistical analysis was performed by testing markers individually. In this article we present some elementary statistical considerations which clearly show that in case of complex traits the approach based on multiple regression or generalized linear models is preferable to multiple testing. We introduce a model selection approach to GWAS based on modifications of Bayesian Information Criterion (BIC) and develop some simple search strategies to deal with the huge number of potential models. Comprehensive simulations based on real SNP data confirm that model selection has larger power than multiple testing to detect causal SNPs in complex models. On the other hand multiple testing has substantial problems with proper ranking of causal SNPs and tends to detect a certain number of false positive SNPs, which are not linked to any of the causal mutations. We show that this behavior is typical in GWAS for complex traits and can be explained by an aggregated influence of many small random sample correlations between genotypes of a SNP under investigation and other causal SNPs. We believe that our findings at least partially explain problems with low power and nonreplicability of results in many real data GWAS. Finally, we discuss the advantages of our model selection approach in the context of real data analysis, where we consider publicly available gene expression data as traits for individuals from the HapMap project.