Source author record

Mihailo Stojnic

Mihailo Stojnic appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

36works
11topics
3close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

36 published item(s)

preprint2026arXiv

Parametric RDT approach to computational gap of symmetric binary perceptron

We study potential presence of statistical-computational gaps (SCG) in symmetric binary perceptrons (SBP) via a parametric utilization of \emph{fully lifted random duality theory} (fl-RDT) [96]. A structural change from decreasingly to arbitrarily ordered $c$-sequence (a key fl-RDT parametric component) is observed on the second lifting level and associated with \emph{satisfiability} ($α_c$) -- \emph{algorithmic} ($α_a$) constraints density threshold change thereby suggesting a potential existence of a nonzero computational gap $SCG=α_c-α_a$. The second level estimate is shown to match the theoretical $α_c$ whereas the $r\rightarrow \infty$ level one is proposed to correspond to $α_a$. For example, for the canonical SBP ($κ=1$ margin) we obtain $α_c\approx 1.8159$ on the second and $α_a\approx 1.6021$ (with converging tendency towards $\sim 1.59$ range) on the seventh level. Our propositions remarkably well concur with recent literature: (i) in [20] local entropy replica approach predicts $α_{LE}\approx 1.58$ as the onset of clustering defragmentation (presumed driving force behind locally improving algorithms failures); (ii) in $α\rightarrow 0$ regime we obtain on the third lifting level $κ\approx 1.2385\sqrt{\frac{α_a}{-\log\left ( α_a \right ) }}$ which qualitatively matches overlap gap property (OGP) based predictions of [43] and identically matches local entropy based predictions of [24]; (iii) $c$-sequence ordering change phenomenology mirrors the one observed in asymmetric binary perceptron (ABP) in [98] and the negative Hopfield model in [100]; and (iv) as in [98,100], we here design a CLuP based algorithm whose practical performance closely matches proposed theoretical predictions.

preprint2023arXiv

Causal Inference (C-inf) -- asymmetric scenario of typical phase transitions

In this paper, we revisit and further explore a mathematically rigorous connection between Causal inference (C-inf) and the Low-rank recovery (LRR) established in [10]. Leveraging the Random duality - Free probability theory (RDT-FPT) connection, we obtain the exact explicit typical C-inf asymmetric phase transitions (PT). We uncover a doubling low-rankness phenomenon, which means that exactly two times larger low rankness is allowed in asymmetric scenarios compared to the symmetric worst case ones considered in [10]. Consequently, the final PT mathematical expressions are as elegant as those obtained in [10], and highlight direct relations between the targeted C-inf matrix low rankness and the time of treatment. Our results have strong implications for applications, where C-inf matrices are not necessarily symmetric.

preprint2023arXiv

Causal Inference (C-inf) -- closed form worst case typical phase transitions

In this paper we establish a mathematically rigorous connection between Causal inference (C-inf) and the low-rank recovery (LRR). Using Random Duality Theory (RDT) concepts developed in [46,48,50] and novel mathematical strategies related to free probability theory, we obtain the exact explicit typical (and achievable) worst case phase transitions (PT). These PT precisely separate scenarios where causal inference via LRR is possible from those where it is not. We supplement our mathematical analysis with numerical experiments that confirm the theoretical predictions of PT phenomena, and further show that the two closely match for fairly small sample sizes. We obtain simple closed form representations for the resulting PTs, which highlight direct relations between the low rankness of the target C-inf matrix and the time of the treatment. Hence, our results can be used to determine the range of C-inf's typical applicability.

preprint2023arXiv

Exact Error in Matrix Completion: Approximately Low-Rank Structures and Missing Blocks

We study the completion of approximately low rank matrices with entries missing not at random (MNAR). In the context of typical large-dimensional statistical settings, we establish a framework for the performance analysis of the nuclear norm minimization ($\ell_1^*$) algorithm. Our framework produces \emph{exact} estimates of the worst-case residual root mean squared error and the associated phase transitions (PT), with both exhibiting remarkably simple characterizations. Our results enable to {\it precisely} quantify the impact of key system parameters, including data heterogeneity, size of the missing block, and deviation from ideal low rankness, on the accuracy of $\ell_1^*$-based matrix completion. To validate our theoretical worst-case RMSE estimates, we conduct numerical simulations, demonstrating close agreement with their numerical counterparts.

preprint2016arXiv

Box constrained $\ell_1$ optimization in random linear systems -- asymptotics

In this paper we consider box constrained adaptations of $\ell_1$ optimization heuristic when applied for solving random linear systems. These are typically employed when on top of being sparse the systems' solutions are also known to be confined in a specific way to an interval on the real axis. Two particular $\ell_1$ adaptations (to which we will refer as the \emph{binary} $\ell_1$ and \emph{box} $\ell_1$) will be discussed in great detail. Many of their properties will be addressed with a special emphasis on the so-called phase transitions (PT) phenomena and the large deviation principles (LDP). We will fully characterize these through two different mathematical approaches, the first one that is purely probabilistic in nature and the second one that connects to high-dimensional geometry. Of particular interest we will find that for many fairly hard mathematical problems a collection of pretty elegant characterizations of their final solutions will turn out to exist.

preprint2016arXiv

Box constrained $\ell_1$ optimization in random linear systems -- finite dimensions

Our companion work \cite{Stojnicl1BnBxasymldp} considers random under-determined linear systems with box-constrained sparse solutions and provides an asymptotic analysis of a couple of modified $\ell_1$ heuristics adjusted to handle such systems (we refer to these modifications of the standard $\ell_1$ as binary and box $\ell_1$). Our earlier work \cite{StojnicISIT2010binary} established that the binary $\ell_1$ does exhibit the so-called phase-transition phenomenon (basically the same phenomenon well-known through earlier considerations to be a key feature of the standard $\ell_1$, see, e.g. \cite{DonohoPol,DonohoUnsigned,StojnicCSetam09,StojnicUpper10}). Moreover, in \cite{StojnicISIT2010binary}, we determined the precise location of the co-called phase-transition (PT) curve. On the other hand, in \cite{Stojnicl1BnBxasymldp} we provide a much deeper understanding of the PTs and do so through a large deviations principles (LDP) type of analysis. In this paper we complement the results of \cite{Stojnicl1BnBxasymldp} by leaving the asymptotic regime naturally assumed in the PT and LDP considerations aside and instead working in a finite dimensional setting. Along the same lines, we provide for both, the binary and the box $\ell_1$, precise finite dimensional analyses and essentially determine their ultimate statistical performance characterizations. On top of that, we explain how the results created here can be utilized in the asymptotic setting, considered in \cite{Stojnicl1BnBxasymldp}, as well. Finally, for the completeness, we also present a collection of results obtained through numerical simulations and observe that they are in a massive agreement with our theoretical calculations.

preprint2016arXiv

Fully bilinear generic and lifted random processes comparisons

In our companion paper \cite{Stojnicgscomp16} we introduce a collection of fairly powerful statistical comparison results. They relate to a general comparison concept and its an upgrade that we call lifting procedure. Here we provide a different generic principle (which we call fully bilinear) that in certain cases turns out to be stronger than the corresponding one from \cite{Stojnicgscomp16}. Moreover, we also show how the principle that we introduce here can also be pushed through the lifting machinery of \cite{Stojnicgscomp16}. Finally, as was the case in \cite{Stojnicgscomp16}, here we also show how the well known Slepian's max and Gordon's minmax comparison principles can be obtained as special cases of the mechanisms that we present here. We also create their lifted upgrades which happen to be stronger than the corresponding ones in \cite{Stojnicgscomp16}. A fairly large collection of results obtained through numerical experiments is also provided. It is observed that these results are in an excellent agreement with what the theory predicts.

preprint2016arXiv

Generic and lifted probabilistic comparisons -- max replaces minmax

In this paper we introduce a collection of powerful statistical comparison results. We first present the results that we obtained while developing a general comparison concept. After that we introduce a separate lifting procedure that is a comparison concept on its own. We then show how in certain scenarios the lifting procedure basically represents a substantial upgrade over the general strategy. We complement the introduced results with a fairly large collection of numerical experiments that are in an overwhelming agreement with what the theory predicts. We also show how many well known comparison results (e.g. Slepian's max and Gordon's minmax principle) can be obtained as special cases. Moreover, it turns out that the minmax principle can be viewed as a single max principle as well. The range of applications is enormous. It starts with revisiting many of the results we created in recent years in various mathematical fields and recognizing that they are fully self-contained as their starting blocks are specialized variants of the concepts introduced here. Further upgrades relate to core comparison extensions on the one side and more practically oriented modifications on the other. Those that we deem the most important we discuss in several separate companion papers to ensure preserving the introductory elegance and simplicity of what is presented here.

preprint2016arXiv

Partial $\ell_1$ optimization in random linear systems -- finite dimensions

In this paper we provide a complementary set of results to those we present in our companion work \cite{Stojnicl1HidParasymldp} regarding the behavior of the so-called partial $\ell_1$ (a variant of the standard $\ell_1$ heuristic often employed for solving under-determined systems of linear equations). As is well known through our earlier works \cite{StojnicICASSP10knownsupp,StojnicTowBettCompSens13}, the partial $\ell_1$ also exhibits the phase-transition (PT) phenomenon, discovered and well understood in the context of the standard $\ell_1$ through Donoho's and our own works \cite{DonohoPol,DonohoUnsigned,StojnicCSetam09,StojnicUpper10}. \cite{Stojnicl1HidParasymldp} goes much further though and, in addition to the determination of the partial $\ell_1$'s phase-transition curves (PT curves) (which had already been done in \cite{StojnicICASSP10knownsupp,StojnicTowBettCompSens13}), provides a substantially deeper understanding of the PT phenomena through a study of the underlying large deviations principles (LDPs). As the PT and LDP phenomena are by their definitions related to large dimensional settings, both sets of our works, \cite{StojnicICASSP10knownsupp,StojnicTowBettCompSens13} and \cite{Stojnicl1HidParasymldp}, consider what is typically called the asymptotic regime. In this paper we move things in a different direction and consider finite dimensional scenarios. Basically, we provide explicit performance characterizations for any given collection of systems/parameters dimensions. We do so for two different variants of the partial $\ell_1$, one that we call exactly the partial $\ell_1$ and another one, possibly a bit more practical, that we call the hidden partial $\ell_1$.

preprint2016arXiv

Partial $\ell_1$ optimization in random linear systems -- phase transitions and large deviations

$\ell_1$ optimization is a well known heuristic often employed for solving various forms of sparse linear problems. In this paper we look at its a variant that we refer to as the \emph{partial} $\ell_1$ and discuss its mathematical properties when used for solving linear under-determined systems of equations. We will focus on large random systems and discuss the phase transition (PT) phenomena and how they connect to the large deviation principles (LDP). Using a variety of probabilistic and geometric techniques that we have developed in recent years we will first present general guidelines that conceptually fully characterize both, the PTs and the LDPs. After that we will put an emphasis on providing a collection of explicit analytical solutions to all of the underlying mathematical problems. As a nice bonus to the developed concepts, the forms of the analytical solutions will, in our view, turn out to be fairly elegant as well.

preprint2016arXiv

Random linear systems with sparse solutions -- asymptotics and large deviations

In this paper we revisit random linear under-determined systems with sparse solutions. We consider $\ell_1$ optimization heuristic known to work very well when used to solve these systems. A collection of fundamental results that relate to its performance analysis in a statistical scenario is presented. We start things off by recalling on now classical phase transition (PT) results that we derived in \cite{StojnicCSetam09,StojnicUpper10}. As these represent the so-called breaking point characterizations, we now complement them by analyzing the behavior in a zone around the breaking points in a sense typically used in the study of the large deviation properties (LDP) in the classical probability theory. After providing a conceptual solution to these problems we attack them on a "hardcore" mathematical level attempting/hoping to be able to obtain explicit solutions as elegant as those we obtained in \cite{StojnicCSetam09,StojnicUpper10} (this time around though, the final characterizations were to be expected to be way more involved than in \cite{StojnicCSetam09,StojnicUpper10}, simply, the ultimate goals are set much higher and their achieving would provide a much richer collection of information about the $\ell_1$'s behavior). Perhaps surprisingly, the final LDP $\ell_1$ characterizations that we obtain happen to match the elegance of the corresponding PT ones from \cite{StojnicCSetam09,StojnicUpper10}. Moreover, as we have done in \cite{StojnicEquiv10}, here we also present a corresponding LDP set of results that can be obtained through an alternative high-dimensional geometry approach. Finally, we also prove that the two types of characterizations, obtained through two substantially different mathematical approaches, match as one would hope that they do.

preprint2016arXiv

Random linear systems with sparse solutions -- finite dimensions

In our companion work \cite{Stojnicl1RegPosasymldp} we revisited random under-determined linear systems with sparse solutions. The main emphasis was on the performance analysis of the $\ell_1$ heuristic in the so-called asymptotic regime, i.e. in the regime where the systems' dimensions are large. Through an earlier sequence of work \cite{DonohoPol,DonohoUnsigned,StojnicCSetam09,StojnicUpper10}, it is now well known that in such a regime the $\ell_1$ exhibits the so-called \emph{phase transition} (PT) phenomenon. \cite{Stojnicl1RegPosasymldp} then went much further and established the so-called \emph{large deviations principle} (LDP) type of behavior that characterizes not only the breaking points of the $\ell_1$'s success but also the behavior in the entire so-called \emph{transition zone} around these points. Both of these concepts, the PTs and the LDPs, are in fact defined so that one can use them to characterize the asymptotic behavior. In this paper we complement the results of \cite{Stojnicl1RegPosasymldp} by providing an exact detailed analysis in the non-asymptotic regime. Of course, not only are the non-asymptotic results complementing those from \cite{Stojnicl1RegPosasymldp}, they actually are the ones that ultimately fully characterize the $\ell_1$'s behavior in the most general sense. We introduce several novel high-dimensional geometry type of strategies that enable us to eventually determine the $\ell_1$'s behavior.

preprint2016arXiv

Random linear under-determined systems with block-sparse solutions -- asymptotics, large deviations, and finite dimensions

In this paper we consider random linear under-determined systems with block-sparse solutions. A standard subvariant of such systems, namely, precisely the same type of systems without additional block structuring requirement, gained a lot of popularity over the last decade. This is of course in first place due to the success in mathematical characterization of an $\ell_1$ optimization technique typically used for solving such systems, initially achieved in \cite{CRT,DOnoho06CS} and later on perfected in \cite{DonohoPol,DonohoUnsigned,StojnicCSetam09,StojnicUpper10}. The success that we achieved in \cite{StojnicCSetam09,StojnicUpper10} characterizing the standard sparse solutions systems, we were then able to replicate in a sequence of papers \cite{StojnicCSetamBlock09,StojnicUpperBlock10,StojnicICASSP09block,StojnicJSTSP09} where instead of the standard $\ell_1$ optimization we utilized its an $\ell_2/\ell_1$ variant as a better fit for systems with block-sparse solutions. All of these results finally settled the so-called threshold/phase transitions phenomena (which naturally assume the asymptotic/large dimensional scenario). Here, in addition to a few novel asymptotic considerations, we also try to raise the level a bit, step a bit away from the asymptotics, and consider the finite dimensions scenarios as well.

preprint2015arXiv

Bounds on restricted isometry constants of random matrices

In this paper we look at isometry properties of random matrices. During the last decade these properties gained a lot attention in a field called compressed sensing in first place due to their initial use in \cite{CRT,CT}. Namely, in \cite{CRT,CT} these quantities were used as a critical tool in providing a rigorous analysis of $\ell_1$ optimization's ability to solve an under-determined system of linear equations with sparse solutions. In such a framework a particular type of isometry, called restricted isometry, plays a key role. One then typically introduces a couple of quantities, called upper and lower restricted isometry constants to characterize the isometry properties of random matrices. Those constants are then usually viewed as mathematical objects of interest and their a precise characterization is desirable. The first estimates of these quantities within compressed sensing were given in \cite{CRT,CT}. As the need for precisely estimating them grew further a finer improvements of these initial estimates were obtained in e.g. \cite{BCTsharp09,BT10}. These are typically obtained through a combination of union-bounding strategy and powerful tail estimates of extreme eigenvalues of Wishart (Gaussian) matrices (see, e.g. \cite{Edelman88}). In this paper we attempt to circumvent such an approach and provide an alternative way to obtain similar estimates.

preprint2015arXiv

Compressed sensing of block-sparse positive vectors

In this paper we revisit one of the classical problems of compressed sensing. Namely, we consider linear under-determined systems with sparse solutions. A substantial success in mathematical characterization of an $\ell_1$ optimization technique typically used for solving such systems has been achieved during the last decade. Seminal works \cite{CRT,DOnoho06CS} showed that the $\ell_1$ can recover a so-called linear sparsity (i.e. solve systems even when the solution has a sparsity linearly proportional to the length of the unknown vector). Later considerations \cite{DonohoPol,DonohoUnsigned} (as well as our own ones \cite{StojnicCSetam09,StojnicUpper10}) provided the precise characterization of this linearity. In this paper we consider the so-called structured version of the above sparsity driven problem. Namely, we view a special case of sparse solutions, the so-called block-sparse solutions. Typically one employs $\ell_2/\ell_1$-optimization as a variant of the standard $\ell_1$ to handle block-sparse case of sparse solution systems. We considered systems with block-sparse solutions in a series of work \cite{StojnicCSetamBlock09,StojnicUpperBlock10,StojnicICASSP09block,StojnicJSTSP09} where we were able to provide precise performance characterizations if the $\ell_2/\ell_1$-optimization similar to those obtained for the standard $\ell_1$ optimization in \cite{StojnicCSetam09,StojnicUpper10}. Here we look at a similar class of systems where on top of being block-sparse the unknown vectors are also known to have components of the same sign. In this paper we slightly adjust $\ell_2/\ell_1$-optimization to account for the known signs and provide a precise performance characterization of such an adjustment.

preprint2015arXiv

Linear under-determined systems with sparse solutions: Redirecting a challenge?

Seminal works \cite{CRT,DonohoUnsigned,DonohoPol} generated a massive interest in studying linear under-determined systems with sparse solutions. In this paper we give a short mathematical overview of what was accomplished in last 10 years in a particular direction of such a studying. We then discuss what we consider were the main challenges in last 10 years and give our own view as to what are the main challenges that lie ahead. Through the presentation we arrive to a point where the following natural rhetoric question arises: is it a time to redirect the main challenges? While we can not provide the answer to such a question we hope that our small discussion will stimulate further considerations in this direction.

preprint2015arXiv

Towards a better compressed sensing

In this paper we look at a well known linear inverse problem that is one of the mathematical cornerstones of the compressed sensing field. In seminal works \cite{CRT,DOnoho06CS} $\ell_1$ optimization and its success when used for recovering sparse solutions of linear inverse problems was considered. Moreover, \cite{CRT,DOnoho06CS} established for the first time in a statistical context that an unknown vector of linear sparsity can be recovered as a known existing solution of an under-determined linear system through $\ell_1$ optimization. In \cite{DonohoPol,DonohoUnsigned} (and later in \cite{StojnicCSetam09,StojnicUpper10}) the precise values of the linear proportionality were established as well. While the typical $\ell_1$ optimization behavior has been essentially settled through the work of \cite{DonohoPol,DonohoUnsigned,StojnicCSetam09,StojnicUpper10}, we in this paper look at possible upgrades of $\ell_1$ optimization. Namely, we look at a couple of algorithms that turn out to be capable of recovering a substantially higher sparsity than the $\ell_1$. However, these algorithms assume a bit of "feedback" to be able to work at full strength. This in turn then translates the original problem of improving upon $\ell_1$ to designing algorithms that would be able to provide output needed to feed the $\ell_1$ upgrades considered in this papers.

preprint2015arXiv

Upper-bounding $\ell_1$-optimization sectional thresholds

In this paper we look at a particular problem related to under-determined linear systems of equations with sparse solutions. $\ell_1$-minimization is a fairly successful polynomial technique that can in certain statistical scenarios find sparse enough solutions of such systems. Barriers of $\ell_1$ performance are typically referred to as its thresholds. Depending if one is interested in a typical or worst case behavior one then distinguishes between the \emph{weak} thresholds that relate to a typical behavior on one side and the \emph{sectional} and \emph{strong} thresholds that relate to the worst case behavior on the other side. Starting with seminal works \cite{CRT,DonohoPol,DOnoho06CS} a substantial progress has been achieved in theoretical characterization of $\ell_1$-minimization statistical thresholds. More precisely, \cite{CRT,DOnoho06CS} presented for the first time linear lower bounds on all of these thresholds. Donoho's work \cite{DonohoPol} (and our own \cite{StojnicCSetam09,StojnicUpper10}) went a bit further and essentially settled the $\ell_1$'s \emph{weak} thresholds. At the same time they also provided fairly good lower bounds on the values on the \emph{sectional} and \emph{strong} thresholds. In this paper, we revisit the \emph{sectional} thresholds and present a simple mechanism that can be used to create solid upper bounds as well. The method we present relies on a seemingly simple but substantial progress we made in studying Hopfield models in \cite{StojnicHopBnds10}.

preprint2013arXiv

A framework to characterize performance of LASSO algorithms

In this paper we consider solving \emph{noisy} under-determined systems of linear equations with sparse solutions. A noiseless equivalent attracted enormous attention in recent years, above all, due to work of \cite{CRT,CanRomTao06,DonohoPol} where it was shown in a statistical and large dimensional context that a sparse unknown vector (of sparsity proportional to the length of the vector) can be recovered from an under-determined system via a simple polynomial $\ell_1$-optimization algorithm. \cite{CanRomTao06} further established that even when the equations are \emph{noisy}, one can, through an SOCP noisy equivalent of $\ell_1$, obtain an approximate solution that is (in an $\ell_2$-norm sense) no further than a constant times the noise from the sparse unknown vector. In our recent works \cite{StojnicCSetam09,StojnicUpper10}, we created a powerful mechanism that helped us characterize exactly the performance of $\ell_1$ optimization in the noiseless case (as shown in \cite{StojnicEquiv10} and as it must be if the axioms of mathematics are well set, the results of \cite{StojnicCSetam09,StojnicUpper10} are in an absolute agreement with the corresponding exact ones from \cite{DonohoPol}). In this paper we design a mechanism, as powerful as those from \cite{StojnicCSetam09,StojnicUpper10}, that can handle the analysis of a LASSO type of algorithm (and many others) that can be (or typically are) used for "solving" noisy under-determined systems. Using the mechanism we then, in a statistical context, compute the exact worst-case $\ell_2$ norm distance between the unknown sparse vector and the approximate one obtained through such a LASSO. The obtained results match the corresponding exact ones obtained in \cite{BayMon10,DonMalMon10}. Moreover, as a by-product of our analysis framework we recognize existence of an SOCP type of algorithm that achieves the same performance.

preprint2013arXiv

A performance analysis framework for SOCP algorithms in noisy compressed sensing

Solving under-determined systems of linear equations with sparse solutions attracted enormous amount of attention in recent years, above all, due to work of \cite{CRT,CanRomTao06,DonohoPol}. In \cite{CRT,CanRomTao06,DonohoPol} it was rigorously shown for the first time that in a statistical and large dimensional context a linear sparsity can be recovered from an under-determined system via a simple polynomial $\ell_1$-optimization algorithm. \cite{CanRomTao06} went even further and established that in \emph{noisy} systems for any linear level of under-determinedness there is again a linear sparsity that can be \emph{approximately} recovered through an SOCP (second order cone programming) noisy equivalent to $\ell_1$. Moreover, the approximate solution is (in an $\ell_2$-norm sense) guaranteed to be no further from the sparse unknown vector than a constant times the noise. In this paper we will also consider solving \emph{noisy} linear systems and present an alternative statistical framework that can be used for their analysis. To demonstrate how the framework works we will show how one can use it to precisely characterize the approximation error of a wide class of SOCP algorithms. We will also show that our theoretical predictions are in a solid agrement with the results one can get through numerical simulations.

preprint2013arXiv

A rigorous geometry-probability equivalence in characterization of $\ell_1$-optimization

In this paper we consider under-determined systems of linear equations that have sparse solutions. This subject attracted enormous amount of interest in recent years primarily due to influential works \cite{CRT,DonohoPol}. In a statistical context it was rigorously established for the first time in \cite{CRT,DonohoPol} that if the number of equations is smaller than but still linearly proportional to the number of unknowns then a sparse vector of sparsity also linearly proportional to the number of unknowns can be recovered through a polynomial $\ell_1$-optimization algorithm (of course, this assuming that such a sparse solution vector exists). Moreover, the geometric approach of \cite{DonohoPol} produced the exact values for the proportionalities in question. In our recent work \cite{StojnicCSetam09} we introduced an alternative statistical approach that produced attainable values of the proportionalities. Those happened to be in an excellent numerical agreement with the ones of \cite{DonohoPol}. In this paper we give a rigorous analytical confirmation that the results of \cite{StojnicCSetam09} indeed match those from \cite{DonohoPol}.

preprint2013arXiv

Another look at the Gardner problem

In this paper we revisit one of the classical perceptron problems from the neural networks and statistical physics. In \cite{Gar88} Gardner presented a neat statistical physics type of approach for analyzing what is now typically referred to as the Gardner problem. The problem amounts to discovering a statistical behavior of a spherical perceptron. Among various quantities \cite{Gar88} determined the so-called storage capacity of the corresponding neural network and analyzed its deviations as various perceptron parameters change. In a more recent work \cite{SchTir02,SchTir03} many of the findings of \cite{Gar88} (obtained on the grounds of the statistical mechanics replica approach) were proven to be mathematically correct. In this paper, we take another look at the Gardner problem and provide a simple alternative framework for its analysis. As a result we reprove many of now known facts and rigorously reestablish a few other results.

preprint2013arXiv

Asymmetric Little model and its ground state energies

In this paper we look at a class of random optimization problems that arise in the forms typically known in statistical physics as Little models. In \cite{BruParRit92} the Little models were studied by means of the well known tool from the statistical physics called the replica theory. A careful consideration produced a physically sound conclusion that the behavior of almost all important features of the Little models essentially resembles the behavior of the corresponding ones of appropriately scaled Sherrington-Kirkpatrick (SK) model. In this paper we revisit the Little models and consider their ground state energies as one of their main features. We then rigorously show that they indeed can be lower-bounded by the corresponding ones related to the SK model. We also provide a mathematically rigorous way to show that the replica symmetric estimate of the ground state energy is in fact a rigorous upper-bound of the ground state energy. Moreover, we then recall on a set of powerful mechanisms we recently designed for a study of the Hopfield models in \cite{StojnicHopBnds10,StojnicMoreSophHopBnds10} and show how one can utilize them to substantially lower the upper-bound that the replica symmetry theory provides.

preprint2013arXiv

Bounding ground state energy of Hopfield models

In this paper we look at a class of random optimization problems that arise in the forms typically known as Hopfield models. We view two scenarios which we term as the positive Hopfield form and the negative Hopfield form. For both of these scenarios we define the binary optimization problems that essentially emulate what would typically be known as the ground state energy of these models. We then present a simple mechanism that can be used to create a set of theoretical rigorous bounds for these energies. In addition to purely theoretical bounds, we also present a couple of fast optimization algorithms that can also be used to provide solid (albeit a bit weaker) algorithmic bounds for the ground state energies.

preprint2013arXiv

Discrete perceptrons

Perceptrons have been known for a long time as a promising tool within the neural networks theory. The analytical treatment for a special class of perceptrons started in seminal work of Gardner \cite{Gar88}. Techniques initially employed to characterize perceptrons relied on a statistical mechanics approach. Many of such predictions obtained in \cite{Gar88} (and in a follow-up \cite{GarDer88}) were later on established rigorously as mathematical facts (see, e.g. \cite{SchTir02,SchTir03,TalBook,StojnicGardGen13,StojnicGardSphNeg13,StojnicGardSphErr13}). These typically related to spherical perceptrons. A lot of work has been done related to various other types of perceptrons. Among the most challenging ones are what we will refer to as the discrete perceptrons. An introductory statistical mechanics treatment of such perceptrons was given in \cite{GutSte90}. Relying on results of \cite{Gar88}, \cite{GutSte90} characterized many of the features of several types of discrete perceptrons. We in this paper, consider a similar subclass of discrete perceptrons and provide a mathematically rigorous set of results related to their performance. As it will turn out, many of the statistical mechanics predictions obtained for discrete predictions will in fact appear as mathematically provable bounds. This will in a way emulate a similar type of behavior we observed in \cite{StojnicGardGen13,StojnicGardSphNeg13,StojnicGardSphErr13} when studying spherical perceptrons.

preprint2013arXiv

Lifting $\ell_1$-optimization strong and sectional thresholds

In this paper we revisit under-determined linear systems of equations with sparse solutions. As is well known, these systems are among core mathematical problems of a very popular compressed sensing field. The popularity of the field as well as a substantial academic interest in linear systems with sparse solutions are in a significant part due to seminal results \cite{CRT,DonohoPol}. Namely, working in a statistical scenario, \cite{CRT,DonohoPol} provided substantial mathematical progress in characterizing relation between the dimensions of the systems and the sparsity of unknown vectors recoverable through a particular polynomial technique called $\ell_1$-minimization. In our own series of work \cite{StojnicCSetam09,StojnicUpper10,StojnicEquiv10} we also provided a collection of mathematical results related to these problems. While, Donoho's work \cite{DonohoPol,DonohoUnsigned} established (and our own work \cite{StojnicCSetam09,StojnicUpper10,StojnicEquiv10} reaffirmed) the typical or the so-called \emph{weak threshold} behavior of $\ell_1$-minimization many important questions remain unanswered. Among the most important ones are those that relate to non-typical or the so-called \emph{strong threshold} behavior. These questions are usually combinatorial in nature and known techniques come up short of providing the exact answers. In this paper we provide a powerful mechanism that that can be used to attack the "tough" scenario, i.e. the \emph{strong threshold} (and its a similar form called \emph{sectional threshold}) of $\ell_1$-minimization.

preprint2013arXiv

Lifting $\ell_q$-optimization thresholds

In this paper we look at a connection between the $\ell_q,0\leq q\leq 1$, optimization and under-determined linear systems of equations with sparse solutions. The case $q=1$, or in other words $\ell_1$ optimization and its a connection with linear systems has been thoroughly studied in last several decades; in fact, especially so during the last decade after the seminal works \cite{CRT,DOnoho06CS} appeared. While current understanding of $\ell_1$ optimization-linear systems connection is fairly known, much less so is the case with a general $\ell_q,0<q<1$, optimization. In our recent work \cite{StojnicLqThrBnds10} we provided a study in this direction. As a result we were able to obtain a collection of lower bounds on various $\ell_q,0\leq q\leq 1$, optimization thresholds. In this paper, we provide a substantial conceptual improvement of the methodology presented in \cite{StojnicLqThrBnds10}. Moreover, the practical results in terms of achievable thresholds are also encouraging. As is usually the case with these and similar problems, the methodology we developed emphasizes their a combinatorial nature and attempts to somehow handle it. Although our results' main contributions should be on a conceptual level, they already give a very strong suggestion that $\ell_q$ optimization can in fact provide a better performance than $\ell_1$, a fact long believed to be true due to a tighter optimization relaxation it provides to the original $\ell_0$ sparsity finding oriented original problem formulation. As such, they in a way give a solid boost to further exploration of the design of the algorithms that would be able to handle $\ell_q,0<q<1$, optimization in a reasonable (if not polynomial) time.

preprint2013arXiv

Lifting/lowering Hopfield models ground state energies

In our recent work \cite{StojnicHopBnds10} we looked at a class of random optimization problems that arise in the forms typically known as Hopfield models. We viewed two scenarios which we termed as the positive Hopfield form and the negative Hopfield form. For both of these scenarios we defined the binary optimization problems whose optimal values essentially emulate what would typically be known as the ground state energy of these models. We then presented a simple mechanisms that can be used to create a set of theoretical rigorous bounds for these energies. In this paper we create a way more powerful set of mechanisms that can substantially improve the simple bounds given in \cite{StojnicHopBnds10}. In fact, the mechanisms we create in this paper are the first set of results that show that convexity type of bounds can be substantially improved in this type of combinatorial problems.

preprint2013arXiv

Meshes that trap random subspaces

In our recent work \cite{StojnicCSetam09,StojnicUpper10} we considered solving under-determined systems of linear equations with sparse solutions. In a large dimensional and statistical context we proved results related to performance of a polynomial $\ell_1$-optimization technique when used for solving such systems. As one of the tools we used a probabilistic result of Gordon \cite{Gordon88}. In this paper we revisit this classic result in its core form and show how it can be reused to in a sense prove its own optimality.

preprint2013arXiv

Negative spherical perceptron

In this paper we consider the classical spherical perceptron problem. This problem and its variants have been studied in a great detail in a broad literature ranging from statistical physics and neural networks to computer science and pure geometry. Among the most well known results are those created using the machinery of statistical physics in \cite{Gar88}. They typically relate to various features ranging from the storage capacity to typical overlap of the optimal configurations and the number of incorrectly stored patterns. In \cite{SchTir02,SchTir03,TalBook} many of the predictions of the statistical mechanics were rigorously shown to be correct. In our own work \cite{StojnicGardGen13} we then presented an alternative way that can be used to study the spherical perceptrons as well. Among other things we reaffirmed many of the results obtained in \cite{SchTir02,SchTir03,TalBook} and thereby confirmed many of the predictions established by the statistical mechanics. Those mostly relate to spherical perceptrons with positive thresholds (which we will typically refer to as the positive spherical perceptrons). In this paper we go a step further and attack the negative counterpart, i.e. the perceptron with negative thresholds. We present a mechanism that can be used to analyze many features of such a model. As a concrete example, we specialize our results for a particular feature, namely the storage capacity. The results we obtain for the storage capacity seem to indicate that the negative case could be more combinatorial in nature and as such a somewhat harder challenge than the positive counterpart.

preprint2013arXiv

Optimality of $\ell_2/\ell_1$-optimization block-length dependent thresholds

The recent work of \cite{CRT,DonohoPol} rigorously proved (in a large dimensional and statistical context) that if the number of equations (measurements in the compressed sensing terminology) in the system is proportional to the length of the unknown vector then there is a sparsity (number of non-zero elements of the unknown vector) also proportional to the length of the unknown vector such that $\ell_1$-optimization algorithm succeeds in solving the system. In more recent papers \cite{StojnicCSetamBlock09,StojnicICASSP09block,StojnicJSTSP09} we considered under-determined systems with the so-called \textbf{block}-sparse solutions. In a large dimensional and statistical context in \cite{StojnicCSetamBlock09} we determined lower bounds on the values of allowable sparsity for any given number (proportional to the length of the unknown vector) of equations such that an $\ell_2/\ell_1$-optimization algorithm succeeds in solving the system. These lower bounds happened to be in a solid numerical agreement with what one can observe through numerical experiments. Here we derive the corresponding upper bounds. Moreover, the upper bounds that we obtain in this paper match the lower bounds from \cite{StojnicCSetamBlock09} and ultimately make them optimal.

preprint2013arXiv

Regularly random duality

In this paper we look at a class of random optimization problems. We discuss ways that can help determine typical behavior of their solutions. When the dimensions of the optimization problems are large such an information often can be obtained without actually solving the original problems. Moreover, we also discover that fairly often one can actually determine many quantities of interest (such as, for example, the typical optimal values of the objective functions) completely analytically. We present a few general ideas and emphasize that the range of applications is enormous.

preprint2013arXiv

Spherical perceptron as a storage memory with limited errors

It has been known for a long time that the classical spherical perceptrons can be used as storage memories. Seminal work of Gardner, \cite{Gar88}, started an analytical study of perceptrons storage abilities. Many of the Gardner's predictions obtained through statistical mechanics tools have been rigorously justified. Among the most important ones are of course the storage capacities. The first rigorous confirmations were obtained in \cite{SchTir02,SchTir03} for the storage capacity of the so-called positive spherical perceptron. These were later reestablished in \cite{TalBook} and a bit more recently in \cite{StojnicGardGen13}. In this paper we consider a variant of the spherical perceptron that operates as a storage memory but allows for a certain fraction of errors. In Gardner's original work the statistical mechanics predictions in this directions were presented sa well. Here, through a mathematically rigorous analysis, we confirm that the Gardner's predictions in this direction are in fact provable upper bounds on the true values of the storage capacity. Moreover, we then present a mechanism that can be used to lower these bounds. Numerical results that we present indicate that the Garnder's storage capacity predictions may, in a fairly wide range of parameters, be not that far away from the true values.

preprint2013arXiv

Under-determined linear systems and $\ell_q$-optimization thresholds

Recent studies of under-determined linear systems of equations with sparse solutions showed a great practical and theoretical efficiency of a particular technique called $\ell_1$-optimization. Seminal works \cite{CRT,DOnoho06CS} rigorously confirmed it for the first time. Namely, \cite{CRT,DOnoho06CS} showed, in a statistical context, that $\ell_1$ technique can recover sparse solutions of under-determined systems even when the sparsity is linearly proportional to the dimension of the system. A followup \cite{DonohoPol} then precisely characterized such a linearity through a geometric approach and a series of work\cite{StojnicCSetam09,StojnicUpper10,StojnicEquiv10} reaffirmed statements of \cite{DonohoPol} through a purely probabilistic approach. A theoretically interesting alternative to $\ell_1$ is a more general version called $\ell_q$ (with an essentially arbitrary $q$). While $\ell_1$ is typically considered as a first available convex relaxation of sparsity norm $\ell_0$, $\ell_q,0\leq q\leq 1$, albeit non-convex, should technically be a tighter relaxation of $\ell_0$. Even though developing polynomial (or close to be polynomial) algorithms for non-convex problems is still in its initial phases one may wonder what would be the limits of an $\ell_q,0\leq q\leq 1$, relaxation even if at some point one can develop algorithms that could handle its non-convexity. A collection of answers to this and a few realted questions is precisely what we present in this paper.

preprint2013arXiv

Upper-bounding $\ell_1$-optimization weak thresholds

In our recent work \cite{StojnicCSetam09} we considered solving under-determined systems of linear equations with sparse solutions. In a large dimensional and statistical context we proved that if the number of equations in the system is proportional to the length of the unknown vector then there is a sparsity (number of non-zero elements of the unknown vector) also proportional to the length of the unknown vector such that a polynomial $\ell_1$-optimization technique succeeds in solving the system. We provided lower bounds on the proportionality constants that are in a solid numerical agreement with what one can observe through numerical experiments. Here we create a mechanism that can be used to derive the upper bounds on the proportionality constants. Moreover, the upper bounds obtained through such a mechanism match the lower bounds from \cite{StojnicCSetam09} and ultimately make the latter ones optimal.

preprint2008arXiv

On the reconstruction of block-sparse signals with an optimal number of measurements

Let A be an M by N matrix (M < N) which is an instance of a real random Gaussian ensemble. In compressed sensing we are interested in finding the sparsest solution to the system of equations A x = y for a given y. In general, whenever the sparsity of x is smaller than half the dimension of y then with overwhelming probability over A the sparsest solution is unique and can be found by an exhaustive search over x with an exponential time complexity for any y. The recent work of Candés, Donoho, and Tao shows that minimization of the L_1 norm of x subject to A x = y results in the sparsest solution provided the sparsity of x, say K, is smaller than a certain threshold for a given number of measurements. Specifically, if the dimension of y approaches the dimension of x, the sparsity of x should be K < 0.239 N. Here, we consider the case where x is d-block sparse, i.e., x consists of n = N / d blocks where each block is either a zero vector or a nonzero vector. Instead of L_1-norm relaxation, we consider the following relaxation min x \| X_1 \|_2 + \| X_2 \|_2 + ... + \| X_n \|_2, subject to A x = y where X_i = (x_{(i-1)d+1}, x_{(i-1)d+2}, ..., x_{i d}) for i = 1,2, ..., N. Our main result is that as n -> \infty, the minimization finds the sparsest solution to Ax = y, with overwhelming probability in A, for any x whose block sparsity is k/n < 1/2 - O(ε), provided M/N > 1 - 1/d, and d = Ω(\log(1/ε)/ε). The relaxation can be solved in polynomial time using semi-definite programming.