Source author record

Rongrong Wang

Rongrong Wang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

14works
11topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

14 published item(s)

preprint2026arXiv

Toward Global Large Language Models in Medicine

Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed the landscape of medicine and holds promise for improving health care quality and expanding access to medical information globally. However, existing LLMs are primarily trained on high-resource languages, limiting their applicability in global medical scenarios. To address this gap, we constructed GlobMed, a large multilingual medical dataset, containing over 500,000 entries spanning 12 languages, including four low-resource languages. Building on this, we established GlobMed-Bench, which systematically assesses 56 state-of-the-art proprietary and open-weight LLMs across multiple multilingual medical tasks, revealing significant performance disparities across languages, particularly for low-resource languages. Additionally, we introduced GlobMed-LLMs, a suite of multilingual medical LLMs trained on GlobMed, with parameters ranging from 1.7B to 8B. GlobMed-LLMs achieved an average performance improvement of over 40% relative to baseline models, with a more than threefold increase in performance on low-resource languages. Together, these resources provide an important foundation for advancing the equitable development and application of LLMs globally, enabling broader language communities to benefit from technological advances.

preprint2022arXiv

Perturbation of invariant subspaces for ill-conditioned eigensystem

Given a diagonalizable matrix $A$, we study the stability of its invariant subspaces when its matrix of eigenvectors is ill-conditioned. Let $\mathcal{X}_1$ be some invariant subspace of $A$ and $X_1$ be the matrix storing the right eigenvectors that spanned $\mathcal{X}_1$. It is generally believed that when the condition number $κ_2(X_1)$ gets large, the corresponding invariant subspace $\mathcal{X}_1$ will become unstable to perturbation. This paper proves that this is not always the case. Specifically, we show that the growth of $κ_2(X_1)$ alone is not enough to destroy the stability. As a direct application, our result ensures that when $A$ gets closer to a Jordan form, one may still estimate its invariant subspaces from the noisy data stably.

preprint2022arXiv

Sigma Delta quantization for images

In signal quantization, it is well-known that introducing adaptivity to quantization schemes can improve their stability and accuracy in quantizing bandlimited signals. However, adaptive quantization has only been designed for one-dimensional signals. The contribution of this paper is two-fold: i). we propose the first family of two-dimensional adaptive quantization schemes that maintain the same mathematical and practical merits as their one-dimensional counterparts, and ii). we show that both the traditional 1-dimensional and the new 2-dimensional quantization schemes can effectively quantize signals with jump discontinuities. These results immediately enable the usage of adaptive quantization on images. Under mild conditions, we show that the adaptivity is able to reduce the reconstruction error of images from the presently best $O(\sqrt P)$ to the much smaller $O(\sqrt s)$, where $s$ is the number of jump discontinuities in the image and $P$ ($P\gg s$) is the total number of samples. This $\sqrt{P/s}$-fold error reduction is achieved via applying a total variation norm regularized decoder, whose formulation is inspired by the mathematical super-resolution theory in the field of compressed sensing. Compared to the super-resolution setting, our error reduction is achieved without requiring adjacent spikes/discontinuities to be well-separated, which ensures its broad scope of application. We numerically demonstrate the efficacy of the new scheme on medical and natural images. We observe that for images with small pixel intensity values, the new method can significantly increase image quality over the state-of-the-art method.

preprint2020arXiv

A Hierarchical User Intention-Habit Extract Network for Credit Loan Overdue Risk Detection

More personal consumer loan products are emerging in mobile banking APP. For ease of use, application process is always simple, which means that few application information is requested for user to fill when applying for a loan, which is not conducive to construct users' credit profile. Thus, the simple application process brings huge challenges to the overdue risk detection, as higher overdue rate will result in greater economic losses to the bank. In this paper, we propose a model named HUIHEN (Hierarchical User Intention-Habit Extract Network) that leverages the users' behavior information in mobile banking APP. Due to the diversity of users' behaviors, we divide behavior sequences into sessions according to the time interval, and use the field-aware method to extract the intra-field information of behaviors. Then, we propose a hierarchical network composed of time-aware GRU and user-item-aware GRU to capture users' short-term intentions and users' long-term habits, which can be regarded as a supplement to user profile. The proposed model can improve the accuracy without increasing the complexity of the original online application process. Experimental results demonstrate the superiority of HUIHEN and show that HUIHEN outperforms other state-of-art models on all datasets.

preprint2016arXiv

From compressed sensing to compressed bit-streams: practical encoders, tractable decoders

Compressed sensing is now established as an effective method for dimension reduction when the underlying signals are sparse or compressible with respect to some suitable basis or frame. One important, yet under-addressed problem regarding the compressive acquisition of analog signals is how to perform quantization. This is directly related to the important issues of how "compressed" compressed sensing is (in terms of the total number of bits one ends up using after acquiring the signal) and ultimately whether compressed sensing can be used to obtain compressed representations of suitable signals. Building on our recent work, we propose a concrete and practicable method for performing "analog-to-information conversion". Following a compressive signal acquisition stage, the proposed method consists of a quantization stage, based on $ΣΔ$ (sigma-delta) quantization, and a subsequent encoding (compression) stage that fits within the framework of compressed sensing seamlessly. We prove that, using this method, we can convert analog compressive samples to compressed digital bitstreams and decode using tractable algorithms based on convex optimization. We prove that the proposed AIC provides a nearly optimal encoding of sparse and compressible signals. Finally, we present numerical experiments illustrating the effectiveness of the proposed analog-to-information converter.

preprint2016arXiv

Sigma Delta quantization with Harmonic frames and partial Fourier ensembles

Sigma Delta quantization, a quantization method which first surfaced in the 1960s, has now been used widely in various digital products such as cameras, cell phones, radars, etc. The method samples an input signal at a rate higher than the Nyquist rate, thus achieves great robustness to quantization noise. Compressed Sensing (CS) is a frugal acquisition method that utilizes the possible sparsity of the signals to reduce the required number of samples for a lossless acquisition. One can deem the reduced number as an effective dimensionality of the set of sparse signals and accordingly, define an effective oversampling rate as the ratio between the actual sampling rate and the effective dimensionality. A natural conjecture is that the error of Sigma Delta quantization, previously shown to decay with the vanilla oversampling rate, should now decay with the effective oversampling rate when carried out in the regime of compressed sensing. Confirming this intuition is one of the main goals in this direction. The study of quantization in CS has so far been limited to proving error convergence results for Gaussian and sub-Gaussian sensing matrices, as the number of bits and/or the number of samples grow to infinity. In this paper, we provide a first result for the more realistic Fourier sensing matrices. The major idea is to randomly permute the Fourier samples before feeding them into the quantizer. We show that the random permutation can effectively increase the low frequency power of the measurements, thus enhance the quality of $ΣΔ$ quantization.

preprint2016arXiv

The gap between the null space property and the restricted isometry property

The null space property (NSP) and the restricted isometry property (RIP) are two properties which have received considerable attention in the compressed sensing literature. As the name suggests, NSP is a property that depends solely on the null space of the measurement procedure and as such, any two matrices which have the same null space will have NSP if either one of them does. On the other hand, RIP is a property of the measurement procedure itself, and given an RIP matrix it is straightforward to construct another matrix with the same null space that is not RIP. %Furthermore, RIP is known to imply NSP and therefore RIP is a strictly stronger assumption than NSP. We say a matrix is RIP-NSP if it has the same null space as an RIP matrix. We show that such matrices can provide robust recovery of compressible signals under Basis pursuit which in many applicable settings is comparable to the guarantee that RIP provides. More importantly, we constructively show that the RIP-NSP is stronger than NSP with the aid of this robust recovery result, which shows that RIP is fundamentally stronger than NSP.

preprint2015arXiv

Quantization of compressive samples with stable and robust recovery

In this paper we study the quantization stage that is implicit in any compressed sensing signal acquisition paradigm. We propose using Sigma-Delta quantization and a subsequent reconstruction scheme based on convex optimization. We prove that the reconstruction error due to quantization decays polynomially in the number of measurements. Our results apply to arbitrary signals, including compressible ones, and account for measurement noise. Additionally, they hold for sub-Gaussian (including Gaussian and Bernoulli) random compressed sensing measurements, as well as for both high bit-depth and coarse quantizers, and they extend to 1-bit quantization. In the noise-free case, when the signal is strictly sparse we prove that by optimizing the order of the quantization scheme one can obtain root-exponential decay in the reconstruction error due to quantization.

preprint2015arXiv

Restricted isometry property of random subdictionaries

We study statistical restricted isometry, a property closely related to sparse signal recovery, of deterministic sensing matrices of size $m \times N$. A matrix is said to have a statistical restricted isometry property (StRIP) of order $k$ if most submatrices with $k$ columns define a near-isometric map of ${\mathbb R}^k$ into ${\mathbb R}^m$. As our main result, we establish sufficient conditions for the StRIP property of a matrix in terms of the mutual coherence and mean square coherence. We show that for many existing deterministic families of sampling matrices, $m=O(k)$ rows suffice for $k$-StRIP, which is an improvement over the known estimates of either $m = Θ(k \log N)$ or $m = Θ(k\log k)$. We also give examples of matrix families that are shown to have the StRIP property using our sufficient conditions.

preprint2014arXiv

Measures of scalability

Scalable frames are frames with the property that the frame vectors can be rescaled resulting in tight frames. However, if a frame is not scalable, one has to aim for an approximate procedure. For this, in this paper we introduce three novel quantitative measures of the closeness to scalability for frames in finite dimensional real Euclidean spaces. Besides the natural measure of scalability given by the distance of a frame to the set of scalable frames, another measure is obtained by optimizing a quadratic functional, while the third is given by the volume of the ellipsoid of minimal volume containing the symmetrized frame. After proving that these measures are equivalent in a certain sense, we establish bounds on the probability of a randomly selected frame to be scalable. In the process, we also derive new necessary and sufficient conditions for a frame to be scalable.

preprint2013arXiv

A null space analysis of the L1 synthesis method in dictionary-based compressed sensing

An interesting topic in compressed sensing aims to recover signals with sparse representations in a dictionary. Recently the performance of the L1-analysis method has been a focus, while some fundamental problems for the L1-synthesis method are still unsolved. For example, what are the conditions for it to stably recover compressible signals under noise? Whether coherent dictionaries allow the existence of sensing matrices that guarantee good performances of the L1-synthesis method? To answer these questions, we build up a framework for the L1-synthesis method. In particular, we propose a dictionary-based null space property DNSP which, to the best of our knowledge, is the first sufficient and necessary condition for the success of L1-synthesis without measurement noise. With this new property, we show that when the dictionary D is full spark, it cannot be too coherent otherwise the method fails for all sensing matrices. We also prove that in the real case, DNSP is equivalent to the stability of L1-synthesis under noise.

preprint2013arXiv

A null space property approach to compressed sensing with frames

An interesting topic in compressive sensing concerns problems of sensing and recovering signals with sparse representations in a dictionary. In this note, we study conditions of sensing matrices A for the L1-synthesis method to accurately recover sparse, or nearly sparse signals in a given dictionary D. In particular, we propose a dictionary based null space property (D-NSP) which, to the best of our knowledge, is the first sufficient and necessary condition for the success of the L1 recovery. This new property is then utilized to detect some of those dictionaries whose sparse families cannot be compressed universally. Moreover, when the dictionary is full spark, we show that AD being NSP, which is well-known to be only sufficient for stable recovery via L1-synthesis method, is indeed necessary as well.

preprint2013arXiv

Random Subdictionaries and Coherence Conditions for Sparse Signal Recovery

The most frequently used condition for sampling matrices employed in compressive sampling is the restricted isometry (RIP) property of the matrix when restricted to sparse signals. At the same time, imposing this condition makes it difficult to find explicit matrices that support recovery of signals from sketches of the optimal (smallest possible)dimension. A number of attempts have been made to relax or replace the RIP property in sparse recovery algorithms. We focus on the relaxation under which the near-isometry property holds for most rather than for all submatrices of the sampling matrix, known as statistical RIP or StRIP condition. We show that sampling matrices of dimensions $m\times N$ with maximum coherence $μ=O((k\log^3 N)^{-1/4})$ and mean square coherence $\bar μ^2=O(1/(k\log N))$ support stable recovery of $k$-sparse signals using Basis Pursuit. These assumptions are satisfied in many examples. As a result, we are able to construct sampling matrices that support recovery with low error for sparsity $k$ higher than $\sqrt m,$ which exceeds the range of parameters of the known classes of RIP matrices.