Source author record

Qiying Wang

Qiying Wang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2026arXiv

MiMo-V2-Flash Technical Report

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that interleaves Sliding Window Attention (SWA) with global attention, with a 128-token sliding window under a 5:1 hybrid ratio. The model is pre-trained on 27 trillion tokens with Multi-Token Prediction (MTP), employing a native 32k context length and subsequently extended to 256k. To efficiently scale post-training compute, MiMo-V2-Flash introduces a novel Multi-Teacher On-Policy Distillation (MOPD) paradigm. In this framework, domain-specialized teachers (e.g., trained via large-scale reinforcement learning) provide dense and token-level reward, enabling the student model to perfectly master teacher expertise. MiMo-V2-Flash rivals top-tier open-weight models such as DeepSeek-V3.2 and Kimi-K2, despite using only 1/2 and 1/3 of their total parameters, respectively. During inference, by repurposing MTP as a draft model for speculative decoding, MiMo-V2-Flash achieves up to 3.6 acceptance length and 2.6x decoding speedup with three MTP layers. We open-source both the model weights and the three-layer MTP weights to foster open research and community collaboration.

preprint2025arXiv

MiMo-Audio: Audio Language Models are Few-Shot Learners

Existing audio language models typically rely on task-specific fine-tuning to accomplish particular audio tasks. In contrast, humans are able to generalize to new audio tasks with only a few examples or simple instructions. GPT-3 has shown that scaling next-token prediction pretraining enables strong generalization capabilities in text, and we believe this paradigm is equally applicable to the audio domain. By scaling MiMo-Audio's pretraining data to over one hundred million of hours, we observe the emergence of few-shot learning capabilities across a diverse set of audio tasks. We develop a systematic evaluation of these capabilities and find that MiMo-Audio-7B-Base achieves SOTA performance on both speech intelligence and audio understanding benchmarks among open-source models. Beyond standard metrics, MiMo-Audio-7B-Base generalizes to tasks absent from its training data, such as voice conversion, style transfer, and speech editing. MiMo-Audio-7B-Base also demonstrates powerful speech continuation capabilities, capable of generating highly realistic talk shows, recitations, livestreaming and debates. At the post-training stage, we curate a diverse instruction-tuning corpus and introduce thinking mechanisms into both audio understanding and generation. MiMo-Audio-7B-Instruct achieves open-source SOTA on audio understanding benchmarks (MMSU, MMAU, MMAR, MMAU-Pro), spoken dialogue benchmarks (Big Bench Audio, MultiChallenge Audio) and instruct-TTS evaluations, approaching or surpassing closed-source models. Model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-Audio.

preprint2020arXiv

Locally trimmed least squares: conventional inference in possibly nonstationary models

A novel IV estimation method, that we term Locally Trimmed LS (LTLS), is developed which yields estimators with (mixed) Gaussian limit distributions in situations where the data may be weakly or strongly persistent. In particular, we allow for nonlinear predictive type of regressions where the regressor can be stationary short/long memory as well as nonstationary long memory process or a nearly integrated array. The resultant t-tests have conventional limit distributions (i.e. N(0; 1)) free of (near to unity and long memory) nuisance parameters. In the case where the regressor is a fractional process, no preliminary estimator for the memory parameter is required. Therefore, the practitioner can conduct inference while being agnostic about the exact dependence structure in the data. The LTLS estimator is obtained by applying certain chronological trimming to the OLS instrument via the utilisation of appropriate kernel functions of time trend variables. The finite sample performance of LTLS based t-tests is investigated with the aid of a simulation experiment. An empirical application to the predictability of stock returns is also provided.

preprint2014arXiv

Uniform convergence rates for a class of martingales with application in non-linear cointegrating regression

For a class of martingales, this paper provides a framework on the uniform consistency with broad applicability. The main condition imposed is only related to the conditional variance of the martingale, which holds true for stationary mixing time series, stationary iterated random function, Harris recurrent Markov chains and $I(1)$ processes with innovations being a linear process. Using the established results, this paper investigates the uniform convergence of the Nadaraya-Watson estimator in a non-linear cointegrating regression model. Our results not only provide sharp convergence rate, but also the optimal range for the uniform convergence to be held. This paper also considers the uniform upper and lower bound estimates for a functional of Harris recurrent Markov chain, which are of independent interests.

preprint2013arXiv

Self-normalized Cramér type moderate deviations for the maximum of sums

Let $X_1,X_2,...$ be independent random variables with zero means and finite variances, and let $S_n=\sum_{i=1}^nX_i$ and $V^2_n=\sum_{i=1}^nX^2_i$. A Cramér type moderate deviation for the maximum of the self-normalized sums $\max_{1\leq k\leq n}S_k/V_n$ is obtained. In particular, for identically distributed $X_1,X_2,...,$ it is proved that $P(\max_{1\leq k\leq n}S_k\geq xV_n)/(1-Φ(x))\rightarrow2$ uniformly for $0<x\leq\mathrm{o}(n^{1/6})$ under the optimal finite third moment of $X_1$.

preprint2012arXiv

A specification test for nonlinear nonstationary models

We provide a limit theory for a general class of kernel smoothed U-statistics that may be used for specification testing in time series regression with nonstationary data. The test framework allows for linear and nonlinear models with endogenous regressors that have autoregressive unit roots or near unit roots. The limit theory for the specification test depends on the self-intersection local time of a Gaussian process. A new weak convergence result is developed for certain partial sums of functions involving nonstationary time series that converges to the intersection local time process. This result is of independent interest and is useful in other applications. Simulations examine the finite sample performance of the test.

preprint2010arXiv

Strong approximations of level exceedences related to multiple hypothesis testing

Particularly in genomics, but also in other fields, it has become commonplace to undertake highly multiple Student's $t$-tests based on relatively small sample sizes. The literature on this topic is continually expanding, but the main approaches used to control the family-wise error rate and false discovery rate are still based on the assumption that the tests are independent. The independence condition is known to be false at the level of the joint distributions of the test statistics, but that does not necessarily mean, for the small significance levels involved in highly multiple hypothesis testing, that the assumption leads to major errors. In this paper, we give conditions under which the assumption of independence is valid. Specifically, we derive a strong approximation that closely links the level exceedences of a dependent ``studentized process'' to those of a process of independent random variables. Via this connection, it can be seen that in high-dimensional, low sample-size cases, provided the sample size diverges faster than the logarithm of the number of tests, the assumption of independent $t$-tests is often justified.