Source author record

Maohua Li

Maohua Li appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

5works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

5 published item(s)

preprint2026arXiv

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training or on heuristic token eviction, creating an undesirable trade-off among efficiency, training cost, and accuracy. In this work, we show that full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation. Our approach is built on three observations: (1) only a small subset of attention heads truly requires full long-context processing; (2) long-range retrieval is governed primarily by a low-dimensional subspace, allowing relevant tokens to be retrieved efficiently with a 16-dimensional indexer; and (3) the useful token budget is strongly query-dependent, making dynamic top-$p$ selection more suitable than fixed top-$k$ sparsification. Based on these insights, we propose RTPurbo, which retains the full KV cache only for retrieval heads and introduces a lightweight token indexer for sparse attention. By exploiting the model's intrinsic sparsity, RTPurbo achieves sparsification with only a few hundred training steps. Experiments on long-context benchmarks and reasoning tasks show that RTPurbo preserves near-lossless accuracy while delivering substantial efficiency gains, including up to a 9.36$\times$ prefill speedup at 1M context and about a 2.01$\times$ decode speedup. These results suggest that strong sparse inference can be obtained from standard full-attention training without expensive native sparse pretraining.

preprint2015arXiv

Ghost symmetry of the discrete KP hierarchy

In this paper, with the help of the $S$ function and ghost symmetry for the discrete KP hierarchy which is a semi-discrete version of the KP hierarchy, the ghost flow on its eigenfunction(adjoint eigenfunction) and the spectral representation of its Baker-Akhiezer function and adjoint Baker-Akhiezer function are derived. From these observations above, some important distinctions between the discrete KP hierarchy and KP hierarchy are shown. Also we give the ghost flow on the tau function and another kind of proof of the ASvM formula of the discrete KP hierarchy.

preprint2014arXiv

The compatibility of additional symmetry and gauge transformations for the constrained discrete Kadomtsev-Petviashvili hierarchy

In this paper, the compatibility between the gauge transformations and the additional symmetry of the constrained discrete Kadomtsev-Petviashvili hierarchy is given, which preserving the form of the additional symmetry of the cdKP hierarchy, up to shifting of the corresponding additional flows by ordinary time flows.

preprint2012arXiv

The Recursion operators of the BKP hierarchy and the CKP Hierarchy

In this paper, under the constraints of the BKP(CKP) hierarchy, a crucial observation is that the odd dynamical variable $u_{2k+1}$ can be explicitly expressed by the even dynamical variable $u_{2k}$ in the Lax operator $L$ through a new operator $B$. Using operator $B$, the essential differences between the BKP hierarchy and the CKP hierarchy are given by the flow equations and the recursion operators under the $(2n+1)$-reduction. The formal formulas of the recursion operators for the BKP and CKP hierarchy under $(2n+1)$-reduction are given. To illustrate this method, the two recursion operators are constructed explicitly for the 3-reduction of the BKP and CKP hierarchies. The $t_7$ flows of $u_2$ are generated from $t_1$ flows by the above recursion operators, which are consistent with the corresponding flows generated by the flow equations under 3-reduction.