Source author record

Ke Cheng

Ke Cheng appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2026arXiv

Differentially Private Subspace Fine-Tuning for Large Language Models

Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differential privacy (DP) offers rigorous privacy guarantees and has been widely adopted in fine-tuning; however, naively injecting noise across the high-dimensional parameter space creates perturbations with large norms, degrading performance and destabilizing training. To address this issue, we propose DP-SFT, a two-stage subspace fine-tuning method that substantially reduces noise magnitude while preserving formal DP guarantees. Our intuition is that, during fine-tuning, significant parameter updates lie within a low-dimensional, task-specific subspace, while other directions change minimally. Hence, we only inject DP noise into this subspace to protect privacy without perturbing irrelevant parameters. In phase one, we identify the subspace by analyzing principal gradient directions to capture task-specific update signals. In phase two, we project full gradients onto this subspace, add DP noise, and map the perturbed gradients back to the original parameter space for model updates, markedly lowering noise impact. Experiments on multiple datasets demonstrate that DP-SFT enhances accuracy and stability under rigorous DP constraints, accelerates convergence, and achieves substantial gains over DP fine-tuning baselines.

preprint2026arXiv

RelayGR: Scaling Long-Sequence Generative Recommendation via Cross-Stage Relay-Race Inference

Real-time recommender systems execute multi-stage cascades (retrieval, pre-processing, fine-grained ranking) under strict tail-latency SLOs, leaving only tens of milliseconds for ranking. Generative recommendation (GR) models can improve quality by consuming long user-behavior sequences, but in production their online sequence length is tightly capped by the ranking-stage P99 budget. We observe that the majority of GR tokens encode user behaviors that are independent of the item candidates, suggesting an opportunity to pre-infer a user-behavior prefix once and reuse it during ranking rather than recomputing it on the critical path. Realizing this idea at industrial scale is non-trivial: the prefix cache must survive across multiple pipeline stages before the final ranking instance is determined, the user population implies cache footprints far beyond a single device, and indiscriminate pre-inference would overload shared resources under high QPS. We present RelayGR, a production system that enables in-HBM relay-race inference for GR. RelayGR selectively pre-infers long-term user prefixes, keeps their KV caches resident in HBM over the request lifecycle, and ensures the subsequent ranking can consume them without remote fetches. RelayGR combines three techniques: 1) a sequence-aware trigger that admits only at-risk requests under a bounded cache footprint and pre-inference load, 2) an affinity-aware router that co-locates cache production and consumption by routing both the auxiliary pre-infer signal and the ranking request to the same instance, and 3) a memory-aware expander that uses server-local DRAM to capture short-term cross-request reuse while avoiding redundant reloads. We implement RelayGR on Huawei Ascend NPUs and evaluate it with real queries. Under a fixed P99 SLO, RelayGR supports up to 1.5$\times$ longer sequences and improves SLO-compliant throughput by up to 3.6$\times$.

preprint2020arXiv

Compact Global Descriptor for Neural Networks

Long-range dependencies modeling, widely used in capturing spatiotemporal correlation, has shown to be effective in CNN dominated computer vision tasks. Yet neither stacks of convolutional operations to enlarge receptive fields nor recent nonlocal modules is computationally efficient. In this paper, we present a generic family of lightweight global descriptors for modeling the interactions between positions across different dimensions (e.g., channels, frames). This descriptor enables subsequent convolutions to access the informative global features with negligible computational complexity and parameters. Benchmark experiments show that the proposed method can complete state-of-the-art long-range mechanisms with a significant reduction in extra computing cost. Code available at https://github.com/HolmesShuan/Compact-Global-Descriptor.

preprint2015arXiv

Study of time resolution by digital methods with DRS4 system

A new Digital Pulse Processing (DPP) system, based on domino ring sampler version 4 (DRS4), with good time resolution for the LaBr3 detectors has been developed and different digital timing analysis methods for processing the detector raw signals are reported. The system, composed of an eight channels DRS4 chip, was used as the readout electronic and acquisition system to process the outputs signals from XP20D0 Photomultiplier Tubes (PMTs). The PMTs were coupled with LaBr3 scintillator and placed on opposite side of the radioactive positron 22Na source for 511keV gama-ray test. By analyzing the raw data acquired by the system, the best coincidence timing resolution is about 194.7ps (FWHM), obtained by the digital constant fraction discrimination (dCFD) method and better than the other digital methods and analysis method based on the conventional analog systems. The results indicate that it is a promising approach to better localize the positron annihilation in the positron emission tomography (PET) with time of flight (TOF) and they are also suitable for the scintillation timing measurement, such as in TOF-DeltaE and TOF-E systems for particle identification, with picosecond accuracy timing measurement. Furthermore, this system is more simple and convenient in comparison with other systems.

preprint2014arXiv

Application of the DRS4 Chip for GHz Waveform Digitizing Circuit

At present, fast waveform digitizing circuit is more and more employed in modern physics experiments for processing the signals from an array detector. A new fast waveform sampling digitizing circuit developed by us is presented in this paper. Different with the traditional waveform digitizing circuit constructed with analog to digital converter(ADC) or time to digital converter(TDC), it is developed based on domino ring sampler(DRS), a switched capacitor array(SCA) chip. A DRS4 chip is used as a core device in our circuit, which has a fast sampling rate up to five gigabit samples per second (GSPS). The circuit has advantages of high resolution, low cost, low power dissipation, high channel density and small size. The quite satisfactory results are acquired by the preliminary performance test of this circuit board. Eight channels can be provided by one board, which has a 1-volt input dynamic range for each channel. The circuit linearity is better than 0.1%, the noise is less than 0.5 mV (root mean square, RMS), and its time resolution is about 50ps. The several boards can be cascaded to construct a multi-board system. The good performances make the circuit board to be used not only for physics experiments, but also for other applications.

preprint2011arXiv

Strong Visible Absorption and Photoluminescence of Titanic Acid Nanotubes by Hydrothermal Method

Titanic acid nanotubes (with a chemical formula H2Ti2O4(OH)2, abbreviated as TANTs) were synthesized by the hydrothermal method using commercial TiO2 nanoparticle powder (P25, Degussa, Germany) including anatase and rutile phase as a starting material. Conversion from nanoparticles to nanotubes was achieved by treating the nanoparticle powder with 10 M NaOH aqueous solution. Absorption and photoluminescence (PL) data indicate that the nanotubes obtained under slow and suitable drying and heating conditions had very strong and stable visible absorption with three peaks at 515, 575, and 675 nm and photoluminescence at room temperature in air.