Source author record

Weifeng Zhao

Weifeng Zhao appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Artificial Intelligence Computation and Language eess.AS math.NA Numerical Analysis physics.comp-ph Sound

Catalog footprint

What is connected

4works

7topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the first expressive end-to-end (E2E) spoken language model (SLM) that goes beyond natural conversation to support both role-playing and singing generation. VITA-QinYu adopts a hybrid speech-text paradigm that extends interleaved text-audio modeling with multi-codebook audio tokens, a design enabling richer paralinguistic representation while preserving a clear separation between modalities to avoid interference. We further develop a comprehensive data generation pipeline to synthesize a total of 15.8K hours of natural conversation, role-playing, and singing data for training. VITA-QinYu demonstrates superior expressiveness, outperforming peer SLMs by 7 percentage points on objective role-playing benchmarks, and surpassing peer models by 0.13 points on a 5-point MOS scale for singing. Simultaneously, it achieves state-of-the-art conversational accuracy and fluency, exceeding prior SLMs by 1.38 and 4.98 percentage points on the C3 and URO benchmarks, respectively. We open-source our code and models and provide an easy-to-use demo with full-stack support for streaming and full-duplex interaction.

preprint2022arXiv

KaraTuner: Towards end to end natural pitch correction for singing voice in karaoke

An automatic pitch correction system typically includes several stages, such as pitch extraction, deviation estimation, pitch shift processing, and cross-fade smoothing. However, designing these components with strategies often requires domain expertise and they are likely to fail on corner cases. In this paper, we present KaraTuner, an end-to-end neural architecture that predicts pitch curve and resynthesizes the singing voice directly from the tuned pitch and vocal spectrum extracted from the original recordings. Several vital technical points have been introduced in KaraTuner to ensure pitch accuracy, pitch naturalness, timbre consistency, and sound quality. A feed-forward Transformer is employed in the pitch predictor to capture longterm dependencies in the vocal spectrum and musical note. We also develop a pitch-controllable vocoder based on a novel source-filter block and the Fre-GAN architecture. KaraTuner obtains a higher preference than the rule-based pitch correction approach through A/B tests, and perceptual experiments show that the proposed vocoder achieves significant advantages in timbre consistency and sound quality compared with the parametric WORLD vocoder, phase vocoder and CLPC vocoder.

preprint2020arXiv

Boundary treatment of high order Runge-Kutta methods for hyperbolic conservation laws

In \cite{ZH2019}, we developed a boundary treatment method for implicit-explicit (IMEX) Runge-Kutta (RK) methods for solving hyperbolic systems with source terms. Since IMEX RK methods include explicit ones as special cases, this boundary treatment method naturally applies to explicit methods as well. In this paper, we examine this boundary treatment method for the case of explicit RK schemes of arbitrary order applied to hyperbolic conservation laws. We show that the method not only preserves the accuracy of explicit RK schemes but also possesses good stability. This compares favourably to the inverse Lax-Wendroff method in \cite{TS2010,TWSN2012} where analysis and numerical experiments have previously verified the presence of order reduction \cite{TS2010,TWSN2012}. In addition, we demonstrate that our method performs well for strong-stability-preserving (SSP) RK schemes involving negative coefficients and downwind spatial discretizations. It is numerically shown that when boundary conditions are present and the proposed boundary treatment is used, that SSP RK schemes with negative coefficients still allow for larger time steps than schemes with all non-negative coefficients. In this regard, our boundary treatment method is an effective supplement to SSP RK schemes with/without negative coefficients for initial-boundary value problems for hyperbolic conservation laws.

preprint2019arXiv

Relaxation-rate formula for the entropic lattice Boltzmann method

An elegant and uniform relaxation-rate formula is presented for the entropic lattice Boltzmann method (ELBM). The formula not only guarantees the discrete time H-theorem at numerical level but also gives full consideration to the consistency with hydrodynamics. With this novel formula, the computational cost of the ELBM is significantly reduced and the method now can be efficiently used for a broad range of hydrodynamics applications including high Renolds number flows. Moreover, we demonstrate that the grid points where flow fields change drastically are effectively marked by the formula.