Source author record

Feng Deng

Feng Deng appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Computer Vision Sound Artificial Intelligence Computation and Language cond-mat.mtrl-sci eess.AS Machine Learning Multimedia physics.chem-ph

Catalog footprint

What is connected

3works

9topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

Apollo: Unified Multi-Task Audio-Video Joint Generation

Audio-video joint generation has progressed rapidly, yet substantial challenges still remain. Non-commercial approaches still suffer audio-visual asynchrony, poor lip-speech alignment, and unimodal degradation, which can be stemmed from weak audio-visual correspondence modeling, limited generalization, and scarce high-quality dense-caption data. To address these issues, we introduce Apollo and delve into three axes--model architecture, training strategy, and data curation. Architecturally, we adopt a single-tower design with unified DiT blocks and an Omni-Full Attention mechanism, achieving tight audio-visual alignment and strong scalability. Training-wise, we adopt a progressive multitask regime--random modality masking to joint optimization across tasks, and a multistage curriculum, yielding robust representations, strengthening A-V aligned world knowledge, and preventing unimodal collapse. For datasets, we present the first large-scale audio-video dataset with dense captions, and introduce a novel automated data-construction pipeline which annotates and filters millions of diverse, high-quality, strictly aligned audio-video-caption triplets. Building on this, Apollo scales to large datasets, delivering high-fidelity, semantically and temporally aligned, instruction-following generation in both joint and unimodal settings while generalizing robustly to out-of-distribution scenarios. Across tasks, it substantially outperforms prior methods by a large margin and achieves performance comparable to Veo 3, offering a unified, scalable path toward next-generation audio-video synthesis.

preprint2021arXiv

SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification

Recently, x-vector has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances. Improvement upon the x-vector has been an active research area, and enormous neural networks have been elaborately designed based on the x-vector, eg, extended TDNN (E-TDNN), factorized TDNN (F-TDNN), and densely connected TDNN (D-TDNN). In this work, we try to identify the optimal architectures from a TDNN based search space employing neural architecture search (NAS), named SpeechNAS. Leveraging the recent advances in the speaker recognition, such as high-order statistics pooling, multi-branch mechanism, D-TDNN and angular additive margin softmax (AAM) loss with a minimum hyper-spherical energy (MHE), SpeechNAS automatically discovers five network architectures, from SpeechNAS-1 to SpeechNAS-5, of various numbers of parameters and GFLOPs on the large-scale text-independent speaker recognition dataset VoxCeleb1. Our derived best neural network achieves an equal error rate (EER) of 1.02% on the standard test set of VoxCeleb1, which surpasses previous TDNN based state-of-the-art approaches by a large margin. Code and trained weights are in https://github.com/wentaozhu/speechnas.git

preprint2016arXiv

Discovery of homogeneously dispersed pentacoordinated Al(V) species on the surface of amorphous silica-alumina

The dispersion and coordination of aluminium species on the surface of silica-alumina based materials are essential for controlling their catalytic activity and selectivity. Al(IV) and Al(VI) are two common coordinations of Al species in the silica network and alumina phase, respectively. Al(V) is rare in nature and was found hitherto only in the alumina phase or interfaces containing alumina, a behavior which negatively affects the dispersion, population, and accessibility of Al(V) species on the silica-alumina surface. This constraint has limited the development of silica-alumina based catalysts, particularly because Al(V) had been confirmed to act as a highly active center for acid reactions and single-atom catalysts. Here, we report the direct observation of high population of homogenously dispersed Al(V) species in amorphous silica-alumina in the absence of any bulk alumina phase, by high resolution TEM/EDX and high magnetic-field MAS NMR. Solid-state 27Al multi-quantum MAS NMR experiments prove unambiguously that most of the Al(V) species formed independently from the alumina phase and are accessible on the surface for guest molecules. These species are mainly transferred to Al(VI) species with partial formation of Al(IV) species after adsorption of water. The NMR chemical shifts and their coordination transformation with and without water adsorption are matching that obtained in DFT calculations of the predicted clusters. The discovery presented in this study not only provides fundamental knowledge of the nature of aluminum coordination, but also paves the way for developing highly efficient catalysts.