Source author record

Zhengyang Zhao

Zhengyang Zhao appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
6topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2026arXiv

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams, whereas existing benchmarks largely fail to evaluate this capability. They restrict videos to short clips, isolate modalities, or reduce questions to one-hop perception. We introduce TraceAV-Bench, the first benchmark to jointly evaluate multi-hop reasoning over long audio-visual trajectories and multimodal hallucination robustness. TraceAV-Bench comprises 2,200 rigorously validated multiple-choice questions over 578 long videos, totaling 339.5 hours, spanning 4 evaluation dimensions and 15 sub-tasks. Each question is grounded in an explicit reasoning chain that averages 3.68 hops across a 15.1-minute temporal span. The dataset is built by a three-step semi-automated pipeline followed by a strict quality assurance process. Evaluation of multiple representative OmniLLMs on TraceAV-Bench reveals that the benchmark poses a persistent challenge across all models, with the strongest closed-source model (Gemini 3.1 Pro) reaching only 68.29% on general tasks, and the best open-source model (Ming-Flash-Omni-2.0) reaching 51.70%, leaving substantial headroom. Moreover, we find that robustness to multimodal hallucination is largely decoupled from general multimodal reasoning performance. We anticipate that TraceAV-Bench will stimulate further research toward OmniLLMs that can reason coherently and faithfully over long-form audio-visual content.

preprint2026arXiv

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

Inference-time harnesses substantially improve large language models on complex reasoning tasks. However, the intrinsic capabilities of the underlying model remain unchanged by the addition of these external workflows. To bridge this gap, we introduce \emph{On-Policy Harness Self-Distillation} (OPHSD), which employs the harness-augmented current model as a teacher for self-distillation, thereby introducing extra supervisory signals from the harness beyond training data. OPHSD internalizes task-specific harness capabilities into the student model, yielding robust generalizability and strong standalone performance across diverse reasoning tasks. Evaluated across draft--verify harness for text classification and plan--solve for mathematical reasoning tasks, OPHSD consistently outperforms strong baselines (e.g., +10.83\% over OPSD on HMMT25). Our analysis further indicates that reattaching the harness during inference yields no additional benefits and can even degrade performance, suggesting that complex harnesses need not always be permanent fixtures; instead, they can serve as temporary training scaffolds whose benefits are permanently fed back into the base model. Our code and training data are available at https://github.com/zzy1127/OPHSD-On-Policy-Harness-Self-Distillation.

preprint2021arXiv

LightXML: Transformer with Dynamic Negative Sampling for High-Performance Extreme Multi-label Text Classification

Extreme Multi-label text Classification (XMC) is a task of finding the most relevant labels from a large label set. Nowadays deep learning-based methods have shown significant success in XMC. However, the existing methods (e.g., AttentionXML and X-Transformer etc) still suffer from 1) combining several models to train and predict for one dataset, and 2) sampling negative labels statically during the process of training label ranking model, which reduces both the efficiency and accuracy of the model. To address the above problems, we proposed LightXML, which adopts end-to-end training and dynamic negative labels sampling. In LightXML, we use generative cooperative networks to recall and rank labels, in which label recalling part generates negative and positive labels, and label ranking part distinguishes positive labels from these labels. Through these networks, negative labels are sampled dynamically during label ranking part training by feeding with the same text representation. Extensive experiments show that LightXML outperforms state-of-the-art methods in five extreme multi-label datasets with much smaller model size and lower computational complexity. In particular, on the Amazon dataset with 670K labels, LightXML can reduce the model size up to 72% compared to AttentionXML.

preprint2016arXiv

Giant Voltage Manipulation of MgO-based Magnetic Tunnel Junctions via Localized Anisotropic Strain: a Potential Pathway to Ultra-Energy-Efficient Memory Technology

Strain-mediated voltage control of magnetization in piezoelectric/ferromagnetic systems is a promising mechanism to implement energy-efficient spintronic memory devices. Here, we demonstrate giant voltage manipulation of MgO magnetic tunnel junctions (MTJ) on a Pb(Mg1/3Nb2/3)0.7Ti0.3O3 (PMN-PT) piezoelectric substrate with (001) orientation. It is found that the magnetic easy axis, switching field, and the tunnel magnetoresistance (TMR) of the MTJ can be efficiently controlled by strain from the underlying piezoelectric layer upon the application of a gate voltage. Repeatable voltage controlled MTJ toggling between high/low-resistance states is demonstrated. More importantly, instead of relying on the intrinsic anisotropy of the piezoelectric substrate to generate the required strain, we utilize anisotropic strain produced using local gating scheme, which is scalable and amenable to practical memory applications. Additionally, the adoption of crystalline MgO-based MTJ on piezoelectric layer lends itself to high TMR in the strain-mediated MRAM devices.

preprint2015arXiv

Spin Hall Switching of the Magnetization in Ta/TbFeCo Structures with Bulk Perpendicular Anisotropy

Spin-orbit torques are studied in Ta/TbFeCo patterned structures with a bulk perpendicular magnetic anisotropy (bulk-PMA) for the first time. The current-induced magnetization switching is investigated in the presence of a perpendicular, longitudinal, or transverse field. In order to rule out Joule heating effect, switching of the magnetization is also demonstrated using current pulses. It is found that the anti-damping torque correlated with spin Hall effect is very strong, and a spin Hall angle of about 0.12 is obtained. The field-like torque related with Rashba effect is negligible in this structure suggesting that the interface play a significant role in Rashba-like torque.

preprint2014arXiv

Room Temperature Spin Pumping in Topological Insulator Bi2Se3

Three-dimensional (3D) topological insulators are known for their strong spin-orbit coupling and the existence of spin-textured topological surface states which could be potentially exploited for spintronics. Here, we investigate spin pumping from a metallic ferromagnet (CoFeB) into a 3D topological insulator (Bi2Se3) and demonstrate successful spin injection from CoFeB into Bi2Se3 and the direct detection of the electromotive force generated by the inverse spin Hal effect (ISHE) at room temperature. The spin pumping, driven by the magnetization dynamics of the metallic ferromagnet, introduces a spin current into the topological insulator layer, resulting in a broadening of the ferromagnetic resonance (FMR) linewidth. We find that the FMR linewidth more than quintuples, the spin mixing conductance can be as large as 3.4*10^20m^-2 and the spin Hall angle can be as large as 0.23 in the Bi2Se3 layer.