Source author record

Zhong Guan

Zhong Guan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

10works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

10 published item(s)

preprint2026arXiv

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training

In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.

preprint2026arXiv

D-VLA: A High-Concurrency Distributed Asynchronous Reinforcement Learning Framework for Vision-Language-Action Models

The rapid evolution of Embodied AI has enabled Vision-Language-Action (VLA) models to excel in multimodal perception and task execution. However, applying Reinforcement Learning (RL) to these massive models in large-scale distributed environments faces severe systemic bottlenecks, primarily due to the resource conflict between high-fidelity physical simulation and the intensive VRAM/bandwidth demands of deep learning. This conflict often leaves overall throughput constrained by execution-phase inefficiencies. To address these challenges, we propose D-VLA, a high-concurrency, low-latency distributed RL framework for large-scale embodied foundation models. D-VLA introduces "Plane Decoupling," physically isolating high-frequency training data from low-frequency weight control to eliminate interference between simulation and optimization. We further design a four-thread asynchronous "Swimlane" pipeline, enabling full parallel overlap of sampling, inference, gradient computation, and parameter distribution. Additionally, a dual-pool VRAM management model and topology-aware replication resolve memory fragmentation and optimize communication efficiency. Experiments on benchmarks like LIBERO show that D-VLA significantly outperforms mainstream RL frameworks in throughput and sampling efficiency for billion-parameter VLA models. In trillion-parameter scalability tests, our framework maintains exceptional stability and linear speedup, providing a robust system for high-performance general-purpose embodied agents.

preprint2026arXiv

Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emph{training--inference discrepancy term} that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emph{policy-staleness term} that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required historical training-side logits, or old logits. This missing-old-logit problem entangles discrepancy repair with staleness correction, breaks the intended semantics of decoupled correction, and makes clipping and masking thresholds interact undesirably. To address this issue, we study both exact and approximate correction routes. We propose three exact old-logit acquisition strategies: snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption, and compare their system trade-offs. From the perspective of approximate correction, we focus on preserving the benefits of decoupled correction through a more appropriate approximate policy when exact old logits cannot be recovered at low cost, without incurring extra system overhead. Following this analysis, we adopt a revised PPO-EWMA method, which achieves significant gains in both training speed and optimization performance.

preprint2026arXiv

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned World Models as generative simulators, enabling policy optimization entirely within "imagination." However, when deployed as simulators for specific environments such as the LIBERO benchmark, existing World Models often suffer from poor generalization and long-horizon error accumulation. During closed-loop rollouts, these models are highly sensitive to initial-state perturbations; minor changes in color, illumination, and other visual factors can trigger cascading hallucinations, leading to severe blurriness or overexposure. Moreover, long-horizon error accumulation further degrades the quality and fidelity of predicted future states. These issues limit the reliability of World Models as simulators. To mitigate these problems, we propose Sword, a robust World Model framework. Our method introduces Structure-Guided Style Augmentation to disentangle the visual textures of interactive environments from task-relevant dynamics, thereby improving generalization. We further propose Dynamic Latent Bootstrapping, which maintains consistency between training and inference while keeping memory consumption low. Extensive experiments on the LIBERO benchmark show that our method significantly outperforms the baseline WoVR in terms of generalization, generation quality, robustness, fidelity, and the success rate of reinforcement-learning post-training for VLA models.

preprint2021arXiv

Maximum Approximate Bernstein Likelihood Estimation of Densities in a Two-sample Semiparametric Model

Maximum likelihood estimators are proposed for the parameters and the densities in a semiparametric density ratio model in which the nonparametric baseline density is approximated by the Bernstein polynomial model. The EM algorithm is used to obtain the maximum approximate Bernstein likelihood estimates. Simulation study shows that the performance of the proposed method is much better than the existing ones. The proposed method is illustrated by real data examples. Some asymptotic results are also presented and proved.

preprint2016arXiv

Enhanced high-order harmonic generation from periodic potentials in inhomogeneous laser fields

We theoretically study high-order harmonic generation (HHG) from solid-phase systems in spatially inhomogeneous strong laser fields originated by resonant plasmons within a metallic nanostructure. The intensity of the second plateau in HHG may be enhanced by two-three orders and be comparable with the intensity of the first plateau. This is due to bigger transition probabilities to higher conduction bands. It provides us a practical way to increase the yields of HHG with laser intensity below the damage threshold. It presents a promising way to triple the range of HHG spectra in experimental measurements. It also allows to generate intense isolated attosecond pulse from solids driven by few-cycle laser fields.

preprint2016arXiv

Improvement of a Theorem of Lorentz (1963) and its Generalization to the Multivariate Case

In this short note we have proved an enhanced version of a theorem of Lorentz [1] and its generalization to the multivariate case which gives a non- uniform estimate of degree of approximation by a polynomial with positive coefficients. The performance of the approximation at the vertices of [0; 1]d is more precisely characterized by the improved result and its multivariate generalization. The latter provides mathematical foundation on which multivariate density approximation by a polynomial with positive coefficients can be established.

preprint2015arXiv

Bernstein Polynomial Model for Grouped Continuous Data

Grouped data are commonly encountered in applications. The Bernstein polynomial model is proposed as an approximate model in this paper for estimating a univariate density function based on grouped data. The coefficients of the Bernstein polynomial, as the mixture proportions of beta distributions, can be estimated using an EM algorithm. The optimal degree of the Bernstein polynomial can be determined using a change-point estimation method. The rate of convergence of the proposed density estimate to the true density is proved to be almost parametric by an acceptance-rejection arguments used in Monte Carlo method. The proposed method is compared with some existing methods in a simulation study and is applied to a real dataset.

preprint2015arXiv

Efficient and Robust Density Estimation Using Bernstein Type Polynomials

Method of parameterizing and smoothing the unknown underling distributions using Bernstein polynomials is proposed, verified and investigated. Any distribution with bounded and smooth enough density can be approximated by the proposed model. The approximating model turns out to be a mixture of beta distributions beta$(i+1, m-i+1)$, $i=0,\ldots, m$, for some optimal degree $m$. A simple change-point estimating method for choosing optimal degree $m$ of the Bernstein polynomials is also presented. The proposed methods give maximum likelihood density estimate which is consistent in $L_2$ distance at an almost parametric rate under some conditions. Simulation study shows that one can benefit from both the smoothness and the accuracy by using the proposed method. The proposed model can also be used to estimate some functional of the unknown distribution such as population mean. As illustration, the proposed methods are applied to three different type data sets including a microarray data set.

preprint2015arXiv

High harmonic generation from periodic potentials driven by few-cycle laser pulses

We investigate the high harmonic generation (HHG) from solids by simulating the dynamics of a single active electron in periodic potentials. The corresponding time-dependent Schrödinger equations (TDSE) are solved numerically by using B-spline basis sets in coordinate space. The energy band structure and wave vectors can be directly retrived from the eigenfunctions. The harmonic spectra obtained agree well with the results simulated by TDSE in $k$ space using Bloch states and show a two-plateau structure. Both of the cutoff energies of the two plateaus in the harmonic spectrum scale linearly with the field strength. We also study HHG driven by intense few-cycle laser pulses and find that the cutoff energy of the harmonic spectrum is as sensitive to the changes of the carrier envelope phase, as to HHG from gas samples, which suggests recollision pictures in HHG as found by recent experiments (Nature {\bf 522}, 462 (2015)).