Source author record

Shengbao Zheng

Shengbao Zheng appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Artificial Intelligence Distributed, Parallel, and Cluster Computing eess.SY Networking and Internet Architecture Systems and Control

Catalog footprint

What is connected

2works

5topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

Collective Communication for 100k+ GPUs

The increasing scale of large language models (LLMs) necessitates highly efficient collective communication frameworks, particularly as training workloads extend to hundreds of thousands of GPUs. Traditional communication methods face significant throughput and latency limitations at this scale, hindering both the development and deployment of state-of-the-art models. This paper presents the NCCLX collective communication framework, developed at Meta, engineered to optimize performance across the full LLM lifecycle, from the synchronous demands of large-scale training to the low-latency requirements of inference. The framework is designed to support complex workloads on clusters exceeding 100,000 GPUs, ensuring reliable, high-throughput, and low-latency data exchange. Empirical evaluation on the Llama4 model demonstrates substantial improvements in communication efficiency. This research contributes a robust solution for enabling the next generation of LLMs to operate at unprecedented scales.

preprint2023arXiv

Modeling and Control of Discrete Event Systems under Joint Sensor-Actuator Cyber Attacks

In this paper, we investigate joint sensor-actuator cyber attacks in discrete event systems. We assume that attackers can attack some sensors and actuators at the same time by altering observations and control commands. Because of the nondeterminism in observation and control caused by cyber attacks, the behavior of the supervised system becomes nondeterministic and may deviate from the safety specification. We define the upper-bound on all possible languages that can be generated by the supervised system to investigate the safety supervisory control problem under cyber attacks. After introducing CA-controllability and CA-observability, we prove that the supervisory control problem under cyber attacks is solvable if and only if the given specification language is CA-controllable and CA-observable. Furthermore, we obtain methods to calculate the state estimates under sensor attacks and to synthesize a state-estimate-based supervisor to achieve a given safety specification under cyber attacks. We further show that of all the solutions, the proposed state-estimate-based supervisor is maximally-permissive.