Source author record

Hongxing Li

Hongxing Li appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2026arXiv

Milestone-Guided Policy Learning for Long-Horizon Language Agents

While long-horizon agentic tasks require language agents to perform dozens of sequential decisions, training such agents with reinforcement learning remains challenging. We identify two root causes: credit misattribution, where correct early actions are penalized due to terminal failures, and sample inefficiency, where scarce successful trajectories result in near-total loss of learning signal. We introduce a milestone-guided policy learning framework, BEACON, that leverages the compositional structure of long-horizon tasks to ensure precise credit assignment. BEACON partitions trajectories at milestone boundaries, applies temporal reward shaping within segments to credit partial progress, and estimates advantages at dual scales to prevent distant failures from corrupting the evaluation of local actions. On ALFWorld, WebShop, and ScienceWorld, BEACON consistently outperforms GRPO and GiGPO. Notably, on long-horizon ALFWorld tasks, BEACON achieves 92.9% success rate, nearly doubling GRPO's 53.5%, while improving effective sample utilization from 23.7% to 82.0%. These results establish milestone-anchored credit assignment as an effective paradigm for training long-horizon language agents. Code is available at https://github.com/ZJU-REAL/BEACON.

preprint2022arXiv

Large Intrinsic Valley Polarization and High Curie Temperature in Stable Two-dimensional Ferrovalley YX$_2$(X=I,Br and Cl)

Ferrovalley materials with spontaneous valley polarization are crucial to valleytronic application. Based on first-principles calculations, we demonstrate that two-dimensional (2D) YX$_2$(X= I, Br,and Cl) in 2H structure constitute a series of promising ferrovalley semiconductors with large spontaneous valley polarization and high Curie temperature. Our calculations reveal that YX$_2$ are dynamically and thermally stable 2D ferromagnetic semiconductors with a Curie temperature above 200 K. Due to the natural noncentrosymmetric structure, intrinsic ferromagnetic order and strong spin orbital coupling, the large spontaneous valley polarizations of 108.98, 57.70 and 22.35 meV are also predicted in single-layer YX$_2$(X = I, Br and Cl),respectively. The anomalous valley Hall effect is also proposed based on the valley contrasting Berry curvature. Moreover, the ferromagnetism and valley polarization are found to be effectively tuning by applying a biaxial strain. Interestingly, the suppressed valley physics of YBr$_2$ and YCl$_2$ can be switched on via applying a moderate compression strain. The present findings promise YX$_2$ as competitive candidates for the further experimental studies and practical applications in valleytronics.

preprint2020arXiv

Spin-dependent Schottky barriers and vacancy-induced spin-selective Ohmic contacts in magnetic vdW heterostructures

The 2D ferromagnets, such as CrX3 (X=Cl, Br and I), have been attracting extensive attentions since they provide novel platforms to fundamental physics and device applications. Integrating CrX3 with other electrodes and substrates is an essential step to their device realization. Therefore, it is important to understand the interfacial properties between CrX3 and other 2D materials. As an illustrative example, we have investigated the heterostructures between CrX3 and graphene (CrX3/Gr) from firstprinciples. We find unique Schottky contacts type with strongly spin-dependent barriers in CrX3/Gr. This can be understood by synergistic effects between the exchange splitting of semiconductor band of CrX3 and interlayer charge transfer. The spinasymmetry of Schottky barriers may result in different tunneling rates of spin-up and down electrons, and then lead to spin-polarized current, namely spin-filter (SF) effect. Moreover, by introducing X vacancy into CrX3/Gr, an Ohmic contact forms in spin-up direction. It may enhance the transport of spin-up electrons, and improve SF effect. Our systematic study reveals the unique interfacial properties of CrX3/Gr, and provides a theoretical view to the understanding and designing of spintronics device based on magnetic vdW heterostructures.

preprint2016arXiv

Trust Exploitation and Attention Competition: A Game Theoretical Model

The proliferation of Social Network Sites (SNSs) has greatly reformed the way of information dissemination, but also provided a new venue for hosts with impure motivations to disseminate malicious information. Social trust is the basis for information dissemination in SNSs. Malicious hosts judiciously and dynamically make the balance between maintaining its social trust and selfishly maximizing its malicious gain over a long time-span. Studying the optimal response strategies for each malicious host could assist to design the best system maneuver so as to achieve the targeted level of overall malicious activities. In this paper, we propose an interaction-based social trust model, and formulate the maximization of long-term malicious gains of multiple competing hosts as a non-cooperative differential game. Through rigorous analysis, optimal response strategies are identified and the best system maneuver mechanism is presented. Extensive numerical studies further verify the analytical results.

preprint2014arXiv

Optimal CSMA-based Wireless Communication with Worst-case Delay and Non-uniform Sizes

Carrier Sense Multiple Access (CSMA) protocols have been shown to reach the full capacity region for data communication in wireless networks, with polynomial complexity. However, current literature achieves the throughput optimality with an exponential delay scaling with the network size, even in a simplified scenario for transmission jobs with uniform sizes. Although CSMA protocols with order-optimal average delay have been proposed for specific topologies, no existing work can provide worst-case delay guarantee for each job in general network settings, not to mention the case when the jobs have non-uniform lengths while the throughput optimality is still targeted. In this paper, we tackle on this issue by proposing a two-timescale CSMA-based data communication protocol with dynamic decisions on rate control, link scheduling, job transmission and dropping in polynomial complexity. Through rigorous analysis, we demonstrate that the proposed protocol can achieve a throughput utility arbitrarily close to its offline optima for jobs with non-uniform sizes and worst-case delay guarantees, with a tradeoff of longer maximum allowable delay.

preprint2013arXiv

Virtual Machine Trading in a Federation of Clouds: Individual Profit and Social Welfare Maximization

By sharing resources among different cloud providers, the paradigm of federated clouds exploits temporal availability of resources and geographical diversity of operational costs for efficient job service. While interoperability issues across different cloud platforms in a cloud federation have been extensively studied, fundamental questions on cloud economics remain: When and how should a cloud trade resources (e.g., virtual machines) with others, such that its net profit is maximized over the long run, while a close-to-optimal social welfare in the entire federation can also be guaranteed? To answer this question, a number of important, inter-related decisions, including job scheduling, server provisioning and resource pricing, should be dynamically and jointly made, while the long-term profit optimality is pursued. In this work, we design efficient algorithms for inter-cloud virtual machine (VM) trading and scheduling in a cloud federation. For VM transactions among clouds, we design a double-auction based mechanism that is strategyproof, individual rational, ex-post budget balanced, and efficient to execute over time. Closely combined with the auction mechanism is a dynamic VM trading and scheduling algorithm, which carefully decides the true valuations of VMs in the auction, optimally schedules stochastic job arrivals with different SLAs onto the VMs, and judiciously turns on and off servers based on the current electricity prices. Through rigorous analysis, we show that each individual cloud, by carrying out the dynamic algorithm in the online double auction, can achieve a time-averaged profit arbitrarily close to the offline optimum. Asymptotic optimality in social welfare is also achieved under homogeneous cloud settings. We carry out trace-driven simulations to examine the effectiveness of our algorithms and the achievable social welfare under heterogeneous cloud settings.