Source author record

Hao Fang

Hao Fang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

30works
17topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

30 published item(s)

preprint2026arXiv

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients

Reinforcement learning with verifiable rewards (RLVR), due to the deterministic verification, becomes a dominant paradigm for enhancing the reasoning ability of large language models (LLMs). The community witnesses the rapid change from the Proximal Policy Optimization (PPO) to Group Relative Policy Optimization (GRPO), in which GRPO reduces the complicated advantage estimation with simple estimation over grouped positive and negative rollouts. However, we note that negative rollouts may admit no gradation of failure severity, and the combinatorial vastness makes penalizing a few sampled negatives unlikely to cover a meaningful reward signal under sparse binary rewards. In this work, we propose Positive-Only Policy Optimization (POPO), a novel RLVR framework in which learning can occur exclusively via online positive rollouts. Specifically, POPO utilizes bounded importance sampling over the positive rollout set. Thus, no disjoint negative rollouts are used for the gradient guidance. We show that implicit negative gradients can emerge naturally through reinforcing the positive probability via rollouts redistribution. Next, POPO stabilizes the policy optimization through two mechanisms. First, it applies a siamese policy network with a momentum-based adaptation law for stabilized policy evolution. Second, we replace the KL-divergence with a bounded similarity penalty term in the siamese representation space. We conduct extensive experiments using publicly available, well-established text-LLM models, e.g., the Qwen family, across all-level mathematical benchmarks. Our experiment demonstrates that POPO achieves performance comparable to, or even superior to GRPO. Notably, we show that POPO can achieve 36.67% in AIME 2025 with Qwen-Math-7B, outperforming GRPO 30.00%. Our ablation and sweep studies further illustrate the necessity and robustness of POPO components.

preprint2026arXiv

Sustainable Intelligence for the Wild: Democratizing Ecological Monitoring via Knowledge-Adaptive Edge Expert Agents

Rapid biodiversity loss underscore the urgency of effective monitoring, yet manual surveys remain resource-intensive. While on-device AI offers a scalable alternative, its performance in the wild is often challenged by environmental variability. Current methods rely heavily on cloud resource, which requires continuous uploading of field data for model retraining. This approach is unsuitable for remote deployments because it consumes limited power and network connectivity. To address these constraints, this research proposes a shift from model adaptation to knowledge adaptation. We introduce an architecture that separates visual perception from reasoning, combining a visual encoder with a dynamic knowledge base. We uses an explicit knowledge base to replace implicitly encoding expert knowledge into model parameters. This method also supports knowledge sustainability by preserving expert insights in a structured form. Through cross-disciplinary collaboration with biologists and Indigenous communities, this work advances ethical AI co-development, fostering responsible and culturally informed ecosystem management.

preprint2024arXiv

Natural Language Decomposition and Interpretation of Complex Utterances

Designing natural language interfaces has historically required collecting supervised data to translate user requests into carefully designed intent representations. This requires enumerating and labeling a long tail of user requests, which is challenging. At the same time, large language models (LLMs) encode knowledge about goals and plans that can help conversational assistants interpret user requests requiring numerous steps to complete. We introduce an approach to handle complex-intent-bearing utterances from a user via a process of hierarchical natural language decomposition and interpretation. Our approach uses a pre-trained language model to decompose a complex utterance into a sequence of simpler natural language steps and interprets each step using the language-to-program model designed for the interface. To test our approach, we collect and release DeCU -- a new NL-to-program benchmark to evaluate Decomposition of Complex Utterances. Experiments show that the proposed approach enables the interpretation of complex utterances with almost no complex training data, while outperforming standard few-shot prompting approaches.

preprint2023arXiv

Application of the correlated B-spline basis functions to the leading relativistic and QED corrections of helium

B-spline functions have been widely used in computational atomic physics. Different from the traditional B-spline basis (a simple product of two B-splines), the recently developed correlated B-spline basis functions(C-BSBF), in which the interelectronic coordinate $r_{12}$ is included explicitly, have greatly improved the computational accuracy of polarizability [S. J. Yang \textit{et al}., Phys. Rev. A \textbf{95}, 062505 (2017)] and bethe logarithm [ S. J. Yang \textit{et al}., Phys. Rev. A \textbf{100}, 042509 (2019)] for singlet states of helium. Here, we report the extension of the C-BSBF to the leading relativistic and QED correction calculations for energy levels of the $1\,^1S$, $2\,^1S$, $2\,^3S$, and $3\,^3S$ states of helium. The relativistic kinetic term $p_{1}^{4}$, contact potential $δ^{3}(r_{1})$, $δ^{3}(r_{12})$ and Araki-Sucher correction $\langle 1/r_{12}^{3} \rangle$ are calculated by using the global operator method, in which $r_{12}^n$ and $r_{12}^n\ln r_{12}$ involved are calculated with the generalization of Laplace's expansions. The obtained values for the ground state are $δE_{rel}/α^{2}=-$1.951 754 7(2) and $δE_{QED}/α^{3}=$57.288 165(2), consistent with previous results, which opens the possibility of calculating higher-order relativistic and QED effects using the C-BSBF.

preprint2023arXiv

Surveillance Face Anti-spoofing

Face Anti-spoofing (FAS) is essential to secure face recognition systems from various physical attacks. However, recent research generally focuses on short-distance applications (i.e., phone unlocking) while lacking consideration of long-distance scenes (i.e., surveillance security checks). In order to promote relevant research and fill this gap in the community, we collect a large-scale Surveillance High-Fidelity Mask (SuHiFiMask) dataset captured under 40 surveillance scenes, which has 101 subjects from different age groups with 232 3D attacks (high-fidelity masks), 200 2D attacks (posters, portraits, and screens), and 2 adversarial attacks. In this scene, low image resolution and noise interference are new challenges faced in surveillance FAS. Together with the SuHiFiMask dataset, we propose a Contrastive Quality-Invariance Learning (CQIL) network to alleviate the performance degradation caused by image quality from three aspects: (1) An Image Quality Variable module (IQV) is introduced to recover image information associated with discrimination by combining the super-resolution network. (2) Using generated sample pairs to simulate quality variance distributions to help contrastive learning strategies obtain robust feature representation under quality variation. (3) A Separate Quality Network (SQN) is designed to learn discriminative features independent of image quality. Finally, a large number of experiments verify the quality of the SuHiFiMask dataset and the superiority of the proposed CQIL.

preprint2022arXiv

A $σ_{2}$ Penrose inequality for conformal asymptotically hyperbolic 4-discs

In this paper, we consider conformal metrics on a unit 4-disc with an asymptotically hyperbolic end and possible isolated conic singularities. We define a mass term of the AH end. If the $σ_{2}$ curvature has lower bound $σ_{2}\geq\frac{3}{2}$, we prove a Penrose type inequality relating the mass and contributions from singularities. We also classify sharp cases, which is the standard hyperbolic 4-space $\mathbb{H}^{4}$ when no singularity occurs. It is worth noting that our curvature condition implies non-positive energy density.

preprint2021arXiv

Task-Oriented Dialogue as Dataflow Synthesis

We describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset and code for replicating experiments are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines.

preprint2020arXiv

Building A User-Centric and Content-Driven Socialbot

To build Sounding Board, we develop a system architecture that is capable of accommodating dialog strategies that we designed for socialbot conversations. The architecture consists of a multi-dimensional language understanding module for analyzing user utterances, a hierarchical dialog management framework for dialog context tracking and complex dialog control, and a language generation process that realizes the response plan and makes adjustments for speech synthesis. Additionally, we construct a new knowledge base to power the socialbot by collecting social chat content from a variety of sources. An important contribution of the system is the synergy between the knowledge base and the dialog management, i.e., the use of a graph structure to organize the knowledge base that makes dialog control very efficient in bringing related content to the discussion. Using the data collected from Sounding Board during the competition, we carry out in-depth analyses of socialbot conversations and user ratings which provide valuable insights in evaluation methods for socialbots. We additionally investigate a new approach for system evaluation and diagnosis that allows scoring individual dialog segments in the conversation. Finally, observing that socialbots suffer from the issue of shallow conversations about topics associated with unstructured data, we study the problem of enabling extended socialbot conversations grounded on a document. To bring together machine reading and dialog control techniques, a graph-based document representation is proposed, together with methods for automatically constructing the graph. Using the graph-based representation, dialog control can be carried out by retrieving nodes or moving along edges in the graph. To illustrate the usage, a mixed-initiative dialog strategy is designed for socialbot conversations on news articles.

preprint2020arXiv

Deepfakes for Medical Video De-Identification: Privacy Protection and Diagnostic Information Preservation

Data sharing for medical research has been difficult as open-sourcing clinical data may violate patient privacy. Traditional methods for face de-identification wipe out facial information entirely, making it impossible to analyze facial behavior. Recent advancements on whole-body keypoints detection also rely on facial input to estimate body keypoints. Both facial and body keypoints are critical in some medical diagnoses, and keypoints invariability after de-identification is of great importance. Here, we propose a solution using deepfake technology, the face swapping technique. While this swapping method has been criticized for invading privacy and portraiture right, it could conversely protect privacy in medical video: patients' faces could be swapped to a proper target face and become unrecognizable. However, it remained an open question that to what extent the swapping de-identification method could affect the automatic detection of body keypoints. In this study, we apply deepfake technology to Parkinson's disease examination videos to de-identify subjects, and quantitatively show that: face-swapping as a de-identification approach is reliable, and it keeps the keypoints almost invariant, significantly better than traditional methods. This study proposes a pipeline for video de-identification and keypoint preservation, clearing up some ethical restrictions for medical data sharing. This work could make open-source high quality medical video datasets more feasible and promote future medical research that benefits our society.

preprint2020arXiv

Solving A Class of Nonsmooth Resource Allocation Problems with Directed Graphs though Distributed Smooth Multi-Proximal Algorithms

In this paper, two distributed multi-proximal primal-dual algorithms are proposed to deal with a class of distributed nonsmooth resource allocation problems. In these problems, the global cost function is the summation of local convex and nonsmooth cost functions, each of which consists of one twice differentiable function and multiple nonsmooth functions. Communication graphs of underling multi-agent systems are directed and strongly connected but not necessarily weighted-balanced. The multi-proximal splitting is designed to deal with the difficulty caused by the unproximable property of the summation of those nonsmooth functions. Moreover, it can also guarantee the smoothness of proposed algorithms. Auxiliary variables in the multi-proximal splitting are introduced to estimate subgradients of nonsmooth functions. Theoretically, the convergence analysis is conducted by employing Lyapunov stability theory and integral input-to-state stability (iISS) theory with respect to set. It shows that proposed algorithms can make states converge to the optimal point that satisfies resource allocation conditions.

preprint2016arXiv

Bi-directional Attention with Agreement for Dependency Parsing

We develop a novel bi-directional attention model for dependency parsing, which learns to agree on headword predictions from the forward and backward parsing directions. The parsing procedure for each direction is formulated as sequentially querying the memory component that stores continuous headword embeddings. The proposed parser makes use of {\it soft} headword embeddings, allowing the model to implicitly capture high-order parsing history without dramatically increasing the computational complexity. We conduct experiments on English, Chinese, and 12 other languages from the CoNLL 2006 shared task, showing that the proposed model achieves state-of-the-art unlabeled attachment scores on 6 languages.

preprint2016arXiv

Learning Latent Local Conversation Modes for Predicting Community Endorsement in Online Discussions

Many social media platforms offer a mechanism for readers to react to comments, both positively and negatively, which in aggregate can be thought of as community endorsement. This paper addresses the problem of predicting community endorsement in online discussions, leveraging both the participant response structure and the text of the comment. The different types of features are integrated in a neural network that uses a novel architecture to learn latent modes of discussion structure that perform as well as deep neural networks but are more interpretable. In addition, the latent modes can be used to weight text features thereby improving prediction accuracy.

preprint2016arXiv

On curvature pinching of conic 2-spheres

We study metrics on conic 2-spheres when no Einstein metrics exist. In particular, when the curvature of a conic metric is positive, we obtain the best curvature pinching constant. We also show that when this best pinching constant is approached, the conic 2-sphere has an explicit Gromov-Hausdorff limit. This is a generalization of the previous results of Chen-Lin and Bartolucci for 2-spheres with one or two conic points.

preprint2016arXiv

Torsion type invariants of singularities

Inspired by the LG/CY correspondence, we study the local index theory of the Schrödinger operator associated to a singularity defined on ${\mathbb C}^n$ by a quasi-homogeneous polynomial $f$. Under some mild assumption on $f$, we show that the small time heat kernel expansion of the corresponding Schrödinger operator exists and is a series of fractional powers of time $t$. Then we prove a local index formula which expresses the Milnor number of $f$ by a Gaussian type integral. Furthermore, the heat kernel expansion provides spectral invariants of $f$. Especially, we define torsion type invariants associated to a singularity. These spectral invariants provide a new direction to study the singularity.

preprint2016arXiv

Volume bounds of conic 2-spheres

We obtain sharp volume bound for a conic 2-sphere in terms of its Gaussian curvature bound. We also give the geometric models realizing the extremal volume. In particular, when the curvature is bounded in absolute value by $1$, we compute the minimal volume of a conic sphere in the sense of Gromov. In order to apply the level set analysis and iso-perimetric inequality as in our previous works, we develop some new analytical tools to treat regions with vanishing curvature.

preprint2015arXiv

From Captions to Visual Concepts and Back

This paper presents a novel approach for automatically generating image descriptions: visual detectors, language models, and multimodal similarity models learnt directly from a dataset of image captions. We use multiple instance learning to train visual detectors for words that commonly occur in captions, including many different parts of speech such as nouns, verbs, and adjectives. The word detector outputs serve as conditional inputs to a maximum-entropy language model. The language model learns from a set of over 400,000 image descriptions to capture the statistics of word usage. We capture global semantics by re-ranking caption candidates using sentence-level features and a deep multimodal similarity model. Our system is state-of-the-art on the official Microsoft COCO benchmark, producing a BLEU-4 score of 29.1%. When human judges compare the system captions to ones written by other people on our held-out test set, the system captions have equal or better quality 34% of the time.

preprint2015arXiv

Language Models for Image Captioning: The Quirks and What Works

Two recent approaches have achieved state-of-the-art results in image captioning. The first uses a pipelined process where a set of candidate words is generated by a convolutional neural network (CNN) trained on images, and then a maximum entropy (ME) language model is used to arrange these words into a coherent sentence. The second uses the penultimate activation layer of the CNN as input to a recurrent neural network (RNN) that then generates the caption sequence. In this paper, we compare the merits of these different language modeling approaches for the first time by using the same state-of-the-art CNN as input. We examine issues in the different approaches, including linguistic irregularities, caption repetition, and data set overlap. By combining key aspects of the ME and RNN methods, we achieve a new record performance over previously published results on the benchmark COCO dataset. However, the gains we see in BLEU do not translate to human judgments.

preprint2015arXiv

Microsoft COCO Captions: Data Collection and Evaluation Server

In this paper we describe the Microsoft COCO Caption dataset and evaluation server. When completed, the dataset will contain over one and a half million captions describing over 330,000 images. For the training and validation images, five independent human generated captions will be provided. To ensure consistency in evaluation of automatic caption generation algorithms, an evaluation server is used. The evaluation server receives candidate captions and scores them using several popular metrics, including BLEU, METEOR, ROUGE and CIDEr. Instructions for using the evaluation server are provided.

preprint2015arXiv

On convergence to a football

We show that spheres of positive constant curvature with $n$ ($n\geq3$) conic points converge to a sphere of positive constant curvature with two conic points (or called an (American) football) in Gromov-Hausdorff topology when the corresponding singular divisors converge to a critical divisor in the sense of Troyanov. We prove this convergence in two different ways. Geometrically, the convergence follows from Luo-Tian's explicit description of conic spheres as boundaries of convex polytopes in $S^{3}$. Analytically, regarding the conformal factors as the singular solutions to the corresponding PDE, we derive the required a priori estimates and convergence result after proper reparametrization.

preprint2015arXiv

Talking to the crowd: What do people react to in online discussions?

This paper addresses the question of how language use affects community reaction to comments in online discussion forums, and the relative importance of the message vs. the messenger. A new comment ranking task is proposed based on community annotated karma in Reddit discussions, which controls for topic and timing of comments. Experimental work with discussion threads from six subreddits shows that the importance of different types of language features varies with the community of interest.

preprint2014arXiv

Performance Limits of Segmented Compressive Sampling: Correlated Samples versus Bits

This paper gives performance limits of the segmented compressive sampling (CS) which collects correlated samples. It is shown that the effect of correlation among samples for the segmented CS can be characterized by a penalty term in the corresponding bounds on the sampling rate. Moreover, this penalty term is vanishing as the signal dimension increases. It means that the performance degradation due to the fixed correlation among samples obtained by the segmented CS (as compared to the standard CS with equivalent size sampling matrix) is negligible for a high-dimensional signal. In combination with the fact that the signal reconstruction quality improves with additional samples obtained by the segmented CS (as compared to the standard CS with sampling matrix of the size given by the number of original uncorrelated samples), the fact that the additional correlated samples also provide new information about a signal is a strong argument for the segmented CS.

preprint2014arXiv

Permutation Enhanced Parallel Reconstruction with A Linear Compressive Sampling Device

In this letter, a permutation enhanced parallel reconstruction architecture for compressive sampling is proposed. In this architecture, a measurement matrix is constructed from a block-diagonal sensing matrix and the sparsifying basis of the target signal. In this way, the projection of the signal onto the sparsifying basis can be divided into several segments and all segments can be reconstructed in parallel. Thus, the computational complexity and the time for reconstruction can be reduced significantly. This feature is especially appealing for big data processing. Furthermore, to reduce the number of measurements needed to achieve the desired reconstruction error performance, permutation is introduced for the projection of the signal. It is shown that the permutation can be performed implicitly by using a pre-designed measurement matrix. Thus, the permutation enhanced parallel reconstruction can be achieved with a linear compressive sampling device.

preprint2013arXiv

Permutation Meets Parallel Compressed Sensing: How to Relax Restricted Isometry Property for 2D Sparse Signals

Traditional compressed sensing considers sampling a 1D signal. For a multidimensional signal, if reshaped into a vector, the required size of the sensing matrix becomes dramatically large, which increases the storage and computational complexity significantly. To solve this problem, we propose to reshape the multidimensional signal into a 2D signal and sample the 2D signal using compressed sensing column by column with the same sensing matrix. It is referred to as parallel compressed sensing, and it has much lower storage and computational complexity. For a given reconstruction performance of parallel compressed sensing, if a so-called acceptable permutation is applied to the 2D signal, we show that the corresponding sensing matrix has a smaller required order of restricted isometry property condition, and thus, storage and computation requirements are further lowered. A zigzag-scan-based permutation, which is shown to be particularly useful for signals satisfying a layer model, is introduced and investigated. As an application of the parallel compressed sensing with the zigzag-scan-based permutation, a video compression scheme is presented. It is shown that the zigzag-scan-based permutation increases the peak signal-to-noise ratio of reconstructed images and video frames.

preprint2013arXiv

The J-flow on Kahler surfaces: a boundary case

We study the J-flow on Kahler surfaces when the Kahler class lies on the boundary of the open cone for which global smooth convergence holds, and satisfies a nonnegativity condition. We obtain a C^0 estimate and show that the J-flow converges smoothly to a singular Kahler metric away from a finite number of curves of negative self-intersection on the surface. We discuss an application to the Mabuchi energy functional on Kahler surfaces with ample canonical bundle.

preprint2012arXiv

A note on renormalized volume functionals

New properties are derived of renormalized volume functionals, which arise as coefficients in the asymptotic expansion of the volume of an asymptotically hyperbolic Einstein (AHE) manifold. A formula is given for the renormalized volume of an even-dimensional AHE manifold in terms of an arbitrary totally geodesic compactification. The second variation of renormalized volume functionals under conformal change is identified, and is used to show that Einstein metrics of nonzero scalar curvature are local extrema.

preprint2010arXiv

On a class of fully nonlinear flow in Kähler geometry

In this paper, we study a class of fully nonlinear metric flow on Kähler manifolds, which includes the J-flow as a special case. We provide a sufficient and necessary condition for the long time convergence of the flow, generalizing the result of Song-Weinkove. As a consequence, under the given condition, we solved the corresponding Euler equation, which is fully nonlinear of Monge-Ampère type. As an application, we also discuss a complex Monge-Ampère type equation including terms of mixed degrees, which was first posed by Chen.