Source author record

Yuan Zuo

Yuan Zuo appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

7works
4topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

7 published item(s)

preprint2026arXiv

Strategy-Aware Optimization Modeling with Reasoning LLMs

Large language models (LLMs) can generate syntactically valid optimization programs, yet often struggle to reliably choose an effective modeling strategy, leading to incorrect formulations and inefficient solver behavior. We propose SAGE, a strategy-aware framework that makes Modeling Strategy explicit in both data construction and post-training. SAGE builds a solver-verified multi-strategy dataset and trains a student model with supervised fine-tuning followed by Segment-Weighted GRPO using a composite reward over format compliance, correctness, and solver efficiency. Across eight benchmarks spanning synthetic and real-world settings, SAGE improves average pass@1 from 72.7 to 80.3 over the strongest open-source baseline. With multiple generations, SAGE discovers more distinct correct formulations and improves component-level diversity at pass@16 by 19-29%. At the largest scale, SAGE produces more compact constraint systems with 14.2% fewer constraints than the baseline, consistent with solver-efficient modeling. Overall, these results show that making Modeling Strategy explicit improves automated optimization modeling. Code is available at https://github.com/rachhhhing/SAGE.

preprint2023arXiv

Language Model as an Annotator: Unsupervised Context-aware Quality Phrase Generation

Phrase mining is a fundamental text mining task that aims to identify quality phrases from context. Nevertheless, the scarcity of extensive gold labels datasets, demanding substantial annotation efforts from experts, renders this task exceptionally challenging. Furthermore, the emerging, infrequent, and domain-specific nature of quality phrases presents further challenges in dealing with this task. In this paper, we propose LMPhrase, a novel unsupervised context-aware quality phrase mining framework built upon large pre-trained language models (LMs). Specifically, we first mine quality phrases as silver labels by employing a parameter-free probing technique called Perturbed Masking on the pre-trained language model BERT (coined as Annotator). In contrast to typical statistic-based or distantly-supervised methods, our silver labels, derived from large pre-trained language models, take into account rich contextual information contained in the LMs. As a result, they bring distinct advantages in preserving informativeness, concordance, and completeness of quality phrases. Secondly, training a discriminative span prediction model heavily relies on massive annotated data and is likely to face the risk of overfitting silver labels. Alternatively, we formalize phrase tagging task as the sequence generation problem by directly fine-tuning on the Sequence-to-Sequence pre-trained language model BART with silver labels (coined as Generator). Finally, we merge the quality phrases from both the Annotator and Generator as the final predictions, considering their complementary nature and distinct characteristics. Extensive experiments show that our LMPhrase consistently outperforms all the existing competitors across two different granularity phrase mining tasks, where each task is tested on two different domain datasets.

preprint2015arXiv

Investigation of practical application for QAM hybrid receiver

We present a quantum receiver for quadrature amplitude modulation (QAM) coherent states discrimination with homodyne-displacement hybrid structure. Our strategy is to carry out two successive measurements on parts of the quantum states. The homodyne result of the first measurement reveals partial information about the state and is forward to a displacement receiver, which finally identifies the input state by using feedback to adjust a reference field. Numerical simulation results show that for 16-QAM, the hybrid receiver could outperform the standard quantum limit (SQL) with a reduced number of codeword interval partitions and on-off detectors, which shows great potential toward implementing the practical application.

preprint2015arXiv

QAM Adaptive Measurements Feedback Quantum Receiver Performance

We theoretically study the quantum receivers with adaptive measurements feedback for discriminating quadrature amplitude modulation (QAM) coherent states in terms of average symbol error rate. For rectangular 16-QAM signal set, with different stages of adaptive measurements, the effects of realistic imperfection parameters including the sub-unity quantum efficiency and the dark counts of on-off detectors, as well as the transmittance of beam splitters and the mode mismatch factor between the signal and local oscillating fields on the symbol error rate are separately investigated through Monte Carlo simulations. Using photon-number-resolving detectors (PNRD) instead of on-off detectors, all the effects on the symbol error rate due to the above four imperfections can be suppressed in a certain degree. The finite resolution and PNR capability of PNRDs are also considered. We find that for currently available technology, the receiver shows a reasonable gain from the standard quantum limit (SQL) with moderate stages.

preprint2014arXiv

Word Network Topic Model: A Simple but General Solution for Short and Imbalanced Texts

The short text has been the prevalent format for information of Internet in recent decades, especially with the development of online social media, whose millions of users generate a vast number of short messages everyday. Although sophisticated signals delivered by the short text make it a promising source for topic modeling, its extreme sparsity and imbalance brings unprecedented challenges to conventional topic models like LDA and its variants. Aiming at presenting a simple but general solution for topic modeling in short texts, we present a word co-occurrence network based model named WNTM to tackle the sparsity and imbalance simultaneously. Different from previous approaches, WNTM models the distribution over topics for each word instead of learning topics for each document, which successfully enhance the semantic density of data space without importing too much time or space complexity. Meanwhile, the rich contextual information preserved in the word-word space also guarantees its sensitivity in identifying rare topics with convincing quality. Furthermore, employing the same Gibbs sampling with LDA makes WNTM easily to be extended to various application scenarios. Extensive validations on both short and normal texts testify the outperformance of WNTM as compared to baseline methods. And finally we also demonstrate its potential in precisely discovering newly emerging topics or unexpected events in Weibo at pretty early stages.

preprint2013arXiv

Suppressing the Errors due to Mode Mismatch for M-ary PSK Quantum Receivers by Photon-Number-Resolving Detector

A M-ary phase shift keying (PSK) quantum receiver consisting of displacement operations, photon counting, and electrical feed-back adaptive measurements is analyzed with a realistic model considering the effects of the sub-unity quantum efficiency and the dark counts of single-photon detectors, as well as the transmittance and the mode mismatch of beam splitters. Among these factors, the mode mismatch has the greatest impact on the error probability of the receiver with on-off detectors. The errors due to mode mismatch can be suppressed effectively by using photon-number-resolving detectors (PNRD) instead of on-off detectors.