Source author record

Jingwen Xu

Jingwen Xu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

8works
8topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

8 published item(s)

preprint2026arXiv

Benchmark^2: Systematic Evaluation of LLM Benchmarks

The rapid proliferation of benchmarks for evaluating large language models (LLMs) has created an urgent need for systematic methods to assess benchmark quality itself. We propose Benchmark^2, a comprehensive framework comprising three complementary metrics: (1) Cross-Benchmark Ranking Consistency, measuring whether a benchmark produces model rankings aligned with peer benchmarks; (2) Discriminability Score, quantifying a benchmark's ability to differentiate between models; and (3) Capability Alignment Deviation, identifying problematic instances where stronger models fail but weaker models succeed within the same model family. We conduct extensive experiments across 15 benchmarks spanning mathematics, reasoning, and knowledge domains, evaluating 11 LLMs across four model families. Our analysis reveals significant quality variations among existing benchmarks and demonstrates that selective benchmark construction based on our metrics can achieve comparable evaluation performance with substantially reduced test sets.

preprint2026arXiv

Controllable Memory Usage: Balancing Anchoring and Innovation in Long-Term Human-Agent Interaction

As LLM-based agents are increasingly used in long-term interactions, cumulative memory is critical for enabling personalization and maintaining stylistic consistency. However, most existing systems adopt an ``all-or-nothing'' approach to memory usage: incorporating all relevant past information can lead to \textit{Memory Anchoring}, where the agent is trapped by past interactions, while excluding memory entirely results in under-utilization and the loss of important interaction history. We show that an agent's reliance on memory can be modeled as an explicit and user-controllable dimension. We first introduce a behavioral metric of memory dependence to quantify the influence of past interactions on current outputs. We then propose \textbf{Stee}rable \textbf{M}emory Agent, \texttt{SteeM}, a framework that allows users to dynamically regulate memory reliance, ranging from a fresh-start mode that promotes innovation to a high-fidelity mode that closely follows interaction history. Experiments across different scenarios demonstrate that our approach consistently outperforms conventional prompting and rigid memory masking strategies, yielding a more nuanced and effective control for personalized human-agent collaboration.

preprint2026arXiv

CSSG: Measuring Code Similarity with Semantic Graphs

Existing code similarity metrics, such as BLEU, CodeBLEU, and TSED, largely rely on surface-level string overlap or abstract syntax tree structures, and often fail to capture deeper semantic relationships between programs.We propose CSSG (Code Similarity using Semantic Graphs), a novel metric that leverages program dependence graphs to explicitly model control dependencies and variable interactions, providing a semantics-aware representation of code.Experiments on the CodeContests+ dataset show that CSSG consistently outperforms existing metrics in distinguishing more similar code from less similar code under both monolingual and cross-lingual settings, demonstrating that dependency-aware graph representations offer a more effective alternative to surface-level or syntax-based similarity measures.

preprint2026arXiv

UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning

User simulators serve as the critical interactive environment for agent post-training, and an ideal user simulator generalizes across domains and proactively engages in negotiation by challenging or bargaining. However, current methods exhibit two issues. They rely on static and context-unaware profiles, necessitating extensive manual redesign for new scenarios, thus limiting generalizability. Moreover, they neglect human strategic thinking, leading to vulnerability to agent manipulation. To address these issues, we propose UserLM-R1, a novel user language model with reasoning capability. Specifically, we first construct comprehensive user profiles with both static roles and dynamic scenario-specific goals for adaptation to diverse scenarios. Then, we propose a goal-driven decision-making policy to generate high-quality rationales before producing responses, and further refine the reasoning and improve strategic capabilities with supervised fine-tuning and multi-reward reinforcement learning. Extensive experimental results demonstrate that UserLM-R1 outperforms competitive baselines, particularly on the more challenging adversarial set.

preprint2016arXiv

Carbon Decorated TiO2 Nanotube Membranes: A Renewable Nanofilter for Size- and Charge Selective Enrichment of Proteins

In this work, we design a TiO2 nanomembrane (TiNM) that can be used as a nanofilter platform for a selective enrichment of specific proteins. After use the photocatalytic properties of TiO2 allow to decompose unwanted remnant on the substrate and thus make the platform reusable. To construct this platform we fabricate a free-standing TiO2 nanotube array and remove the bottom oxide to form a both-end open TiNM. By pyrolysis of the natural tube wall contamination (C/TiNM), the walls become decorated with graphitic carbon patches. Owing to the large surface area, the amphiphilic nature and the charge adjustable character, this C/TiNM can be used to extract and enrich hydrophobic and charged biomolecules from solutions. Using human serum albumin (HSA) as a model protein as well as protein mixtures, we show that the composite membrane exhibits a highly enhanced loading capacity and protein selectivity and is reusable after a short UV treatment.

preprint2016arXiv

Graphitic C3N4 Sensitized TiO2 Nanotube Layers: A Visible Light Activated Efficient Antimicrobial Platform

In this work, we introduce a facile procedure to graft a thin graphitic C3N4 (g-C3N4) layer on aligned TiO2 nanotube arrays (TiNT) by one-step chemical vapor deposition (CVD) approach. This provides a platform to enhance the visible-light response of TiO2 nanotubes for antimicrobial applications. The formed g- C3N4/TiNT binary nanocomposite exhibits excellent bactericidal efficiency against E. coli as a visiblelight activated antibacterial coating.

preprint2016arXiv

Momentum mapping of continuum electron wave packet interference

We analyze the two-dimensional photoelectrons momentum distribution of Ar atom ionized by midinfrared laser pulses and mainly concentrate on the energy range below 2Up. By using a generalized quantum trajectory Monte Carlo (GQTMC) simulation and comparing with the numerical solution of time-dependent Schrodinger equation (TDSE), we show that in the deep tunneling regime, the rescattered electron trajectories plays unimportant role and the interplay between the intracycle and inter-cycle results in a ring-like interference pattern. The ring-like interference pattern will mask the holographic interference structure in the low longitudinal momentum region. When the nonadiabatic tunneling contributes significantly to ionization, i.e., the Keldysh parameter 1, the contribution of the rescattered electron trajectories become large, thus holographic interference pattern can be clearly observed. Our results help paving the way for gaining physical insight into ultrafast electron dynamic process with attosecond temporal resolution.

preprint2016arXiv

Visible Light Triggered Drug Release from TiO2 Nanotube Arrays: A Novel Controllable Antibacterial Platform

In this work, we use a double-layered stack of TiO2 nanotubes (TiNTs) to construct a visible-light triggered drug delivery system. Key for visible-light drug release is a hydrophobic cap on the nanotubes containing Au nanoparticles (AuNPs). The AuNPs allow for a photocatalytic scission of the hydrophobic chain under visible light. To demonstrate the principle, we loaded antibiotic (ampicillin sodium (AMP)) in the lower part of the TiO2 nanotube stack, triggered visible light induced release, and carried out antibacterial studies. The release from the platform becomes most controllable if the drug is silane-grafted in hydrophilic bottom layer for drug storage. Thus visible-light photocatalysis can also determine the release kinetics of the active drug from the nanotube wall.