Source author record

Zhe Fu

Zhe Fu appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

9works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

9 published item(s)

preprint2026arXiv

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-thought prompting, have achieved considerable success on foundational reasoning tasks. However, this success is heavily contingent upon extensive human-annotated demonstrations, and models' capabilities are still insufficient for more complex problems. Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories. The proposed RL framework facilitates the emergent development of advanced reasoning patterns, such as self-reflection, verification, and dynamic strategy adaptation. Consequently, the trained model achieves superior performance on verifiable tasks such as mathematics, coding competitions, and STEM fields, surpassing its counterparts trained via conventional supervised learning on human demonstrations. Moreover, the emergent reasoning patterns exhibited by these large-scale models can be systematically harnessed to guide and enhance the reasoning capabilities of smaller models.

preprint2026arXiv

Supervised and Unsupervised Neural Network Solver for First Order Hyperbolic Nonlinear PDEs

We present a neural network-based method for learning scalar hyperbolic conservation laws. Our method replaces the traditional numerical flux in finite volume schemes with a trainable neural network while preserving the conservative structure of the scheme. The model can be trained both in a supervised setting with efficiently generated synthetic data or in an unsupervised manner, leveraging the weak formulation of the partial differential equation. We provide theoretical results that our model can perform arbitrarily well, and provide associated upper bounds on neural network size. Extensive experiments demonstrate that our method often outperforms efficient schemes such as Godunov's scheme, WENO, and Discontinuous Galerkin for comparable computational budgets. Finally, we demonstrate the effectiveness of our method on a traffic prediction task, leveraging field experimental highway data from the Berkeley DeepDrive drone dataset.

preprint2024arXiv

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

The rapid development of open-source large language models (LLMs) has been truly remarkable. However, the scaling law described in previous literature presents varying conclusions, which casts a dark cloud over scaling LLMs. We delve into the study of scaling laws and present our distinctive findings that facilitate scaling of large scale models in two commonly used open-source configurations, 7B and 67B. Guided by the scaling laws, we introduce DeepSeek LLM, a project dedicated to advancing open-source language models with a long-term perspective. To support the pre-training phase, we have developed a dataset that currently consists of 2 trillion tokens and is continuously expanding. We further conduct supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) on DeepSeek LLM Base models, resulting in the creation of DeepSeek Chat models. Our evaluation results demonstrate that DeepSeek LLM 67B surpasses LLaMA-2 70B on various benchmarks, particularly in the domains of code, mathematics, and reasoning. Furthermore, open-ended evaluations reveal that DeepSeek LLM 67B Chat exhibits superior performance compared to GPT-3.5.

preprint2022arXiv

Learning energy-efficient driving behaviors by imitating experts

The rise of vehicle automation has generated significant interest in the potential role of future automated vehicles (AVs). In particular, in highly dense traffic settings, AVs are expected to serve as congestion-dampeners, mitigating the presence of instabilities that arise from various sources. However, in many applications, such maneuvers rely heavily on non-local sensing or coordination by interacting AVs, thereby rendering their adaptation to real-world settings a particularly difficult challenge. To address this challenge, this paper examines the role of imitation learning in bridging the gap between such control strategies and realistic limitations in communication and sensing. Treating one such controller as an "expert", we demonstrate that imitation learning can succeed in deriving policies that, if adopted by 5% of vehicles, may boost the energy-efficiency of networks with varying traffic conditions by 15% using only local observations. Results and code are available online at https://sites.google.com/view/il-traffic/home.

preprint2019arXiv

The three-state Potts model on the centered triangular lattice

We study phase transitions of the Potts model on the centered-triangular lattice with two types of couplings, namely $K$ between neighboring triangular sites, and $J$ between the centered and the triangular sites. Results are obtained by means of a finite-size analysis based on numerical transfer matrix calculations and Monte Carlo simulations. Our investigation covers the whole $(K, J)$ phase diagram, but we find that most of the interesting physics applies to the antiferromagnetic case $K<0$, where the model is geometrically frustrated. In particular, we find that there are, for all finite $J$, two transitions when K is varied. Their critical properties are explored. In the limits $J\to \pm \infty$ we find algebraic phases with infinite-order transitions to the ferromagnetic phase.

preprint2016arXiv

Special transitions in an O($n$) loop model with an Ising-like constraint

We investigate the O($n$) nonintersecting loop model on the square lattice under the constraint that the loops consist of ninety-degree bends only. The model is governed by the loop weight $n$, a weight $x$ for each vertex of the lattice visited once by a loop, and a weight $z$ for each vertex visited twice by a loop. We explore the $(x,z)$ phase diagram for some values of $n$. For $0<n<1$, the diagram has the same topology as the generic O($n$) phase diagram with $n<2$, with a first-order line when $z$ starts to dominate, and an O($n$)-like transition when $x$ starts to dominate. Both lines meet in an exactly solved higher critical point. For $n>1$, the O($n$)-like transition line appears to be absent. Thus, for $z=0$, the $(n,x)$ phase diagram displays a line of phase transitions for $n\le 1$. The line ends at $n=1$ in an infinite-order transition. We determine the conformal anomaly and the critical exponents along this line. These results agree accurately with a recent proposal for the universal classification of this type of model, at least in most of the range $-1 \leq n \leq 1$. We also determine the exponent describing crossover to the generic O($n$) universality class, by introducing topological defects associated with the introduction of `straight' vertices violating the ninety-degree-bend rule. These results are obtained by means of transfer-matrix calculations and finite-size scaling.

preprint2013arXiv

Ising-like transitions in the O($n$) loop model on the square lattice

We explore the phase diagram of the O($n$) loop model on the square lattice in the $(x,n)$ plane, where $x$ is the weight of a lattice edge covered by a loop. These results are based on transfer-matrix calculations and finite-size scaling. We express the correlation length associated with the staggered loop density in the transfer-matrix eigenvalues. The finite-size data for this correlation length, combined with the scaling formula, reveal the location of critical lines in the diagram. For $n>>2$ we find Ising-like phase transitions associated with the onset of a checkerboard-like ordering of the elementary loops, i.e., the smallest possible loops, with the size of an elementary face, which cover precisely one half of the faces of the square lattice at the maximum loop density. In this respect, the ordered state resembles that of the hard-square lattice gas with nearest-neighbor exclusion, and the finiteness of $n$ represents a softening of its particle-particle potentials. We also determine critical points in the range $-2\leq n\leq 2$. It is found that the topology of the phase diagram depends on the set of allowed vertices of the loop model. Depending on the choice of this set, the $n>2$ transition may continue into the dense phase of the $n \leq 2$ loop model, or continue as a line of $n \leq 2$ O($n$) multicritical points.

preprint2012arXiv

Exact critical points of the O($n$) loop model on the martini and the 3-12 lattices

We derive the exact critical line of the O($n$) loop model on the martini lattice as a function of the loop weight $n$.A finite-size scaling analysis based on transfer matrix calculations is also performed.The numerical results coincide with the theoretical predictions with an accuracy up to 9 decimal places. In the limit $n\to 0$, this gives the exact connective constant $μ=1.7505645579...$ of self-avoiding walks on the martini lattice. Using similar numerical methods, we also study the O($n$) loop model on the 3-12 lattice. We obtain similarly precise agreement with the exact critical points given by Batchelor [J. Stat. Phys. 92, 1203 (1998)].

preprint2010arXiv

Critical frontier for the Potts and percolation models on triangular-type and kagome-type lattices II: Numerical analysis

In a recent paper (arXiv:0911.2514), one of us (FYW) considered the Potts model and bond and site percolation on two general classes of two-dimensional lattices, the triangular-type and kagome-type lattices, and obtained closed-form expressions for the critical frontier with applications to various lattice models. For the triangular-type lattices Wu's result is exact, and for the kagome-type lattices Wu's expression is under a homogeneity assumption. The purpose of the present paper is two-fold: First, an essential step in Wu's analysis is the derivation of lattice-dependent constants $A, B, C$ for various lattice models, a process which can be tedious. We present here a derivation of these constants for subnet networks using a computer algorithm. Secondly, by means of a finite-size scaling analysis based on numerical transfer matrix calculations, we deduce critical properties and critical thresholds of various models and assess the accuracy of the homogeneity assumption. Specifically, we analyze the $q$-state Potts model and the bond percolation on the 3-12 and kagome-type subnet lattices $(n\times n):(n\times n)$, $n\leq 4$, for which the exact solution is not known. To calibrate the accuracy of the finite-size procedure, we apply the same numerical analysis to models for which the exact critical frontiers are known. The comparison of numerical and exact results shows that our numerical determination of critical thresholds is accurate to 7 or 8 significant digits. This in turn infers that the homogeneity assumption determines critical frontiers with an accuracy of 5 decimal places or higher. Finally, we also obtained the exact percolation thresholds for site percolation on kagome-type subnet lattices $(1\times 1):(n\times n)$ for $1\leq n \leq 6$.