Source author record

David Wang

David Wang appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Computer Vision eess.IV Machine Learning math.NT physics.app-ph physics.chem-ph Robotics

Catalog footprint

What is connected

5works

7topics

4close collaborators

Actions

Connect this record

Open graph Browse works

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

preprint2026arXiv

STEP3-VL-10B Technical Report

We present STEP3-VL-10B, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. STEP3-VL-10B is realized through two strategic shifts: first, a unified, fully unfrozen pre-training strategy on 1.2T multimodal tokens that integrates a language-aligned Perception Encoder with a Qwen3-8B decoder to establish intrinsic vision-language synergy; and second, a scaled post-training pipeline featuring over 1k iterations of reinforcement learning. Crucially, we implement Parallel Coordinated Reasoning (PaCoRe) to scale test-time compute, allocating resources to scalable perceptual reasoning that explores and synthesizes diverse visual hypotheses. Consequently, despite its compact 10B footprint, STEP3-VL-10B rivals or surpasses models 10$\times$-20$\times$ larger (e.g., GLM-4.6V-106B, Qwen3-VL-235B) and top-tier proprietary flagships like Gemini 2.5 Pro and Seed-1.5-VL. Delivering best-in-class performance, it records 92.2% on MMBench and 80.11% on MMMU, while excelling in complex reasoning with 94.43% on AIME2025 and 75.95% on MathVision. We release the full model suite to provide the community with a powerful, efficient, and reproducible baseline.

preprint2020arXiv

SK-Net: Deep Learning on Point Cloud via End-to-end Discovery of Spatial Keypoints

Since the PointNet was proposed, deep learning on point cloud has been the concentration of intense 3D research. However, existing point-based methods usually are not adequate to extract the local features and the spatial pattern of a point cloud for further shape understanding. This paper presents an end-to-end framework, SK-Net, to jointly optimize the inference of spatial keypoint with the learning of feature representation of a point cloud for a specific point cloud task. One key process of SK-Net is the generation of spatial keypoints (Skeypoints). It is jointly conducted by two proposed regulating losses and a task objective function without knowledge of Skeypoint location annotations and proposals. Specifically, our Skeypoints are not sensitive to the location consistency but are acutely aware of shape. Another key process of SK-Net is the extraction of the local structure of Skeypoints (detail feature) and the local spatial pattern of normalized Skeypoints (pattern feature). This process generates a comprehensive representation, pattern-detail (PD) feature, which comprises the local detail information of a point cloud and reveals its spatial pattern through the part district reconstruction on normalized Skeypoints. Consequently, our network is prompted to effectively understand the correlation between different regions of a point cloud and integrate contextual information of the point cloud. In point cloud tasks, such as classification and segmentation, our proposed method performs better than or comparable with the state-of-the-art approaches. We also present an ablation study to demonstrate the advantages of SK-Net.

preprint2019arXiv

Design Principles for Self-forming Interfaces Enabling Stable Lithium Metal Anodes

The path toward Li-ion batteries with higher energy-densities will likely involve use of thin lithium metal (Li) anode (<50 $μ$m in thickness), whose cyclability today remains limited by dendrite formation and low Coulombic efficiency. Previous studies have shown that the solid-electrolyte-interface (SEI) of Li metal plays a crucial role in Li electrodeposition and stripping. However, design rules for optimal SEIs on lithium metal are not well-established. Here, using integrated experimental and modeling studies on a series of structurally-similar SEI-modifying compounds as model systems, we reveal the relationship between SEI compositions, Li deposition morphology and coulombic efficiency, and identify two key descriptors (ionicity and compactness) for high performance SEIs through integrated experimental and modeling studies. Using this understanding, we design a highly ionic and compact SEI that shows excellent cycling performance in LiCoO$_2$-Li full cells at practical current densities. Our results provide guidance for the rational selection and optimization of SEI modifiers to further improve Li metal anodes.

preprint2019arXiv

Mechanical Search: Multi-Step Retrieval of a Target Object Occluded by Clutter

When operating in unstructured environments such as warehouses, homes, and retail centers, robots are frequently required to interactively search for and retrieve specific objects from cluttered bins, shelves, or tables. Mechanical Search describes the class of tasks where the goal is to locate and extract a known target object. In this paper, we formalize Mechanical Search and study a version where distractor objects are heaped over the target object in a bin. The robot uses an RGBD perception system and control policies to iteratively select, parameterize, and perform one of 3 actions -- push, suction, grasp -- until the target object is extracted, or either a time limit is exceeded, or no high confidence push or grasp is available. We present a study of 5 algorithmic policies for mechanical search, with 15,000 simulated trials and 300 physical trials for heaps ranging from 10 to 20 objects. Results suggest that success can be achieved in this long-horizon task with algorithmic policies in over 95% of instances and that the number of actions required scales approximately linearly with the size of the heap. Code and supplementary material can be found at http://ai.stanford.edu/mech-search .

preprint2018arXiv

Some Families of Super Congruences Involving Alternating Multiple Harmonic Sums

Let $p$ be a prime. In this short note we study some families of super congruences involving the following alternating sums \begin{equation*} \sum_{\substack{j_1+j_2+\cdots+j_n=2 p^r p\nmid j_1 j_2 \cdots j_n}} \frac{(-1)^{j_1+\cdots+j_b}}{j_1\cdots j_n} \pmod{p^r}, \end{equation*} which extend similar statements proved by Shen and Cai who treated the cases when $n=4,5$.