Source author record

Andrew Chi-Chih Yao

Andrew Chi-Chih Yao appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

4works
5topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

4 published item(s)

preprint2026arXiv

Tensor Product Attention Is All You Need

Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this paper, we propose Tensor Product Attention (TPA), a novel attention mechanism that uses tensor decompositions to represent queries, keys, and values compactly, substantially shrinking the KV cache size at inference time. By factorizing these representations into contextual low-rank components and seamlessly integrating with Rotary Position Embedding (RoPE), TPA achieves improved model quality alongside memory efficiency. Based on TPA, we introduce the Tensor ProducT ATTenTion Transformer (T6), a new model architecture for sequence modeling. Through extensive empirical evaluation on language modeling tasks, we demonstrate that T6 surpasses or matches the performance of standard Transformer baselines including Multi-Head Attention (MHA), Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and Multi-Head Latent Attention (MLA) across various metrics, including perplexity and a range of established evaluation benchmarks. Notably, TPA's memory efficiency and computational efficiency at decoding stage enables processing longer sequences under fixed resource constraints, addressing a critical scalability challenge in modern language models. Project Page: https://github.com/tensorgi/TPA.

preprint2016arXiv

On Solutions for the Maximum Revenue Multi-item Auction under Dominant-Strategy and Bayesian Implementations

Very few exact solutions are known for the monopolist's $k$-item $n$-buyer maximum revenue problem with additive valuation in which $k, n >1$ and the buyers $i$ have independent private distributions $F^j_i$ on items $j$. In this paper we derive exact formulas for the maximum revenue when $k=2$ and $F^j_i$ are any IID distributions on support of size 2, for both the dominant-strategy (DIC) and the Bayesian (BIC) implementations. The formulas lead to the simple characterization that, the two implementations have identical maximum revenue if and only if selling-separately is optimal for the distribution. Our results also give the first demonstration, in this setting, of revenue gaps between the two implementations. For instance, if $k=n=2$ and $Pr\{X_F=1\}=Pr\{X_F=2\}=\frac{1}{2}$, then the maximum revenue in the Bayesian implementation exceeds that in the dominant-strategy by exactly $2\%$; the same gap exists for the continuous uniform distribution $X_F$ over $[a, a+1]\cup[2a, 2a+1]$ for all large $a$.

preprint2014arXiv

An n-to-1 Bidder Reduction for Multi-item Auctions and its Applications

In this paper, we introduce a novel approach for reducing the $k$-item $n$-bidder auction with additive valuation to $k$-item $1$-bidder auctions. This approach, called the \emph{Best-Guess} reduction, can be applied to address several central questions in optimal revenue auction theory such as the power of randomization, and Bayesian versus dominant-strategy implementations. First, when the items have independent valuation distributions, we present a deterministic mechanism called {\it Deterministic Best-Guess} that yields at least a constant fraction of the optimal revenue by any randomized mechanism. Second, if all the $nk$ valuation random variables are independent, the optimal revenue achievable in {\it dominant strategy incentive compatibility} (DSIC) is shown to be at least a constant fraction of that achievable in {\it Bayesian incentive compatibility} (BIC). Third, when all the $nk$ values are identically distributed according to a common one-dimensional distribution $F$, the optimal revenue is shown to be expressible in the closed form $Θ(k(r+\int_0^{mr} (1-F(x)^n) \ud x))$ where $r= sup_{x\geq 0} \, x(1 - F(x)^n)$ and $m=\lceil k/n\rceil$; this revenue is achievable by a simple mechanism called \emph{2nd-Price Bundling}. All our results apply to arbitrary distributions, regular or irregular.

preprint2014arXiv

Quantum replication at the Heisenberg limit

No process in nature can perfectly clone an arbitrary quantum state. But is it possible to engineer processes that replicate quantum information with vanishingly small error? Here we demonstrate the possibility of probabilistic super-replication phenomena where N equally prepared quantum clocks are transformed into a much larger number of M nearly perfect replicas, with an error that rapidly vanishes whenever M is small compared to the square of N. The quadratic replication rate is the ultimate limit imposed by Quantum Mechanics to the proliferation of information and is fundamentally linked with the Heisenberg limit of quantum metrology.