Source author record

Rui Han

Rui Han appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

35works
21topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

35 published item(s)

preprint2026arXiv

LLM-Guided Quantified SMT Solving over Uninterpreted Functions

Quantified formulas with Uninterpreted Functions (UFs) over non-linear real arithmetic pose fundamental challenges for Satisfiability Modulo Theories (SMT) solving. Traditional quantifier instantiation methods struggle because they lack semantic understanding of UF constraints, forcing them to search through unbounded solution spaces with limited guidance. We present AquaForte, a framework that leverages Large Language Models to provide semantic guidance for UF instantiation by generating instantiated candidates for function definitions that satisfy the constraints, thereby significantly reducing the search space and complexity for solvers. Our approach preprocesses formulas through constraint separation, uses structured prompts to extract mathematical reasoning from LLMs, and integrates the results with traditional SMT algorithms through adaptive instantiation. AquaForte maintains soundness through systematic validation: LLM-guided instantiations yielding SAT solve the original problem, while UNSAT results generate exclusion clauses for iterative refinement. Completeness is preserved by fallback to traditional solvers augmented with learned constraints. Experimental evaluation on SMT-COMP benchmarks demonstrates that AquaForte solves numerous instances where state-of-the-art solvers like Z3 and CVC5 timeout, with particular effectiveness on satisfiable formulas. Our work shows that LLMs can provide valuable mathematical intuition for symbolic reasoning, establishing a new paradigm for SMT constraint solving.

preprint2026arXiv

On the small denominator problem for generalized Minkowski--Funk transforms

Rubin's generalized Minkowski--Funk transforms $M_t^α$ on the sphere $\mathbb{S}^n$ give rise, for irrational radii $t=\cos(βπ)$, to a small denominator problem governed by the asymptotic behavior of their spectral multipliers. We show that for Lebesgue-almost every $β$ the corresponding two-sine small divisor inequality has infinitely many solutions, and deduce that $(M_t^α)^{-1}$ is not bounded from $\tilde{H}^{s+ρ+1}(\mathbb{S}^n)$ to $H^s(\mathbb{S}^n)$ in the non-critical case $ρ\neq 0,1$. In the critical cases $ρ\in\{0,1\}$ we prove Rubin's Conjectures 4.4 and 4.7 on the failure of endpoint Sobolev regularity for the inverse transforms.

preprint2025arXiv

Influence of ambient temperature on cavitation bubble dynamics

We investigate the influence of ambient temperature on the dynamics of spark-generated cavitation bubbles over a broad temperature range of 23 to 90$^\circ \text{C}$. Increasing temperature, the attenuation of collapse intensity of a bubble in a free field is quantitatively characterised through the Rayleigh factor, minimum bubble volume, and maximum collapse velocity. In scenarios where the bubble is initiated near a rigid boundary, this temperature-dependent weakening effect manifests further as a reduction in jet velocity and bubble migration. Additionally, our findings demonstrate that when ambient temperature exceeds 70$^\circ \text{C}$, secondary cavitation forms near the bubble surface around the moment of maximum bubble expansion, followed by coalescence-induced surface wrinkles. These perturbations trigger Rayleigh-Taylor instability and enhance bubble fission. We determine the internal gas pressure of the bubble at its maximum expansion via the Rayleigh-Plesset equation with the input of bubble radius from experimental measurements. It reveals that the secondary cavitation is derived from the gas pressure descending below the saturated vapor pressure, which provides nucleation-favorable conditions. This study sheds light on the physics behind erosion mitigation in high-temperature fluids from the perspective of cavitation bubble dynamics.

preprint2024arXiv

Cavitation bubble dynamics inside a droplet suspended in a different host fluid

In this paper, we present a theoretical, experimental, and numerical study of the dynamics of cavitation bubbles inside a droplet suspended in another host fluid. On the theoretical side, we provided a modified Rayleigh collapse time and natural frequency for spherical bubbles in our particular context, characterized by the density ratio between the two liquids and the bubble-to-droplet size ratio. Regarding the experimental aspect, experiments were carried out for laser-induced cavitation bubbles inside oil-in-water (O/W) or water-in-oil (W/O) droplets. Two distinct fluid-mixing mechanisms were unveiled in the two systems, respectively. In the case of O/W droplets, a liquid jet emerges around the end of the bubble collapse phase, effectively penetrating the droplet interface. We offer a detailed analysis of the criteria governing jet penetration, involving the standoff parameter and impact velocity of the bubble jet on the droplet surface. Conversely, in the scenario involving W/O droplets, the bubble traverses the droplet interior, inducing global motion and eventually leading to droplet pinch-off when the local Weber number exceeds a critical value. This phenomenon is elucidated through the equilibrium between interfacial and kinetic energies. Lastly, our boundary integral model faithfully reproduces the essential physics of nonspherical bubble dynamics observed in the experiments. We conduct a parametric study spanning a wide parameter space to investigate bubble-droplet interactions. The insights from this study could serve as a valuable reference for practical applications in the field of ultrasonic emulsification, pharmacy, etc.

preprint2023arXiv

Avila's acceleration via zeros of determinants, and applications to Schrödinger cocycles

In this paper we give a characterization of Avila's quantized acceleration of the Lyapunov exponent via the number of zeros of the Dirichlet determinants in finite volume. As applications, we prove $β$-Hölder continuity of the integrated density of states for supercritical quasi-periodic Schrödinger operators restricted to the $\ell$-th stratum, for any $β<(2(\ell-1))^{-1}$ and $\ell\ge2$. We establish Anderson localization for all Diophantine frequencies for the operator with even analytic potential function on the first supercritical stratum, which has positive measure if it is nonempty.

preprint2022arXiv

Atrial Fibrillation Detection Using Weight-Pruned, Log-Quantised Convolutional Neural Networks

Deep neural networks (DNN) are a promising tool in medical applications. However, the implementation of complex DNNs on battery-powered devices is challenging due to high energy costs for communication. In this work, a convolutional neural network model is developed for detecting atrial fibrillation from electrocardiogram (ECG) signals. The model demonstrates high performance despite being trained on limited, variable-length input data. Weight pruning and logarithmic quantisation are combined to introduce sparsity and reduce model size, which can be exploited for reduced data movement and lower computational complexity. The final model achieved a 91.1% model compression ratio while maintaining high model accuracy of 91.7% and less than 1% loss.

preprint2022arXiv

Decay of multi-point correlation functions in $\mathbb{Z}^d$

We prove multi-point correlation bounds in $\mathbb{Z}^d$ for arbitrary $d\geq 1$ with symmetrized distances, answering open questions proposed by Sims-Warzel \cite{SW} and Aza-Bru-Siqueira Pedra \cite{ABP}. As applications, we prove multi-point correlation bounds for the Ising model on $\mathbb{Z}^d$, and multi-point dynamical localization in expectation for uniformly localized disordered systems, which provides the first examples of this conjectured phenomenon by Bravyi-König \cite{BK}.

preprint2021arXiv

Density of states and Delocalization for discrete magnetic random Schrödinger operators

We study discrete magnetic random Schrödinger operators on the square and honeycomb lattice. For the non-random magnetic operator on the hexagonal lattice with any rational magnetic flux, we show that the middle two dispersion surfaces exhibit Dirac cones. We then derive an asymptotic expansion for the density of states on the honeycomb lattice for oscillations of arbitrary rational magnetic flux. This allows us, as a corollary, to rigorously study the quantum Hall effect and conclude dynamical delocalization close to the conical point under disorder. We obtain similar results for the discrete random Schrödinger operator on the $\mathbb Z^2$-lattice with weak magnetic fields, close to the bottom and top of its spectrum.

preprint2021arXiv

Universally Optimal Verification of Entangled States with Nondemolition Measurements

The efficient and reliable characterization of quantum states plays a vital role in most, if not all, quantum information processing tasks. In this work, we present a universally optimal protocol for verifying entangled states by employing the so-called quantum nondemolition measurements, such that the verification efficiency is equivalent to that of the optimal global strategy. Instead of being probabilistic as the standard verification strategies, our protocol is constructed sequentially, which is thus more favorable for experimental realizations. In addition, the target states are preserved in the protocol after each measurement, so can be reused in any subsequent tasks. We demonstrate the power of our protocol for the optimal verification of Bell states, arbitrary two-qubit pure states, and stabilizer states. We also prove that our protocol is able to perform tasks including fidelity estimation and state preparation.

preprint2020arXiv

Averages Along the Primes: Improving and Sparse Bounds

Consider averages along the prime integers $ \mathbb P $ given by \begin{equation*} \mathcal{A}_N f (x) = N ^{-1} \sum_{ p \in \mathbb P \;:\; p\leq N} (\log p) f (x-p). \end{equation*} These averages satisfy a uniform scale-free $ \ell ^{p}$-improving estimate. For all $ 1< p < 2$, there is a constant $ C_p$ so that for all integer $ N$ and functions $ f$ supported on $ [0,N]$, there holds \begin{equation*} N ^{-1/p' }\lVert \mathcal{A}_N f\rVert_{\ell^{p'}} \leq C_p N ^{- 1/p} \lVert f\rVert_{\ell^p}. \end{equation*} The maximal function $ \mathcal{A}^{\ast} f =\sup_{N} \lvert \mathcal{A}_N f \rvert$ satisfies $ (p,p)$ sparse bounds for all $ 1< p < 2$. The latter are the natural variants of the scale-free bounds. As a corollary, $ \mathcal{A}^{\ast} $ is bounded on $ \ell ^{p} (w)$, for all weights $ w$ in the Muckenhoupt $A_p$ class. No prior weighted inequalities for $ \mathcal{A}^{\ast} $ were known.

preprint2020arXiv

ExpertNet: Adversarial Learning and Recovery Against Noisy Labels

Today's available datasets in the wild, e.g., from social media and open platforms, present tremendous opportunities and challenges for deep learning, as there is a significant portion of tagged images, but often with noisy, i.e. erroneous, labels. Recent studies improve the robustness of deep models against noisy labels without the knowledge of true labels. In this paper, we advocate to derive a stronger classifier which proactively makes use of the noisy labels in addition to the original images - turning noisy labels into learning features. To such an end, we propose a novel framework, ExpertNet, composed of Amateur and Expert, which iteratively learn from each other. Amateur is a regular image classifier trained by the feedback of Expert, which imitates how human experts would correct the predicted labels from Amateur using the noise pattern learnt from the knowledge of both the noisy and ground truth labels. The trained Amateur and Expert proactively leverage the images and their noisy labels to infer image classes. Our empirical evaluations on noisy versions of CIFAR-10, CIFAR-100 and real-world data of Clothing1M show that the proposed model can achieve robust classification against a wide range of noise ratios and with as little as 20-50% training data, compared to state-of-the-art deep models that solely focus on distilling the impact of noisy labels.

preprint2020arXiv

Honeycomb structures in magnetic fields

We consider reduced-dimensionality models of honeycomb lattices in magnetic fields and report results about the spectrum, the density of states, self-similarity, and metal/insulator transitions under disorder. We perform a spectral analysis by which we discover a fractal Cantor spectrum for irrational magnetic flux through a honeycomb, prove the existence of zero energy Dirac cones for each rational flux, obtain an explicit expansion of the density of states near the conical points, and show the existence of mobility edges under Anderson-type disorder. Our results give a precise description of de Haas-van Alphen and Quantum Hall effects, and provide quantitative estimates on transport properties. In particular, our findings explain experimentally observed asymmetry phenomena by going beyond the perfect cone approximation.

preprint2020arXiv

Improving estimates for discrete polynomial averages

For a polynomial $P$ mapping the integers into the integers, define an averaging operator $A_{N} f(x):=\frac{1}{N}\sum_{k=1}^N f(x+P(k))$ acting on functions on the integers. We prove sufficient conditions for the $\ell^{p}$-improving inequality \begin{equation*} \|A_N f\|_{\ell^q(\mathbb{Z})} \lesssim_{P,p,q} N^{-d(\frac{1}{p}-\frac{1}{q})} \|f\|_{\ell^p(\mathbb{Z})}, \qquad N \in\mathbb{N}, \end{equation*} where $1\leq p \leq q \leq \infty$. For a range of quadratic polynomials, the inequalities established are sharp, up to the boundary of the allowed pairs of $(p,q)$. For degree three and higher, the inequalities are close to being sharp. In the quadratic case, we appeal to discrete fractional integrals as studied by Stein and Wainger. In the higher degree case, we appeal to the Vinogradov Mean Value Theorem, recently established by Bourgain, Demeter, and Guth.

preprint2020arXiv

Large deviation estimates and Hölder regularity of the Lyapunov exponents for quasi-periodic Schrödinger cocycles

We consider one-dimensional quasi-periodic Schrödinger operators with analytic potentials. In the positive Lyapunov exponent regime, we prove large deviation estimates which lead to optimal Hölder continuity of the Lyapunov exponents and the integrated density of states, in both small Lyapunov exponent and large coupling regimes. Our results cover all the Diophantine frequencies and some Liouville frequencies.

preprint2020arXiv

Mass Yields of Fission Fragment of Pt to Ra Isotopes

An effective Fourier nuclear shape parametrization which describes well the most relevant degrees of freedom on the way to fission is used to construct a 3D collective model. The potential energy surface is evaluated within the macroscopic-microscopic approach based on the Lublin-Strasbourg Drop (LSD) macroscopic energy and Yukawa-folded single particle potential. A phenomenological inertia parameter is used to describe the kinetic properties of the fissioning system. The fission fragment mass yields are obtained by using an approximate solution of the underlying Hamiltonian. The predicted mass fragmentations for even-even Pt to Ra isotopes are compared with available experimental data. Their main characteristics are well reproduced when the neck rupture probability dependent on the neck radius is introduced.

preprint2019arXiv

Extracting unambiguous information from a single qubit by sequential observers

In a recent paper [Phys. Rev. Lett. 111, 100501 (2013)], a scheme was proposed where subsequent observers can extract unambiguous information about the initial state of a qubit, with finite joint probability of success. Here, we generalize the problem for arbitrary preparation probabilities (arbitrary priors). We discuss two different schemes: one where only the joint probability of success is maximized and another where, in addition, the joint probability of failure is also minimized. We also derive the mutual information for these schemes and show that there are some parameter regions for the scheme without minimizing the joint failure probability where, even though the joint success probability is maximum, no information is actually transmitted by Alice.

preprint2019arXiv

Low-$n$ global ideal MHD instabilities in CFETR baseline scenario

This article reports an evaluation on the linear ideal magnetohydrodynamic (MHD) stability of the China Fusion Engineering Test Reactor (CFETR) baseline scenario for various first-wall locations. The initial-value code NIMROD and eigen-value code AEGIS are employed in this analysis. A good agreement is achieved between two codes in the growth rates of $n=1-10$ ideal MHD modes for various locations of the perfect conducting first-wall. The higher-$n$ modes are dominated by ballooning modes and localized in the pedestal region, while the lower-$n$ modes have more prominent external kink components and broader mode profiles. The influences of plasma-vacuum profile and wall shape are also examined using NIMROD. In presence of resistive wall, the low-$n$ ideal MHD instabilities are further studied using AEGIS. For the designed first-wall location, the $n = 1$ resistive wall mode (RWM) is found unstable, which could be fully stabilized by uniform toroidal rotation above 2.9\% core Alfvén speed.

preprint2018arXiv

A higher dimensional Bourgain-Dyatlov fractal uncertainty principle

We establish a version of the fractal uncertainty principle, obtained by Bourgain and Dyatlov in 2016, in higher dimensions. The Fourier support is limited to sets $Y\subset \mathbb{R}^d$ which can be covered by finitely many products of $δ$-regular sets in one dimension, but relative to arbitrary axes. Our results remain true if $Y$ is distorted by diffeomorphisms. Our method combines the original approach by Bourgain and Dyatlov, in the more quantitative 2017 rendition by Jin and Zhang, with Cartan set techniques.

preprint2017arXiv

The Helstrom measurement: A nondestructive implementation

We discuss a novel implementation of the minimum error state discrimination measurement, originally introduced by Helstrom. In this implementation, instead of performing the optimal projective measurement directly on the system, it is first entangled to an ancillary system and the measurement is performed on the ancilla. We show that, by an appropriate choice of the entanglement transformation, the Helstrom bound can be attained. The advantage of this approach is twofold. First, it provides a novel implementation when the optimal projective measurement cannot be directly performed. For example, in the case of continuous variable states (binary and N phase-shifted coherent signals), the available detection methods, photon counting and homodyning, are insufficient to perform the required cat-state projection. In the case of symmetric states, the square-root measurement is optimal, but it is not easy to perform directly for more than two states. Our approach provides a feasible alternative in both cases. Second, the measurement is non-destructive from the point of view of the original system and one has a certain amount of freedom in designing the post-measurement state, which can then be processed further.

preprint2016arXiv

AccuracyTrader: Accuracy-aware Approximate Processing for Low Tail Latency and High Result Accuracy in Cloud Online Services

Modern latency-critical online services such as search engines often process requests by consulting large input data spanning massive parallel components. Hence the tail latency of these components determines the service latency. To trade off result accuracy for tail latency reduction, existing techniques use the components responding before a specified deadline to produce approximate results. However, they may skip a large proportion of components when load gets heavier, thus incurring large accuracy losses. This paper presents AccuracyTrader that produces approximate results with small accuracy losses while maintaining low tail latency. AccuracyTrader aggregates information of input data on each component to create a small synopsis, thus enabling all components producing initial results quickly using their synopses. AccuracyTrader also uses synopses to identify the parts of input data most related to arbitrary requests' result accuracy, thus first using these parts to improve the produced results in order to minimize accuracy losses. We evaluated AccuracyTrader using workloads in real services. The results show: (i) AccuracyTrader reduces tail latency by over 40 times with accuracy losses of less than 7% compared to existing exact processing techniques; (ii) when using the same latency, AccuracyTrader reduces accuracy losses by over 13 times comparing to existing approximate processing techniques.

preprint2015arXiv

Benchmarking Big Data Systems: State-of-the-Art and Future Directions

The great prosperity of big data systems such as Hadoop in recent years makes the benchmarking of these systems become crucial for both research and industry communities. The complexity, diversity, and rapid evolution of big data systems gives rise to various new challenges about how we design generators to produce data with the 4V properties (i.e. volume, velocity, variety and veracity), as well as implement application-specific but still comprehensive workloads. However, most of the existing big data benchmarks can be described as attempts to solve specific problems in benchmarking systems. This article investigates the state-of-the-art in benchmarking big data systems along with the future challenges to be addressed to realize a successful and efficient benchmark.

preprint2015arXiv

BigDataBench-MT: A Benchmark Tool for Generating Realistic Mixed Data Center Workloads

Long-running service workloads (e.g. web search engine) and short-term data analysis workloads (e.g. Hadoop MapReduce jobs) co-locate in today's data centers. Developing realistic benchmarks to reflect such practical scenario of mixed workload is a key problem to produce trustworthy results when evaluating and comparing data center systems. This requires using actual workloads as well as guaranteeing their submissions to follow patterns hidden in real-world traces. However, existing benchmarks either generate actual workloads based on probability models, or replay real-world workload traces using basic I/O operations. To fill this gap, we propose a benchmark tool that is a first step towards generating a mix of actual service and data analysis workloads on the basis of real workload traces. Our tool includes a combiner that enables the replaying of actual workloads according to the workload traces, and a multi-tenant generator that flexibly scales the workloads up and down according to users' requirements. Based on this, our demo illustrates the workload customization and generation process using a visual interface. The proposed tool, called BigDataBench-MT, is a multi-tenant version of our comprehensive benchmark suite BigDataBench and it is publicly available from http://prof.ict.ac.cn/BigDataBench/multi-tenancyversion/.

preprint2015arXiv

Characterization and Architectural Implications of Big Data Workloads

Big data areas are expanding in a fast way in terms of increasing workloads and runtime systems, and this situation imposes a serious challenge to workload characterization, which is the foundation of innovative system and architecture design. The previous major efforts on big data benchmarking either propose a comprehensive but a large amount of workloads, or only select a few workloads according to so-called popularity, which may lead to partial or even biased observations. In this paper, on the basis of a comprehensive big data benchmark suite---BigDataBench, we reduced 77 workloads to 17 representative workloads from a micro-architectural perspective. On a typical state-of-practice platform---Intel Xeon E5645, we compare the representative big data workloads with SPECINT, SPECCFP, PARSEC, CloudSuite and HPCC. After a comprehensive workload characterization, we have the following observations. First, the big data workloads are data movement dominated computing with more branch operations, taking up to 92% percentage in terms of instruction mix, which places them in a different class from Desktop (SPEC CPU2006), CMP (PARSEC), HPC (HPCC) workloads. Second, corroborating the previous work, Hadoop and Spark based big data workloads have higher front-end stalls. Comparing with the traditional workloads i. e. PARSEC, the big data workloads have larger instructions footprint. But we also note that, in addition to varied instruction-level parallelism, there are significant disparities of front-end efficiencies among different big data workloads. Third, we found complex software stacks that fail to use state-of-practise processors efficiently are one of the main factors leading to high front-end stalls. For the same workloads, the L1I cache miss rates have one order of magnitude differences among diverse implementations with different software stacks.

preprint2015arXiv

Investigations into Elasticity in Cloud Computing

The pay-as-you-go model supported by existing cloud infrastructure providers is appealing to most application service providers to deliver their applications in the cloud. Within this context, elasticity of applications has become one of the most important features in cloud computing. This elasticity enables real-time acquisition/release of compute resources to meet application performance demands. In this thesis we investigate the problem of delivering cost-effective elasticity services for cloud applications. Traditionally, the application level elasticity addresses the question of how to scale applications up and down to meet their performance requirements, but does not adequately address issues relating to minimising the costs of using the service. With this current limitation in mind, we propose a scaling approach that makes use of cost-aware criteria to detect the bottlenecks within multi-tier cloud applications, and scale these applications only at bottleneck tiers to reduce the costs incurred by consuming cloud infrastructure resources. Our approach is generic for a wide class of multi-tier applications, and we demonstrate its effectiveness by studying the behaviour of an example electronic commerce site application. Furthermore, we consider the characteristics of the algorithm for implementing the business logic of cloud applications, and investigate the elasticity at the algorithm level: when dealing with large-scale data under resource and time constraints, the algorithm's output should be elastic with respect to the resource consumed. We propose a novel framework to guide the development of elastic algorithms that adapt to the available budget while guaranteeing the quality of output result, e.g. prediction accuracy for classification tasks, improves monotonically with the used budget.

preprint2015arXiv

PCS: Predictive Component-level Scheduling for Reducing Tail Latency in Cloud Online Services

Modern latency-critical online services often rely on composing results from a large number of server components. Hence the tail latency (e.g. the 99th percentile of response time), rather than the average, of these components determines the overall service performance. When hosted on a cloud environment, the components of a service typically co-locate with short batch jobs to increase machine utilizations, and share and contend resources such as caches and I/O bandwidths with them. The highly dynamic nature of batch jobs in terms of their workload types and input sizes causes continuously changing performance interference to individual components, hence leading to their latency variability and high tail latency. However, existing techniques either ignore such fine-grained component latency variability when managing service performance, or rely on executing redundant requests to reduce the tail latency, which adversely deteriorate the service performance when load gets heavier. In this paper, we propose PCS, a predictive and component-level scheduling framework to reduce tail latency for large-scale, parallel online services. It uses an analytical performance model to simultaneously predict the component latency and the overall service performance on different nodes. Based on the predicted performance, the scheduler identifies straggling components and conducts near-optimal component-node allocations to adapt to the changing performance interferences from batch jobs. We demonstrate that, using realistic workloads, the proposed scheduler reduces the component tail latency by an average of 67.05\% and the average overall service latency by 64.16\% compared with the state-of-the-art techniques on reducing tail latency.

preprint2014arXiv

BDGS: A Scalable Big Data Generator Suite in Big Data Benchmarking

Data generation is a key issue in big data benchmarking that aims to generate application-specific data sets to meet the 4V requirements of big data. Specifically, big data generators need to generate scalable data (Volume) of different types (Variety) under controllable generation rates (Velocity) while keeping the important characteristics of raw data (Veracity). This gives rise to various new challenges about how we design generators efficiently and successfully. To date, most existing techniques can only generate limited types of data and support specific big data systems such as Hadoop. Hence we develop a tool, called Big Data Generator Suite (BDGS), to efficiently generate scalable big data while employing data models derived from real data to preserve data veracity. The effectiveness of BDGS is demonstrated by developing six data generators covering three representative data types (structured, semi-structured and unstructured) and three data sources (text, graph, and table data).

preprint2014arXiv

Characterizing and Subsetting Big Data Workloads

Big data benchmark suites must include a diversity of data and workloads to be useful in fairly evaluating big data systems and architectures. However, using truly comprehensive benchmarks poses great challenges for the architecture community. First, we need to thoroughly understand the behaviors of a variety of workloads. Second, our usual simulation-based research methods become prohibitively expensive for big data. As big data is an emerging field, more and more software stacks are being proposed to facilitate the development of big data applications, which aggravates hese challenges. In this paper, we first use Principle Component Analysis (PCA) to identify the most important characteristics from 45 metrics to characterize big data workloads from BigDataBench, a comprehensive big data benchmark suite. Second, we apply a clustering technique to the principle components obtained from the PCA to investigate the similarity among big data workloads, and we verify the importance of including different software stacks for big data benchmarking. Third, we select seven representative big data workloads by removing redundant ones and release the BigDataBench simulation version, which is publicly available from http://prof.ict.ac.cn/BigDataBench/simulatorversion/.

preprint2014arXiv

Neutron Time-Of-Flight Spectrometer Based on HIRFL for Studies of Spallation Reactions Related to ADS Project

A Neutron Time-Of-Flight (NTOF) spectrometer based on Heavy Ion Research Facility in Lanzhou (HIRFL) is developed for studies of neutron production of proton induced spallation reactions related to the ADS project. After the presentation of comparisons between calculated spallation neutron production double-differential cross sections and the available experimental one, a detailed description of NTOF spectrometer is given. Test beam results show that the spectrometer works well and data analysis procedures are established. The comparisons of the test beam neutron spectra with those of GEANT4 simulations are presented.

preprint2014arXiv

On Big Data Benchmarking

Big data systems address the challenges of capturing, storing, managing, analyzing, and visualizing big data. Within this context, developing benchmarks to evaluate and compare big data systems has become an active topic for both research and industry communities. To date, most of the state-of-the-art big data benchmarks are designed for specific types of systems. Based on our experience, however, we argue that considering the complexity, diversity, and rapid evolution of big data systems, for the sake of fairness, big data benchmarks must include diversity of data and workloads. Given this motivation, in this paper, we first propose the key requirements and challenges in developing big data benchmarks from the perspectives of generating data with 4V properties (i.e. volume, velocity, variety and veracity) of big data, as well as generating tests with comprehensive workloads for big data systems. We then present the methodology on big data benchmarking designed to address these challenges. Next, the state-of-the-art are summarized and compared, following by our vision for future research directions.

preprint2013arXiv

Beyond adiabatic elimination: A hierarchy of approximations for multi-photon processes

In multi-level systems, the commonly used adiabatic elimination is a method for approximating the dynamics of the system by eliminating irrelevant, non-resonantly coupled levels. This procedure is, however, somewhat ambiguous and it is not clear how to improve on it systematically. We use an integro-differential equation for the probability amplitudes of the levels of interest, which is equivalent to the original Schrodinger equation for all probability amplitudes. In conjunction with a Markov approximation, the integro-differential equation is then used to generate a hierarchy of approximations, in which the zeroth order is the adiabatic-elimination approximation. It works well with a proper choice of interaction picture; the procedure suggests criteria for optimizing this choice. The first-order approximation in the hierarchy provides significant improvements over standard adiabatic elimination, without much increase in complexity, and is furthermore not so sensitive to the choice of interaction picture. We illustrate these points with several examples.

preprint2012arXiv

Raman transitions without adiabatic elimination: A simple and accurate treatment

Driven Raman processes --- nearly resonant two-photon transitions through an intermediate state that is non-resonantly coupled and does not acquire a sizeable population --- are commonly treated with a simplified description in which the intermediate state is removed by adiabatic elimination. While the adiabatic-elimination approximation is reliable when the detuning of the intermediate state is quite large, it cannot be trusted in other situations, and it does not allow one to estimate the population in the eliminated state. We introduce an alternative method that keeps all states in the description, without increasing the complexity by much. An integro-differential equation of Lippmann-Schwinger type generates a hierarchy of approximations, but very accurate results are already obtained in the lowest order.