Source author record

M. A. Clark

M. A. Clark appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

32works
10topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

32 published item(s)

preprint2022arXiv

Detailed analysis of excited state systematics in a lattice QCD calculation of $g_A$

Excited state contamination remains one of the most challenging sources of systematic uncertainty to control in lattice QCD calculations of nucleon matrix elements and form factors: early time separations are contaminated by excited states and late times suffer from an exponentially bad signal-to-noise problem. High-statistics calculations at large time separations $\gtrsim1$ fm are commonly used to combat these issues. In this work, focusing on $g_A$, we explore the alternative strategy of utilizing a large number of relatively low-statistics calculations at short to medium time separations (0.2--1 fm), combined with a multi-state analysis. On an ensemble with a pion mass of approximately 310 MeV and a lattice spacing of approximately 0.09 fm, we find this provides a more robust and economical method of quantifying and controlling the excited state systematic uncertainty. A quantitative separation of various types of excited states enables the identification of the transition matrix elements as the dominant contamination. The excited state contamination of the Feynman-Hellmann correlation function is found to reduce to the 1% level at approximately 1 fm while for the more standard three-point functions, this does not occur until after 2 fm. Critical to our findings is the use of a global minimization, rather than fixing the spectrum from the two-point functions and using them as input to the three-point analysis. We find that the ground state parameters determined in such a global analysis are stable against variations in the excited state model, the number of excited states, and the truncation of early-time or late-time numerical data.

preprint2022arXiv

QED with massive photons for precision physics: zero modes and first result for the hadron spectrum

The current precision reached by lattice QCD calculations of low-energy hadronic observables, requires not only the introduction of electromagnetic corrections, but also control over all the potential systematic uncertainties introduced by the lattice version of QED. Introducing a massive photon as an infrared regulator in lattice QED, provides a well defined theory, dubbed QEDM, amenable to numerical evaluation [arXiv:1507.08916]. The photon mass is removed through extrapolation. In this contribution we scrutinise aspects of QEDM such as the presence and fate of the zero modes contributions and we describe the determination of the photon mass corrections in finite and infinite volume. We demonstrate that the required extrapolations are well controlled using numerical data obtained on two ensembles which only differ in volume.

preprint2022arXiv

The hyperon spectrum from lattice QCD

Hyperon decays present a promising alternative for extracting $\vert V_{us} \vert$ from lattice QCD combined with experimental measurements. Currently $\vert V_{us} \vert$ is determined from the kaon decay widths and a lattice calculation of the associated form factor. In this proceeding, I will present preliminary work on a lattice determination of the hyperon mass spectrum. I will additionally summarize future goals in which we will calculate the hyperon transition matrix elements, which will provide an alternative means for accessing $\vert V_{us} \vert$. This work is based on a particular formulation of SU(2) chiral perturbation theory for hyperons; determining the extent to which this effective field theory converges is instrumental in understanding the limits of its predictive power, especially since some hyperonic observables are difficult to calculate near the physical pion mass (e.g., hyperon-to-nucleon form factors), and thus the use of heavier than physical pion masses is likely to yield more precise results when combined with extrapolations to the physical point.}

preprint2021arXiv

Two-nucleon S-wave interactions at the $SU(3)$ flavor-symmetric point with $m_{ud}\simeq m_s^{\rm phys}$: a first lattice QCD calculation with the stochastic Laplacian Heaviside method

We report on the first application of the stochastic Laplacian Heaviside method for computing multi-particle interactions with lattice QCD to the two-nucleon system. Like the Laplacian Heaviside method, this method allows for the construction of interpolating operators which can be used to construct a positive definite set of two-nucleon correlation functions, unlike nearly all other applications of lattice QCD to two nucleons in the literature. It also allows for a variational analysis in which optimal linear combinations of the interpolating operators are formed that couple predominantly to the eigenstates of the system. Utilizing such methods has become of paramount importance in order to help resolve the discrepancy in the literature on whether two nucleons in either isospin channel form a bound state at pion masses heavier than physical, with the discrepancy persisting even in the $SU(3)$-flavor symmetric point with all quark masses near the physical strange quark mass. This is the first in a series of papers aimed at resolving this discrepancy. In the present work, we employ the stochastic Laplacian Heaviside method without a hexaquark operator in the basis at a lattice spacing of $a\sim0.086$~fm, lattice volume of $L=48a\simeq4.1$~fm and pion mass $m_π\simeq714$ MeV. With this setup, the observed spectrum of two-nucleon energy levels strongly disfavors the presence of a bound state in either the deuteron or dineutron channel.

preprint2020arXiv

$F_K / F_π$ from Möbius domain-wall fermions solved on gradient-flowed HISQ ensembles

We report the results of a lattice quantum chromodynamics calculation of $F_K/F_π$ using Möbius domain-wall fermions computed on gradient-flowed $N_f=2+1+1$ highly-improved staggered quark (HISQ) ensembles. The calculation is performed with five values of the pion mass ranging from $130 \lesssim m_π\lesssim 400$ MeV, four lattice spacings of $a\sim 0.15, 0.12, 0.09$ and $0.06$ fm and multiple values of the lattice volume. The interpolation/extrapolation to the physical pion and kaon mass point, the continuum, and infinite volume limits are performed with a variety of different extrapolation functions utilizing both the relevant mixed-action effective field theory expressions as well as discretization-enhanced continuum chiral perturbation theory formulas. We find that the $a\sim0.06$ fm ensemble is helpful, but not necessary to achieve a subpercent determination of $F_K/F_π$. We also include an estimate of the strong isospin breaking corrections and arrive at a final result of $F_{K^\pm}/F_{π^\pm} = 1.1942(45)$ with all sources of statistical and systematic uncertainty included. This is consistent with the Flavour Lattice Averaging Group average value, providing an important benchmark for our lattice action. Combining our result with experimental measurements of the pion and kaon leptonic decays leads to a determination of $|V_{us}|/|V_{ud}| = 0.2311(10)$.

preprint2019arXiv

Lattice QCD Determination of $g_A$

The nucleon axial coupling, $g_A$, is a fundamental property of protons and neutrons, dictating the strength with which the weak axial current of the Standard Model couples to nucleons, and hence, the lifetime of a free neutron. The prominence of $g_A$ in nuclear physics has made it a benchmark quantity with which to calibrate lattice QCD calculations of nucleon structure and more complex calculations of electroweak matrix elements in one and few nucleon systems. There were a number of significant challenges in determining $g_A$, notably the notorious exponentially-bad signal-to-noise problem and the requirement for hundreds of thousands of stochastic samples, that rendered this goal more difficult to obtain than originally thought. I will describe the use of an unconventional computation method, coupled with "ludicrously'" fast GPU code, access to publicly available lattice QCD configurations from MILC and access to leadership computing that have allowed these challenges to be overcome resulting in a determination of $g_A$ with 1% precision and all sources of systematic uncertainty controlled. I will discuss the implications of these results for the convergence of $SU(2)$ Chiral Perturbation theory for nucleons, as well as prospects for further improvements to $g_A$ (sub-percent precision, for which we have preliminary results) which is part of a more comprehensive application of lattice QCD to nuclear physics. This is particularly exciting in light of the new CORAL supercomputers coming online, Sierra and Summit, for which our lattice QCD codes achieve a machine-to-machine speed up over Titan of an order of magnitude.

preprint2018arXiv

Simulating the weak death of the neutron in a femtoscale universe with near-Exascale computing

The fundamental particle theory called Quantum Chromodynamics (QCD) dictates everything about protons and neutrons, from their intrinsic properties to interactions that bind them into atomic nuclei. Quantities that cannot be fully resolved through experiment, such as the neutron lifetime (whose precise value is important for the existence of light-atomic elements that make the sun shine and life possible), may be understood through numerical solutions to QCD. We directly solve QCD using Lattice Gauge Theory and calculate nuclear observables such as neutron lifetime. We have developed an improved algorithm that exponentially decreases the time-to solution and applied it on the new CORAL supercomputers, Sierra and Summit. We use run-time autotuning to distribute GPU resources, achieving 20% performance at low node count. We also developed optimal application mapping through a job manager, which allows CPU and GPU jobs to be interleaved, yielding 15% of peak performance when deployed across large fractions of CORAL.

preprint2016arXiv

Accelerating Lattice QCD Multigrid on GPUs Using Fine-Grained Parallelization

The past decade has witnessed a dramatic acceleration of lattice quantum chromodynamics calculations in nuclear and particle physics. This has been due to both significant progress in accelerating the iterative linear solvers using multi-grid algorithms, and due to the throughput improvements brought by GPUs. Deploying hierarchical algorithms optimally on GPUs is non-trivial owing to the lack of parallelism on the coarse grids, and as such, these advances have not proved multiplicative. Using the QUDA library, we demonstrate that by exposing all sources of parallelism that the underlying stencil problem possesses, and through appropriate mapping of this parallelism to the GPU architecture, we can achieve high efficiency even for the coarsest of grids. Results are presented for the Wilson-Clover discretization, where we demonstrate up to 10x speedup over present state-of-the-art GPU-accelerated methods on Titan. Finally, we look to the future, and consider the software implications of our findings.

preprint2016arXiv

Neutrinoless double beta decay from lattice QCD

While the discovery of non-zero neutrino masses is one of the most important accomplishments by physicists in the past century, it is still unknown how and in what form these masses arise. Lepton number-violating neutrinoless double beta decay is a natural consequence of Majorana neutrinos and many BSM theories, and many experimental efforts are involved in the search for these processes. Understanding how neutrinoless double beta decay would manifest in nuclear environments is key for understanding any observed signals. In these proceedings we present an overview of a set of one- and two-body matrix elements relevant for experimental searches for neutrinoless double beta decay, describe the role of lattice QCD calculations, and present preliminary lattice QCD results.

preprint2015arXiv

Digital Signal Processing using Stream High Performance Computing: A 512-input Broadband Correlator for Radio Astronomy

A "large-N" correlator that makes use of Field Programmable Gate Arrays and Graphics Processing Units has been deployed as the digital signal processing system for the Long Wavelength Array station at Owens Valley Radio Observatory (LWA-OV), to enable the Large Aperture Experiment to Detect the Dark Ages (LEDA). The system samples a ~100MHz baseband and processes signals from 512 antennas (256 dual polarization) over a ~58MHz instantaneous sub-band, achieving 16.8Tops/s and 0.236 Tbit/s throughput in a 9kW envelope and single rack footprint. The output data rate is 260MB/s for 9 second time averaging of cross-power and 1 second averaging of total-power data. At deployment, the LWA-OV correlator was the largest in production in terms of N and is the third largest in terms of complex multiply accumulations, after the Very Large Array and Atacama Large Millimeter Array. The correlator's comparatively fast development time and low cost establish a practical foundation for the scalability of a modular, heterogeneous, computing architecture.

preprint2015arXiv

Optimizing performance per watt on GPUs in High Performance Computing: temperature, frequency and voltage effects

The magnitude of the real-time digital signal processing challenge attached to large radio astronomical antenna arrays motivates use of high performance computing (HPC) systems. The need for high power efficiency (performance per watt) at remote observatory sites parallels that in HPC broadly, where efficiency is an emerging critical metric. We investigate how the performance per watt of graphics processing units (GPUs) is affected by temperature, core clock frequency and voltage. Our results highlight how the underlying physical processes that govern transistor operation affect power efficiency. In particular, we show experimentally that GPU power consumption grows non-linearly with both temperature and supply voltage, as predicted by physical transistor models. We show lowering GPU supply voltage and increasing clock frequency while maintaining a low die temperature increases the power efficiency of an NVIDIA K20 GPU by up to 37-48% over default settings when running xGPU, a compute-bound code used in radio astronomy. We discuss how temperature-aware power models could be used to reduce power consumption for future HPC installations. Automatic temperature-aware and application-dependent voltage and frequency scaling (T-DVFS and A-DVFS) may provide a mechanism to achieve better power efficiency for a wider range of codes running on GPUs

preprint2015arXiv

The Murchison Widefield Array Correlator

The Murchison Widefield Array (MWA) is a Square Kilometre Array (SKA) Precursor. The telescope is located at the Murchison Radio--astronomy Observatory (MRO) in Western Australia (WA). The MWA consists of 4096 dipoles arranged into 128 dual polarisation aperture arrays forming a connected element interferometer that cross-correlates signals from all 256 inputs. A hybrid approach to the correlation task is employed, with some processing stages being performed by bespoke hardware, based on Field Programmable Gate Arrays (FPGAs), and others by Graphics Processing Units (GPUs) housed in general purpose rack mounted servers. The correlation capability required is approximately 8 TFLOPS (Tera FLoating point Operations Per Second). The MWA has commenced operations and the correlator is generating 8.3 TB/day of correlation products, that are subsequently transferred 700 km from the MRO to Perth (WA) in real-time for storage and offline processing. In this paper we outline the correlator design, signal path, and processing elements and present the data format for the internal and external interfaces.

preprint2014arXiv

A Scalable Hybrid FPGA/GPU FX Correlator

Radio astronomical imaging arrays comprising large numbers of antennas, O(10^2-10^3) have posed a signal processing challenge because of the required O(N^2) cross correlation of signals from each antenna and requisite signal routing. This motivated the implementation of a Packetized Correlator architecture that applies Field Programmable Gate Arrays (FPGAs) to the O(N) "F-stage" transforming time domain to frequency domain data, and Graphics Processing Units (GPUs) to the O(N^2) "X-stage" performing an outer product among spectra for each antenna. The design is readily scalable to at least O(10^3) antennas. Fringes, visibility amplitudes and sky image results obtained during field testing are presented.

preprint2012arXiv

Multigrid Algorithms for Domain-Wall Fermions

We describe an adaptive multigrid algorithm for solving inverses of the domain-wall fermion operator. Our multigrid algorithm uses an adaptive projection of near-null vectors of the domain-wall operator onto coarser four-dimensional lattices. This extension of multigrid techniques to a chiral fermion action will greatly reduce overall computation cost, and the elimination of the fifth dimension in the coarse space reduces the relative cost of using chiral fermions compared to discarding this symmetry. We demonstrate near-elimination of critical slowing as the quark mass is reduced and small volume dependence, which may be suppressed by taking advantage of the recursive nature of the algorithm.

preprint2012arXiv

Shadow Hamiltonians, Poisson Brackets, and Gauge Theories

Numerical lattice gauge theory computations to generate gauge field configurations including the effects of dynamical fermions are usually carried out using algorithms that require the molecular dynamics evolution of gauge fields using symplectic integrators. Sophisticated integrators are in common use but are hard to optimise, and force-gradient integrators show promise especially for large lattice volumes. We explain why symplectic integrators lead to very efficient Monte Carlo algorithms because they exactly conserve a shadow Hamiltonian. The shadow Hamiltonian may be expanded in terms of Poisson brackets, and can be used to optimize the integrators. We show how this may be done for gauge theories by extending the formulation of Hamiltonian mechanics on Lie groups to include Poisson brackets and shadows, and by giving a general method for the practical computation of forces, force-gradients, and Poisson brackets for gauge theories.

preprint2011arXiv

Accelerating Radio Astronomy Cross-Correlation with Graphics Processing Units

We present a highly parallel implementation of the cross-correlation of time-series data using graphics processing units (GPUs), which is scalable to hundreds of independent inputs and suitable for the processing of signals from "Large-N" arrays of many radio antennas. The computational part of the algorithm, the X-engine, is implementated efficiently on Nvidia's Fermi architecture, sustaining up to 79% of the peak single precision floating-point throughput. We compare performance obtained for hardware- and software-managed caches, observing significantly better performance for the latter. The high performance reported involves use of a multi-level data tiling strategy in memory and use of a pipelined algorithm with simultaneous computation and transfer of data from host to device memory. The speed of code development, flexibility, and low cost of the GPU implementations compared to ASIC and FPGA implementations have the potential to greatly shorten the cycle of correlator development and deployment, for cases where some power consumption penalty can be tolerated.

preprint2011arXiv

First Spectroscopic Imaging Observations of the Sun at Low Radio Frequencies with the Murchison Widefield Array Prototype

We present the first spectroscopic images of solar radio transients from the prototype for the Murchison Widefield Array (MWA), observed on 2010 March 27. Our observations span the instantaneous frequency band 170.9-201.6 MHz. Though our observing period is characterized as a period of `low' to `medium' activity, one broadband emission feature and numerous short-lived, narrowband, non-thermal emission features are evident. Our data represent a significant advance in low radio frequency solar imaging, enabling us to follow the spatial, spectral, and temporal evolution of events simultaneously and in unprecedented detail. The rich variety of features seen here reaffirms the coronal diagnostic capability of low radio frequency emission and provides an early glimpse of the nature of radio observations that will become available as the next generation of low frequency radio interferometers come on-line over the next few years.

preprint2011arXiv

Improving dynamical lattice QCD simulations through integrator tuning using Poisson brackets and a force-gradient integrator

We show how the integrators used for the molecular dynamics step of the Hybrid Monte Carlo algorithm can be further improved. These integrators not only approximately conserve some Hamiltonian $H$ but conserve exactly a nearby shadow Hamiltonian $\tilde{H}$. This property allows for a new tuning method of the molecular dynamics integrator and also allows for a new class of integrators (force-gradient integrators) which is expected to reduce significantly the computational cost of future large-scale gauge field ensemble generation.

preprint2011arXiv

Scaling Lattice QCD beyond 100 GPUs

Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success to the post-Monte Carlo "analysis" phase which accounts for a substantial fraction of the workload in a typical LQCD calculation, the initial Monte Carlo "gauge field generation" phase requires capability-level supercomputing, corresponding to O(100) GPUs or more. Such strong scaling has not been previously achieved. In this contribution, we demonstrate that using a multi-dimensional parallelization strategy and a domain-decomposed preconditioner allows us to scale into this regime. We present results for two popular discretizations of the Dirac operator, Wilson-clover and improved staggered, employing up to 256 GPUs on the Edge cluster at Lawrence Livermore National Laboratory.

preprint2010arXiv

Adaptive multigrid algorithm for the lattice Wilson-Dirac operator

We present an adaptive multigrid solver for application to the non-Hermitian Wilson-Dirac system of QCD. The key components leading to the success of our proposed algorithm are the use of an adaptive projection onto coarse grids that preserves the near null space of the system matrix together with a simplified form of the correction based on the so-called gamma_5-Hermitian symmetry of the Dirac operator. We demonstrate that the algorithm nearly eliminates critical slowing down in the chiral limit and that it has weak dependence on the lattice volume.

preprint2010arXiv

Enabling a High Throughput Real Time Data Pipeline for a Large Radio Telescope Array with GPUs

The Murchison Widefield Array (MWA) is a next-generation radio telescope currently under construction in the remote Western Australia Outback. Raw data will be generated continuously at 5GiB/s, grouped into 8s cadences. This high throughput motivates the development of on-site, real time processing and reduction in preference to archiving, transport and off-line processing. Each batch of 8s data must be completely reduced before the next batch arrives. Maintaining real time operation will require a sustained performance of around 2.5TFLOP/s (including convolutions, FFTs, interpolations and matrix multiplications). We describe a scalable heterogeneous computing pipeline implementation, exploiting both the high computing density and FLOP-per-Watt ratio of modern GPUs. The architecture is highly parallel within and across nodes, with all major processing elements performed by GPUs. Necessary scatter-gather operations along the pipeline are loosely synchronized between the nodes hosting the GPUs. The MWA will be a frontier scientific instrument and a pathfinder for planned peta- and exascale facilities.

preprint2010arXiv

Interferometric imaging with the 32 element Murchison Wide-field Array

The Murchison Wide-field Array (MWA) is a low frequency radio telescope, currently under construction, intended to search for the spectral signature of the epoch of re-ionisation (EOR) and to probe the structure of the solar corona. Sited in Western Australia, the full MWA will comprise 8192 dipoles grouped into 512 tiles, and be capable of imaging the sky south of 40 degree declination, from 80 MHz to 300 MHz with an instantaneous field of view that is tens of degrees wide and a resolution of a few arcminutes. A 32-station prototype of the MWA has been recently commissioned and a set of observations taken that exercise the whole acquisition and processing pipeline. We present Stokes I, Q, and U images from two ~4 hour integrations of a field 20 degrees wide centered on Pictoris A. These images demonstrate the capacity and stability of a real-time calibration and imaging technique employing the weighted addition of warped snapshots to counter extreme wide field imaging distortions.

preprint2010arXiv

Multigrid solver for clover fermions

We present an adaptive multigrid Dirac solver developed for Wilson clover fermions which offers order-of-magnitude reductions in solution time compared to conventional Krylov solvers. The solver incorporates even-odd preconditioning and mixed precision to solve the Dirac equation to double precision accuracy and shows only a mild increase in time to solution for decreasing quark mass. We show actual time to solution on production lattices in comparison to conventional Krylov solvers and will also discuss the setup process and its relative cost to the total solution time.

preprint2010arXiv

The removal of critical slowing down

We present promising initial results of our adaptive multigrid solver developed for application directly to the non-Hermitian Wilson-Dirac system in 4 dimensions, as opposed to the solver developed in [1] for the corresponding normal equations. The key behind the success of this algorithm is the use of an adaptive projection onto coarse grids that preserves the near null space of the system matrix. We demonstrate that the resulting algorithm has weak dependence on the gauge coupling and exhibits extremely mild critical slowing down in the chiral limit.

preprint2009arXiv

Force Gradient Integrators

We present initial results of the use of Force Gradient integrators for lattice field theories. These promise to give significant performance improvements, especially for light fermions and large lattices. Our results show that this is indeed the case, indicating a speed-up of more than a factor of two, which is expected to increase as the integration step size becomes smaller for larger lattices and smaller fermion masses.

preprint2009arXiv

QCD on GPUs: cost effective supercomputing

The exponential growth of floating point power in graphics processing units (GPUs), together with their low cost, has given rise to an attractive platform upon which to deploy lattice QCD calculations. GPUs are essentially many (O(100)) core chips, that are programmed using a massively threaded environment, and so are representative of the future of high performance computing (HPC). The large ratio of raw floating point operations per second to memory bandwidth that is characteristic of GPUs necessitates that unique algorithmic design choices are made to harness their full potential. We review the progress to date in using GPUs for large scale calculations, and contrast GPUs against more traditional HPC architectures

preprint2009arXiv

Solving Lattice QCD systems of equations using mixed precision solvers on GPUs

Modern graphics hardware is designed for highly parallel numerical tasks and promises significant cost and performance benefits for many scientific applications. One such application is lattice quantum chromodyamics (lattice QCD), where the main computational challenge is to efficiently solve the discretized Dirac equation in the presence of an SU(3) gauge field. Using NVIDIA's CUDA platform we have implemented a Wilson-Dirac sparse matrix-vector product that performs at up to 40 Gflops, 135 Gflops and 212 Gflops for double, single and half precision respectively on NVIDIA's GeForce GTX 280 GPU. We have developed a new mixed precision approach for Krylov solvers using reliable updates which allows for full double precision accuracy while using only single or half precision arithmetic for the bulk of the computation. The resulting BiCGstab and CG solvers run in excess of 100 Gflops and, in terms of iterations until convergence, perform better than the usual defect-correction approach for mixed precision.

preprint2008arXiv

Tuning HMC using Poisson brackets

We discuss how the integrators used for the Hybrid Monte Carlo (HMC) algorithm not only approximately conserve some Hamiltonian $H$ but exactly conserve a nearby shadow Hamiltonian (\tilde H), and how the difference $ΔH \equiv \tilde H - H $ may be expressed as an expansion in Poisson brackets. By measuring average values of these Poisson brackets over the equilibrium distribution $\propto e^{-H}$ generated by HMC we can find the optimal integrator parameters from a single simulation. We show that a good way of doing this in practice is to minimize the variance of $ΔH$ rather than its magnitude, as has been previously suggested. Some details of how to compute Poisson brackets for gauge and fermion fields, and for nested and force gradient integrators are also presented.

preprint2007arXiv

2+1 flavor domain wall QCD on a (2 fm)^3 lattice: light meson spectroscopy with Ls = 16

We present results for light meson masses and pseudoscalar decay constants from the first of a series of lattice calculations with 2+1 dynamical flavors of domain wall fermions and the Iwasaki gauge action. The work reported here was done at a fixed lattice spacing of about 0.12 fm on a 16^3\times32 lattice, which amounts to a spatial volume of (2 fm)^3 in physical units. The number of sites in the fifth dimension is 16, which gives m_{res} = 0.00308(4) in these simulations. Three values of input light sea quark masses, m_l^{sea} \approx 0.85 m_s, 0.59 m_s and 0.33 m_s were used to allow for extrapolations to the physical light quark limit, whilst the heavier sea quark mass was fixed to approximately the physical strange quark mass m_s. The exact rational hybrid Monte Carlo algorithm was used to evaluate the fractional powers of the fermion determinants in the ensemble generation. We have found that f_π= 127(4) MeV, f_K = 157(5) MeV and f_K/f_π= 1.24(2), where the errors are statistical only, which are in good agreement with the experimental values.