Source author record

Karthee Sivalingam

Karthee Sivalingam appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

6works
7topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

6 published item(s)

preprint2020arXiv

Optimising AI Training Deployments using Graph Compilers and Containers

Artificial Intelligence (AI) applications based on Deep Neural Networks (DNN) or Deep Learning (DL) have become popular due to their success in solving problems likeimage analysis and speech recognition. Training a DNN is computationally intensive and High Performance Computing(HPC) has been a key driver in AI growth. Virtualisation and container technology have led to the convergence of cloud and HPC infrastructure. These infrastructures with diverse hardware increase the complexity of deploying and optimising AI training workloads. AI training deployments in HPC or cloud can be optimised with target-specific libraries, graph compilers, andby improving data movement or IO. Graph compilers aim to optimise the execution of a DNN graph by generating an optimised code for a target hardware/backend. As part of SODALITE (a Horizon 2020 project), MODAK tool is developed to optimise application deployment in software defined infrastructures. Using input from the data scientist and performance modelling, MODAK maps optimal application parameters to a target infrastructure and builds an optimised container. In this paper, we introduce MODAK and review container technologies and graph compilers for AI. We illustrate optimisation of AI training deployments using graph compilers and Singularity containers. Evaluation using MNIST-CNN and ResNet50 training workloads shows that custom built optimised containers outperform the official images from DockerHub. We also found that the performance of graph compilers depends on the target hardware and the complexity of the neural network.

preprint2016arXiv

Variance reduction with practical all-to-all lattice propagators

We discuss all-to-all quark propagator techniques in two (related) contexts within Lattice QCD: the computation of closed quark propagators, and applications to the so-called "eye diagrams" appearing in the computation of non-leptonic kaon decay amplitudes. Combinations of low-mode averaging and diluted stochastic volume sources that yield optimal signal-to-noise ratios for the latter problem are developed. We also apply a recently proposed probing algorithm to compute directly the diagonal of the inverse Dirac operator, and compare its performance with that of stochastic methods. At fixed computational cost the two procedures yield comparable signal-to-noise ratios, but probing has practical advantages which make it a promising tool for a wide range of applications in Lattice QCD.

preprint2015arXiv

Performance analysis and Optimisation of the Met Unified Model on a Cray XC30

The Unified Model (UM) code supports simulation of weather, climate and earth system processes. It is primarily developed by the UK Met Office, but in recent years a wider community of users and developers have grown around the code. Here we present results from the optimisation work carried out by the UK National Centre for Atmospheric Science (NCAS) for a high resolution configuration (N512 $\approx$ 25km) on the UK ARCHER supercomputer, a Cray XC-30. On ARCHER, we use Cray Performance Analysis Tools (CrayPAT) to analyse the performance of UM and then Cray Reveal to identify and parallelise serial loops using OpenMP directives. We compare performance of the optimised version at a range of scales, and with a range of optimisations, including altered MPI rank placement, and addition of OpenMP directives. It is seen that improvements in MPI configuration yield performance improvements of between 5 and 12\%, and the added OpenMP directives yield an additional 5-16\% speedup. We also identify further code optimisations which could yield yet greater improvement in performance. We note that speedup gained using addition of OpenMP directives does not result in improved performance on the IBM Power platform where much of the code has been developed. This suggests that performance gains on future heterogeneous architectures will be hard to port. Nonetheless, it is clear that the investment of months in analysis and optimisation has yielded performance gains that correspond to the saving of tens of millions of core-hours on current climate projects.

preprint2014arXiv

Clover Action for Blue Gene-Q and Iterative solvers for DWF

In Lattice QCD, a major challenge in simulating physical quarks is the computational complexity of these simulations. In this proceeding, we describe the optimisation of Clover fermion action for Blue gene-Q architecture and how different iterative solvers behave for Domain Wall Fermion action. We find that the optimised Clover term achieved a maximum efficiency of 29.1% and 20.2% for single and double precision respectively for iterative Conjugate Gradient solver. For Domain Wall Fermion action (DWF) we found that Modified Conjugate Residual(MCR) as the most efficient solver compared to CG and GCR. We have developed a new multi-shift MCR algorithm that is 18.5% faster compared to multi-shift CG for the evaluation of rational functions in RHMC.

preprint2013arXiv

The kaon semileptonic form factor with near physical domain wall quarks

We present a new calculation of the K->pi semileptonic form factor at zero momentum transfer in domain wall lattice QCD with Nf=2+1 dynamical quark flavours. By using partially twisted boundary conditions we simulate directly at the phenomenologically relevant point of zero momentum transfer. We perform a joint analysis for all available ensembles which include three different lattice spacings (a=0.09-0.14fm), large physical volumes (m_pi*L>3.9) and pion masses as low as 171 MeV. The comprehensive set of simulation points allows for a detailed study of systematic effects leading to the prediction f+(0)=0.9670(20)(+18/-46), where the first error is statistical and the second error systematic. The result allows us to extract the CKM-matrix element |Vus|=0.2237(+13/-8) and confirm first-row CKM-unitarity in the Standard Model at the sub per mille level.

preprint2012arXiv

Kaon semileptonic decays near the physical point

The CKM matrix element $|V_{us}|$ can be extracted from the experimental measurement of semileptonic $K\toπ$ decays. The determination depends on theory input for the corresponding vector form factor in QCD. We present a preliminary update on our efforts to compute it in $N_f=2+1$ lattice QCD using domain wall fermions for several lattice spacings and with a lightest pion mass of about $170\,\mathrm{MeV}$. By using partially twisted boundary conditions we avoid systematic errors associated with an interpolation of the form factor in momentum-transfer, while simulated pion masses near the physical point reduce the systematic error due to the chiral extrapolation.