Source author record

Wes Armour

Wes Armour appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

11works
9topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

11 published item(s)

preprint2026arXiv

Runtime-Orchestrated Second-Order Optimization for Scalable LLM Training

Second-order methods offer an attractive path toward more sample-efficient LLM training, but their practical use is often blocked by the systems cost of maintaining and updating large matrix-based optimizer states. We introduce \textbf{Asteria}, a runtime system designed to remove this bottleneck by separating second-order optimization logic from the critical GPU training path. Rather than keeping all preconditioner state on the accelerator, Asteria dynamically distributes optimizer state across GPU memory, CPU memory, and optional NVMe storage according to architectural constraints and runtime pressure. It further uses training hooks to prepare shadow states in advance, allowing expensive inverse-root computations to proceed asynchronously on the host while GPU computation continues. For distributed training, Asteria employs a bounded-staleness protocol that limits synchronization frequency while preserving optimizer effectiveness through topology-aware coordination. We evaluate Asteria on both memory-constrained and distributed training settings. On a DGX Spark platform with a single GB10 GPU and 128GB unified memory, Asteria supports second-order training for a 1B-parameter language model. On multi-node GH200 systems, it lowers visible optimizer overhead, reduces recurring latency spikes, accelerates convergence in wall-clock time, and maintains the optimization advantages of SOAP and KL-Shampoo in a 7B-parameter language model. Our results suggest that second-order LLM training can be made practical not by simplifying the optimizer alone, but by rethinking how optimizer state, background computation, and distributed synchronization are managed at the runtime level.

preprint2026arXiv

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models

Autoregressive next-token training offers a unified formulation for image generation and text understanding, but it also creates strong modality competition that destabilizes optimization and limits large-batch scaling. We show that first-order optimizers such as AdamW are vulnerable to cross-modality gradient heterogeneity, while second-order preconditioning, particularly SOAP, provides a more stable basis for multimodal alignment. Building on this insight, we propose \emph{ML-FOP-SOAP}, a second-order optimization framework with Multi-Level Variance Correction. Our Fisher-Orthogonal Projection suppresses variance-induced modality conflicts, reducing the trade-off between visual generation and textual understanding. To make this practical under large gradient accumulation, we introduce a hierarchical folding strategy that captures fine-grained variance with low micro-step overhead. Experiments on Janus and Emu3 show consistent gains across both modalities and stable training at batch size 8192. Compared with AdamW, our method improves sample efficiency by up to $1.4\times$ and accelerates wall-clock training by up to $1.5\times$, offering a robust optimizer for scaling multimodal foundation models.

preprint2016arXiv

A polyphase filter for many-core architectures

In this article we discuss our implementation of a polyphase filter for real-time data processing in radio astronomy. We describe in detail our implementation of the polyphase filter algorithm and its behaviour on three generations of NVIDIA GPU cards, on dual Intel Xeon CPUs and the Intel Xeon Phi (Knights Corner) platforms. All of our implementations aim to exploit the potential for data reuse that the algorithm offers. Our GPU implementations explore two different methods for achieving this, the first makes use of L1/Texture cache, the second uses shared memory. We discuss the usability of each of our implementations along with their behaviours. We measure performance in execution time, which is a critical factor for real-time systems, we also present results in terms of bandwidth (GB/s), compute (GFlop/s) and type conversions (GTc/s). We include a presentation of our results in terms of the sample rate which can be processed in real-time by a chosen platform, which more intuitively describes the expected performance in a signal processing setting. Our findings show that, for the GPUs considered, the performance of our polyphase filter when using lower precision input data is limited by type conversions rather than device bandwidth. We compare these results to an implementation on the Xeon Phi. We show that our Xeon Phi implementation has a performance that is 1.47x to 1.95x greater than our CPU implementation, however is not insufficient to compete with the performance of GPUs. We conclude with a comparison of our best performing code to two other implementations of the polyphase filter, showing that our implementation is faster in nearly all cases. This work forms part of the Astro-Accelerate project, a many-core accelerated real-time data processing library for digital signal processing of time-domain radio astronomy data.

preprint2016arXiv

Commissioning of ALFABURST: initial tests and results

Fast Radio Bursts (FRBs) are apparently one-time, relatively bright radio pulses that have been observed in recent years. The origin of FRBs is currently unknown and many instruments are being built to detect more of these bursts to better characterize their physical properties and identify the source population. ALFABURST is one such instrument. ALFABURST takes advantage of the 7-beam Arecibo L-band Feed Array (ALFA) receiver on the 305-m Arecibo Radio Telescope in Puerto Rico, to detect FRBs in real-time at L-band (1.4 GHz). We present the results of recent on-sky tests and observations undertaken during the commissioning phase of the instrument. ALFABURST is now available for commensal observations with other ALFA projects.

preprint2015arXiv

ALFABURST: A realtime fast radio burst monitor for the Arecibo telescope

Fast radio bursts (FRBs) constitute an emerging class of fast radio transient whose origin continues to be a mystery. Realizing the importance of increasing coverage of the search parameter space, we have designed, built, and deployed a realtime monitor for FRBs at the 305-m Arecibo radio telescope. Named 'ALFABURST', it is a commensal instrument that is triggered whenever the 1.4 GHz seven-beam Arecibo $L$-Band Feed Array (ALFA) receiver commences operation. The ongoing commensal survey we are conducting using ALFABURST has an instantaneous field of view of 0.02 sq. deg. within the FWHM of the beams, with the realtime software configurable to use up to 300 MHz of bandwidth. We search for FRBs with dispersion measure up to 2560 cm$^{-3}$ pc and pulse widths ranging from 0.128 ms to 16.384 ms. Commissioning observations performed over the past few months have demonstrated the capability of the instrument in detecting single pulses from known pulsars. In this paper, I describe the instrument and the associated survey.

preprint2015arXiv

Graphene as a Lattice Field Theory

We introduce effective field theories for the electronic properties of graphene in terms of relativistic fermions propagating in 2+1 dimensions, and outline how strong inter-electron interactions may be modelled by numerical simulation of a lattice field theory. For strong enough coupling an insulating state can form via condensation of particle-hole pairs, and it is demonstrated that this is a theoretical possibility for monolayer graphene. For bilayer graphene the effect of an interlayer bias voltage can be modelled by the introduction of a chemical potential (akin to isopsin chemical potential in QCD) with no accompanying sign problem; simulations reveal the presence of strong interactions among the residual degrees of freedom at the resulting Fermi surface, which is disrupted by an excitonic condensate. We also present preliminary results for the quasiparticle dispersion, which permit direct estimates of both the Fermi momentum and the induced gap.

preprint2015arXiv

Strong Interaction Effects at a Fermi Surface in a Model for Voltage-Biased Bilayer Graphene

Monte Carlo simulation of a 2+1 dimensional model of voltage-biased bilayer graphene, consisting of relativistic fermions with chemical potential mu coupled to charged excitations with opposite sign on each layer, has exposed non-canonical scaling of bulk observables near a quantum critical point found at strong coupling. We present a calculation of the quasiparticle dispersion relation E(k) as a function of exciton source j in the same system, employing partially twisted boundary conditions to boost the number of available momentum modes. The Fermi momentum k_F and superfluid gap Delta are extracted in the limit j tends to zero for three different values of mu, and support a strongly interacting scenario at the Fermi surface with Delta of order O(mu). We propose an explanation for the observation mu < k_F in terms of a dynamical critical exponent z < 1.

preprint2014arXiv

The Implementation of a Real-Time Polyphase Filter

In this article we study the suitability of dierent computational accelerators for the task of real-time data processing. The algorithm used for comparison is the polyphase filter, a standard tool in signal processing and a well established algorithm. We measure performance in FLOPs and execution time, which is a critical factor for real-time systems. For our real-time studies we have chosen a data rate of 6.5GB/s, which is the estimated data rate for a single channel on the SKAs Low Frequency Aperture Array. Our findings how that GPUs are the most likely candidate for real-time data processing. GPUs are better in both performance and power consumption.

preprint2013arXiv

Monte Carlo Study of Strongly-Interacting Degenerate Fermions: a Model for Voltage-Biased Bilayer Graphene

We formulate a model of N_f=4 flavors of relativistic fermion in 2+1d in the presence of a chemical potential mu coupled to two flavor doublets with opposite sign, akin to isopsin chemical potential in QCD. This is argued to be an effective theory for low energy electronic excitations in bilayer graphene, in which an applied voltage between the layers ensures equal populations of particles on one layer and holes on the other. The model is then reformulated on a spacetime lattice using staggered fermions, and in the absence of a sign problem, simulated using an orthodox hybrid Monte Carlo algorithm. With the coupling strength chosen to be close to a quantum critical point believed to exist for N_f<N_fc\approx4.8, it is found that there is a region below saturation where both the carrier density and a particle-hole "excitonic" condensate scale anomalously with increasing mu, much more rapidly that the corresponding quantities in free field theory, while the conventional chiral condensate is strongly suppressed. The corresponding ground state is speculated to be a strongly-correlated degenerate fermion system, with a remnant Fermi surface distorted by a superfluid excitonic condensate. The model thus shows qualitatively different behaviour to any model with mu=/=0 previously studied by lattice simulation.

preprint2012arXiv

Observations of transients and pulsars with LOFAR international stations

The LOw FRequency ARray - LOFAR is a new radio telescope that is moving the science of radio pulsars and transients into a new phase. Its design places emphasis on digital hardware and flexible software instead of mechanical solutions. LOFAR observes at radio frequencies between 10 and 240 MHz where radio pulsars and many transients are expected to be brightest. Radio frequency signals emitted from these objects allow us to study the intrinsic pulsar emission and phenomena such as propagation effects through the interstellar medium. The design of LOFAR allows independent use of its stations to conduct observations of known bright objects, or wide field monitoring of transient events. One such combined software/hardware solution is called the Advanced Radio Transient Event Monitor and Identification System (ARTEMIS). It is a backend for both targeted observations and real-time searches for millisecond radio transients which uses Graphical Processing Unit (GPU) technology to remove interstellar dispersion and detect millisecond radio bursts from astronomical sources in real-time using a single LOFAR station.

preprint2009arXiv

Monte Carlo Simulation of the Semimetal-Insulator Phase Transition in Monolayer Graphene

A 2+1 dimensional fermion field theory is proposed as a model for the low-energy electronic excitations in monolayer graphene. The model consists of N=2 four-component Dirac fermions moving in the plane and interacting via a contact interaction between charge densities. For strong couplings there is a continuous transition to a Mott insulting phase. We present results of an extensive numerical study of the model's critical region, including the order parameter, its associated susceptibility, and for the first time the quasiparticle propagator. The data enables an extraction of the critical exponents at the transition, including the dynamical critical exponent, which are hypothesised to be universal features of a quantum critical point. The relation of our model with others in the literature is discussed, along with the implications for physical graphene following from our value of the critical coupling.