Source author record

Raghunathan Ramakrishnan

Raghunathan Ramakrishnan appears in the imported research catalog. Authorship, coauthor and topic links are available while profile ownership is still unclaimed.

ResearcherUnclaimed source record

Catalog footprint

What is connected

14works
3topics
4close collaborators

Actions

Connect this record

Log in to claim

Research graph

See the researcher in context

Open full explorer

Inspect adjacent papers, topics, institutions and collaborators without losing the researcher page.

Building this map preview

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

14 published item(s)

preprint2025arXiv

Machine-Learned Potentials for Solvation Modeling

Solvent environments play a central role in determining molecular structure, energetics, reactivity, and interfacial phenomena. However, modeling solvation from first principles remains difficult due to the complex interplay of interactions and unfavorable computational scaling of first-principles treatment with system size. Machine-learned potentials (MLPs) have recently emerged as efficient surrogates for quantum chemistry methods, offering first-principles accuracy at greatly reduced computational cost. MLPs approximate the underlying potential energy surface, enabling efficient computation of energies and forces in solvated systems, and are capable of accounting for effects such as hydrogen bonding, long-range polarization, and conformational changes. This review surveys the development and application of MLPs in solvation modeling. We summarize the theoretical basis of MLP-based energy and force predictions and present a classification of MLPs based on training targets, model types, and design choices related to architectures, descriptors, and training protocols. Integration into established solvation workflows is discussed, with case studies spanning small molecules, interfaces, and reactive systems. We conclude by outlining open challenges and future directions toward transferable, robust, and physically grounded MLPs for solvation-aware atomistic modeling.

preprint2022arXiv

Resolution-vs.-Accuracy Dilemma in Machine Learning Modeling of Electronic Excitation Spectra

In this study, we explore the potential of machine learning for modeling molecular electronic spectral intensities as a continuous function in a given wavelength range. Since presently available chemical space datasets provide excitation energies and corresponding oscillator strengths for only a few valence transitions, here, we present a new dataset -- \bigqm -- with 12,880 molecules containing up to 7 CONF atoms and report ground state and excited state properties. A publicly accessible web-based data-mining platform is presented to facilitate on-the-fly screening of several molecular properties including harmonic vibrational and electronic spectra. We present all singlet electronic transitions from the ground state calculated using the time-dependent density functional theory framework with the $ω$B97XD exchange-correlation functional and a diffuse-function augmented basis set. The resulting spectra predominantly span the X-ray to deep-UV region (10--120 nm). To compare the target spectra with predictions based on small basis sets, we bin spectral intensities and show good agreement is obtained only at the expense of the resolution. Compared to this, machine learning models with latest structural representations trained directly using $<10 \%$ of the target data recover the spectra of the remaining molecules with better accuracies at a desirable $<1$ nm wavelength resolution.

preprint2021arXiv

Data-Driven Modeling of S0 -> S1 Excitation Energy in the BODIPY Chemical Space: High-Throughput Computation, Quantum Machine Learning, and Inverse Design

Derivatives of BODIPY are popular fluorophores due to their synthetic feasibility, structural rigidity, high quantum yield, and tunable spectroscopic properties. While the characteristic absorption maximum of BODIPY is at 2.5 eV, combinations of functional groups and substitution sites can shift the peak position by +/- 1 eV. Time-dependent long-range corrected hybrid density functional methods can model the lowest excitation energies offering a semi-quantitative precision of +/- 0.3 eV. Alas, the chemical space of BODIPYs stemming from combinatorial introduction of -- even a few dozen -- substituents is too large for brute-force high-throughput modeling. To navigate this vast space, we select 77,412 molecules and train a kernel-based quantum machine learning model providing < 2% hold-out error. Further reuse of the results presented here to navigate the entire BODIPY universe comprising over 253 giga (253 x 10^9) molecules is demonstrated by inverse-designing candidates with desired target excitation energies.

preprint2021arXiv

High-Throughput Design of Peierls and Charge Density Wave Phases in Q1D Organometallic Materials

Soft-phonon modes of an undistorted phase encode a material's preference for symmetry lowering. However, the evidence is sparse for the relationship between an unstable phonon wavevector's reciprocal and the number of formula units in the stable distorted phase. This "1/q*-criterion" holds great potential for the first-principles design of materials, especially in low-dimension. We validate the approach on the Q1D materials space containing 1199 ring-metal units and identify candidates that are stable in undistorted (1 unit), Peierls (2 units), charge density wave (3-5 units), or long wave (>5 units) phases. We highlight materials exhibiting gap-opening as well as an uncommon gap-closing Peierls transition, and discuss an example case stabilized as a charge density wave insulator. We present the data generated for this study through an interactive publicly accessible Big Data analytics platform (http://moldis.tifrh.res.in/data/rmq1d) facilitating limitless and seamless data-mining explorations.

preprint2020arXiv

Charge-Transfer Selectivity and Quantum Interference in Real-Time Electron Dynamics: Gaining Insights from Time-Dependent Configuration Interaction Simulations

Many-electron wavepacket dynamics based on time-dependent configuration interaction (TDCI) is a numerically rigorous approach to quantitatively model electron-transfer across molecular junctions. TDCI simulations of cyanobenzene thiolates---para- and meta-linked to an acceptor gold atom---show donor states \emph{conjugating} with the benzene $π$-network to allow better through-molecule electron migration in the para isomer compared to the meta counterpart. For dynamics involving \emph{non-conjugating} states, we find electron-injection to stem exclusively from distance-dependent non-resonant quantum mechanical tunneling, in which case the meta isomer exhibits better dynamics. Computed trend in donor-to-acceptor net-electron transfer through differently linked azulene bridges agrees with the trend seen in low-bias conductivity measurements. Disruption of $π$-conjugation has been shown to be the cause of diminished electron-injection through the 1,3-azulene, a pathological case for graph-based diagnosis of destructive quantum interference. Furthermore, we demonstrate quantum interference of many-electron wavefunctions to drive para- vs. meta- selectivity in the coherent evolution of superposed $π$(CN)- and $σ$(NC-C)-type wavepackets. Analyses reveal that in the para-linked benzene, $σ$ and $π$ MOs localized at the donor terminal are \emph{in-phase} leading to constructive interference of electron density distribution while phase-flip of one of the MOs in the meta isomer results in destructive interference. These findings suggest that \emph{a priori} detection of orbital phase-flip and quantum coherence conditions can aid in molecular device design strategies.

preprint2020arXiv

Critical Benchmarking of the G4(MP2) Model, the Correlation Consistent Composite Approach and Popular Density Functional Approximations on a Probabilistically Pruned Benchmark Dataset of Formation Enthalpies

First-principles calculation of the standard formation enthalpy, $ΔH_f^\circ$ (298K), in such large scale as required by chemical space explorations, is amenable only with density functional approximations (DFAs) and some composite wave function theories (cWFTs). Alas, the accuracies of popular range-separated hybrid, `rung-4' DFAs, and cWFTs that offer the best accuracy-vs.-cost trade-off have as yet been established only for datasets predominantly comprising small molecules, hence, their transferability to larger datasets remains vague. In this study, we present an extended benchmark dataset of over 1600 values of $ΔH_f^\circ$ for structurally and electronically diverse molecules. We apply quartile-ranking based on boundary-corrected kernel density estimation to filter outliers and arrive at Probabilistically Pruned Enthalpies of 1694 compounds (PPE1694). For this dataset, we rank the prediction accuracies of G4, G4(MP2), ccCA, CBS-QB3 and 23 popular DFAs using conventional and probabilistic error metrics. We discuss systematic prediction errors and highlight the role an empirical higher-level correction (HLC) plays in the G4(MP2) model. Furthermore, we comment on uncertainties associated with the reference empirical data for atoms and the systematic errors stemming from these that grow with the molecular size. We believe these findings to aid in identifying meaningful application domains for quantum thermochemical methods.

preprint2020arXiv

Quantum-chemistry-aided identification, synthesis and experimental validation of model systems for conformationally controlled reaction studies: Separation of the conformers of 2,3-dibromobuta-1,3-diene in the gas phase

The Diels-Alder cycloaddition, in which a diene reacts with a dienophile to form a cyclic compound, counts among the most important tools in organic synthesis. Achieving a precise understanding of its mechanistic details on the quantum level requires new experimental and theoretical methods. Here, we present an experimental approach that separates different diene conformers in a molecular beam as a prerequisite for the investigation of their individual cycloaddition reaction kinetics and dynamics under single-collision conditions in the gas phase. A low- and high-level quantum-chemistry-based screening of more than one hundred dienes identified 2,3-dibromobutadiene (DBB) as an optimal candidate for efficient separation of its gauche and s-trans conformers by electrostatic deflection. A preparation method for DBB was developed which enabled the generation of dense molecular beams of this compound. The theoretical predictions of the molecular properties of DBB were validated by the successful separation of the conformers in the molecular beam. A marked difference in photofragment ion yields of the two conformers upon femtosecond-laser pulse ionization was observed, pointing at a pronounced conformer-specific fragmentation dynamics of ionized DBB. Our work sets the stage for a rigorous examination of mechanistic models of cycloaddition reactions under controlled conditions in the gas phase.

preprint2016arXiv

Fast and accurate predictions of covalent bonds in chemical space

We assess the predictive accuracy of perturbation theory based estimates of changes in covalent bonding due to linear alchemical interpolations among molecules. We have investigated $σ$ bonding to hydrogen, as well as $σ$ and $π$ bonding between main-group elements, occurring in small sets of iso-valence-electronic molecular species with elements drawn from second to fourth rows in the $p$-block of the periodic table. Numerical evidence suggests that first order estimates of covalent bonding potentials can achieve chemical accuracy if (i) the alchemical interpolation is vertical (fixed geometry), (ii) involves molecules containing elements in the third and fourth row of the periodic table, and (iii) a reference geometry is optimized. In this case, changes in the bonding potential become near-linear in coupling parameter, resulting in analytical predictions with very high accuracy ($\sim$1 kcal/mol). Second order estimates deteriorate the prediction. If initial and final molecules differ not only in composition but also in geometry, all estimates become substantially worse, with second order being slightly more accurate than first order. The independent particle approximation to the second order perturbation performs poorly when compared to the coupled perturbed or finite difference approach. Taylor series expansions up to fourth order of the potential energy curve of highly symmetric systems indicate a finite radius of convergence, as illustrated for the alchemical stretching of H$_2^+$. Numerical results are presented for covalent bonds to hydrogen in 12 molecules with 8 valence electrons; (ii) main-group single bonds in 9 molecules with 14 valence electrons; (iii) main-group double bonds in 9 molecules with 12 valence electrons; (iv) main-group triple bonds in 9 molecules with 10 valence electrons; (v) H$_2^+$ single bond with 1 electron.

preprint2016arXiv

Genetic optimization of training sets for improved machine learning models of molecular properties

The training of molecular models of quantum mechanical properties based on statistical machine learning requires large datasets which exemplify the map from chemical structure to molecular property. Intelligent a priori selection of training examples is often difficult or impossible to achieve as prior knowledge may be sparse or unavailable. Ordinarily representative selection of training molecules from such datasets is achieved through random sampling. We use genetic algorithms for the optimization of training set composition consisting of tens of thousands of small organic molecules. The resulting machine learning models are considerably more accurate with respect to small randomly selected training sets: mean absolute errors for out-of-sample predictions are reduced to ~25% for enthalpies, free energies, and zero-point vibrational energy, to ~50% for heat-capacity, electron-spread, and polarizability, and by more than ~20% for electronic properties such as frontier orbital eigenvalues or dipole-moments. We discuss and present optimized training sets consisting of 10 molecular classes for all molecular properties studied. We show that these classes can be used to design improved training sets for the generation of machine learning models of the same properties in similar but unrelated molecular sets.

preprint2015arXiv

Electronic Spectra from TDDFT and Machine Learning in Chemical Space

Due to its favorable computational efficiency time-dependent (TD) density functional theory (DFT) enables the prediction of electronic spectra in a high-throughput manner across chemical space. Its predictions, however, can be quite inaccurate. We resolve this issue with machine learning models trained on deviations of reference second-order approximate coupled-cluster singles and doubles (CC2) spectra from TDDFT counterparts, or even from DFT gap. We applied this approach to low-lying singlet-singlet vertical electronic spectra of over 20 thousand synthetically feasible small organic molecules with up to eight CONF atoms. The prediction errors decay monotonously as a function of training set size. For a training set of 10 thousand molecules, CC2 excitation energies can be reproduced to within $\pm$0.1 eV for the remaining molecules. Analysis of our spectral database via chromophore counting suggests that even higher accuracies can be achieved. Based on the evidence collected, we discuss open challenges associated with data-driven modeling of high-lying spectra, and transition intensities.

preprint2015arXiv

Fourier series of atomic radial distribution functions: A molecular fingerprint for machine learning models of quantum chemical properties

We introduce a fingerprint representation of molecules based on a Fourier series of atomic radial distribution functions. This fingerprint is unique (except for chirality), continuous, and differentiable with respect to atomic coordinates and nuclear charges. It is invariant with respect to translation, rotation, and nuclear permutation, and requires no pre-conceived knowledge about chemical bonding, topology, or electronic orbitals. As such it meets many important criteria for a good molecular representation, suggesting its usefulness for machine learning models of molecular properties trained across chemical compound space. To assess the performance of this new descriptor we have trained machine learning models of molecular enthalpies of atomization for training sets with up to 10k organic molecules, drawn at random from a published set of 134k organic molecules. We validate the descriptor on all remaining molecules of the 134k set. For a training set of 5k molecules the fingerprint descriptor achieves a mean absolute error of 8.0 kcal/mol, respectively. This is slightly worse than the performance attained using the Coulomb matrix, another popular alternative, reaching 6.2 kcal/mol for the same training and test sets.

preprint2015arXiv

Machine Learning for Quantum Mechanical Properties of Atoms in Molecules

We introduce machine learning models of quantum mechanical observables of atoms in molecules. Instant out-of-sample predictions for proton and carbon nuclear chemical shifts, atomic core level excitations, and forces on atoms reach accuracies on par with density functional theory reference. Locality is exploited within non-linear regression via local atom-centered coordinate systems. The approach is validated on a diverse set of 9k small organic molecules. Linear scaling of computational cost in system size is demonstrated for saturated polymers with up to sub-mesoscale lengths.

preprint2015arXiv

Many Molecular Properties from One Kernel in Chemical Space

We introduce property-independent kernels for machine learning modeling of arbitrarily many molecular properties. The kernels encode molecular structures for training sets of varying size, as well as similarity measures sufficiently diffuse in chemical space to sample over all training molecules. Corresponding molecular reference properties provided, they enable the instantaneous generation of ML models which can systematically be improved through the addition of more data. This idea is exemplified for single kernel based modeling of internal energy, enthalpy, free energy, heat capacity, polarizability, electronic spread, zero-point vibrational energy, energies of frontier orbitals, HOMO-LUMO gap, and the highest fundamental vibrational wavenumber. Models of these properties are trained and tested using 112 kilo organic molecules of similar size. Resulting models are discussed as well as the kernels' use for generating and using other property models.