Researcher profile

Teresa Head-Gordon

Teresa Head-Gordon contributes to research discovery and scholarly infrastructure.

ResearcherAffiliation not importedOpen to collaborate

Trust snapshot

Quick read

Trust 21 - EmergingVerification L1Unclaimed author
8works
0followers
7topics
4close collaborators

Actions

Decide how to stay connected

Follow researcher0

Identity and collaboration

How to connect with this researcher

Claiming links this public author record to a researcher profile and unlocks direct collaboration workflows.

Log in to claim

Direct collaboration

Open a focused conversation when the fit is right

Claim this author entity first to unlock direct invitations.

Research graph

See the researcher in context

Open full explorer

Inspect adjacent work, topics, institutions and collaborators without jumping out to a separate graph page.

Building this graph slice

BZPEER is loading the nearby papers, people, topics and institutions for this page.

Published work

8 published item(s)

preprint2026arXiv

Leak Proof PDBBind: A Reorganized Dataset of Protein-Ligand Complexes for More Generalizable Binding Affinity Prediction

The majority of machine learning scoring functions used in drug discovery for predicting protein-ligand binding poses and affinities have been trained on the PDBBind dataset. However, it is unclear whether these new scoring functions are actually an improvement over traditional models since often the training and test sets are cross-contaminated with proteins and ligands with high similarity, and hence they may not perform comparably well in binding prediction of unrelated protein-ligand complexes. In this work we have carefully prepared a new split of the PDBBind data set to control for data leakage, defined as proteins and ligands with high sequence and structural similarity. The resulting leak-proof (LP)-PDBBind data is used to retrain four popular SFs: AutoDock Vina, Random Forest (RF)-Score, InteractionGraphNet (IGN), and DeepDTA, to better test their capabilities when applied to new protein-ligand complexes. In particular we have formulated a new independent data set, BDB2020+, by matching high quality binding free energies from BindingDB with co-crystalized ligand-protein complexes from the PDB that have been deposited since 2020. Based on all the benchmark results, the retrained models using LP-PDBBind consistently perform better, with IGN especially being recommended for scoring and ranking applications for new protein-ligand systems.

preprint2022arXiv

Learning Correlations between Internal Coordinates to improve 3D Cartesian Coordinates for Proteins

We consider a generic representation problem of internal coordinates (bond lengths, valence angles, and dihedral angles) and their transformation to 3-dimensional Cartesian coordinates of a biomolecule. We show that the internal-to-Cartesian process relies on correctly predicting chemically subtle correlations among the internal coordinates themselves, and learning these correlations increases the fidelity of the Cartesian representation. This general problem has been solved with machine learning for proteins, but with appropriately formulated data is extensible to any type of chain biomolecule including RNA, DNA, and lipids. We show that the internal-to-Cartesian process relies on correctly predicting chemically subtle correlations among the internal coordinates themselves, and learning these correlations increases the fidelity of the Cartesian representation. We developed a machine learning algorithm, Int2Cart, to predict bond lengths and bond angles from backbone torsion angles and residue types of a protein, and allows reconstruction of protein structures better than using fixed bond lengths and bond angles, or a static library method that relies on backbone torsion angles and residue types on a single residue. The Int2Cart algorithm has been implemented as an individual python package at https://github.com/THGLab/int2cart.

preprint2020arXiv

An Isolated Water Droplet in the Aqueous Solution of a Supramolecular Tetrahedral Cage

Water under nanoconfinement at ambient conditions has exhibited low-dimensional ice formation and liquid-solid phase transitions, but with structural and dynamical signatures which map onto known regions of waters phase diagram. Using THz absorption spectroscopy and ab initio molecular dynamics, we have investigated the ambient water confined in a supramolecular tetrahedral assembly, and determined that a distinct network of 9-10 water molecules is present within the nanocavity of the host. The low-frequency absorption spectrum and theoretical analysis of the water in the $Ga_4$$L_6$$^{-12}$ host demonstrate that the structure and dynamics of the encapsulated droplet is distinct from any known phase of water. A further inference is that the release of the highly unusual encapsulated water droplet creates a strong thermodynamic driver for the high affinity binding of guests in aqueous solution for the $Ga_4$$L_6$$^{-12}$ supramolecular construct.

preprint2020arXiv

Convergence of Stochastic-extended Lagrangian molecular dynamics method for polarizable force field simulation

Extended Lagrangian molecular dynamics (XLMD) is a general method for performing molecular dynamics simulations using quantum and classical many-body potentials. Recently several new XLMD schemes have been proposed and tested on several classes of many-body polarization models such as induced dipoles or Drude charges, by creating an auxiliary set of these same degrees of freedom that are reversibly integrated through time. This gives rise to a singularly perturbed Hamiltonian system that provides a good approximation to the time evolution of the real mutual polarization field. To further improve upon the accuracy of the XLMD dynamics, and to potentially extend it to other many-body potentials, we introduce a stochastic modification which leads to a set of singularly perturbed Langevin equations with degenerate noise. We prove that the resulting Stochastic-XLMD converges to the accurate dynamics, and the convergence rate is both optimal and is independent of the accuracy of the initial polarization field. We carefully study the scaling of the damping factor and numerical noise for efficient numerical simulation for Stochastic-XLMD, and we demonstrate the effectiveness of the method for model polarizable force field systems.

preprint2020arXiv

Learning to Make Chemical Predictions: the Interplay of Feature Representation, Data, and Machine Learning Algorithms

Recently supervised machine learning has been ascending in providing new predictive approaches for chemical, biological and materials sciences applications. In this Perspective we focus on the interplay of machine learning algorithm with the chemically motivated descriptors and the size and type of data sets needed for molecular property prediction. Using Nuclear Magnetic Resonance chemical shift prediction as an example, we demonstrate that success is predicated on the choice of feature extracted or real-space representations of chemical structures, whether the molecular property data is abundant and/or experimentally or computationally derived, and how these together will influence the correct choice of popular machine learning algorithms drawn from deep learning, random forests, or kernel methods.

preprint2020arXiv

Stochastic Constrained Extended System Dynamics for Solving Charge Equilibration Models

We present a new stochastic extended Lagrangian solution to charge equilibration that eliminates self-consistent field (SCF) calculations, eliminating the computational bottleneck in solving the many-body solution with standard SCF solvers. By formulating both charges and chemical potential as latent variables, and introducing a holonomic constraint that satisfies charge conservation, the SC-XLMD method accurately reproduces structural, thermodynamic, and dynamics properties using ReaxFF, and shows excellent weak- and strong-scaling performance in the LAMMPS molecular simulation package.

preprint2020arXiv

Strong Anisotropy in Liquid Water upon Librational Excitation using Terahertz Laser Fields

Tracking the excitation of water molecules in the homogeneous liquid is challenging due to the ultrafast dissipation of rotational excitation energy through the hydrogen-bonded network. Here we demonstrate strong transient anisotropy of liquid water through librational excitation using single-color pump-probe experiments at 12.3 THz. We deduce a third order response of chi^3 exceeding previously reported values in the optical range by three orders of magnitude. Using a theory that replaces the nonlinear response with a material response property amenable to molecular dynamics simulation, we show that the rotationally damped motion of water molecules in the librational band is resonantly driven at this frequency, which could explain the enhancement of the anisotropy in the liquid by the external Terahertz field. By addition of salt (MgSO4), the hydration water is instead dominated by the local electric field of the ions, resulting in reduction of water molecules that can be dynamically perturbed by THz pulses.

preprint2019arXiv

Extended Experimental Inferential Structure Determination Method for Evaluating the Structural Ensembles of Disordered Protein States

Characterization of proteins with intrinsic or unfolded state disorder comprises a new frontier in structural biology, requiring the characterization of diverse and dynamic structural ensembles. We introduce a comprehensive Bayesian framework, the Extended Experimental Inferential Structure Determination (X-EISD) method, that calculates the maximum log-likelihood of a protein structural ensemble by accounting for the uncertainties of a wide range of experimental data and back-calculation models from structures, including NMR chemical shifts, J-couplings, Nuclear Overhauser Effects, paramagnetic relaxation enhancements, residual dipolar couplings, and hydrodynamic radii, single molecule fluorescence Förster resonance energy transfer efficiencies and small angle X-ray scattering intensity curves. We apply X-EISD to the drkN SH3 unfolded state domain and show that certain experimental data types are more influential than others for both eliminating structural ensemble models, while also finding equally probable disordered ensembles that have alternative structural properties that will stimulate further experiments to discriminate between them.