Center for Molecular Biophysics
The Center for Molecular Biophysics (CMB) at Oak Ridge National Laboratory (ORNL) has been advancing molecular science since 2006, with support from ORNL and the University of Tennessee.
CMB conducts research at the intersection of biology, chemistry, physics, computation, and neutron sciences, using high-performance computer simulations and biophysical experiments to study the structure and function of biologically relevant molecular systems.
A significant area of research at CMB focuses on subsurface biogeochemistry and environmental science, particularly understanding how bacteria interact with contaminants like mercury. Scientists investigate bacterial mercury-resistance proteins and the catalytic mechanisms of enzymes that degrade mercury, with the goal of developing new methods to reduce mercury pollution.
CMB is also a leader in bioenergy and biomass research, working to improve biofuel production by studying the physical and chemical properties of lignin, cellulose, and biomass-degrading microbes. Researchers explore the role of hydrogen bonding in cellulose breakdown, examine the catalytic mechanisms of cellulose-degrading enzymes, and investigate how lignin and cellulose interact to find more efficient ways to process biomass.

Scientists use computational and neutron-based techniques to study protein folding, dynamics, and function, examining structured folding pathways, sugar recognition by ricin-like domains, and the identification of mercury methylation genes and proteins. These studies provide valuable insights into molecular mechanisms underlying environmental and biological processes.
CMB harnesses ORNL’s supercomputing resources to perform large-scale simulations of biomolecular systems. Researchers conduct multimillion-atom simulations of biomass breakdown and use rapid ligand docking to accelerate drug discovery. The center also integrates neutron scattering with computational modeling to examine protein interactions at the atomic level. This approach helps researchers study biomolecular dynamics and understand how biomembranes are structured and how lipids move within them.
CMB helps scientists bridge molecular insights with larger biological systems, contributing to advancements in precision medicine, synthetic biology, and bioengineering. By combining high-performance computing, neutron science, and experimental biophysics, researchers can continue to push the boundaries of molecular science and enhance scientific understanding of biomolecular systems.
Explore Our Research
Visualization of Solvent Disruption of Biomass and Biomembrane
Lignocellulosic biomass is recalcitrant to deconstruction and saccharification due to its fundamental molecular architecture and multicomponent laminate composition. A fundamental understanding of the structural changes and associations that occur at the molecular level during biosynthesis, deconstruction, and hydrolysis of biomass is essential for improving processing and conversion methods for lignocellulose-based fuels and products. The objective of this research is to develop and demonstrate a combined neutron scattering and computer simulation technology for multiple-length scale, real-time imaging of biomass during pretreatment and enzymatic hydrolysis.
Biomass Solvent Pretreatment
Production of ethanol by bioconversion of lignocellulosic biomass requires biomass pretreatment to increase the enzymatic digestibility of cellulose. Three major pretreatment regimes have been investigated, involving: aqueous solutions; co-solvents of water and organic solvents, such as THF; ionic liquids. We describe how those three pretreatment regimes change the structure of biomass components: cellulose, lignin and hemicellulose.
Biomass Dynamics
The full utilization of plant biomass for the production of energy and novel materials often involves high temperature heating/cooling cycles. High temperature is applied to soften lignin by enhancing its underlying atomic dynamics. Moreover, the hydration of the lignin is different in those process: biomass having a higher water content than isolated lignin has. However, a molecular-level description of the dependence on temperature and hydration of lignin dynamics is lacking. Molecular dynamics simulations were combined with neutron scattering and dielectric spectroscopy experiments to probe the dependence of lignin dynamics on hydration and thermal history. Hydration was found to always make lignin more dynamic. Further, at any given temperature and hydration during heating/cooling cycle, lignin was found to be more dynamic upon cooling than upon heating. Syringyl lignin units were found to be more dynamic than Guaiacyl, and aliphatic chains more dynamic than the aromatic rings.
Efflux Pump
The problem of antibiotic resistance is an emerging public health threat. Many bacteria are resistant to not just one antibiotic, but many different antibiotics. In fact, some bacteria are resistant to all currently available antibiotics. One of the primary causes of resistance in bacteria are multiprotein complexes known as efflux pumps. When an antibiotic enters a bacterial cell, the efflux pump kicks it back out preventing the antibiotic from fulfilling its function of killing bacteria and healing the infection. At CMB in collaboration with JC Gumbart at Georgia Tech, Helen Zgurskaya at Oklahoma University, and John Walker at St. Louis University School of Medicine, we are using three-dimensional models of each component and the full complex of the efflux pump in Escherichia coli to discover new small molecules that can inhibit the pump, design new molecules from discovered scaffolds, and gain an atomistic understanding of the assembly and transport mechanism.
Histone Deacetylase 4
Histone deacetylase 4 (HDAC4) is a member of the histone deacetylase family of enzymes and is believed to function as a scaffold for the recruitment of other multiprotein complexes like Nuclear receptor co-repressor (NCoR). However, it has been reported that HDAC4 is overexpressed in esophageal squamous cell carcinoma (ESCC) and prostate cancer cells. All the known HDAC4 inhibitors till date are non-selective binders and cause toxicity. Hence to effectively prevent HDAC4 activity in cancer growth, there is a need to identify selective inhibitors. In our study, we have used a novel strategy to inhibit HDAC4 by targeting the interface of HDAC4-NCoR instead of the catalytic pocket. Targeting a protein-protein interface (PPI) has been challenging because in most cases, they lack cavities for small molecules to bind or are intrinsically disordered. However, using molecular dynamics (MD) simulations, We have identified novel cavities at the PPI which were not evident in the crystal structures. To find potential inhibitors, we have performed a high-throughput virtual screening of small molecules to target these new pockets.
Adhesin Protein
Infective Endocarditis (IE). IE is a life-threatening cardiovascular infection associated with significant morbidity and mortality. We are working on adhesin proteins that mediate microbe attachment to human platelets via sialoglycans thus enabling pathogenesis. IE kills thousands of people in the United States alone but there are no known anti-adhesion molecules to target the disease early. Our major goal is thus to find small molecules which inhibit by binding at the sialoglycan binding pocket (competitive inhibitors) or bind to an allosteric pocket (allosteric inhibitors). We have performed first-round of virtual screening to identify novel competitive inhibitors which will be tested by our collaborators at Vanderbilt University.
AKRAS
Normal KRAS proteins have a vital role in cell division, differentiation and in tissue signaling. However, dysfunctional KRAS proteins can lead to the development of malignant tumors like colorectal and lung cancers. Despite its role in cancer development, there is no effective drug against KRAS at present. Hence, there is a need to perform a high-throughput study to identify novel KRAS inhibitors. On this project, The work is done by Rupesh Agarwal, Wade Seifert and Dr. Micholas D. Smith. We want to identify novel scaffolds of small molecules which bind to allosteric sites of KRAS rather than at native substrate binding site. Our plan is to perform a high-thoughput virtual screening of 16 million small molecules present in the ZINC database on an ensemble of KRAS structures obtained from MD simulations and use this to identify a common pharmacophore or scaffold between the hits.
FGF-23
Patients with excess fibroblast growth factor-23 (FGF-23) in their circulation have metabolic deficiencies that contribute to chronic kidney disease and the bone-softening disorder rickets. We used computationally based screening (ensemble docking and virtual high-throughput screening) to identify a compound predicted to bind the FGF-23. When tested experimentally, the compound was selective for FGF-23 compared to other FGF family ligands and prevented FGF signaling in cultured kidney cells.
Novel Strategies
Developing a quantitative understanding of the structure, function, and dynamics of protein-protein, protein-ligand, and protein-nucleic acid interactions, and how such interactions modulate protein function, is at the heart of a molecular-level understanding of disease and how to target mutant or otherwise pathogenic proteins with drugs that have a low side-effect profile. Towards this end, computational strategies to predict high-specificity binding of drug ligands to proteins and strategies to simulate how the binding of such ligands modulates protein function are being developed, with a specific focus on targeted cancer therapy as a replacement for chemotherapy. Specifically, machine learning methods using a range of physicochemical features will be applied to the development of highly accurate scoring functions for prediction of drug ligand binding sites. After this stage, advanced sampling methods such as the Markov state model and features calculated therefrom, such as mutual information among protein residues, or direct observation of a modulatory effect, will be used to assess allostery, that is, the ability of a binding event at one site to modulate the active site function. Targeting allosteric sites holds the promise of greater specificity and a lower side effect profile compared with the standard approach of targeted cancer therapy involving inhibition of the active site.
Quantum Chemical Calculation of pKas of Environmentally Relevant Functional Groups: Carboxylic Acids, Amines and Thiols in Aqueous Solution
Developing accurate quantum chemical approaches for calculating pKas is of broad interest. Useful accuracy can be obtained by using density functional theory (DFT) in combination with a polarizable continuum solvent model. However, some classes of molecules present problems for this approach, yielding errors greater than 5 pK units. Various methods have been developed to improve the accuracy of the combined strategy. These methods perform well but either do not generalize or introduce additional degrees of freedom, increasing the computational cost. The Solvation Model based on Density (SMD) has emerged as one of the most commonly used continuum solvent models. Nevertheless, for some classes of organic compounds, e.g., thiols, the pKas calculated with the original SMD model show errors of 6–10 pK units, and we traced these errors to inaccuracies in the solvation free energies of the anions. To improve the accuracy of pKas calculated with DFT and the SMD model, we developed a scaled solvent-accessible surface approach for constructing the solute–solvent boundary. By using a “direct” approach, in which all quantities are computed in the presence of the continuum solvent, the use of thermodynamic cycles is avoided. Furthermore, no explicit water molecules are required. Three benchmark data sets of experimentally measured pKa values, including 28 carboxylic acids, 10 aliphatic amines, and 45 thiols, were used to assess the optimized SMD model, which we call SMD with a scaled solvent-accessible surface (SMDsSAS). Of the methods tested, the M06-2X density functional approximation, 6-31+G(d,p) basis set, and SMDsSAS solvent model provided the most accurate pKas for each set, yielding mean unsigned errors of 0.9, 0.4, and 0.5 pK units, respectively, for carboxylic acids, aliphatic amines, and thiols. This approach is therefore useful for efficiently calculating the pKas of environmentally relevant functional groups.
Methylmercury Speciation and Dimethylmercury Production in Sulfidic Solutions
Decades of research have been devoted to understanding methylmercury formation and degradation, but little is known about the formation of dimethylmercury in aquatic systems. In collaboration with Andrew Graham at Grinnell College and Niri Govind at the Environmental Molecular Sciences Laboratory (EMSL) at Pacific Northwest National Laboratory (PNNL), we combined complementary experimental and computational approaches to examine methylmercury speciation and dimethylmercury formation in sulfidic solutions. We found that bis(methylmercuric) sulfide, (CH3Hg)2S, decomposes slowly to form dimethylmercury and mercury sulfide (HgS). Using relativistic quantum chemical calculations, we proposed a mechanism for the reaction that is in excellent agreement with the experimental kinetics measurements. In natural systems with relatively high methylmercury/sulfide ratios (e.g., at the oxic/anoxic interface), highly toxic dimethylmercury may be produced.
Identification of Mercury and Dissolved Organic Matter Complexes Using Ultrahigh-Resolution Mass Spectrometry
The chemical speciation and bioavailability of mercury (Hg) is markedly influenced by its complexation with natural dissolved organic matter (DOM) in aquatic environments. To date, however, analytical methodologies capable of identifying specific Hg-DOM complexes are scarce. In collaboration with Baohua Gu, who led the work, we identified candidate molecules with the observed chemical formula (C8H17N2S4Hg+) and then used density functional theory (DFT) to predict their Hg-binding properties and solution-phase structures. Our prediction that the most prevalent complex is a HgII bis(N,N-diethyldithiocarbamate) complex was subsequently verified with mass spectrometry experiments.
Passive Permeability of Mercury and Methylmercury Complexes Through a Bacterial Cytoplasmic Membrane
Cellular uptake and export are important steps in the biotransformation of mercury by microorganisms. However, the mechanisms of transport across biological membranes remain unclear. To identify biophysical features that govern Hg uptake, we performed extensive molecular dynamics simulations of the passive permeation of HgII and methylmercury complexes with thiolate ligands through a model bacterial cytoplasmic membrane. We found that favorable interactions with carbonyl and tail groups of phospholipids stabilize Hg-containing solutes in the tail–head interface region of the membrane.
Site-directed Mutagenesis of HgcA and HgcB Reveals Amino Acid Residues Important for Mercury Methylation
We collaborated with Judy Wall’s group at the University of Missouri, who performed site-directed mutagenesis and protein sequence truncation, to determine amino acid residues in HgcA and HgcB that are important for mercury methylation. Mutations of conserved amino acids near the strictly conserved Cys residue, Cys93, in HgcA had varying impacts on mercury methylation capacity but showed that the structure of the putative ‘cap helix’ region harboring Cys93 is crucial for methylation. In the ferredoxin HgcB, only one of two conserved cysteines at the C-terminus is necessary for methylation. An additional, strictly conserved Cys in HgcB was found to be essential for methylation. This study supports the previously predicted importance of Cys93 in HgcA for methylation of mercury and reveals additional residues in HgcA and HgcB that facilitate the production of this neurotoxin.
X-ray Structure of a Hg2+ Complex of Mercuric Reductase and QM/MM study of Hg2+ Transfer Between the C-terminal and Buried Catalytic Site Cysteine Pairs
In collaboration with Sue Miller at UCSF, we combined X-ray crystallography and QM/MM simulation to understand how the bacterial enzyme mercuric ion reductase, MerA, shuttles Hg2+ ions into its active site, where it is reduced to less toxic Hg0. We found that Hg2+ transfer in MerA involves a balanced competition between Hg2+ and H+ for pairs of cysteine thiolates. Our simulations revealed at atomic detail how MerA, a key component of bacterial mercury resistance, orchestrates the handling of toxic Hg2+ through successive binding by and transfer among pairs of cysteines.
Why Mercury Prefers Soft Ligands
Defining the factors that determine the relative affinities of different ligands for Hg2+ is critical to understanding its speciation, transformation, and bioaccumulation in the environment. We used quantum chemistry to dissect the relative binding free energies for a series of inorganic anion complexes of Hg2+. By comparing Hg2+−ligand interactions in the gaseous and aqueous phases, we showed that differences in interactions with a few, local water molecules leads to a clear periodic trend within the chalcogenide and halide groups and recovers the well-known experimentally observed preference of Hg2+ for soft ligands such as thiols. This approach established a strong basis for understanding Hg speciation in the biosphere with quantum chemical calculations.
The Genetic Basis for Bacterial Mercury Methylation
Methylmercury is a potent neurotoxin produced in natural environments from inorganic mercury by anaerobic bacteria. However, the genes and proteins responsible remained elusive for decades. We discovered a two-gene cluster, hgcAB, required for mercury methylation by Desulfovibrio desulfuricans ND132 and Geobacter sulfurreducens PCA. In either bacterium, deletion of hgcA, hgcB, or both genes abolishes mercury methylation. These genes encode a corrinoid protein, HgcA, and a ferredoxin, HgcB, consistent with roles as a methyl carrier and an electron donor required for corrinoid cofactor reduction, respectively. Among bacteria and archaea with sequenced genomes, gene orthologs are present in confirmed methylators but absent in nonmethylators, suggesting a common mercury methylation pathway in all methylating bacteria and archaea sequenced to date.
Mechanism of Hg-C Protonolysis in the Organomercurial Lyase MerB
The bacterial organomercurial lyase, MerB, catalyzes the demethylation of methylmercury. Using a DFT-based quantum chemical cluster approach, we investigated two proposed reaction mechanisms of MerB. The calculations ruled out a direct protonation mechanism in which a cysteine residue delivers a catalytic proton to the methyl leaving group. Instead, the calculations support a two-step mechanism in which two Cys thiolates coordinate methylmercury and an Asp protonates the methyl leaving group to cleave the Hg-C bond and release CH4. Natural population analysis revealed that MerB lowers the activation free energy by redistributing electron density into the leaving group and away from the catalytic proton.
Plant Cell Wall
Terrestrial plants harness solar energy and remove carbon dioxide from the atmosphere, transforming it into plant cell walls composed of energy-rich, renewable and complex biomaterial, also known as cellulosic biomass or lignocellulose. We are partners of the Center for Lignocellulose structure and Formation that is focused on discovering how plants make these complex, hierarchically-structured biomaterials.
Protein Flexibility
Our understanding of the structure:function paradigm for proteins has shifted in recent years, recognizing that disordered proteins and flexible assemblies are highly abundant in cells and play central roles in biological function. These systems are particularly challenging because they cannot be studied using widely available structural biology techniques such as crystallography, nuclear magnetic resonance spectroscopy or cryo- electron microscopy. We are developing new integrative approaches and technologies are needed to characterize disordered and flexible systems. Small-angle neutron scattering (SANS) is ideally suited to studies of flexible and disordered proteins, and large dynamic complexes. The specific labeling that is observable with neutrons and enabled by deuteration enhances the visibility of specific parts of complex and dynamic assemblies. Combining SANS with computational modelling can thus provide unique structural information that is unattainable by other means.
Developing physical models of disordered proteins and flexible dynamic assemblies is uniquely challenging because these systems typically exist as an ensemble of conformational states. We address the two main computational challenges facing molecular dynamics (MD) simulations: obtaining adequate sampling and employing reliable force fields. Running scalable MD codes that harvest the immense computational power of ORNL provide enhanced sampling. The MD models are validated by calculating SANS structure factors and comparing directly to experiment. Integration of experiment and MD simulation involves the use of simulation to guide experimental design, as well as SANS data to refine the simulation results.
GROMACS: Atomic Simulations on Massively Parallel Machines
The most highly used MD program in recent years is the GROMACS molecular simulation toolkit. GROMACS is cited by ~2,500 scientific articles per year, a rate that has been rapidly increasing recently. GROMACS is known for its high performance and its efficient, effective use of system resources, but new, massively parallel HPC (high-performance computing) architectures present a continuing challenge. At CMB, we are helping to meet this challenge. Specifically, we are working to upgrade GROMACS’ threading framework to keep all CPU cores busy by overlapping computing and communication. This requires careful rescheduling of threads to different tasks during data transfers, which was previously time that threads spent idle. This must be done while avoiding thread-management overhead, which can erode performance gains. Finally, all this has to be done within the narrow two millisecond window when GROMACS computes atom positions, forces, etc. for the next simulation step. To meet this goal, we are developing a new threading framework, called STS (static thread scheduler) (https://github.com/eblen/sts) and integrating it into GROMACS. Funding provided through the INTEL(R) IPCC Program.
Scaling of Multimillion-Atom MD Simulation on Petascale Supercomputers
The performance of multimillion-atom MD simulations on petascale supercomputers is limited by the most commonly used electrostatic method: PME. Particle mesh ewald (PME) requires two 3D fast fourier transforms (FFT) and these in turn require global communication. These global communications are very time consuming on massively parallel supercomputers, such as Jaguar at ORNL. Our recent paper shows that for the chosen test system (lignocellulosic biomass) the reaction field (RF) method produces very accurate results when compared to results obtained using PME. As seen in the graph, a simulation using RF shows significant better scaling than one of the same system using PME.
AdaptiveMD: HPC Adaptive Sampling Workflow Package
We are working on a software package that automates biomolecular dynamics simulation workflows. With AdaptiveMD, you can create workflows that iteratively restart a swarm of simulation replicates by sampling initial states from a Markov State Model of pre-existing data. Look at the github page for instructions on how to install and use the package.
Contact
