A protein’s function is determined by its structure, and structure — the way a protein folds — is determined by its sequence of amino acids, the building blocks of proteins.
Many methods for designing novel proteins, including examples that could bind to a disease-causing molecule in our cells, involve a two-step process: The structure comes first, and then a machine-learning framework generates a repertoire of sequences that could potentially adopt that structure.
In nature, many different amino acid sequences can fold into the same structure. At the same time, one amino acid sequence can potentially adopt different structures depending on the protein’s flexibility or a functional trigger. Therefore, when researchers use artificial intelligence to design new proteins, the challenge is to guide AI to “see” that there are many potentially useful answers — that many sequences can adopt the same fold
“For years, the field has measured success by asking whether a model can reproduce the protein sequence that evolution happened to select — our work shows that this isn’t the best metric for protein design,” says Amy E. Keating, Department of Biology head, Jay A. Stein (1968) Professor of Biology, professor of biological engineering, and senior author of a paper recently published in PNAS.
PottsMPNN, a new machine-learning framework developed in the Department of Biology, incorporates the physical principles that govern protein structure and stability, improving sequence generation and the ability to predict how mutations will affect a protein’s stability. In other words, the model has a better understanding of the sequence-energy landscape, meaning the relationship between the identity of each amino acid and the stability of the protein.
Source link







