Bioinformatics Notes

Structural Bioinformatics and Protein Structure Prediction

Lecture slides and notes for Structural Bioinformatics and Protein Structure Prediction in Bioinformatics Notes by Md Ahbab. 13 pages.

Document Info: 13 pages · PDF

Structural Bioinformatics and Protein Structure Prediction, first page preview

Content Preview

CSE 4893: Introduction to Bioinformatics Module 10: Structural Bioinformatics and Protein Structure Prediction Abstract These notes introduce structural bioinformatics, the study of how a protein sequence becomes athree dimensional shape and how computers help us read that shape. We start with amino acid chemistry,the four levels of structure and the backbone angles that geometry allows. We then survey how structuresare measured, how they are stored, how they are compared and how they are predicted, from early propensityrules to today’s deep learning systems. Practical topics include root mean square deviation, the templatemodelling score, comparative modelling, coevolution signals, docking and molecular mechanics energy. Themodule closes with protein design and the duties that come with it, followed by worked exercises, keywordsand links for further study. Contents 1 Learning Objectives 1 2 Protein Architecture 2 3 Geometry and Folding 2 4 Experimental Structure Determination 3 5 The Protein Data Bank 4 5.1 Anatomy of an entry . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 5.2 PDB format against mmCIF . . . . . . . . . . . . . . . . . . . . . .

CSE 4893 Introduction to Bioinformatics Module 10 2. Explain backbone geometry, read a Ramachandran plot, and state Anfinsen’s hypothesis and Levinthal’s paradox. 3. Compare X-ray crystallography, nuclear magnetic resonance and cryogenic electron microscopy, and judge the quality of a deposited structure. 4. Compute and interpret RMSD and the template modelling score after an optimal superpo- sition. 5. Choose an appropriate prediction route, from comparative modelling to deep learning predictors, and read its confidence scores honestly. 2 Protein Architecture A protein is a linear polymer of amino acids joined by peptide bonds. Every residue shares the samebackbone atoms, N, Cα and C, and differs only in its side chain. Because the side chain decides whether aresidue prefers water or prefers the oily interior, the sequence quietly encodes the shape. Table 1: The twenty standard amino acids grouped by side chain property, with three letter and one letter codes. Group Members (name, three letter code, one letter code) Nonpolar aliphatic Glycine Gly G; Alanine Ala A; Valine Val V; Leucine Leu L; Isoleucine Ile I; Proline Pro P; Methionine Met M Aromatic Phenylalanine Phe F; Tyrosine

CSE 4893 Introduction to Bioinformatics Module 10 −180 −90 0 90 180 −180 −90 0 90 180 beta and polyproline right handed alpha left handed alpha ϕ (degrees) ψ (degrees) Figure 2: A Ramachandran plot standing in for a real survey of high resolution structures. Each dot is one residue.The shaded blocks mark the sterically allowed regions; the empty space is forbidden by atomic clashes. Anfinsen’s hypothesis For many small single domain proteins, the native structure is the thermodynamic minimum of the freeenergy under physiological conditions. In other words, all the information needed to fold is already in thesequence, which is precisely why prediction from sequence alone is a sensible goal. Levinthal’s paradox points out that folding cannot be a blind search. If each of 100 residues had onlythree accessible backbone states, the chain would have 3100 ≈5 × 1047 conformations; sampling themat 10−13 seconds each would take longer than the age of the universe, yet real proteins fold in milliseconds.The resolution is that the energy landscape is not flat but funnel shaped, so partly correct structures arealready downhill. free energy conformational entropy, many states to few unfolded ens