Bioinformatics Notes
Lecture slides and notes for Pairwise Sequence Alignment and Dynamic Programming in Bioinformatics Notes by Md Ahbab. 11 pages.

CSE 4893 Introduction to Bioinformatics Module 2 Introduction to Bioinformatics Module 2: Pairwise Sequence Alignment and Dynamic Programming Abstract These notes introduce pairwise sequence alignment, the task of lining up two DNA or protein sequences so that related positions sit in the same column. We start with homology, similarity and identity, then build a scoring model from log odds substitution scores and gap penalties. We derive the Needleman and Wunsch recurrence for global alignment, the Smith and Waterman recurrence for local alignment, and the Gotoh formulation for affine gaps. Worked matrices, pseudocode, tool commands and short exercises support hands on practice. Placeholder sections are left open so the module can grow with the class. Read the algorithms with pencil and paper nearby. Contents 1 Learning Objectives 1 2 Homology, Similarity and Identity 2 3 Scoring Models 2 4 Global Alignment: Needleman and Wunsch 3 5 Local Alignment: Smith and Waterman 4 6 Affine Gaps: The Gotoh Formulation 5 7 PLACEHOLDER: Additional Formulas 5 8 Algorithms 6 9 Practical Considerations 6 10 Tools 7 11 Landmark Paper Summaries 8 12 Worked Exercises 9 13 Keywords 9 14 Useful Links 10
CSE 4893 Introduction to Bioinformatics Module 2 gene A human gene A mouse gene B human gene B mouse ancestral gene duplication speciation speciation Figure 1: One duplication followed by one speciation. Gene A in human and gene A in mouse are ortho- logues. Gene A and gene B in the same genome are paralogues. 5. Explain why affine gap costs are more realistic, and state the three coupled recurrences of the Gotoh formulation. 2 Homology, Similarity and Identity Alignment is a way of asking a historical question with a numerical answer. Before we score anything, we need three words to be crisp. Homology A yes or no statement. Two sequences are homologous if they descend from a common ancestral sequence. There is no such thing as ninety percent homologous. Similarity A measurable score. It counts how alike two sequences look under a chosen scoring model, including conservative substitutions. Identity The percentage of aligned columns that hold exactly the same residue. It depends on the alignment, so it is only meaningful when the alignment length is reported with it. Homologous genes come in two flavours. Orthologues are copies separated by a speciation event, so the human and mouse
CSE 4893 Introduction to Bioinformatics Module 2 Table 1: How the two matrix families are built and used. Feature PAM BLOSUM Source data Global alignments of closely re- lated proteins Ungapped blocks of local alignments Model Extrapolated by matrix powers Counted directly at each clustering level Number meaning Higher number means more diver- gence Higher number means less divergence Typical use PAM70 for close pairs, PAM250 for distant pairs BLOSUM80 for close pairs, BLO- SUM45 for distant pairs Default in tools Rare today BLOSUM62 in BLASTP and most aligners Table 2: Common default gap settings. Values are penalties, so they are subtracted from the score. Setting Matrix Open Extend BLASTP default BLOSUM62 11 1 EMBOSS needle default BLOSUM62 10.0 0.5 Distant protein pairs BLOSUM45 14 2 Close protein pairs BLOSUM80 10 1 +11 and the cysteine to cysteine score is +9, because those residues are rare and rarely replaced. Conservative swaps stay positive, for example serine to threonine is +1 and leucine to isoleucine is +2, while asparagine to tryptophan is −4. Scores in the matrix run from +11 down to −4. 3.2 PAM and BLOSUM PAM matrices, from Dayhoff and colleagues in 1978, start fro