Bioinformatics Notes

Multiple Sequence Alignment and Molecular Phylogenetics

Lecture slides and notes for Multiple Sequence Alignment and Molecular Phylogenetics in Bioinformatics Notes by Md Ahbab. 15 pages.

Document Info: 15 pages ยท PDF

Multiple Sequence Alignment and Molecular Phylogenetics, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 4 CSE 4893: Introduction to Bioinformatics Module 4: Multiple Sequence Alignment and Molecular Phylogenetics Abstract These notes introduce the two ideas that carry most of comparative genomics: aligning many sequences at once, and turning that alignment into a tree. We start from the sum of pairs objec- tive, explain why exact alignment of many sequences is out of reach, and follow the progressive, iterative and consistency strategies used by current tools. We then measure conservation with information content, and move to phylogenetics: Newick notation, corrected distances, neigh- bour joining, parsimony, likelihood with the standard substitution models, Bayesian inference and bootstrap support. Worked numeric examples, pseudocode, common pitfalls and a keyword appendix support laboratory practice and self study. Contents 1 Learning Objectives 2 2 Multiple Sequence Alignment 2 2.1 The Objective Function . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2.2 Progressive Alignment . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2.3 Iterative and Consistency Methods . . . .

CSE 4893 Introduction to Bioinformatics Module 4 1 Learning Objectives After working through this module a student should be able to: 1. State the sum of pairs objective for a multiple alignment and explain why exact multi way dynamic programming is impractical. 2. Run and compare progressive, iterative and consistency based aligners, and judge when each is appropriate. 3. Read and write Newick strings, and interpret branch lengths, rooting and support values on a tree. 4. Compute corrected pairwise distances, build a neighbour joining tree by hand, and score a topology under maximum parsimony. 5. Explain maximum likelihood and Bayesian inference, including substitution models, rate heterogeneity and convergence diagnostics. 2 Multiple Sequence Alignment A multiple sequence alignment (MSA) arranges three or more sequences into a matrix so that residues in the same column are hypotheses of shared ancestry. Everything downstream, from conserved motif discovery to tree inference, inherits the errors of this matrix, so an MSA is best treated as an estimate with uncertainty rather than as data. 2.1 The Objective Function The most common score for an alignment A of k sequences is the sum

CSE 4893 Introduction to Bioinformatics Module 4 B C D A 0.11 0.42 0.45 B 0.40 0.44 C 0.09 A B C D 1. align A with B 2. align C with D 3. align the two profiles Stage 1: pairwise distances Stage 2: guide tree Stage 3: merge order Figure 1: The three stages of progressive alignment: distances, guide tree, then profile merging in tree order. 2.3 Iterative and Consistency Methods Iterative refinement repeatedly splits a finished alignment into two groups, realigns the two pro- files, and keeps the change if the objective improves. Consistency methods take a different route: they collect evidence from all pairwise alignments, and reward a column only if third sequences agree with it, so the score of matching residue x in sequence i with residue y in sequence j is boosted when many intermediate sequences support that pairing. Table 1: Three widely used aligners, with the release current at the time of writing. Tool Strategy Scales to Best use MUSCLE v5.3 Progressive with probabilistic consistency, plus alignment ensembles for replicate MSAs Tens of thousands of sequences Protein families where you want to test how alignment uncertainty affects the tree T-Coffee Consistency from a librar