Bioinformatics Notes
Lecture slides and notes for Database Search Heuristics BLAST, FASTA and Profile Methods in Bioinformatics Notes by Md Ahbab. 12 pages.

CSE 4893 Introduction to Bioinformatics Module 3 Introduction to Bioinformatics Module 3: Database Search Heuristics Abstract Exhaustive dynamic programming compares a query with every residue of a database and is fartoo slow for present day sequence collections. This module explains the heuristics that make database searchpractical. We follow FASTA from k-tuple lookup through diagonal scoring to banded rescoring, then BLAST fromword neighbourhoods and the two hit rule to gapped exten- sion. We derive the Karlin and Altschul statistics thatturn a raw score into a bit score, an expect value and a probability, and we work a full numeric example.Position specific scoring matrices and profile hidden Markov models extend the same idea to whole families.We close with filtering, common pitfalls, and the modern fast search tools. Contents 1 Learning Objectives 1 2 Why Heuristics 2 3 The FASTA Algorithm 2 4 BLAST 2 4.1 Seed and Extend . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 4.2 The BLAST Family . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 5 Statistics of Local Alignment 3 6 Position Specific Scoring and P
CSE 4893 Introduction to Bioinformatics Module 3 3. Describe the BLAST seed and extend strategy, including the threshold T, HSPs, the two hit rule and gapped extension. 4. Compute bit scores, expect values and p values from the Karlin and Altschul formulation, and interpret them. 5. Build a position specific scoring matrix and explain how a profile hidden Markov model generalises it. 2 Why Heuristics The Smith and Waterman algorithm gives the optimal local alignment, but it fills one cell for every pair ofresidues, so a query of length m against a database of n residues costs O(mn) time. That cost is fine for twosequences and hopeless for a database, because n now runs into the hundreds of billions of residues. Aheuristic keeps the same scoring model but refuses to look at regions that cannot possibly contain a goodalignment. It trades a small, measurable loss of sensitivity for a very large gain in speed.Table 1 gives rough orders of magnitude for a query of m = 1000 residues on a single core that fillsabout 108 cells per second. The numbers are indicative only, but the shape of the growth is the lesson. Table 1: Rough single core running time of exhaustive dynamic programming for
CSE 4893 Introduction to Bioinformatics Module 3 database sequence, position j query, position i grey dots: isolated hits, discarded orange dots: hits sharing one diagonal d = j −i, joined and rescored in a band Figure 1: FASTA diagonal hit diagram. Hits that fall on a common diagonal support one ungapped alignment, so onlythose regions are passed to banded dynamic programming. T produces many words,which is slow but sensitive, while a high T produces few words and misses weak similarity. For proteins thedefaults are w = 3 with the BLOSUM62 matrix. For nucleotides the default megablast task uses w = 28, thetraditional -task blastn uses w = 11, and blastn-short uses w = 7.High scoring segment pairs. A seed is extended in both directions without gaps for as long as the runningscore stays within X of the best score seen so far, a rule called X drop. The maximal scoring ungapped pairof segments found this way is a high scoring segment pair, or HSP.The two hit rule. Gapped BLAST only triggers an extension when two non overlapping word hits appear onthe same diagonal within a window of A residues, typically A = 40. Most isolated hits are noise, so this singlerule removes the great majori