Bioinformatics Notes

Database Search Heuristics BLAST, FASTA and Profile Methods

Lecture slides and notes for Database Search Heuristics BLAST, FASTA and Profile Methods in Bioinformatics Notes by Md Ahbab. 12 pages.

Document Info: 12 pages · PDF

Database Search Heuristics BLAST, FASTA and Profile Methods, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 3 Introduction to Bioinformatics Module 3: Database Search Heuristics Abstract Exhaustive dynamic programming compares a query with every residue of a database and is fartoo slow for present day sequence collections. This module explains the heuristics that make database searchpractical. We follow FASTA from k-tuple lookup through diagonal scoring to banded rescoring, then BLAST fromword neighbourhoods and the two hit rule to gapped exten- sion. We derive the Karlin and Altschul statistics thatturn a raw score into a bit score, an expect value and a probability, and we work a full numeric example.Position specific scoring matrices and profile hidden Markov models extend the same idea to whole families.We close with filtering, common pitfalls, and the modern fast search tools. Contents 1 Learning Objectives 1 2 Why Heuristics 2 3 The FASTA Algorithm 2 4 BLAST 2 4.1 Seed and Extend . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 4.2 The BLAST Family . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 5 Statistics of Local Alignment 3 6 Position Specific Scoring and P

CSE 4893 Introduction to Bioinformatics Module 3 3. Describe the BLAST seed and extend strategy, including the threshold T, HSPs, the two hit rule and gapped extension. 4. Compute bit scores, expect values and p values from the Karlin and Altschul formulation, and interpret them. 5. Build a position specific scoring matrix and explain how a profile hidden Markov model generalises it. 2 Why Heuristics The Smith and Waterman algorithm gives the optimal local alignment, but it fills one cell for every pair ofresidues, so a query of length m against a database of n residues costs O(mn) time. That cost is fine for twosequences and hopeless for a database, because n now runs into the hundreds of billions of residues. Aheuristic keeps the same scoring model but refuses to look at regions that cannot possibly contain a goodalignment. It trades a small, measurable loss of sensitivity for a very large gain in speed.Table 1 gives rough orders of magnitude for a query of m = 1000 residues on a single core that fillsabout 108 cells per second. The numbers are indicative only, but the shape of the growth is the lesson. Table 1: Rough single core running time of exhaustive dynamic programming for

CSE 4893 Introduction to Bioinformatics Module 3 database sequence, position j query, position i grey dots: isolated hits, discarded orange dots: hits sharing one diagonal d = j −i, joined and rescored in a band Figure 1: FASTA diagonal hit diagram. Hits that fall on a common diagonal support one ungapped alignment, so onlythose regions are passed to banded dynamic programming. T produces many words,which is slow but sensitive, while a high T produces few words and misses weak similarity. For proteins thedefaults are w = 3 with the BLOSUM62 matrix. For nucleotides the default megablast task uses w = 28, thetraditional -task blastn uses w = 11, and blastn-short uses w = 7.High scoring segment pairs. A seed is extended in both directions without gaps for as long as the runningscore stays within X of the best score seen so far, a rule called X drop. The maximal scoring ungapped pairof segments found this way is a high scoring segment pair, or HSP.The two hit rule. Gapped BLAST only triggers an extension when two non overlapping word hits appear onthe same diagonal within a window of A residues, typically A = 40. Most isolated hits are noise, so this singlerule removes the great majori