Bioinformatics Notes
Lecture slides and notes for Genome Sequencing, Assembly and Read Mapping in Bioinformatics Notes by Md Ahbab. 12 pages.

CSE 4893: Introduction to Bioinformatics Module 5: Genome Sequencing, Assembly and Read Mapping Abstract These subsidiary notes cover how a genome is read, rebuilt and searched. We start withthe chem- istry and error profile of the main sequencing platforms, then move to read quality controland to the coverage theory that tells us how much data is enough. We estimate genome size fromk-mer spectra, compare overlap layout consensus assembly with de Bruijn graph assembly, and learn toread the graph pathologies that break contiguity. The second half covers read mapping: suffix arrays,the Burrows Wheeler transform, the FM index, the seed chain extend strategy, and the SAM format.Worked numbers, algorithms, exercises and a keyword list support self study. Contents 1 Learning Objectives 1 2 Sequencing Technologies 2 3 Reads and Quality Control 2 4 Coverage Theory 3 5 Genome Size from k-mer Spectra 4 6 Assembly 4 7 Read Mapping 6 8 PLACEHOLDER: Additional Formulas 8 9 Algorithms 8 10 Reference Genomes and Coordinates 8 11 Landmark Paper Summaries 9 12 Worked Exercises 11 13 Keywords 11 14 Useful Links 12 1 Learning Objectives 1. Compare the major sequencing platforms by read length, raw ac
CSE 4893 Introduction to Bioinformatics Module 5 2 Sequencing Technologies Every platform trades three things against each other: read length, raw accuracy andcost per base. No platform wins on all three, so serious projects often combine two of them.Short Table 1: Sequencing platforms at a glance. Figures follow current vendor specification sheets andshould be rechecked every year. Platform Chemistry Read length Raw accuracy Throughput Typical use Sanger Capillary chain termination 500 to 1000 bp Q40 to Q50 Under 100 kb per run Single amplicons, clinical confirmation Illumina Reversible terminator sequencing by synthesis 2 x 50 to 2 x 300 bp 85 to 90 percentof bases at Q30 or better Up to about 16 Tb per dual flow cell run on NovaSeq X Plus Resequencing, RNA counting, variant calling Ion Torrent Semiconductor pH detection 200 to 600 bp About Q20, indel errors in homopoly- mers Tens of Gb per run Fast targeted panels and amplicons PacBio HiFi Circular consensus on SMRT cells (Revio, Vega) 15 to 25 kb Median read Q30 orbetter About 100 to 120 Gb per SMRT cell De novo assembly, phasing, methylation Oxford Nanopore Ionic current through a protein pore, R10.4.1 10 kb to above 100 kb ul
CSE 4893 Introduction to Bioinformatics Module 5 Over trimming Aggressive quality trimming looks tidy in the report and quietly damages the analysis. Very shorttrimmed reads map ambiguously, uneven trimming biases coverage at fragment ends, and assemblers losethe overlaps they depend on. Always trim adapters, trim quality gently (a thresh- old near Q20 with aminimum length near 50 bp for 150 bp reads), and remember that mappers such as BWA-MEM already softclip poor ends for you. 4 Coverage Theory Lander and Waterman modelled shotgun sequencing as random sampling of read start points along thegenome. Three results follow at once. C = LN G P0 = e−C E[contigs] = N e−C (2) C coverage depth, the average number of reads covering a base, quoted as 30×. L read length in bases; for paired data use the sequenced bases per pair. N number of reads produced. G genome size in bases. P0 probability that a base is covered by no read, so P0G is the expected number of uncov- ered bases and Ne−C is the expected number of contigs, since each contig ends where a read is followed by a gap. Worked example. Take a bacterium with G = 5,000,000 bp and L = 150 bp. • At C = 10: N = CG/L = 10 × 5,000,000/150 =