Bioinformatics Notes
Lecture slides and notes for Master Appendix in Bioinformatics Notes by Md Ahbab. 21 pages.

CSE 4893 Introduction to Bioinformatics Master Appendix CSE 4893: Introduction to Bioinformatics Master Appendix: Glossary, Formulas, Tools and Papers How to use this appendix This appendix gathers every keyword, formula, tool and landmark paper from thethirteen modules into one place. Each item carries a small coloured module tag,so you can always trace it back to the lecture it came from. Read Appendix Awhen a word is unfamiliar. Read Appendix B when you need a formula and itssymbols. Use Appendix C to pick software, and Appendix D to find the originalpaper. Appendices E to H support lab work and last minute revision. Contents 1 Appendix A: Keyword Glossary 2 2 Appendix B: Master Formula Sheet 14 2.1 M1: Foundations of Bioinformatics and Biological Databases . . . . . . . . . . . . . . 14 2.2 M2: Pairwise Sequence Alignment and Dynamic Programming . . . . . . . . . . . . . 14 2.3 M3: Database Search Heuristics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.4 M4: Multiple Sequence Alignment and Molecular Phylogenetics . . . . . . . . . . . . . 15 2.5 M5: Genome Sequencing, Assembly and Read Mapping . . . . . . . . . . . . . . . . . 16 2.6 M6: Gene Prediction, G
CSE 4893 Introduction to Bioinformatics Master Appendix 1 Appendix A: Keyword Glossary 16S rRNA gene M13 A short bacterial gene that changes slowly and is used as a barcode to tell microbial species apart. See also: Amplicon sequencing, Taxonomy. Ab initio gene prediction M6 Finding genes using only the DNA sequence and a statistical model, with no help from known transcripts. See also: Hidden Markov model, Open reading frame. Accession number M1 A stable identifier given to a record in a public database so anyone can retrieve exactly the same entry later. See also: Ensembl, UniProt. Adapter trimming M5 Cutting the short synthetic sequences added during library preparation off the ends of reads before analysis. See also: Trimming, Quality control. Adjusted p value M7 A p value that has been corrected for the fact that thousands of genes were tested at once. See also: False discovery rate, Multiple testing. Affine gap penalty M2 A gap cost with a large charge for opening a gap and a smaller charge for each extra gap position, which favours few long gaps. See also: Gap penalty, Gotoh algorithm. Allele M9 One of the alternative versions of a sequence found at the same position in diff
CSE 4893 Introduction to Bioinformatics Master Appendix identifier, scRNA-seq. Base calling M5 Turning the raw signal from a sequencing machine into letters of DNA with a confidence value for each. See also: Phred score, FASTQ. Basic Local Alignment Search Tool M3 A fast heuristic search that finds short strong local matches between a query and a whole database. See also: Seed and extend, E value. Batch effect M12 An unwanted systematic difference between samples caused by when or how they were processed rather than by biology. See also: Normalisation, Harmony. Bayesian inference M4 A way of combining prior belief with data to give a probability for each possible answer. See also: Likelihood, Maximum likelihood tree. BED M6 A simple text format listing genome intervals with a chromosome, a start and an end. See also: GFF3, Genome browser. Beta sheet M10 A flat pleated protein shape formed when several stretches of backbone lie side by side. See also: Alpha helix, Secondary structure. Betweenness centrality M8 A score for how often a node sits on the shortest path between other nodes, marking it as a bridge. See also: Hub, Node. Bin M13 A group of assembled contigs from a mixed samp