Bioinformatics Notes

Foundations of Bioinformatics and Biological Databases

Lecture slides and notes for Foundations of Bioinformatics and Biological Databases in Bioinformatics Notes by Md Ahbab. 11 pages.

Document Info: 11 pages · PDF

Foundations of Bioinformatics and Biological Databases, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 1 CSE 4893: Introduction to Bioinformatics Module 1: Foundations of Bioinformatics and Biological Databases Abstract These notes introduce bioinformatics as a working discipline. We start with what the field is,how it grew out of protein sequence collections in the 1960s and 1970s, and why the sheer volume ofsequencing data made computation unavoidable. We then review just enough molecular biology to reada data file with understanding: base pairing, the central dogma, the genetic code, gene anatomy andprotein structure. The core of the module is the public database landscape, the file formats thatconnect tools together, and the small amount of mathematics you need every day, such as Phred qualityscores, GC content and Shannon entropy. We close with reproducible working practice, data ethics,landmark papers, worked exercises and a keyword list for revision. Contents 1 Learning Objectives 2 2 What Bioinformatics Is 2 3 Molecular Biology in Brief 2 3.1 DNA and Base Pairing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 3.2 The Central Dogma . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

CSE 4893 Introduction to Bioinformatics Module 1 1 Learning Objectives LO1. Define bioinformatics, name its three pillars, and explain why data growth created the field. LO2. Describe the central dogma, the genetic code and gene anatomy well enough to interpret sequence records. LO3. Navigate the major public databases and state, for each, its content and level of curation. LO4. Read and write the common file formats, and convert quality scores, GC content and entropy by hand. LO5. Apply reproducible and ethical working practice, including version control, containers and the FAIR principles. 2 What Bioinformatics Is Definition Bioinformatics is the design and use of computational methods to store, retrieve, analyse and- interpret biological data, mostly sequences, structures and measurements of molecular activity. The field rests on three pillars. The first is biology, which supplies the question and themeaning. The second is computer science, which supplies data structures, algorithms andengineering. The third is statistics, which tells us whether a pattern is real or is noise dressedup as a result. Weak- ness in any one pillar produces confident nonsense.The word itself was coine

CSE 4893 Introduction to Bioinformatics Module 1 DNA pre-mRNA mRNA Protein transcription splicing translation replication reverse transcription Figure 1: The central dogma pipeline, with the reverse transcription shortcut shown dashed. 3.3 The Genetic Code Translation reads mRNA in non-overlapping triplets called codons. There are 43 = 64 codons for 20 standard amino acids plus stop, so the code is degenerate, meaning several codons encode the same amino acid. AUG both starts translation and encodes methionine. UAA, UAG and UGA are stop codons. Degeneracy sits mostly in the third codon position, which is why synonymous changes there often leave the protein unchanged. A reading frame is the choice of starting offset, so a double-stranded region has six frames in total. 3.4 Gene Anatomy A eukaryotic gene is more than its coding sequence. Upstream lies a promoter that recruits the transcription machinery. The transcript begins with a 5’ untranslated region, then alter- nates exons, which are retained, with introns, which are removed by splicing. It ends with a 3’ untranslated region and a polyadenylation signal. Choosing different exon combinations is alternative splicing, and it lets