Bioinformatics Notes
Lecture slides and notes for Frontiers and Applied Bioinformatics in Bioinformatics Notes by Md Ahbab. 13 pages.

CSE 4893 Introduction to Bioinformatics Module 13 CSE 4893: Introduction to Bioinformatics Module 13: Frontiers and Applied Bioinformatics Abstract Module 13 is the capstone of this course. It asks a simple question: once you canalign, as- semble, annotate and model sequence data, what do you actually do with it? We tourseven applied frontiers, namely metagenomics and the microbiome, clinical and precisiongenomics, cancer genomics, immunoinformatics and drug discovery, genome editing analysis,agricultural and environmental genomics, and genomic surveillance for public health. We thenlook at the plumbing that makes such work reproducible, and at the ethics and policy that makeit legiti- mate. The treatment is deliberately broad and practical, with compact tables, workednumbers, and small algorithms you can implement in a week. Contents 1 Learning Objectives 2 2 Metagenomics and the Microbiome 2 3 Clinical and Precision Genomics 3 4 Cancer Genomics 3 5 Immunoinformatics and Drug Discovery 4 6 Genome Editing Analysis 5 7 Agricultural and Environmental Genomics 6 8 Genomic Surveillance and Public Health 6 9 Reproducible Infrastructure 7 10 Ethics, Law and Policy 8 11 PLACEHOLDER: Addit
CSE 4893 Introduction to Bioinformatics Module 13 1. Learning Objectives By the end of this module you should be able to: 1. Choose between amplicon and shotgun metagenomics for a stated question, andcompute and interpret alpha and beta diversity from a feature table. 2. Walk a clinical case from sample to report, applying variant interpretation tiers and phar- macogenomic guidance to a prescribing decision. 3. Separate somatic from germline findings, distinguish drivers from passengers, andexplain how mutational signatures are extracted from a 96 class matrix. 4. Design and score guide RNAs, and quantify editing outcomes including base and primeediting, while naming the caveats of each tool. 5. Build a reproducible workflow in a containerised pipeline manager, estimate its cost, andargue its ethical and legal position before data are collected. 2. Metagenomics and the Microbiome A metagenome is the pooled genetic material of a community, read without culturing it. Twode- signs dominate. Amplicon sequencing reads one marker gene, usually 16S rRNA for bacteria,ITS for fungi, and is cheap enough for hundreds of samples. Shotgun metagenomics readseverything present and therefore sees
CSE 4893 Introduction to Bioinformatics Module 13 Worked community example. Sample J holds 100 reads: A 40, B 30, C 20, D 10, sop = (0.4, 0.3, 0.2, 0.1).H = −(0.4 ln 0.4+0.3 ln 0.3+0.2 ln 0.2+0.1 ln 0.1) = 0.3665+0.3612+0.3219+ 0.2303 = 1.280.D = 0.16 + 0.09 + 0.04 + 0.01 = 0.30, so 1 −D = 0.70 and 1/D = 3.33 effective taxa.Sample K holds A 10, B 30, C 40, D 20. ThenBCJK = (30 + 0 + 20 + 10)/200 = 60/200 = 0.30. The two samples share most of their mass but therank order of A and C has flipped, which is exactly what a moderate Bray Curtis value describes. Pitfall. Sequencing gives compositional data, not absolute counts. Uneven depth inflatesrich- ness, so rarefy or use a compositional normalisation before comparing samples, and never read araw count as a cell concentration. 3. Clinical and Precision Genomics 3.1 The diagnostic workflow A clinical genome is a chain of custody, not just a FASTQ file: 1. Consent, phenotyping with HPO terms, and sample accessioning. 2. Sequencing, usually a gene panel, exome or genome, with quality gates on coverage. 3. Alignment, variant calling, and joint genotyping against a reference build. 4. Filtering by inheritance model, population frequency an