Bioinformatics Notes
Lecture slides and notes for Population Genomics, Variant Analysis and Genome Wide Association Studies in Bioinformatics Notes by Md Ahbab. 14 pages.

CSE 4893: Introduction to Bioinformatics Module 9: Population Genomics, Variant Analysis and GWAS Abstract These subsidiary notes introduce the study of genetic variation across populations and the statistics that link that variation to traits. We move from the classes of variation carried by a single human genome, through variant calling with the GATK best practices workflow and the VCF format, to the population genetic ideas of Hardy Weinberg equilibrium, linkage dise- quilibrium and population structure. We then build a genome wide association study step by step: quality control, imputation, association testing, reading Manhattan and quantile quantile plots, and mixed model tools for large biobanks. Fine mapping, polygenic risk scores, equity and data governance close the module, with worked exercises, algorithms and a keyword glossary for revision. Contents 1 Learning Objectives 1 2 Genetic Variation 2 3 Variant Calling and the VCF Format 2 4 Population Genetics Foundations 3 4.1 Hardy Weinberg . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 4.2 Linkage Disequilibrium . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .
CSE 4893 Introduction to Bioinformatics Module 9 3. Apply Hardy Weinberg equilibrium, linkage disequilibrium and population structure measures to real genotype counts. 4. Design a genome wide association study, from quality control and imputation to association testing and interpretation. 5. Judge the promise and the limits of polygenic risk scores, and discuss the ethics of genomic data across ancestries. 2 Genetic Variation Two unrelated people differ at roughly one base in a thousand, yet that small fraction is enough to shape height, drug response and disease risk. A single genome sequenced at 30 × depth, as in the 1000 Genomes high coverage resource of 3,202 samples on GRCh38, yields about four to five million variant sites. Most are common and shared, and only a handful are private. Reference frequency resources such as gnomAD v4.1, which aggregates 730,947 exomes and 76,215 genomes, let us ask at once whether a variant is ordinary or genuinely rare. Table 1: Classes of human genetic variation and how each one is usually found. Class Size range Typical count per genome Detection method Single nucleotide variant 1 bp 4.1 to 5.0 million Short read pileup and local assembly, GAT
CSE 4893 Introduction to Bioinformatics Module 9 FASTQ reads BWA-MEM alignment Mark duplicates Base quality recalibration HaplotypeCaller GVCF mode GenomicsDB Import GenotypeGVCFs joint calling VQSR or hard filters Annotate with VEP or bcftools Analysis ready VCF Figure 1: The GATK best practices path for germline short variants. The grey block is data pre-processing, the tinted block is discovery and joint genotyping. Table 2: The eight fixed columns of a VCF record. Field Content Notes for the reader CHROM Contig name Must match the reference dictionary, for example chr20 on GRCh38 POS 1 based position For indels the position of the base before the event ID Variant identifier dbSNPrs number if known, otherwise a dot REF Reference allele Taken from the reference FASTA, never from the reads ALT Alternate alleles Comma separated when the site is multiallelic QUAL Phred confidence Confidence that a variant exists at this site, not that a genotype is right FILTER Filter status PASS, or the name of every filter the site failed INFO Site level annotations Semicolon separated key value pairs such as DP, AF, AC, AN Pitfall: filters change your answer quietly A depth cut, a GQ cut or a swi