Global reference mapping of human transcription factor footprints
NATURE
Authors: Vierstra, Jeff; Lazar, John; Sandstrom, Richard; Halow, Jessica; Lee, Kristen; Bates, Daniel; Diegel, Morgan; Dunn, Douglas; Neri, Fidencio; Haugen, Eric; Rynes, Eric; Reynolds, Alex; Nelson, Jemma; Johnson, Audra; Frerker, Mark; Buckley, Michael; Kaul, Rajinder; Meuleman, Wouter; Stamatoyannopoulos, John A.
Abstract
A high-density DNase I cleavage map from 243 human cell and tissue types provides a genome-wide, nucleotide-resolution map of human transcription factor footprints. Combinatorial binding of transcription factors to regulatory DNA underpins gene regulation in all organisms. Genetic variation in regulatory regions has been connected with diseases and diverse phenotypic traits(1), but it remains challenging to distinguish variants that affect regulatory function(2). Genomic DNase I footprinting enables the quantitative, nucleotide-resolution delineation of sites of transcription factor occupancy within native chromatin(3-6). However, only a small fraction of such sites have been precisely resolved on the human genome sequence(6). Here, to enable comprehensive mapping of transcription factor footprints, we produced high-density DNase I cleavage maps from 243 human cell and tissue types and states and integrated these data to delineate about 4.5 million compact genomic elements that encode transcription factor occupancy at nucleotide resolution. We map the fine-scale structure within about 1.6 million DNase I-hypersensitive sites and show that the overwhelming majority are populated by well-spaced sites of single transcription factor-DNA interaction. Cell-context-dependentcis-regulation is chiefly executed by wholesale modulation of accessibility at regulatory DNA rather than by differential transcription factor occupancy within accessible elements. We also show that the enrichment of genetic variants associated with diseases or phenotypic traits in regulatory regions(1,7)is almost entirely attributable to variants within footprints, and that functional variants that affect transcription factor occupancy are nearly evenly partitioned between loss- and gain-of-function alleles. Unexpectedly, we find increased density of human genetic variation within transcription factor footprints, revealing an unappreciated driver ofcis-regulatory evolution. Our results provide a framework for both global and nucleotide-precision analyses of gene regulatory mechanisms and functional genetic variation.
Comprehensive functional annotation of susceptibility variants identifies genetic heterogeneity between lung adenocarcinoma and squamous cell carcinoma
FRONTIERS OF MEDICINE
Authors: Qin, Na; Li, Yuancheng; Wang, Cheng; Zhu, Meng; Dai, Juncheng; Hong, Tongtong; Albanes, Demetrius; Lam, Stephen; Tardon, Adonina; Chen, Chu; Goodman, Gary; Bojesen, Stig E.; Landi, Maria Teresa; Johansson, Mattias; Risch, Angela; Wichmann, H-Erich; Bickeboller, Heike; Rennert, Gadi; Arnold, Susanne; Brennan, Paul; Field, John K.; Shete, Sanjay; Le Marchand, Loic; Melander, Olle; Brunnstrom, Hans; Liu, Geoffrey; Hung, Rayjean J.; Andrew, Angeline; Kiemeney, Lambertus A.; Zienolddiny, Shan; Grankvist, Kjell; Johansson, Mikael; Caporaso, Neil; Woll, Penella; Lazarus, Philip; Schabath, Matthew B.; Aldrich, Melinda C.; Stevens, Victoria L.; Jin, Guangfu; Christiani, David C.; Hu, Zhibin; Amos, Christopher I.; Ma, Hongxia; Shen, Hongbing
Abstract
Although genome-wide association studies have identified more than eighty genetic variants associated with non-small cell lung cancer (NSCLC) risk, biological mechanisms of these variants remain largely unknown. By integrating a large-scale genotype data of 15 581 lung adenocarcinoma (AD) cases, 8350 squamous cell carcinoma (SqCC) cases, and 27 355 controls, as well as multiple transcriptome and epigenomic databases, we conducted histology-specific meta-analyses and functional annotations of both reported and novel susceptibility variants. We identified 3064 credible risk variants for NSCLC, which were overrepresented in enhancer-like and promoter-like histone modification peaks as well as DNase I hypersensitive sites. Transcription factor enrichment analysis revealed that USF1 was AD-specific while CREB1 was SqCC-specific. Functional annotation and gene-based analysis implicated 894 target genes, including 274 specifics for AD and 123 for SqCC, which were overrepresented in somatic driver genes (ER = 1.95,P= 0.005). Pathway enrichment analysis and Gene-Set Enrichment Analysis revealed that AD genes were primarily involved in immune-related pathways, while SqCC genes were homologous recombination deficiency related. Our results illustrate the molecular basis of both well-studied and new susceptibility loci of NSCLC, providing not only novel insights into the genetic heterogeneity between AD and SqCC but also a set of plausible gene targets for post-GWAS functional experiments.