CBG.

02: Genomes ON THE IMMORALITY OF TELEVISION SETS:
“FUNCTION” IN THE HUMAN GENOME
ACCORDING TO THE EVOLUTION-FREE GOSPEL OF
ENCODE PROJECT WRITE EULOGY FOR JUNK ENCODE (Dan Graur)
DNA
- less than 10% of the genome is evolutionarily
conserved through purifying selection
- human DNA has 3 billion bases - according to ENCODE, a biological function can be
- instead of the expected 100,000 genes, the initial maintained indefinitely without selection, which implies that at
analysis found about 35,000 and that number has since been
least 80-10=70% of the genome is perfectly invulnerable to
whittled down to about 21,000
deleterious mutations, either because no mutation can ever
- 80% of the human genome serves some purpose,
occur in these “functional” regions or because no mutation in
biochemically speaking: these regions can be deleterious
- specify landing spots for proteins that influence - problems in ENCODE logic:
gene activity o seldom used “causal role” definition of
- strands of RNA with myriad roles biological function and then applying it inconsistently to
- places where chemical modifications serve to
different biochemical properties
silence stretches of our chromosomes
o logical fallacy “affirming the consequent”
- a gene’s regulation is far more complex than
o failing to appreciate the crucial difference
previously thought, being influenced by multiple stretches of between “junk DNA” and “garbage DNA”
regulatory DNA located both near and far from the gene o using analytical methods that yield biased
itself and by strands of RNA not translated into proteins errors and inflate estimates of functionality
(=noncoding RNA) o favouring statistical sensitivity over specificity
- 11,224 DNA stretches are classified as pseudogenes, o emphasizing statistical significance rather than
“dead” genes now known to be active in some cell types or
the magnitude of the effect
individuals
- in biology, there are 2 main concepts of function:
- there are many “genes” out there in which DNA o “selected effect”: function of a trait is the
codes for RNA, not a protein, as the end product effect for which it was selected, or by which it
- various cell genes home in on different cell is maintained
compartments, as if they have fixed addresses where they o “causal role”: historical and non-evolutionary:
operate: some go to the nucleus, some to the nucleolus and for a trait Q to have a “causal role” function, G,
some to the cytoplasm
it is necessary and sufficient that Q performs G


GINGERAS: the fundamental unit of the genome and the
(ex:) TATAAA – maintained by natural selection to bind a
basic unit of heredity should be the transcript – the piece of transcription factor; a mutated sequence, resembling this one,
RNA decoded from DNA- and not the gene also binds the transcription factor, but does not result in
transcription (no adaptive or maladaptive consequence); hence,
- 5% of the human genome is conserved across the second sequence has no selected effect function, but its
mammals
causal role function is to bind a transcription factor
- DNA’s bases function in gene regulation through

their interactions with transcription factors and other
- from an evolutionary viewpoint, a function can be
proteins; assigned to a DNA sequence if and only if it is possible to
- 8% of the genome falls within a transcription factor destroy it; unless a genomic functionality is actively protected by
binding site, a percentage that is expected to double once selection, it will accumulate deleterious mutations and will cease
more transcription factors have been tested to be functional
- the fact that sometimes it is difficult to identify
selection should never be used as a justification to ignore
selection altogether in assigning functionality to parts of the
human genome
- the surest indicator of the existence of a genomic
function is that losing it has some phenotypic consequence for
the organism
- functional regions of the genome should evolve more
slowly and be more conserved among species than non-
functional ones
- Ward and Kellis confirmed that approx.. 5% of the
genome is interspecifically conserved and an additional 4% of
the human genome in under selection
- According to ENCODE:
o 74.7% of the genome is transcribed
o 56.1% is associated with modified histones

” Where they against excess genome is extremely efficient due to the meet. that are recognised by a group of transcription factors called coding genes GATA2. pseudogenes (up to 1/10 transcribed. ENCODE is working form the genome out. .517 protein. mammalian conservation suggests that aprox. suggesting that some loss of constraint 60 percent more likely to lie within functional. Are there hotspots? Are there SNPs that .” says Birney. the diseases that people have done GWAS studies for. transcribed.2% is found in open-chromatin areas regions. .” says Birney. lack coding went: Yes!” potential due to the presence of disruptive mutations. classes of sequences that are known to be abundantly one of those too good to be true moments. hence. suggesting that they may of disease. misconceptions in common objections to “junk DNA” biologists had on their radar. less than 2% of the histone modifications may have . suggesting recent loss in function and activity. BRENNER: differentiated between “junk DNA” and “garbage DNA”. biochemically active. 2003. and many of them are new. however. and the fact that such hotspots that are worth looking into. So far. and . as well as being useless. This suggests o 8. and because the IN HUMANS FOR RECENTLY ACQUIRED molecular process generating extra DNA outpaces those getting REGULATORTY FUNCTIONS rid of it. Others are head-scratchers. Imagine a massive table. these also show higher primate divergence relative compared to random SNPs. human constraint correlates with mammalian conservation. o 15. it’s a new lead to follow up o the belief that evolution can always can rid of on. a type of bowel disorder. conserved across mammals. non-coding predates human-macaque divergence . nematode Caenorhabditis elegans has 20.” says Birney. Of these. They found that just 12 percent of the SNPs lie regions. but more than 80% is transcribed. I was in the room [when they got the result] and I . the team have identified 400 enormous effective population sizes. generation time are correlated with 50 and 100 were predictable. a substantial factor of human constraint lies outside mammalian-conserved regions . 5% of “PURPOSE IS THE ONLY THING EVOLUTION CANNOT PROVIDE” the human genome is conserved due to noncoding and regulatory roles.6% consists of methylated CpG dinucleotides different genes. introns (some human introns harbour regulatory the top are all the possible cell types and transcription sequences TISHKOFF. Across .” In other words. 2006 as well as sequences that produce factors (proteins that control how genes are activated) in the small RNA molecules (HIROSE. evolve . attempting to understand the genetic basis constraint than ancestral repeats. indifferent DNA refers to DNA sites that are functional. geneticists have run a . mRNA splice sites and regulatory elements. within protein-coding areas. Take Crohn’s disease. and provides many fresh leads for . “Suddenly we’ve o a lack of knowledge of he original and correct made an unbiased association between a disease and a piece sense of the term of basic biology. Down the left side are all very rapidly and are mostly subject to no functional constraint). For the last decade. selection GWAS studies are working from disease in. although only 5% of the human genome is . Lots. transcription is fundamentally a stochastic process understanding how they affect our risk of disease. The ENCODE team have mapped all of these to activity show reduced human constraint relative to active their data. but are typically devoid of function: “Literally. The something to do with function team found five SNPs that increase the risk of Crohn’s.“We’re now working with lots of different disease o the belief that “future potential” constitutes “a biologists looking at their data sets. there is interest. is deleterious. while . Some of the rest make intuitive genome size sense. regions that do not overlap with active ENCODE seemingly endless stream of “genome-wide association elements and inactive chromatin states show lower studies” (GWAS). mobile elements correspond to both? Yes. “That wasn’t something that the Crohn’s disease . They have thrown up a long list of SNPs – variants provide a more accurate neutral reference than repeats that at specific DNA letters—that correlate with the risk of can have exapted functions different conditions. and is subject to additional elements evolve neutrally or confer a lineage- purifying selection specific fitness advantage . 2004)) ENCODE study. mammalian conserved regions lacking ENCODE . the excess DNA in our genome is junk and it is there EVIDENCE OF ABUNDANT PURIFYING SELECTION because it is harmless. between replication time and. in the majority of known bacterial species. ENCODE: THE ROUGH GUIDE TO THE HUMAN similar selective pressures act in humans and across GENOME mammals . They also showed that . raising the question of whether the deletion of these sites. especially in promoters and enhancers. “It was . . or associated with chromatin states suggestive of regulatory functions ENCODE HYPE . the nunfunctional DNA . a substantially larger portion is but show no evidence of selection against point mutations.5% binds transcription factors that many of these variants are controlling the activity of o 4. bound by A SLIGHTLY DIFFERENT RESPONSE TO TODAY’S a regulator. ZHOU. the disease-associated ones are to active regions. “In some function” sense.

as proposed during mammalian radiation . nucleosomes protect the delicate strands (=gene deserts) can be well tolerated by any organism from physical damage . . the homozygous deletion mice for both deletions were viable . in MMU3 desert.243 human-mouse conserved non-coding elements . rate 1:2:1 (wild-type : heterozygous : mutant homozygous) . the job of the nucleosome is paradoxical. phenotypic parameters measures in the homozygous deletion mice. encoded in 3 billion bp of DNA . nearly RESULT IN VIABLE MICE all of our cells) contain a copy of this genome. on the other hand. a collection of repar enzymes corrects mammalian genomes not corresponding to protein coding chemical changes inflicted on the strands by environmental sequences remains largely undetermined insults . this is usually related to the removal of genes with redundancy elsewhere in the genome . genome-wide association studies suggest that 85% of disease-associated variants are noncoding. nucleosome must be stable. compared with controls: o post-natal survival rates for 25 weeks o measurable growth retardation o clinical chemistry tests (general and specific plasma parameters) o morphological abnormalities o abnormal growth o tissue degeneration o organ mass was similar in both groups of deletion mice and their wild-type littermates . the method by which nucleosomes solve these opposed needs is not well understood. the boundaries for the deletions permitted proximate regulatory elements nearby the flanking genes to remain intact . each of our cells (or more correctly. a fraction similar to the proportion of human constraint that we estimate lies outside protein-coding regions. almost half of human constraint lies outside NUCLEOSOME mammalian-conserved regions. whereas regulatory constraint is primarily lineage-specific. some large-scale deletions of the non-coding DNA . although gene inactivation can sometimes fail to result in detectable phenotype. even though the strength of human constraint is higher in conserved elements . both to transcribe mRNA for building new proteins and to replicate the DNA when the cell divides. the deletions weren’t lethal in embryons because of approx. reduced in the brain sheltering structures that compact the DNA and keep it from . forming tight. the heterozygous mice appeared phenotypically normal . nucleosome must be labile enough promoter to allow the information in the DNA to be used. protein-coding constraint occurs primarily in conserved regions. quantitative assays revealed detectable alterations in levels requiring it to perform 2 opposite functions simultaneously: of expression: Prkacb reduced in the heart and Rpp30 on one hand. polymerases must be allowed access to the DNA. the functional importance of the roughly 98% of the . deletion of a gene desert mapping to mouse chromosome 3 and in chromosome 19 (with no evidence of transcription) à contain 1. this suggests that mutations outside conserved elements play important roles in both human evolution and disease MEGABASE DELETIONS OF GENE DESERTS . but may involve a partial unfolding of the DNA from around the . molecular level impact: only 2 out of the 108 . beta-galactosidase expression harm.

ment length. with their subtly . of the chromosomes. largest known human gene is dystrophin 2. as the information in the of “foreign” DNA can be inserted into them. for instance. the histone proteins are perfectly designed for their jobs. HUMAN GENOME PROJECT (TO KNOW OURSELEVES) genetic linkage map: based on careful analyses of human inheritance patterns.000 bp restriction fragments. or YAC. in this way. the human genome is not so very different from physical maps: distances between features are measured not that of chimpanzees or mice. chromosome 1 has the most genes (3. the average gene is 3. the tails extend outward from the compact of human DNA. reaching out to neighbouring nucleosomes and reinserted into a yeast cell. with distances measured in centimorgans (=measure of recombination frequency) à the closer 2 genes are on a single chromosome the less likely they are to get split up during genetic recombination à when they are close enough that the chances of being separated are only 1/100 they are said to be separated by a distance of 1 centimorgan .some is then nucleosome.” The classic example is the cloning vector. .some. when digested with a particular restriction enzyme. DNA from the same . the cell makes containing a copy. some regions of the genome resist cloning in YACs and others are prone to rearrangements (PRIMER ON MOLECULAR GENETICS) . can yield dissimilar sets of functions. restriction enzyme: cleave dsDNA molecules at . whereby the DNA is read inserted DNA is replicated along with the rest of the vector as . the surface of the octamer is decorated with positively charged AA. or from bacteriophages (viruslike parasites of bacteria). often by a single base . but in “real” physical units (base pairs) many common elements with the genome of lowly fruit fly . so much so that histones are nearly identical in all non-bacterial organisms.mosomes constructed from yeast or bacterial genomic DNA. of the same fragment of human particular genes more accessible to polymerases. allowing DNA their particular information to be copied and used to build new proteins . that interact strongly with the negatively-charged phosphate groups of the DNA. even slight modifications can be lethal.cules derived from bacteria genome where individuals differ in their DNA sequence. single nucleotide polymorphism (SNP) are sites in a which may be circular DNA mole. the histone proteins. is constructed by assembling the essential genes that they store: each nucleosome is composed of 8 functional parts of a nat.4 mil. as if it were part of the yeast’s normal complement regulatory enzymes that chemically modifies these tails to of chromosomes. sequences that mark the ends encircled by 2 loops of DNA. which comprises nearly a quarter of their .a third necessary tool is some means of DNA chromosome has the fewest (344) “amplification. this serves to glue the DNA strand to the protein core. The characteristic all these vectors share is that fragments . Bp genomic region of 2 different people. which reproduces the YAC during binding them tightly together. each weaken their interactions. The result is a colony of yeast cells. however. A yeast artificial chromo. and sequences required for chromosome are not completely globular like most other proteins à they separa- have long tails. to 250 million bp then.168) and Y . one loop at a time. usually 4 or 6 nucleotides long molecule of DNA that ranges in length from about 50 million . each chromosome is a physically separated specific recognition sites. identical segments of human DNA yield identical sets of . on the other hand.ural yeast chromosome—DNA “histone” proteins bundled tightly together at the centre.nucleosome. scientists believe that human genome has at least 10 mil SNPs . nucleosomes also modify the activity of the the host reproduces itself. or clone.tion during cell division—then splicing in a frag. it indicates for each chromosome the whereabouts of genes or other “heritable markers”. or artificial chro. which then produce different patterns when dynamics sorted according to size . sequences that initiate replication. but they shed light on chromosome structure and fragments. and it even shares in genetic terms. repeat sequences are thought to have no direct different genomic sequences. This engineered chromo. . the nucleus contains cell division.

Sign up to vote on this title
UsefulNot useful