Regulatory sequence

A regulatory sequence is a segment of a nucleic acid molecule which is capable of increasing or decreasing the expression of specific genes within an organism. Regulation of gene expression is an essential feature of all living organisms and viruses.

Description

Regulatory sequence

/silencer

5'UTR

3'UTR

/silencer

Core

Start

Stop

Terminator

Transcription

DNA

Exon

Intron

Post-transcriptional
modification

Pre-
mRNA

Protein coding region

5'cap

Poly-A tail

Translation

Mature
mRNA

Protein

The structure of a

3' untranslated regions (blue) regulate translation into the final protein product.^[1]

Polycistronic operon

Regulatory sequence

Enhancer

/silencer

Operator

Promoter

5'UTR

ORF

UTR

3'UTR

Start

Stop

Terminator

Transcription

DNA

RBS

Protein coding region

mRNA

Translation

Protein

The structure of a

transcription of the gene into an mRNA. The mRNA untranslated regions (blue) regulate translation into the final protein products.^[1]

In

repressors, or both. Repressors often act by preventing RNA polymerase from forming a productive complex with the transcriptional initiation region (promoter), while activators facilitate formation of a productive complex. Furthermore, DNA motifs have been shown to be predictive of epigenomic modifications, suggesting that transcription factors play a role in regulating the epigenome.^[2]

In

riboswitches

.

Activation and implementation

A regulatory DNA sequence does not regulate unless it is activated. Different regulatory sequences are activated and then implement their regulation by different mechanisms.

Enhancer activation and implementation

Expression of genes in mammals can be upregulated when signals are transmitted to the promoters associated with the genes.

transcription factor proteins have a leading role in the regulation of gene expression.^[5]

Enhancers are sequences of the genome that are major gene-regulatory elements. Enhancers control cell-type-specific gene expression programs, most often by looping through long distances to come in physical proximity with the promoters of their target genes.^[6] In a study of brain cortical neurons, 24,937 loops were found, bringing enhancers to promoters.^[3] Multiple enhancers, each often at tens or hundred of thousands of nucleotides distant from their target genes, loop to their target gene promoters and coordinate with each other to control expression of their common target gene.^[6]

The schematic illustration in this section shows an enhancer looping around to come into close physical proximity with the promoter of a target gene. The loop is stabilized by a dimer of a connector protein (e.g. dimer of CTCF or YY1), with one member of the dimer anchored to its binding motif on the enhancer and the other member anchored to its binding motif on the promoter (represented by the red zigzags in the illustration).^[7] Several cell function specific transcription factor proteins (in 2018 Lambert et al. indicated there were about 1,600 transcription factors in a human cell^[8]) generally bind to specific motifs on an enhancer^[9] and a small combination of these enhancer-bound transcription factors, when brought close to a promoter by a DNA loop, govern the level of transcription of the target gene. Mediator (coactivator) (a complex usually consisting of about 26 proteins in an interacting structure) communicates regulatory signals from enhancer DNA-bound transcription factors directly to the RNA polymerase II (RNAP II) enzyme bound to the promoter.^[10]

Enhancers, when active, are generally transcribed from both strands of DNA with RNA polymerases acting in two different directions, producing two eRNAs as illustrated in the Figure.^[11] An inactive enhancer may be bound by an inactive transcription factor. Phosphorylation of the transcription factor may activate it and that activated transcription factor may then activate the enhancer to which it is bound (see small red star representing phosphorylation of a transcription factor bound to an enhancer in the illustration).^[12] An activated enhancer begins transcription of its RNA before activating a promoter to initiate transcription of messenger RNA from its target gene.^[13]

CpG island methylation and demethylation

CpG sites). About 28 million CpG dinucleotides occur in the human genome.^[14] In most tissues of mammals, on average, 70% to 80% of CpG cytosines are methylated (forming 5-methyl-CpG, or 5-mCpG).^[15] Methylated cytosines within CpG sequences often occur in groups, called CpG islands. About 59% of promoter sequences have a CpG island while only about 6% of enhancer sequences have a CpG island.^[16] CpG islands constitute regulatory sequences, since if CpG islands are methylated in the promoter of a gene this can reduce or silence gene expression.^[17]

DNA methylation regulates gene expression through interaction with methyl binding domain (MBD) proteins, such as MeCP2, MBD1 and MBD2. These MBD proteins bind most strongly to highly methylated CpG islands.^[18] These MBD proteins have both a methyl-CpG-binding domain and a transcriptional repression domain.^[18] They bind to methylated DNA and guide or direct protein complexes with chromatin remodeling and/or histone modifying activity to methylated CpG islands. MBD proteins generally repress local chromatin by means such as catalyzing the introduction of repressive histone marks or creating an overall repressive chromatin environment through nucleosome remodeling and chromatin reorganization.^[18]

Transcription factors are proteins that bind to specific DNA sequences in order to regulate the expression of a given gene. The binding sequence for a transcription factor in DNA is usually about 10 or 11 nucleotides long. There are approximately 1,400 different transcription factors encoded in the human genome and they constitute about 6% of all human protein coding genes.^[19] About 94% of transcription factor binding sites that are associated with signal-responsive genes occur in enhancers while only about 6% of such sites occur in promoters.^[9]

EGR1 is a transcription factor important for regulation of methylation of CpG islands. An EGR1 transcription factor binding site is frequently located in enhancer or promoter sequences.^[20] There are about 12,000 binding sites for EGR1 in the mammalian genome and about half of EGR1 binding sites are located in promoters and half in enhancers.^[20] The binding of EGR1 to its target DNA binding site is insensitive to cytosine methylation in the DNA.^[20]

While only small amounts of EGR1 protein are detectable in cells that are un-stimulated, EGR1 translation into protein at one hour after stimulation is markedly elevated.^[21] Expression of EGR1 in various types of cells can be stimulated by growth factors, neurotransmitters, hormones, stress and injury.^[21] In the brain, when neurons are activated, EGR1 proteins are upregulated, and they bind to (recruit) pre-existing TET1 enzymes, which are highly expressed in neurons. TET enzymes can catalyze demethylation of 5-methylcytosine. When EGR1 transcription factors bring TET1 enzymes to EGR1 binding sites in promoters, the TET enzymes can demethylate the methylated CpG islands at those promoters. Upon demethylation, these promoters can then initiate transcription of their target genes. Hundreds of genes in neurons are differentially expressed after neuron activation through EGR1 recruitment of TET1 to methylated regulatory sequences in their promoters.^[20]

Activation by double- or single-strand breaks

About 600 regulatory sequences in promoters and about 800 regulatory sequences in enhancers appear to depend on double-strand breaks initiated by topoisomerase 2β (TOP2B) for activation.^[22]^[23] The induction of particular double-strand breaks is specific with respect to the inducing signal. When neurons are activated in vitro, just 22 TOP2B-induced double-strand breaks occur in their genomes.^[24] However, when contextual fear conditioning is carried out in a mouse, this conditioning causes hundreds of gene-associated DSBs in the medial prefrontal cortex and hippocampus, which are important for learning and memory.^[25]

Such TOP2B-induced double-strand breaks are accompanied by at least four enzymes of the non-homologous end joining (NHEJ) DNA repair pathway (DNA-PKcs, KU70, KU80 and DNA LIGASE IV) (see figure). These enzymes repair the double-strand breaks within about 15 minutes to 2 hours.^[24]^[26] The double-strand breaks in the promoter are thus associated with TOP2B and at least these four repair enzymes. These proteins are present simultaneously on a single promoter nucleosome (there are about 147 nucleotides in the DNA sequence wrapped around a single nucleosome) located near the transcription start site of their target gene.^[26]

The double-strand break introduced by TOP2B apparently frees the part of the promoter at an RNA polymerase–bound transcription start site to physically move to its associated enhancer. This allows the enhancer, with its bound transcription factors and mediator proteins, to directly interact with the RNA polymerase that had been paused at the transcription start site to start transcription.^[24]^[10]

Similarly, topoisomerase I (TOP1) enzymes appear to be located at many enhancers, and those enhancers become activated when TOP1 introduces a single-strand break.

RAD50 and ATR.^[27]

Examples

Genomes can be analyzed systematically to identify regulatory regions.[28] Conserved non-coding sequences often contain regulatory regions, and so they are often the subject of these analyses.

CAAT box
CCAAT box
Operator (biology)
Pribnow box
TATA box
SECIS element, mRNA
Polyadenylation signal, mRNA
A-box
Z-box
C-box
E-box
G-box

Insulin gene

Regulatory sequences for the

insulin gene are:^[29]

A5

Z

negative regulatory element (NRE)^[30]

C2

E2

A3

cAMP response element

A2

CAAT enhancer binding

References

^
ISSN 2002-4436
.

^ Whitaker JW, Zhao Chen, Wei Wang. (2014) Predicting the Human Epigenome from DNA Motifs. Nature Methods. doi:10.1038/nmeth.3065
^
PMID 32451484
.

PMID 33102493
.

S2CID 205485256
.

^
S2CID 152283312
.

PMID 29224777
.

PMID 29425488
.

^
PMID 29987030
.

^
PMID 25693131
.

PMID 29378788
.

PMID 12514134
.

PMID 32810208
.

PMID 26932361
.

PMID 15177689
.

PMID 32338759
.

PMID 11782440
.

^
PMID 25927341
.

S2CID 3207586
.

^
PMID 31467272
.

^
PMID 19374776
.

S2CID 159041612
.

PMID 32029477
.

^
PMID 26052046
.

PMID 34197463
.

^
S2CID 206508330
.

^
PMID 25619691
.

PMID 15699025
.

PMID 11914736
.

PMID 17150186
.

External links

ORegAnno - Open Regulatory Annotation Database

ReMap - database of transcriptional regulators

Regulatory sequences
General

CAAT box

CCAAT box

Pribnow box

TATA box

SECIS element

A-box

Z-box

C-box

E-box

G-box

Insulin gene

A5

Z

negative regulatory element

C2

E2

A3

cAMP response element

A2

CAAT enhancer binding
(CEB)

C1

E1

ILPR

Retrieved from "https://en.wikipedia.org/w/index.php?title=Regulatory_sequence&oldid=1188122190"

[ShafeeLowe2017-1] 
ISSN 2002-4436
.

[2] Whitaker JW, Zhao Chen, Wei Wang. (2014) Predicting the Human Epigenome from DNA Motifs. Nature Methods. doi:10.1038/nmeth.3065

[Beagan-3] 
PMID 32451484
.

[pmid33102493-4] PMID 33102493
.

[pmid22868264-5] S2CID 205485256
.

[Schoenfelder-6] 
S2CID 152283312
.

[pmid29224777-7] PMID 29224777
.

[pmid29425488-8] PMID 29425488
.

[pmid29987030-9] 
PMID 29987030
.

[Allen2015-10] 
PMID 25693131
.

[pmid29378788-11] PMID 29378788
.

[pmid12514134-12] PMID 12514134
.

[pmid32810208-13] PMID 32810208
.

[pmid26932361-14] PMID 26932361
.

[pmid15177689-15] PMID 15177689
.

[pmid32338759-16] PMID 32338759
.

[pmid11782440-17] PMID 11782440
.

[Du-18] 
PMID 25927341
.

[pmid19274049-19] S2CID 3207586
.

[SunZ-20] 
PMID 31467272
.

[Kubosaki-21] 
PMID 19374776
.

[pmid31110352-22] S2CID 159041612
.

[pmid32029477-23] PMID 32029477
.

[Madabhushi-24] 
PMID 26052046
.

[25] PMID 34197463
.

[Ju-26] 
S2CID 206508330
.

[Puc-27] 
PMID 25619691
.

[Stepanova_2005-28] PMID 15699025
.

[Melloul_2002-29] PMID 11914736
.

[30] PMID 17150186
.

[1]

[2]

[5]

[6]

[3]

[7]

[8]

[9]

[10]

[11]

[12]

[13]

[14]

[15]

[16]

[17]

[18]

[19]

[20]

[21]

[22]

[23]

[24]

[25]

[26]

[27]

[29]

Description

Activation and implementation

Enhancer activation and implementation

CpG island methylation and demethylation

Activation by double- or single-strand breaks

Examples

Insulin gene

See also

References

External links