PMID 6158093 — Structure and in vitro transcription of human globin genes.
good_results R=5256w / 36¶ | figs=31 Alex
TITLE
[1] 9w Structure and in vitro Transcription of Human Globin Genes
ABSTRACT
[1] 264w The recent application of molecular cloning procedures to the human globin gene family has led to significant ad- vances in the understanding of the struc- ture, chromosomal arrangement, and evolution of these genes [see references in (1-4)]. Individual members of this rel-1-globin genes (a-and 18-thalassemia) or the switch from fetal to adult globin gene expression. Thus, human globin genes provide a model system for studying the molecular genetics of eukaryotic gene regulation and the molecular basis of hu- man genetic disease.Summary. The a-like and 13-like subunits of human hemoglobin are encoded by a small family of genes that are differentially expressed during development. Through the use of molecular cloning procedures, each member of this gene family has been isolated and extensively characterized. Although the a-like and 13-like globin genes are located on different chromosomes, both sets of genes are arranged in closely linked clusters. In both clusters, each of the genes is transcribed from the same DNA strand, and the genes are arranged in the order of their expressions during development. Structural comparisons of immediately adjacent genes within each cluster have provided evidence for the occurrence of gene duplication and correction during evolution and have led to the discovery of pseudogenes, genes that have acquired numerous mutations that prevent their normal expression. Recently, in vivo and in vitro systems for studying the expression of cloned eukaryotic genes have been developed as a means of identifying DNA sequences that are necessary for normal gene func- tion. This article describes the application of an in vitro transcription procedure to the study of human globin gene expression.
RESULTS
[1] 55w The known a-like and 18-like globins are listed in Table 1. The earliest embryonic hemoglobin tetramer, Hb(Gower 1), consists of e (fl-like) and C (a-like) polypeptide chains. Beginning at approximately 8 weeks of gestation, the embryonic chains are gradually replaced by the adult a-globin chain and two different fetal 1-like chains, designated Gy and -y.
[2] 140w The y chains differ only in the presence of glycine or alanine, respectively, at position 136. During the transition period between embryonic and fetal development, Hb(Gower 2) (a2E2) and Hb(Portland) (;2Va) are detected. HbF (a2y2) eventually becomes the predominant Hb tetramer throughout the remainder of fetal life. Beginning just before birth, the y-globin chains are gradually replaced by the adult ,8-globin and & globin polypeptides. At 6 months after birth, 97 to 98 percent of the hemoglobin is HbA (a2A2), while HbA2 (a282) ac- counts for approximately 2 percent. Small amounts of HbF (1 percent) are al- so found in adult peripheral blood. The site of erythropoiesis changes from the yolk sac in the early embryo, to the developing liver, spleen, and bone marrow in the fetus, and finally to the bone marrow in adults [see references in (3, 4)].
[3] 58w In summary, the a-like and ,8-like globin gene families have coordinated programs for differential gene expression. The primary difference between the two gene families is that two switches in gene expression (embryonic to fetal to adult) are observed for the 18-like genes, while a single switch results in the turnoffof em- bryonic C-globin production early in fetal life.
[4] 199w The entire ,8-globin (5) and a-globin (6) gene clusters have been isolated in sets of overlapping bacteriophage recombi- nants, which were obtained from librar- ies of random, high-molecular-weight, human DNA (7,8). The linkage arrangement of the human a-like and ,8-like globin gene clusters, which was established by genomic blotting (9, 10) and molecu- lar cloning (5,6,8,11,12) experiments, is shown in Fig. 1. Although the size of the two gene clusters differs by almost a factor of 2, the genes in both clusters are arranged on the chromosome in the or- der of their expression during development. A similar pattern of gene organization has been found in the rabbit (13,14) 0036-8075/8010919-1329$02.00/0 Copyright X) 1980 AAAS The linkage arrangement of human (3-like and a-like globin genes. The top line shows the relative locations of the five functional (3-like globin genes and the two (-like sequences, which do not correspond to any known (-globin polypeptide chain. The bottom line shows the map of the four functional a-like globin genes and the a-globin pseudogene al1. The mRNA coding and intervening sequences are designated by filled and open rectangles, respectively. The direction of transcription of all the globin genes is indicated by the arrow.
[5] 118w and mouse (15) (3-like globin gene clusters. All of the known human globin polypeptides can be accounted for by the genes shown in Fig. 1. However, both gene clusters contain additional sequences that are detected by globin gene hybridization probes but cannot be identified with any of the known globin poly- peptides. Three genes, designated t,3l, %P,82 (5), and fal (6, 16), fall into this cat- egory. Structural analysis of the ial gene indicates that it is a pseudogene (a gene that displays significant homology to a functional gene but has mutations that prevent its expression) (16). Similar pseudogenes have been identified at cor- responding positions in the rabbit (13, 17) and mouse (15) (3-like globin gene clusters.
[6] 258w Globin gene fine structure. Since the discovery of an intervening sequencc (intron) in the rabbit (3-globin (18) and the mouse a-globin ( 19) and (3-globin (20) genes, two introns have been identified in all of the functional globin genes thus far studied [see references in (2,21)]. In particular, the five qwressed human (- like globin genes amc jerrupted by two introns at identica Jecations: the first, 122 to 130 base pairs (bp) in length, is lo- cated between codons 30 and 31; and the second, 850 to 900 bp, is between codons 104 and 105 (2) (Fig. 2). Similarly, the lo- cations of introns in the human (22) and mouse (23) a-globin genes are identical and are analogous to the positions of in- trons in (3-like globin genes (Fig. 2). In the case of the mouse (24), rabbit (13, 25), and human (26,27) ,8-globin genes, both the messenger RNA (mRNA) cod- ing sequences (exons) and intron se- quences are transcribed to produce a detectable nuclear mRNA precursor which is processed or spliced to give a mature globin mRNA. Furthermore, the 5' ends of the mouse and rabbit ,(-globin nuclear precursors are coterminal with mature mRNA (25, 28). These nuclear precursors could be the primary globin mRNA transcript. In this case, the sequence encoding the 5' end of the mature mRNA would correspond to the tran- scriptional initiation site. Alternatively, these nuclear transcripts may represent intermediates in the mRNA maturation. In this case, the transcriptional initiation site would be located proximal to the 5' end of the mRNA sequence.
[7] 156w The complete nucleotide sequences of the five human (8-like globin genes have been determined (29,30), and a detailed comparison of these and other mamma- lian globin gene sequences has been pre- sented (2). This sequence comparison re- vealed interesting sequence homologies in regions that are potentially involved in globin gene transcription and splicing. In particular, alignment of the sequences on the 5' sides of the human (8-like globin genes revealed two blocks of sequence homology, which are present in analo- gous positions adjacent to most eu- karyotic genes (31). The first homology block is an AT-rich (A, adenine; T, thymine) sequence originally identified in the Drosophila histone gene cluster called the Hogness box (32). A com- parison of a number of different (3-like globin genes has revealed that the AT- rich sequence CATAAA (C, cytosine) is found 31 + 1 bp on the 5' side of the mRNA capping sites, but that the se- aS23 HbA
[8] 177w C2y2 (Portland) 1330 quence shared by all of the (8-like genes is ATA. This sequence was therefore designated the ATA box (2). The second homology block (designated the CCAAT box) is located 77 ± 10 bp on the 5' side of each gene. A possible role of these se- quences in transcriptional initiation, RNA processing, or both, has been dis- cussed (2). Previous comparisons of noncoding sequence in human, mouse, and rabbit ,8like globin genes indicated that these re- gions diverged by deletion and addition as well as by simple base substitution (2,33,34). Examination of the nucleotide sequences surrounding putative deletion sites suggests that short (two to eight nu- cleotide sequences) and direct repeats are involved in the generation of dele- tions. This pattern is remarkably similar to that observed for preferred deletion sites (hot spots) in the lac i gene ofEsch- erichia coli (35). A model for the in- volvement of short, direct repeat se- quences in the generation of deletions in the noncoding regions of (8-like globin genes during evolution has been pro- posed (2).
[9] 177w A common feature of globin gene clus- ters is the occurrence of two immediately adjacent genes, which are coordinately expressed during a given develop- mental stage. Examples of this are the human 8-I3, G>y A^y, al-a2, and 1-P2 glo- bin gene pairs. The 8 and ( genes are highly homologous in the coding regions, but the noncoding sequences within and surrounding the two genes have diverged considerably (2). Extensive divergence of noncoding regions has also been ob- served in some other closely linked, coordinately expressed globin gene pairs (13,15,33). In contrast, the two members of the Gy_A-y gene pair are virtually identical to one another throughout their coding, intervening, 4pd flanking se- quences (30). Although the nucleotide sequences of linked human a-globin genes have not yet been determined, re- striction mapping and heteroduplex anal- ysis of the al and a2 genes indicate that the sequences within and flanking these two genes are virtually identical. Each a- globin gene is located within an approximately 4-kbp (kilobase pair) region of homology interrupted by two small regions of nonhomology (6).
[10] 155w The extensive sequence homology within and flanking the (jyA-y and al-a2 gene pairs appears to be the product of a mechanism for gene matching during evolution (6, 30, 36). Based on the nearly SCIENCE, VOL. 209 identical distribution of restriction sites surrounding the a-globin genes in a num- ber of primate species, it has been sug- gested that the a-globin gene duplication occurred before the time of primate di- vergence (36). Differences between the a-globin amino acid sequences of various primate species are consistent with sequence drift following primate diver- gence (37). However, intraspecies com- parisons show much less divergence, in- dicating that the a-globin genes within a species have been corrected against one another. Maintenance of homology among a family of evolving genes within a species has been termed "concerted" evolution (36). Gene conversion and ex- pansion-contraction of gene number by homologous but unequal crossing-over have been proposed as mechanisms for concerted evolution (38).
[11] 285w The precise end points of the a-globin gene duplication unit have been located by nucleotide sequence analysis (16). The left end point of the duplication is located immediately adjacent to the putative poly(A) (polyadenylate) addi- tion site of 4al, while the right end point is found 15 bp on the 3' side of the poly(A) addition site of al. The 15-bp se- quence on the 3' side of the poly(A) addi- tion sites of pal, a2, and al consists of a repeated pentanucleotide (GCCTG) (G, guanine), separated by TGTGT. The occurrence of this sequence in all three genes and its location with respect to the end points of the a-globin gene duplication suggest that this sequence might be associated with the mechanism by which the genes were duplicated or cor- rected (16). Zimmer et al. (36) and Lauer et al. ( 6) have proposed a model for a- gene correction that involves inter- chromosomal, unequal crossing-over events. Proudfoot and Maniatis (16) have suggested that the pentanucleotide re- peat acts as a boundary or terminator for the recombination event. Evidence that a-globin gene sequence matching could occur by expansion and contraction of gene number by unequal crossing-over is provided by the frequent occurrence of chromosomes containing one (39) or three (40) adult a-globin genes in some human populations (6,36). The chromo- some containing only one a-globin gene is associated with the common form of a- thalassemia, designated a-thalassemia 2. Comparison of the end points of the dele- tion associated with this disorder, with the location of blocks of homologous se- quence within the al-a2 gene duplication, strongly suggests that the dele- tion results from unequal crossing-over between homologous sequences (6). In- terestingly, deletions that are in-
[12] 136w 19 SEPTEMBER 1980 0 200 400 600 800 1000 1200 a _ 31 32 99 100 141 a I _ Fig. 2. The fine structure of a-an genes. The canonical structures for a-like and 3-like globin genes are dr proximate scale. Solid and open bc sent coding (exon) and noncoding 4 quences, respectively. The a-like gl contain introns of approximately c bp, located between codons 31 and and 100, respectively. The ,8-like gl contain introns of approximately and 850 to 900 bp, located between and 31, and 104 and 105, respectiv distinguishable from those foi thalassemia 2 occur in the cle cluster during propagation in E Analysis of the complete r sequences of cloned Giy and Ay led to formulation of a spec chromosomal gene conversion explain sequence matching linked genes (30). The Gy and A)
[13] 63w one chromosome are identical gion on the 5' side of the cen large intron, yet show greater d on the 3' side of that position. tion of the boundary between served and divergent regions r block of "simple sequence" D (TG) (30). Slightom et al. (30) posed that this simple sequenc4 ferred site for initiation of recoi events that lead to unidirecti4 conversion.
[14] 138w The a-and 1&Globin Pseudogen As was mentioned above, gl sequences that cannot be ident known globin polypeptide ch; been detected in several mamm cies (5,6,13,(15)(16)(17)41). Nucl quence analysis of a rabbit ,3 ps (132) (17), a human a pseudog Fig. 1) ( 16), a mouse 18 pseudol and a mouse a pseudogene (41) onstrated a variety of struc ferences between each gene an tional counterpart. The human quences (4,31 and tp,82, Fig. 1, not yet been extensively chara Each pseudogene that has 1 lyzed exhibits 75 to 80 percent homology when compared with sponding normal gene (excludi side of the mouse pseudogei is not homologous to the adult gene) (15). None of these pse can encode a functional glo peptide due to the presence of letions or insertions that result 1400 1600 tions of the translational reading frame.
[15] 86w In many cases, these frameshifts lead to the presence of in-phase termination co- 105 146 dons. In addition, one or more of the in- tron-exon junctions of rabbit *,32, mouse id,-globin 13H3, and upal are different from the se- the human quence common to splicing junctions in rawn to apglobin genes and all expressed genes [)xes repre-(intron) se-studied to date (42). Thus, even if these lobin genes pseudogenes are transcribed, it is unlike-95 and 125 ly that they would normally produce an 132, and 99 mRNA.
[16] 434w It is interesting to note that in all of the codons 30 mammalian globin gene clusters thus far rely. characterized, a pseudogene is found between the embryonic (or fetal) genes and the adult genes. It is possible that und in a-pseudogenes have some as yet uniden- ned gene tified function in globin gene clusters. Al- coli (6). ternatively, pseudogenes may be the iucleotide products of gene duplication and sub- genes has sequent sequence divergence (16, 17). ific intra-The variation in human a-globin gene model to number observed in present-day popubetween lations and the location of qial within the y genes on a-like globin gene cluster are consistent in the re-with the latter possibility. As shown in lter of the Fig. 2, 4sal, a2, and al are separated livergence from each other by approximately 4 kbp, Examina-which is the size of the al-a2 duplication the con-unit noted above. The nucleotide se-*evealed a quence of ial indicates that it is a-like INA polyrather than c-like (16). It therefore seems have propossible that 4al was once part of a set e is a pre-of three functional a-globin genes. mbination A novel mouse a-globin pseudogene onal gene has recently been described (41). As in the case of the pseudogenes described above, the mouse a-globin pseudogene has frameshift mutations that would re. es sult in premature translatipnal termi- nation. However, unlike the other lobin gene pseudogenes, the mouse gene is missing tified with both introns. The mechanism by which ains have this pseudogene arose and its location ialian spe-with respect to the normal a-globin gene [eotide se-cluster is unknown. eudogene en.e (vial, gene (15), Repetitive Sequences in Globin has dem-Gene Clusters tural difd its func-Cross-hybridization experiments be- 13-like setween the intragenic sequences of the (5) have human ,8-gene clusters (5) and a-gene Lcterized. clusters (43,47) revealed a nonglobin re- been ana-peat sequence that is interspersed within sequence the globin gene clusters and also repeatits corre-ed many times in the human genome. ing the 5' Nucleotide sequence analysis of the re- ne, which petitive sequences within the /8-globin mouse 18 gene cluster (44) indicates that they are udogenes members of a particular repeat sequence bin poly-family, the Alu family, which is reiter- small de-ated approximately 300,000 times in the in altera-human genome (45). In addition, the re- peats are transcribed in vitro by RNA polymerase III (46,47). They show se- quence homology with an abundant class of small nuclear RNA's (44,48) and with double-stranded, heterogeneous, nuclear RNA (5, 44, 47, 49). Finally, the repeats contain a sequence that is homologous to a sequence found near the replication A 1450 (39
[17] 29w origin of SV40, polyoma, and BK DNA tumor viruses (44). A similar set of repet- itive sequences has been identified with a cluster of rabbit ,8-like globin genes (50).
[18] 559w At present, there is little information regarding the expression or function of these interesting repetitive elements in vivo. 3. In vitro transcription of human /-like and a-like globin genes. The human globin genes were originally isolated as lambda recombinants and subsequently were subcloned into pBR322 as follows: A, Pst 4.4 kbp; 8, Pst 2.3 kbp; Gy, Pst 4.0 kbp; E, Bam 0.7 kbp; al, Sst 4.3 kbp; *al, Barn-Hind III 7.3 kbp; C1, RI (linker)-Bam 4.7 kbp (5-7, 12). Subcloned DNA's were digested with various restriction enzymes chosen to cleave the gene sequence at a specific position. In some cases, these digests were purified by phenol extraction and ethanol precipitation and then added directly to in vitro transcription reactions. Alternatively, the globin genes containing DNA fragments were isolated by horizontal agarose gel electrophoresis followed by electroelution into 10 mM tris-Cl,pH 8.0, 0.1M NaCl. The eluant was passed over a 0.5-ml DEAE- Sephadex column (equilibrated with 0. IM tris-Cl, pH 8.0, 0. IM KCI). After extensive washing with equilibration buffer, the column was eluted with O. iM KCI (0.3 ml), then 0.4M KCI (0.3 ml), and finally 0.6M KCI (0.6 ml). The 0.6M KCI wash was collected in 50-Al amounts which were assayed for DNA content by agarose gel electrophoresis. The DNA obtained was precipitated twice with ethanol and washed with 70 percent ethanol. Whole cell in vitro transcription extracts were prepared according to Manley et al. (54). In vitro transcription reactions were as described (54) with some modifications. Plasmid DNA digests (50 Lg) were used in reactions but only 1 to 5 Ag of purified DNA fragments. A ribonuclease inhibitor, ribonucleoside-vanadyl complex (73), was added at 10 mM to the deoxyribonuclease step after the in vitro transcription reaction. The RNA purification was simplified to two extractions with phenol and chloroform and one extraction with chloroform followed by one ethanol precipitation with carrier transfer RNA (tRNA) and sodium acetate (0.25M). Finally, the RNA was dissolved in aqueous 5 mM methyl mercury and run on 2 percent agarose, 5 mM methyl mercury gels (74) with an Alu restriction enzyme digest of pBR322 (75) as size markers. After electrophoresis, the gel was soaked in 0.5M ammonium acetate to inactivate the methyl mercury, stained with ethidium bromide, and photographed under ultraviolet light. The presence of a discrete 18S RNA band was indicative of a fully recovered, undegraded RNA sample. Finally, the gel was dried and autoradiographed in a cassette with a preflashed film and intensifying screen at -70°C (76). Exposures of 6 hours were normally sufficient. (A) In vitro transcripts at the /3-globin gene truncated as indicated in (D). (B) In vitro transcripts of the /-like globin genes e, Gy, and 8 truncated as indicated. (C) In vitro transcripts of the a-like globin genes C1, afal, and a2 trun- cated as indicated. The numbers adjacent to the arrows indicate the size of marked band which, in each case, agreed with the predicted transcript sizes. However, in the case of; 1, the 5' end of the gene has not been identified by sequence analysis. Since in many cases the different lanes were derived from different experiments, the intensities of the various globin gene transcripts cannot be directly compared. However, the ,8RI, /3Bam, and SRI came from the same fractiona- tion and exposure as did qal and a. These particular band intensities arc therefore comparable.
[19] 7w In vitro Transcription of Human Globin Genes
[20] 91w Although the structural studies de- scribed above have provided much use- ful information, the identification of reg- ulatory sequences, such as transcriptional initiation sites, binding sites for regulatory proteins, and RNA processing sites, require the use of in vivo or in vitro assays for gene expression. In vivo assays include the use of the DNA-mediated gene transfer procedure (51) and SV40 vector systems (52). The recent development of cell-free extracts for RNA polymerase II-dependent transcription of cloned eukaryotic genes provides an in vitro approach to the study of globin gene expression.
[21] 131w Two in vitro transcription systems have been described. One system consists of a cytoplasmic extract, which requires the addition of purified RNA polymerase II for activity (53), while the other system consists of a concentrated whole cell extract with endogenous RNA polymerase II activity (54, 55). In both systems, specific transcription of adenovirus genes was demonstrated by the fact that the capped 5' terminus of the in vitro transcript is indistinguishable from that found in vivo (53,54). The general applicability of these in vitro transcription. systems was recently demonstrated by specific transcription of the chicken con- albumin and ovalbumin genes (56) and the mouse ,8-globin gene (57). We report the results of an in vitro transcription study of human globin genes for which a whole cell extract procedure was used (54).
[22] 168w Analysis of embryonic, fetal, and adult globin gene transcripts. We used a truncated template assay (53) to deter- mine whether individual globin genes can function as templates for specific in vitro transcription. The principle of this assay is illustrated in Fig. 3D. A DNA fragment containing the human ,8-globin gene is digested with a restriction en- zyme that recognizes one or more sites within the gene. For example, if the hu- man ,3-globin gene is digested with Eco RI, Ban HI, or Mbo 1, transcripts of approximately 1450, 480, and 320 nu- cleotides, respectively, should be detect- ed in vitro if transcription begins near the mRNA capping site of the gene. Such transcripts are in fact made when Eco RI-, Bam HI-, or Mbo 1-truncated ,3-globin gene fragments are added to in vitro transcription extracts (Fig. 3A). Similarly, in vitro transcripts of the ex- pected size are observed when the hu- man globin genes E, Gy, and 8, are ana- lyzed (Fig. 3B). The efficiency of in vitro
[23] 13w SCIENCE, VOL. 209 0-o" A 480 " 320 '11. 1-1 , .; 19
[24] 79w transcription of the e-and 'y-globin genes approximately the same as that of the ,8-globin gene. In contrast, it appears that the 8-globin gene is somewhat less efficiently transcribed. A quantitative analysis of this consistently observed difference is in progress. An analysis of the in vitro transcription of the human a- like globin genes is shown in Fig. 3C. The embryonic 1l and the adult ca2 globin genes are transcribed with an efficiency comparable to that of the ,3-like genes.
[25] 13w The pseudogene, tal, is also accurately .ranscribed, although at a lower efficien- cy.
[26] 91w We can make two conclusions on the basis of these results. (i) All of the hu- man globin genes assayed appear to comprise individual transcription units. (ii) With the exception of the 8.globin gene, the embryonic, fetal, and adult genes are transcribed with roughly equal efficiencies in an extract prepared from HeLa cells, which do not ordinarily ex- press globin genes. Thus, the interaction of different globin promoters with RNA polymerase II in vitro is approximately the same, and the mechanisms that medi- ate tissue-specific transcription do not operate in vitro.
[27] 199w Analysis ofthe 5' end of f&globin RNA transcribed in vitro. The results of the experiments shown in Fig. 3, A to C, suggest that the 5' ends of in vitro globin gene transcripts are near to their respec- tive mRNA capping sites. To define these 5' ends precisely, we used two dif- ferent procedures. First, DNA fragments of 50 to 100 bp were isolated from the first exon of 8-, e-, and a-globin genes (Fig. 4B). These fragments were end-labeled, strand-separated, and used in pri- mer extension experiments (see legend to Fig. 4). As shown in Fig. 4B, exten- sion of the (3-, E-, and a-globin gene exon 1 primers with reverse transcriptase should produce DNA fragments of 115, 130, and 75 nucleotides, respectively, when an mRNA template is used. Fragments of exactly these sizes are ob- served when either mRNA or in vitro transcripts of the three globin genes are used as templates (Fig. 3A). We con- clude, therefore, that the 5' ends of the in vitro (8-, e-, and a-globin transcripts are coterminal with their mRNA capping sites. This conclusion was confirmed and extended by structural analysis of the 5' end of the (3-globin in vitro transcript.
[28] 100w The ,3-globin RNA synthesized in vitro contains the same 5' cap structure as au- thentic (3-globin mRNA. This was shown by analyzing the ribonuclease Ti oligonucleotides that bound to dihydroxyl- boryl cellulose (an affinity column that binds 3' hydroxyl groups in RNA, in- 130 transcripts. Primer extension tech- 5' 3# nique: A novel method was used to a L L.,J/-7 Exon 1 define the 5' termini of in vitro tran-Primer scripts (IVT). Short DNA fragments, mRNA or IVT _ 50 to 100 nucleotides in length, were RT product < + X isolated from the first exon of the f8-, 75
[29] 424w e-, and a-globin genes. The (t-, e-, and a-gene primers were Hinf-Hph, Eco RII-Mbo II, and Hinf-Hae III DNA fragments, re- spectively. In each case the primer was end-labeled by partially filling the 5' sticky end; the E. coli DNA polymerase I A fragment (77) and high specific activity (a32P)-labeled nucleotide triphosphate were used. The end-labeled fragment strands were readily separated on 12 percent polyacrylamide 7M urea gels after being denatured by boiling in formamide since the two strands are of different lengths. The labeled antisense single strand DNA fragment was then annealed in O.1M NaCl to either mRNA or unlabeled IVT RNA (100°C, 5 minutes; 60"C, 1 hour). The annealed mixture was then added directly to a reverse transcriptase (RT) reaction with all four unlabeled deoxyribonucleoside triphosphates (78). An equal volume of formamide was added to each reaction and, after denaturation by boiling, direct fractionation on poly- acrylamide 7M urea gels was carried out. Maxam and Gilbert sequence ladders (79) were run on the same gel to determine precisely the sizes of primer extension products. (A) The primer extension data on the (-, e-, and a-globin genes. The a and (8 mRNA's were obtained from purified human globin mRNA while e mRNA was obtained from total poly(A) RNA of hemin- induced K562 cells (80). The colinear mRNA and IVT extension products are denoted by their sizes (deduced from the sequence ladder, which is not shown) and by arrows. The complex set of bands close to the original primer position (denoted by arrows and size) result from modifica- tion of excess unannealed primer by reverse transcriptase. The strong intermediate band in the ft primer mRNA slot probably derives from premature termination due to secondary or tertiary structure in the mRNA template. (B) Line diagrams illustrating the different components in- volved in the (3, e, and a primer extension experiments. (C) Dihydroxylboryl cellulose-selected Ti oligonucleotides synthesized in vitro. The 100-,ul reaction mixtures contain pBR-,8, Pst 4.4 kbp DNA which had been digested with Bam HI as template, and 600 ,Ci of either [a-32P]GTP (Cl) or ATP (C2). After a 1-hour incubation (as above), RNA was purified, digested with ribnu- clease Tl, and selected on columns of dihydroxylboryl cellulose. Oligonucleotides were then purified and fractionated in two dimensions (81). The arrows at the bottom of the figure indicate the direction of electrophoresis (horizontal) and homochromatography (vertical). The solid ar- row in each panel indicates the oligonucleotide that nearest neighbor analysis indicates has a sequence identical to the capped TI oligonucleotide found in authentic human ,-globin mRNA.
[30] 246w The dashed arrow indicates an oligonucleotide that appears to contain the same sequence but may differ in its methylation pattern (see legend to Table 2). Table 2. Nearest neighbor analysis of Ti that bind to dihydroxylboryl celiulose. The oligonucleotides shown in Fig. 4C as well as those obtained from reaction mixtures that contained [a-32P]UTP or CTP as a labeled precursor were fractionated in two dimensions by electrophoresis and homochromatography (81). The TI oligonucleotides were recovered and digested with ribonuclease T2; the products were fractionated by electrophoresis on DEAE paper at pH 3.5 (81) and quantitated by liquid scintillation counting. Three lines of evidence suggest that bound TI oligonucleotides (denoted by solid arrows in Fig. 4C) contained a 5' cap structure: (i) the presence of a T2-resistant moiety, X; (ii) 2 moles of [a-32P]phosphate are incorporated into X when [a-32P]ATP is used as the label, an indication that the A residue adjacent to the cap is methylated; and (iii) further analysis of X with the use of ribonuclease P1 provided the sequence expected for the 5' end of f-globin mRNA (data not shown). Similar analysis of the oligonucleotides, indicated by the dashed arrows in Fig. 4C, is consistent with their containing the same primary structure. It appears, though, to contain a cap 2 structure (that is, the base of the second nucleotide from the cap, a C residue, is also methylated at its 2' position). However, an insufficient quantity of radioactivity in this oligonucleotide prevented further analysis.
[31] 8w Ribonuclease T2 digestion products (32p counts per minute)
[32] 22w nucleotide x A + G C U U 5 48 98 C 56 60 12 G 50 42 A 115 15 9
[33] 251w cluding cap structures containing 3' hydroxyl) by standard RNA fractionation procedures. This method was previously used to analyze the 5' structure of adenovirus 2-specific RNA synthesized in vitro (54). RNA, labeled in four separate in vitro transcription reactions with different (a-32P)-labeled ribonucleoside tri- phosphates, was extracted and analyzed as described in the legend to Fig. 4. The TI ribonuclease oligonucleotides that bound to the affinity column were separated by two-dimensional fractionations and further analyzed by ribonuclease T2 digestion. The results obtained for one oligonucleotide are shown in Fig. 4C and in Table 2. The nearest neighbor fre- quencies obtained, together with the identification of a T2-resistant moiety, are consistent with this oligonucleotide having the same sequence and structure as the in vivo TI oligonucleotide found at the 5' terminus of human ,B-globin mRNA (58), GpppA(m)CATTTG(C). Fur- ther evidence that this oligonucleotide contains a cap structure is described in the legend to Table 2. In vitro transcription of abnormal globin genes. Low levels of mRNA have been associated with the mutant /3-globin genes of patients with (3-thalassemia. The most frequently occurring mutations in /3-globin gene expression are called ,(3- and f3-thalassemia. The f8+-thalassemia is characterized by reduced levels of 8- globin polypeptide production, while in ,B°-thalassemia there is no detectable f3- globin synthesis (3). Recent experiments indicate that in the cases studied thus far, the molecular defect in f3-thalas- semia is abnormal processing of the 38- globin mRNA precursor (26,27). In con- trast, the molecular defects in 830-thalassemia are quite heterogeneous (59).
[34] 102w A comparison of the in vitro transcription products of /3-globin genes isolated from individuals with,B-or 3O-thalassemia is presented in Fig. 5. Consistent with the possibility that the defect in ,8/thalassemia is in mRNA processing rather than transcription, a (83 gene (60) is transcribed with the same efficiency as the normal /8-globin gene. We have also examined the in vitro transcription of f8- globin genes isolated from individuals In vitro transcription of f3-thalassemic globin genes. In vitro transcripts (IVT) ob- tained from the /3-globin genes of two pa- tients, one (A) with f+-thalassemia (60} and the other (B) with /3"-thalassemia (61). The ,8+
[35] 21w IVT was run alongside a control ,8 IVT. Simi- larly, the' , IVT was run alongside 8 and e IVT controls.
[36] 67w with a type of (-thalassemia in which trace amounts of f8-globin mRNA are produced. As shown in Fig. 5B, the efficiency of in vitro transcription of a recently isolated 130-thalassemic gene of this type (61) appears to be the same as that of the normal gene. Thus, as in the case of the 83-thalassemia gene, the in vitro assay does not reveal an obvious defect in transcription.
CONCL
[1] 73w Advances in the molecular biology of human globin genes in conjunction with the information provided by clinical in- vestigations of inherited disorders in glo- bin gene expression make it possible to study the molecular genetics of globin gene regulation. For example, cloned hybridization probes, containing specific regions of the a-or f3-globin gene clus- ters have been used to map deletions that are associated with certain types of a-or f3-thalassemias [see references in (1)].
[2] 202w The most interesting conclusion derived from these studies is that deletions, which alter the normal pattern of globin gene expression during development, map many kilobase pairs away from the genes that are affected. One interpretation of this observation is that glo- bin gene clusters consist of one or more functional domains and that deletions within these domains alter the chromo- some structure and thereby affect the normal pattern of differential globin gene expression (1,10,70). A correlation be- tween gene activity and alterations in chromosome structure is suggested by the observation that actively transcribed genes are more sensitive to deoxyribonuclease I digestion than nontranscribed genes (71). Recently Stalder et al. (72) have shown that deoxyribonuclease I sensitivity extends for many kilobase pairs on either side of active chicken glo- bin genes. This suggests that globin gene activation is associated with a structural alteration of a large region of chromo- somal DNA. Further development of in vitro and in vivo systems for studying globin gene expression, the isolation and structural characterization of abnormal globin genes, and the application of site-direct- ed mutagenesis procedures to normal globin genes should make it possible to identify sequences involved in the regu- lation of human globin gene expressioni.
UNMAPPED
[1] 126w In vitro transcription systems have been used in conjunction with molecular cloning and in vitro mutagenesis procedures as a means of defining RNA polymerase II initiation (promoter) sites. The most thoroughly studied example'is the late promoter of adenovirus 2, the promoter for genes that are expressed late in adenovirus lytic growth. The DNA sequence corresponding to the 5' end of the in vivo late gene transcript and the structure of the capped 5' end of this transcript has been determined (62). In vitro transcription of a cloned DNA fragment containing the late promoter yields a transcript whose capped 5' end is iden- tical to that of the in vivo transcript (53,54). An a-amanitin inhibition experimentdemonstrated that this specific in vitro transcript is RNA polymerase Il-dependent (53,54).
[2] 250w The sequences necessary for specific in vitro transcription of late gene pro- moter fragments have been identified by in vitro mutagenesis experiments. As in the case of globin genes, the sequence ATA is located 31 nucleotides on the 5' side of the mRNA capping site of the adeno 2 late promoter. This sequence was shown to be necessary and sufficient for synthesis of normal levels of specific transcripts by construction and in vitro analysis of deletions encompassing the region from -52 to +33 nucleotides from the mRNA capping site (63). More ex- tensive deletion analysis demonstrated that the ATA box is necessary but not sufficient to support optimal levels of in vitro transcription (64). For example, de- letion of the 5' region to nucleotide -47 leads to a threefold reduction in pro- moter activity. Deletion of all sequences on the 3' side of nucleotide +2 leads to a twofold stimulation of in vitro transcription. However, further deletions to nu- cleotide -2 result in an approximately threefold reduction in RNA synthesis. Thus, although the mRNA capping site is required for optimal levels of in vitro transcription, it does not appear to be necessary for the production of a normal transcript. the basis of the analysis of scription products of deleted adenovirus late promoter DNA fragments; they used the in vitro transcription system de- scribed by Weil et al. (53). A similar analysis of the chicken conalbumin gene has shown that the ATA box is required for specific in vitro transcription (56,65).
[3] 126w In contrast to the examples cited above, the ATA box is not required for specific in vitro transcription of the SV40 late region. This region lacks the ATA sequence but is efficiently transcribed in vivo as well as in vitro (63). Specific tran- scripts are also detected when the SV40 early region, which does contain an ATA box, is examined in vitro. Not all promoters that are known to be active in vivo are also active in vitro. For example, the adenovirus early region II promoter does not function in vitro even though transcripts originating from this site are detected in vivo. It is interesting to note that an ATA box is not located on the immediate 5' side of the apparent transcription start point (66).
[4] 104w An ATA sequence does not appear to be necessary for the accurate transcription of cloned eukaryotic genes in the Xenopus oocyte injection system. Gros- schedl and Birnsteil (67) reported the faithful synthesis of sea urchin histone mRNA in this system. The deletion of 5' flanking sequence including the ATA box does not prevent mRNA synthesis. Similarly, Wickens et al. ( 68) have re- ported the synthesis of chicken ovalbu- min protein in oocytes following the in- jection of either intact genes or a deleted gene fragment that lacks the 5' flanking region and a part of the 5' noncoding re- gion and first intron.
[5] 149w Experiments to identify the nucleotide sequence required for accurate in vitro transcription of the human /3-globin gene are in progress (69). Our preliminary re- sults indicate that the CCAAT box is not necessary for normal f3-globin gene in vitro transcription while a region of sequence including the ATA box appears to be essential. As shown in Fig. 3, we have reproducibly found that the 8-globin gene and *al-globin pseudogene are transcribed at lower efficiencies when compared to the 18and a-globin genes, respectively. Figure 6 compares the 5' end regions of these four genes and illus- trates the sequence differences found in 8 and tial at normally conserved posi- tions (2). Clearly, any one of these dif- ferences could account for lower effi- ciency of in vitro transcription. Work is in progress to define which sequences in 8 and *a l are responsible for the low lev- els of transcription.
[6] 69w In summary, the analysis of the struc- sequences found in all the expressed globin genes. The number of base pairs between homolo- gous sequences is indicated. In order to align the 5' flanking sequences of the 8 and thal genes with other globin genes, it was necessary to introduce the deletions in the region between the CCAAT and ATA box. These putative deletions are indicated by the crosshatched boxes.
[7] 104w ture and function of a number of normal and mutant ribonuclease II polymerase promoters have provided the first indications of the sequences essential for eukaryotic gene transcription. It is not surprising that the initial series of experi- ments revealed that different promoters behave differently when assayed by in vitro transcription. A more detailed analysis of mutagenized DNA templates as well as protein factors affecting transcription will be necessary to achieve a better understanding of eukaryotic pro- moters. In addition, it is essential to compare the behavior of normal and mu- tated promoter sequences in a variety of in vitro and in vivo assay systems.