Part 1

PART 1: Basic Concepts — Introduction to Genetics and Genomics

← Back to Genetics Contents
Ch1 — Introduction (3) Ch2 — Introduction to the Human Genome (17) Ch3 — The Human Genome — Gene Structure and Function (22)

Introduction

1/42
Introduction THE BIRTH AND DEVELOPMENT OF GENETICS AND GENOMICS It may surprise students today to learn that an apprecia…
Ch1 — Segment 1
Introduction THE BIRTH AND DEVELOPMENT OF GENETICS AND GENOMICS It may surprise students today to learn that an appreciation of the role of genetics in medicine dates back to the recognition by Archibald Garrod and others, in the early 20th century, that Mendel’s laws of inheritance could explain the recurrence of certain clinical disorders in families.引言 遗传学与基因组学的诞生与发展 当今的学生可能会惊讶地发现,对遗传学在医学中作用的认识可追溯到20世纪初Archibald Garrod等人认识到孟德尔遗传定律可以解释某些临床疾病在家族中的复发。
During the ensuing years, with developments in molecular biology, the field of medical genetics grew from a small clinical subspecialty concerned with a handful of rare hereditary disorders to a recognized medical specialty whose concepts and approaches are integral to diagnosis and management across the spectrum of disease.在随后的岁月中,随着分子生物学的发展,医学遗传学领域从仅关注少数罕见遗传性疾病的小型临床亚专科,发展成为一门公认的医学专业,其概念和方法贯穿于各种疾病的诊断和管理中。
At the dawn of the 21st century, the international Human Genome Project generated a virtually complete sequence of human DNA—our genome (the suffix -ome coming from the Greek for “all” or “complete”).在21世纪之初,国际人类基因组计划生成了近乎完整的人类DNA序列——我们的基因组(后缀-ome源自希腊语,意为“全部”或“完整”)。
This now serves as the foundation of efforts to catalogue all human genes (both DNA and RNA based), understand their structure and regulation, determine the extent of their variation in different populations, and uncover how genetic variation contributes to disease.这现已作为编录所有人类基因(包括基于DNA和RNA的基因)、理解其结构与调控、确定其在不同人群中的变异程度,以及揭示遗传变异如何导致疾病等工作的基础。
Driven by advances in technology and analytics and the ensuing databases, the human genome of any individual can now be studied in its entirety (genomics) rather than one gene at a time (genetics).在技术、分析方法的进步以及随之而来的数据库的推动下,现在可以研究任何一个个体的整个人类基因组(基因组学),而非一次研究一个基因(遗传学)。
This paradigm shift of testing “genome first”—compared to gene (or gene panel) or phenotype first—arises often throughout the remaining chapters in this book.这种“先测基因组”而非先测基因(或基因组合)或表型的范式转变,在本书后续章节中经常出现。
In the research realm there are already more than 1 million human genome sequences completed; we predict by the next edition of this book, or sooner, whole genome sequencing will be fully adopted as the first-tier diagnostic test for most heritable conditions.在研究领域,已完成超过100万个人类基因组序列;我们预测,到本书下一版或更早,全基因组测序将被完全采纳为大多数遗传性疾病的一线诊断检测。
The Practice of Genetics The medical geneticist is usually a physician who works as part of a team of health care providers, including many other physicians, nurses, laboratory geneticists, genetic counselors, and (more recently) genome informaticists, to evaluate patients for possible hereditary diseases.遗传学实践 医学遗传学家通常是作为医疗团队一员工作的医生,该团队包括许多其他医生、护士、实验室遗传学家、遗传咨询师以及(近来)基因组信息学专家,共同评估患者可能存在的遗传性疾病。
They characterize the patient’s illness through careful history taking and physical examination, assess possible modes of inheritance, arrange for diagnostic testing, develop treatment and surveillance plans, and participate in outreach to other family members at risk for the disorder.他们通过仔细采集病史和体格检查来刻画患者的疾病特征,评估可能的遗传方式,安排诊断性检测,制定治疗和监测计划,并参与对存在该疾病风险的其他家庭成员的延伸服务。
However, genetic principles and approaches are not restricted to any one medical specialty or subspecialty; they permeate many, and perhaps all, areas of medicine.然而,遗传学原理和方法并不局限于任何单一的医学专科或亚专科;它们渗透到许多、甚至可能所有医学领域。
We provide 49 case studies of how genetics and genomics are applied to medicine today.我们提供了49个案例研究,展示遗传学和基因组学如何应用于当今医学。
For example, A psychiatrist, developmental pediatrician, or medical geneticist evaluates a child with autism; microarray or genome sequencing reveals the 16p11. 2 microdeletion associated with a syndrome exhibiting variable expressivity with diverse outcomes (Case 5).例如,精神科医生、发育儿科医生或医学遗传学家评估一名患有自闭症的儿童;微阵列或基因组测序揭示与一种具有可变表达性和不同结局的综合征相关的16p11.2微缺失(案例5)。
A genetic counselor specializing in hereditary breast cancer offers education, testing, interpretation, and support to a young woman with a family history of hereditary breast and ovarian cancer (Case 7).一名专攻遗传性乳腺癌的遗传咨询师向一位有遗传性乳腺癌和卵巢癌家族史的年轻女性提供教育、检测、解读和支持(案例7)。
A hematologist combines family and medical history with gene panel sequencing of a young adult with deep venous thrombosis to assess the benefits and risks of initiating and maintaining anticoagulant therapy (Case 46).一名血液科医生将家族史和病史与对一名患有深静脉血栓形成的年轻成人的基因组合测序相结合,以评估启动和维持抗凝治疗的获益与风险(案例46)。
An immunologist suspecting a rare congenital disorder, including antibody deficiency, uses genome sequencing to reveal compound pathogenic splicing variants in the RNU4ATAC small nuclear RNA gene, thus confirming a diagnosis of Roifman syndrome (Case 35).一名怀疑存在罕见先天性疾病(包括抗体缺乏)的免疫学家使用基因组测序揭示RNU4ATAC小核RNA基因中的复合致病性剪接变异,从而确认了Roifman综合征的诊断(案例35)。
A teenager with a diagnosis of type 1 diabetes was not doing well on insulin treatment.一名被诊断为1型糖尿病的青少年在接受胰岛素治疗时效果不佳。
His mother felt that diabetes “ran in the family,” and the local medical team referred them to Genetics, where a multigene sequencing panel for monogenic diabetes revealed a variant consistent with maturity-onset diabetes of the young (MODY).他的母亲认为糖尿病“有家族史”,当地医疗团队将他们转诊至遗传学科,在那里针对单基因糖尿病的多基因测序组合揭示了一个与青年起病的成熟型糖尿病(MODY)一致的变异。
With the revised diagnosis, the young man’s insulin treatment was discontinued in favor of other low-dose medication, and he was soon back on the soccer field (Case 15).随着诊断的修正,该年轻男性的胰岛素治疗被停用,改用其他低剂量药物,他很快重返足球场(案例15)。
2/42
Categories of Genetic Disease Virtually any disease is the result of the combined action of genes and environment, but t…
Ch1 — Segment 2
Categories of Genetic Disease Virtually any disease is the result of the combined action of genes and environment, but the relative effect of the genetic component may be large or small.遗传病的分类 几乎任何疾病都是基因与环境共同作用的结果,但遗传因素的相对效应可大可小。
Among disorders caused wholly or partly by genetic factors, three main types are recognized: chromosome disorders, single-gene disorders, and multifactorial disorders.在完全或部分由遗传因素引起的疾病中,公认有三种主要类型:染色体病、单基因病和多因素病。
At the time of writing this chapter (1 May 2022) 7160 phenotypes were catalogued as caused by variants in 4629 genes.在撰写本章时(2022年5月1日),已有7160种表型被归类为由4629个基因的变异引起。
Many others remain to be defined.还有许多其他表型尚待明确。
There have also been phenomenal advances in cataloguing genetic variation in online databases, to be used in genotype and phenotype correlations; we point to many of the mostused ones throughout the book.在在线数据库中编目遗传变异以用于基因型-表型关联方面也取得了显著进展;我们在全书中指出了许多最常用的数据库。
In chromosome disorders, the defect is due not to a single alteration in the genetic blueprint but to a change in dosage of genes located on entire chromosomes or chromosome segments.在染色体病中,缺陷并非源于遗传蓝图中的单一改变,而是由于整条染色体或染色体片段上基因剂量的变化。
For example, an extra copy (trisomy) of chromosome 21 underlies a specific disorder, Down syndrome, even though no individual gene on that chromosome is abnormal.例如,21号染色体的额外拷贝(三体)是特定疾病——唐氏综合征的基础,尽管该染色体上没有任何单个基因异常。
Duplication or deletion of smaller segments of chromosomes—called copy number variations (CNVs), ranging in size from submicroscopic to a few percent of a chromosome’s length—can cause complex birth defects such as 22q11. 2 deletion syndrome or isolated phenotypes such as autism without obvious physical abnormalities.较小染色体片段的重复或缺失——称为拷贝数变异(CNV),大小从亚显微水平到染色体长度的百分之几不等——可导致复杂的出生缺陷,如22q11.2缺失综合征,或孤立表型,如无明显躯体异常的孤独症。
Along with 22q11. 2 deletion syndrome, a list of identified recurrent genomic disorders is growing.除22q11.2缺失综合征外,已确定的复发性基因组疾病的列表正在增长。
These arise frequently due to illegitimate recombination catalyzed by flanking duplicated repeat segments; the resulting CNVs involving genes are associated with multiple phenotypes (see Chapter 6).这些疾病常由侧翼重复片段催化的非法重组引起;由此产生的涉及基因的CNV与多种表型相关(见第6章)。
As a group, chromosome disorders are common, with a prevalence of 3% in liveborn infants and accounting for approximately half of all spontaneous losses within the first trimester of pregnancy.作为一类疾病,染色体病很常见,在活产婴儿中患病率为3%,约占妊娠早期所有自然流产的一半。
These types of disorders are discussed in Chapter 6.这些类型的疾病在第6章中讨论。
Monogenic disorders are caused by pathogenic mutations in individual genes.单基因病由单个基因中的致病性突变引起。
The resulting variants may be present on both chromosomes of a pair (one of paternal and one of maternal origin) or on only one chromosome of a pair (matched with a normal copy of that gene on the other copy of that chromosome).由此产生的变异可能存在于一对染色体的两条上(分别来自父源和母源),或仅存在于一对染色体中的一条上(另一条染色体上为该基因的正常拷贝)。
Singlegene defects, also called mendelian conditions, often cause diseases that follow one of the classic inheritance patterns in families (autosomal recessive, autosomal dominant, or X linked).单基因缺陷,也称为孟德尔遗传病,常导致遵循经典遗传模式(常染色体隐性、常染色体显性或X连锁)的疾病。
In a few cases, the mutation occurs in the mitochondrial rather than in the nuclear genome.少数情况下,突变发生在线粒体基因组而非核基因组中。
In any case, the cause is a critical error in the genetic information carried by a single gene.无论如何,病因都是单个基因所携带遗传信息中的一个关键错误。
Monogenic disorders, such as cystic fibrosis (Case 12), Huntington ­disease (Case 24), or Marfan syndrome (Case 30), usually exhibit obvious and characteristic pedigree patterns.单基因病,如囊性纤维化(病例12)、亨廷顿病(病例24)或马凡综合征(病例30),通常表现出明显且特征性的系谱模式。
Most such defects are rare, with a frequency that may be as high as 1 in 500 to 1000 individuals but is usually much less.大多数此类缺陷是罕见的,其频率可能高达每500至1000人中1例,但通常低得多。
Although individually rare, monogenic disorders as a group are responsible for a significant proportion of disease and death.尽管每种单基因病罕见,但作为一类疾病,它们导致了相当比例的疾病和死亡。
Overall, the incidence of serious monogenic disorders in the pediatric population has been estimated to be approximately 1 per 300 liveborn infants; over an entire lifetime, the prevalence of these disorders is 1 in 50.总体而言,在儿科人群中,严重单基因病的发病率估计约为每300名活产婴儿中1例;在整个生命周期中,这些疾病的患病率为1/50。
They are discussed in Chapter 7.这些在第7章中讨论。
Multifactorial disease with complex inheritance encompasses most diseases in which there is a genetic contribution.具有复杂遗传模式的多因素病涵盖了大多数有遗传因素参与的疾病。
The category is characterized by increased incidence of disease in identical twins or other close relatives of affected individuals, compared to that in the general population, yet the family history does not fit the inheritance patterns typical of monogenic disorders.该类别的特点是,与一般人群相比,患者同卵双胞胎或其他近亲的疾病发病率升高,但家族史不符合单基因病典型的遗传模式。
Multifactorial diseases include congenital malformations—such as Hirschsprung disease, cleft lip and palate, or congenital heart defects—as well as many common disorders of adult life—such as breast and ovarian cancer (Case 7), diabetes and inflammatory bowel disease (Case 11).多因素病包括先天性畸形——如先天性巨结肠、唇腭裂或先天性心脏缺陷——以及许多成人常见疾病——如乳腺癌和卵巢癌(病例7)、糖尿病和炎症性肠病(病例11)。
There appears to be no single pathogenic genetic variant in many of these conditions.在这些疾病中,许多似乎不存在单一的致病性遗传变异。
Rather, disease results from the combined impact of variants in many different genes; each variant may cause, protect from, or predispose to a serious defect, often in concert with or triggered by environmental factors.相反,疾病源于许多不同基因中变异的联合影响;每个变异可能导致、保护或易感于严重缺陷,且常与环境因素协同或由环境因素触发。
Estimates of the impact of multifactorial disease range from 5% in the pediatric population to more than 60% in the entire population.对多因素病影响的估计范围从儿科人群的5%到整个人群的60%以上。
Aspects of qualitative and quantitative traits, relative risks, and heritability are the subjects of Chapter 9.定性性状和定量性状、相对风险及遗传度等方面是第9章的主题。
Another remarkable expansion is our understanding of the regulation of gene expression and the role of ­epigenetics in rare monogenic conditions, in common complex diseases, and in cancer (see Chapter 8).另一个显著扩展是我们对基因表达调控以及表观遗传学在罕见单基因病、常见复杂疾病和癌症中作用的理解(见第8章)。
EVOLUTION OF GENETICS IN MEDICINE In the 6 years since the previous edition of this book, our understanding of the role of genetics in medicine has increased dramatically.医学遗传学的演变 自本书上一版出版以来的6年间,我们对遗传学在医学中作用的理解显著增加。
Sometimes genetic interpretations have become clearer with enhanced knowledge, but not always.有时随着知识的增加,遗传学解释变得更清晰,但并非总是如此。
Terms such as variant of uncertain significance, variable expression, penetrance, and pleiotropy prevail, reflecting the residual challenges.如意义不明确变异、可变表达、外显率和多效性等术语普遍存在,反映了尚未解决的挑战。
There are still roughly 20,000 protein-coding genes, but many more of their transcriptional isoforms are now characterized, and an equal number of noncoding RNA genes and regulatory elements must also be considered.仍有大约20,000个蛋白质编码基因,但它们的许多转录异构体现在已被鉴定,并且同样数量的非编码RNA基因和调控元件也必须纳入考虑。
To capture all of these, the medical geneticist now regularly relies on genome-wide testing, sometimes sequencing of the whole genome.为了捕获所有这些,医学遗传学家现在常规依赖全基因组检测,有时进行全基因组测序。
The genome reference build (most recent sequence) used for comparison has already changed twice in less than a decade.用于比较的参考基因组版本(最新序列)在不到十年内已更改了两次。
Indeed, these newer reference genomes and the accompanying population genomic databases better capture the extent of sequence-level and structural variation used in medical genetic interpretations; we discuss their use.事实上,这些较新的参考基因组及配套的人群基因组数据库更好地捕捉了医学遗传学解释中所用的序列水平和结构变异的范围;我们讨论了它们的应用。
3/42
Introduction 3 As an adaptation to changing norms, we have endeavored to update certain terminology for this edition.
Ch1 — Segment 3
Introduction 3 As an adaptation to changing norms, we have endeavored to update certain terminology for this edition.引言3 为适应不断变化的标准,我们努力在本版中更新了某些术语。
The rationale is that the terms mutation and polymorphism have gradually taken on meanings that can be ambiguous, value laden, or confusing.其理由是,“突变”和“多态性”这两个术语逐渐带上了可能含糊、带有价值观色彩或令人困惑的含义。
Instead, variant is more neutral and can be readily modified for precision and clarity.相反,“变异”更为中性,并且可以方便地通过修饰语来达到精确和清晰。
Thus we generally limit use of mutation to a genetic change that gives rise to heritable variation or to the process by which such change arises.因此,我们通常将“突变”的使用限制为引起可遗传变异的遗传改变,或此类改变发生的过程。
Polymorphism refers to having two or more alleles at one locus, each with appreciable frequency.“多态性”指在一个基因座上存在两个或更多等位基因,每个等位基因具有显著频率。
The outcome of mutation is a variant, which may have any number of characteristics, such as pathogenic (disease causing), loss of function, exonic, common, of uncertain significance, etc.突变的结果是产生一个变异,该变异可能具有多种特征,例如致病性(引起疾病)、功能丧失、外显子区、常见、意义未明等。
Judicious use of relevant modifiers for variant is an essential part of the transition; some existing nomenclature remains entrenched.明智地使用变异的修饰语是这一转变的关键部分;一些现有命名法仍根深蒂固。
We have also tried to emphasize the need for dignity, inclusion, respect, and privacy in genetics practice.我们也试图强调在遗传学实践中需要维护尊严、包容、尊重和隐私。
To such ends, we have (respectively) updated picture formats of individuals with genetic conditions, undertaken new discussion of the importance of studies of diverse populations, been mindful of our word choices in reference to individuals and groups, and advocated proper delivery of genetic information through genetic counseling.为此,我们分别更新了患有遗传状况个体的图片格式,重新讨论了多样化人群研究的重要性,注意了我们在提及个体和群体时的措辞选择,并倡导通过遗传咨询恰当传递遗传信息。
This introduction to the ever-evolving language and concepts of human and medical genetics, along with appreciation of the genetic and genomic perspectives on health and disease, will serve to frame the lifelong ­learning that comprises every health professional’s career.这个对人类和医学遗传学不断演变的语言和概念的介绍,以及对健康和疾病的遗传与基因组学视角的理解,将有助于构建构成每位卫生专业人员职业生涯的终身学习框架。

Introduction to the Human Genome

4/42
Introduction to the Human Genome Understanding the organization, variation, and transmission of the human genome is cent…
Ch2 — Segment 4
Introduction to the Human Genome Understanding the organization, variation, and transmission of the human genome is central to appreciating the role of genetics in medicine, as well as the emerging principles of genomic and individualized medicine.人类基因组导论:理解人类基因组的组织、变异和传递,对于认识遗传学在医学中的作用以及基因组医学和个体化医学的新兴原理至关重要。
With the availability of the sequence of the human genome and a growing awareness of the role of genome variation in disease, it is now possible to begin to interpret the impact of that variation on human health on a broad scale.随着人类基因组序列的获取以及对基因组变异在疾病中作用的认识日益加深,如今已有可能开始广泛解读这种变异对人类健康的影响。
The comparison of individual genomes underscores the first major take-home lesson of this book—every individual has a unique constitution of gene products, produced in response to the combined inputs of the genome sequence and one's particular set of environmental exposures and experiences.个体基因组的比较强调了本书的第一个重要结论——每个个体拥有独特的基因产物组成,这些产物是基因组序列与个体特定环境暴露和经历的综合输入所产生的结果。
As pointed out in the previous chapter, this realization reflects what Garrod termed “chemical individuality” over a century ago and provides a conceptual foundation for the practice of genomic and individualized medicine.正如前一章所指出的,这一认识反映了加罗德在一个多世纪前所称的“化学个体性”,并为基因组医学和个体化医学的实践提供了概念基础。
Advances in genome technology and the resulting explosion in knowledge and information stemming from the Human Genome Project are thus playing an increasingly transformational role in integrating and applying concepts and discoveries in genetics to the practice of medicine.基因组技术的进步以及由此带来的人类基因组计划知识和信息的爆炸性增长,因此在将遗传学概念和发现整合并应用于医学实践中发挥着越来越具有变革性的作用。
THE HUMAN GENOME AND THE CHROMOSOMAL BASIS OF HEREDITY Appreciation of the importance of genetics to medicine requires an understanding of the nature of the hereditary material, how it is packaged into the human genome, and how it is transmitted from cell to cell during cell division and from generation to generation during reproduction.人类基因组与遗传的染色体基础:认识遗传学对医学的重要性,需要理解遗传物质的本质、如何包装成人类基因组、如何在细胞分裂过程中在细胞间传递,以及在生殖过程中世代相传。
The human genome consists of large amounts of the chemical deoxyribonucleic acid (DNA) that contains within its structure the genetic information needed to specify all aspects of embryogenesis, development, growth, metabolism, and ­reproduction— essentially all aspects of what makes a human being a functional organism.人类基因组由大量的化学物质脱氧核糖核酸(DNA)组成,其结构内包含指定胚胎发生、发育、生长、代谢和生殖所有方面所需的遗传信息——本质上涵盖了使人类成为功能有机体的所有方面。
Every nucleated cell in the body carries its own copy of the human genome, which contains, depending on how one defines the term, approximately 20,000 to 50,000 genes (see 1).体内每个有核细胞都携带自身的人类基因组拷贝,该基因组包含大约20,000到50,000个基因(具体取决于该术语的定义)(参见1)。
Genes, which at this point we consider simply and most broadly as functional units of genetic information, are encoded in the DNA of the genome, organized into a number of rod-shaped organelles called chromosomes in the nucleus of each cell.基因,在此我们简单且最广泛地将其视为遗传信息的功能单位,编码于基因组的DNA中,并组织成每个细胞核内称为染色体的多个杆状细胞器。
The influence of genes and genetics on states of health and disease is profound, and its roots are found in the information encoded in the DNA that makes up the human genome.基因和遗传学对健康与疾病状态的影响深远,其根源存在于构成人类基因组的DNA所编码的信息之中。
Each species has a characteristic chromosome complement (karyotype) in terms of the number, morphology, and content of the chromosomes that make up its genome.每个物种在构成其基因组的染色体的数量、形态和内容方面都具有特征性的染色体组型(核型)。
The genes are in linear order along the chromosomes, each gene having a precise position or locus.基因沿染色体线性排列,每个基因具有精确的位置或位点。
A gene map is the map of the genomic location of the genes and is characteristic of each species and the individuals within a species.基因图是基因在基因组位置的图谱,每个物种及物种内的个体都具有其特征。
The study of chromosomes, their structure, and their inheritance is called cytogenetics.对染色体及其结构和遗传的研究称为细胞遗传学。
The science of human 1 CHROMOSOME AND GENOME ANALYSIS IN CLINICAL MEDICINE Chromosome and genome analysis has become important diagnostic procedures in clinical medicine.人类科学1:临床医学中的染色体和基因组分析——染色体和基因组分析已成为临床医学中重要的诊断程序。
As described more fully in subsequent chapters, these applications include the following: Clinical diagnosis.如后续章节更详细所述,这些应用包括以下方面:临床诊断。
Numerous medical conditions, including some that are common, are associated with changes in chromosome number or structure and require chromosome or genome analysis for diagnosis and genetic counseling (see Chapters 5 and 6).许多医学状况,包括一些常见疾病,与染色体数目或结构的改变有关,需要染色体或基因组分析以进行诊断和遗传咨询(见第5章和第6章)。
Disease gene identification.疾病基因鉴定。
A major goal of medical genetics and genomics today is the identification of the role of specific genes in health and disease.当今医学遗传学和基因组学的一个主要目标是识别特定基因在健康和疾病中的作用。
This topic is referred to repeatedly but is discussed in detail in Chapter 11.这一主题被反复提及,但在第11章中进行了详细讨论。
Cancer genomics.癌症基因组学。
Genomic and chromosomal changes in somatic cells are involved in the initiation and progression of many types of cancer (see Chapter 16).体细胞中的基因组和染色体改变参与多种癌症的发生和进展(见第16章)。
Disease treatment.疾病治疗。
Understanding the exact molecular basis of each individual’s monogenic disorder allows for targeted therapies to address the condition (see Chapter 14).理解每个个体单基因疾病的确切分子基础,使得能够针对该疾病进行靶向治疗(见第14章)。
Prenatal diagnosis.产前诊断。
Chromosome and genome analysis are essential procedures in prenatal diagnosis (see Chapter 18).染色体和基因组分析是产前诊断中必不可少的程序(见第18章)。
5/42
Somatic cell Mitochondrial chromosomes Human Genome Sequence... ...
Ch2 — Segment 5
Somatic cell Mitochondrial chromosomes Human Genome Sequence... ...体细胞线粒体染色体人类基因组序列......
Nuclear chromosomes cytogenetics dates from 1956, when it was first established that the normal human chromosome number is 46.核染色体细胞遗传学始于1956年,当时首次确定正常人类染色体数目为46条。
Since that time, much has been learned about human chromosomes, their normal structure and composition, and the identity of the genes that they contain, as well as their numerous and varied abnormalities.自那时起,人们对人类染色体、其正常结构与组成、所含基因的特性,以及其众多且多样的异常有了大量了解。
With the exception of cells that develop into gametes (the germline), all cells that contribute to one’s body are called somatic cells (soma, “body”).除了发育成配子的细胞(生殖系)外,构成身体的所有细胞都称为体细胞(soma意为“身体”)。
The genome contained in the nucleus of human somatic cells consists of 46 chromosomes, made up of 24 different types and arranged in 23 pairs .人类体细胞核中所含的基因组由46条染色体组成,分为24种不同类型,排列成23对。
Of those 23 pairs, 22 are alike in males and females and are called autosomes, originally numbered in order of their apparent size from the largest to the smallest.在这23对中,有22对在男性和女性中相同,称为常染色体,最初按其表观大小从最大到最小排序编号。
The remaining pair comprises the two different types of sex chromosomes: an X and a Y chromosome in males and two X chromosomes in females.剩余的一对由两种不同类型的性染色体组成:男性为一条X和一条Y染色体,女性为两条X染色体。
Central to the concept of the human genome, each chromosome carries a different subset of genes arranged linearly along its DNA (see 2).人类基因组概念的核心是,每条染色体携带一组不同的基因,这些基因沿其DNA线性排列(见图2)。
Members of a pair of chromosomes (referred to as homologous chromosomes or homologues) carry matching genetic information; that is, they typically have the same genes in the same order.一对染色体中的两条(称为同源染色体)携带匹配的遗传信息;也就是说,它们通常以相同顺序拥有相同的基因。
At any specific locus, however, the homologues either may be identical or may vary slightly in sequence; these different forms of a gene are called alleles.然而,在任何特定基因座上,同源染色体在序列上可能相同,也可能略有差异;基因的这些不同形式称为等位基因。
One member of each pair of chromosomes is inherited from the father, the other from the mother.每对染色体中的一条来自父亲,另一条来自母亲。
Normally, the members of a pair of autosomes are microscopically indistinguishable from each other.正常情况下,一对常染色体的两条在显微镜下无法区分。
In females, the sex chromosomes, the two X chromosomes, are likewise largely indistinguishable.在女性中,性染色体(两条X染色体)也基本无法区分。
In males, however, the sex chromosomes differ.然而,在男性中,性染色体不同。
One is an X, identical to the Xs of the female, inherited by a male from his mother and transmitted to his daughters; the other, the Y chromosome, is inherited from his father and transmitted to his sons.一条是X染色体,与女性的X染色体相同,由男性从其母亲遗传并传递给其女儿;另一条是Y染色体,从其父亲遗传并传递给其儿子。
In Chapter 6, as we explore the chromosomal and genomic basis of disease, we will look at some exceptions to the simple and almost universal rule that human biological females are XX and human biological males are XY.在第6章中,当我们探讨疾病的染色体和基因组基础时,我们将看到一些例外情况,这些例外挑战了人类生物学女性为XX、生物学男性为XY这一简单且近乎普遍的规律。
In addition to the nuclear genome, a small but important part of the human genome resides in mitochondria in the cytoplasm .除核基因组外,人类基因组中一个虽小但重要的部分位于细胞质内的线粒体中。
The mitochondrial chromosome, to be described later in this chapter, has a number of unusual features that distinguish it from the rest of the human genome.线粒体染色体(将在本章后面描述)具有许多不寻常的特征,使其与人类基因组的其余部分区别开来。
6/42
Introduction to the Human Genome 7 DNA Structure: A Brief Review Before the organization of the human genome and its chr…
Ch2 — Segment 6
Introduction to the Human Genome 7 DNA Structure: A Brief Review Before the organization of the human genome and its chromosomes are considered in detail, it is necessary to review the nature of the DNA that makes up the genome.人类基因组导论 7 DNA结构:简要回顾 在详细探讨人类基因组及其染色体的组织之前,有必要回顾构成基因组的DNA的性质。
DNA is a polymeric nucleic acid macromolecule composed of three types of units: a five-carbon sugar, deoxyribose; a nitrogen-containing base; and a phosphate group .DNA是一种聚合核酸大分子,由三种单元组成:五碳糖(脱氧核糖)、含氮碱基和磷酸基团。
The bases are of two types, purines and pyrimidines.碱基分为两种类型:嘌呤和嘧啶。
In DNA, there are two purine bases, adenine Cytosine (C) Guanine (G) Base O Phosphate Deoxyribose O O O O _ P CH2 C OH H H H H C C H C 5' 3' Adenine (A) Thymine (T) N HC Purines Pyrimidines NH2 N C C C N N CH H O N C C C N N CH H C H2N HN _ O O CH3 CH C C C N H HN NH2 O CH C CH C N H N Each of the four bases bonds with deoxyribose (through the nitrogen shown in magenta) forming a nucleoside which bonds with a phosphate group to form the corresponding nucleotides..在DNA中,有两种嘌呤碱基(腺嘌呤和鸟嘌呤)和两种嘧啶碱基(胞嘧啶和胸腺嘧啶)。四种碱基各自与脱氧核糖(通过品红色显示的氮原子)结合形成核苷,核苷再与磷酸基团结合形成相应的核苷酸。
GENES IN THE HUMAN GENOME What is a gene?人类基因组中的基因 什么是基因?
And how many genes do we have?我们有多少个基因?
These questions are more difficult to answer than they might seem.这些问题比表面看起来更难回答。
The word gene, first introduced in 1908, has been used in many different contexts since the essential features of heritable “unit characters” were first outlined by Mendel over 150 years ago.“基因”一词最早于1908年提出,自150多年前孟德尔首次概述可遗传“单位性状”的基本特征以来,该词已在许多不同语境中使用。
To physicians (and indeed to Mendel and other early geneticists), a gene can be defined by its observable impact on an organism and on its statistically determined transmission from generation to generation.对医生(实际上对孟德尔和其他早期遗传学家而言)来说,基因可以通过其对生物体的可观察影响及其统计学确定的代际传递来定义。
To medical geneticists, a gene is recognized clinically in the context of an observable variant that leads to a characteristic clinical condition, and today we recognize over 7000 such conditions (see Chapter 7).对医学遗传学家而言,基因在临床上是通过导致特征性临床状况的可观察变异来识别的,如今我们已识别出7000多种此类状况(见第7章)。
The Human Genome Project provided a more systematic basis for delineating human genes, relying on DNA sequence analysis rather than clinical acumen and family studies alone; indeed, this was one of the most compelling rationales for initiating the project in the late 1980s.人类基因组计划为描绘人类基因提供了更系统的基础,其依据是DNA序列分析,而非仅依靠临床判断和家族研究;事实上,这正是20世纪80年代末启动该计划最令人信服的理由之一。
However, even with the finished sequence product in 2003, it was apparent that our ability to recognize features of the sequence that point to the existence or identity of a gene was sorely lacking.然而,即使到2003年完成序列产品时,我们识别序列中指向基因存在或身份的性状的能力仍然明显不足。
Interpreting the human genome sequence and relating its variation to human biology in both health and disease is thus an ongoing challenge for biomedical research.因此,解读人类基因组序列并将其变异与人类健康与疾病生物学联系起来,仍然是生物医学研究面临的持续挑战。
Although the ultimate catalogue of human genes remains an elusive target, we recognize two general types of gene, those whose product is a protein and those whose product is a functional RNA.尽管人类基因的最终目录仍然是一个难以捉摸的目标,但我们认识到两大类基因:产物为蛋白质的基因和产物为功能性RNA的基因。
The number of protein-coding genes—recognized by features in the genome that will be discussed in is estimated to be somewhere between 20,000 and 25,000.蛋白质编码基因的数量——通过基因组中的特征(将在后续讨论)识别——估计在20,000到25,000之间。
In this book, we typically use approximately 20,000 as the number, and the reader should recognize that this is both imprecise and perhaps an underestimate.在本书中,我们通常使用约20,000作为该数值,读者应认识到这既不精确,可能还是低估。
In addition, however, it has been clear for several decades that the ultimate product of some genes is not a protein at all, but an RNA transcribed from the DNA sequence.然而,此外,几十年来我们已经清楚,某些基因的最终产物根本不是蛋白质,而是从DNA序列转录而来的RNA。
There are many different types of such RNA genes (typically called noncoding genes to distinguish them from protein-coding genes), and it is currently estimated that there are at least another 20,000 to 25,000 noncoding RNA genes around the human genome.这类RNA基因有许多不同类型(通常称为非编码基因,以区别于蛋白质编码基因),目前估计人类基因组中至少还有20,000到25,000个非编码RNA基因。
Thus overall—and depending on what one means by the term—the total number of genes in the human genome is on the order of approximately 20,000 to 50,000.因此总体而言——取决于该术语的含义——人类基因组中的基因总数大约在20,000到50,000的数量级。
However, the reader will appreciate that this remains a moving target, subject to evolving definitions, increases in technological capabilities and analytical precision, advances in informatics and digital medicine, and more complete genome anno­tation.然而,读者会理解这仍然是一个动态目标,受定义演变、技术能力和分析精度的提高、信息学和数字医学的进步以及更完整的基因组注释的影响。
Recent efforts to complete the sequencing of the highly repetitive heterochromatin segments of chromosomes in the centromeres, on the short arms of the acrocentric chromosomes, in segmental duplications and in subtelomeric and telomeric regions, accounting for 8% of the human genome, has identified 200 million bp of DNA, and thousands of new genes, hundreds of them protein coding.近期完成对染色体高度重复异染色质片段(包括着丝粒、近端着丝粒染色体短臂、节段重复区以及亚端粒和端粒区域,占人类基因组的8%)测序的努力,已鉴定出2亿碱基对的DNA,以及数千个新基因,其中数百个为蛋白质编码基因。
7/42
(A) and guanine (G), and two pyrimidine bases, thymine (T) and cytosine (C).
Ch2 — Segment 7
(A) and guanine (G), and two pyrimidine bases, thymine (T) and cytosine (C).(A)和鸟嘌呤(G),以及两种嘧啶碱基:胸腺嘧啶(T)和胞嘧啶(C)。
Nucleotides, each composed of a base, a phosphate, and a sugar moiety, polymerize into long polynucleotide chains held together by 5′-3′ phosphodiester bonds formed between adjacent deoxyribose units .核苷酸由碱基、磷酸基团和糖基组成,它们聚合成长的多核苷酸链,链间通过相邻脱氧核糖单元之间形成的5'-3'磷酸二酯键连接。
In the human genome, these polynucleotide chains exist in the form of a double helix that can be hundreds of millions of nucleotides long in the case of the largest human chromosomes.在人类基因组中,这些多核苷酸链以双螺旋形式存在,对于最大的人类染色体而言,其长度可达数亿个核苷酸。
The anatomic structure of DNA carries the chemical information that allows the exact transmission of genetic information from one cell to its daughter cells and from one generation to the next.DNA的解剖结构携带着化学信息,使得遗传信息能够从亲代细胞精确传递给子代细胞,并世代相传。
At the same time, the primary structure of DNA specifies the amino acid sequences of the polypeptide chains of proteins, as described in the next chapter.同时,DNA的一级结构决定了蛋白质多肽链的氨基酸序列,这将在下一章中描述。
DNA has elegant features that give it these properties.DNA具有赋予其这些特性的精巧特征。
The native state of DNA, as elucidated by James Watson and Francis Crick in 1953, is a double helix .DNA的天然状态是双螺旋结构,这一结构由詹姆斯·沃森和弗朗西斯·克里克于1953年阐明。
The helical structure resembles a right-handed spiral staircase in which its two polynucleotide chains run in opposite directions, held together by hydrogen bonds between pairs of bases: T of one chain paired with A of the other, and G with C.螺旋结构类似于右旋螺旋楼梯,其两条多核苷酸链以相反方向延伸,通过碱基对之间的氢键连接:一条链上的T与另一条链上的A配对,G与C配对。
The specific nature of the genetic information encoded in the human genome lies in the sequence of Cs, As, Gs, and Ts on the two strands of the double helix along each of the chromosomes, both in the nucleus and in mitochondria .人类基因组中编码的遗传信息的特异性在于每条染色体(包括细胞核和线粒体中的染色体)上双螺旋两条链上的C、A、G和T的序列。
Because of the complementary nature of the two strands of DNA, knowledge of the sequence of nucleotide bases on one strand automatically allows one to determine the sequence of bases on the other strand.由于DNA两条链的互补性质,已知一条链上的核苷酸碱基序列即可自动确定另一条链上的碱基序列。
The double-stranded structure of DNA molecules allows them to replicate precisely by separation of the two strands, followed by synthesis of two new complementary strands, in accordance with the sequence of the original template strands .DNA分子的双链结构使其能够通过两条链的分离,随后根据原始模板链的序列合成两条新的互补链,从而精确复制。
Similarly, when necessary, the base complementarity allows efficient and correct repair of damaged DNA molecules.同样,必要时碱基互补性允许对受损DNA分子进行高效且正确的修复。
Structure of Human Chromosomes The composition of genes in the human genome, as well as the determinants of their expression, is specified in the 5' 5' 3' 3' C C G G G C T A Base 3 O H2C C OH H H H H C C H C 5' 3' O O O O P _ Base 2 O H2C C H H H H C C H C 5' 3' O O O O P _ Base 1 O H2C C H H H H C C H C 5' 3' O O O O P _ _ 3' end 5' end Hydrogen bonds 3. 4 Å 34 Å 20 Å A B A portion of a DNA polynucleotide chain, showing the 5′-3′ phosphodiester bonds that link adjacent nucleotides.人类染色体的结构 人类基因组中基因的组成及其表达的决定因素在5' 5' 3' 3' C C G G G C T A 碱基 3 O H2C C OH H H H H C C H C 5' 3' O O O O P _ 碱基2 O H2C C H H H H C C H C 5' 3' O O O O P _ 碱基1 O H2C C H H H H C C H C 5' 3' O O O O P _ _ 3'端 5'端 氢键 3.4 Å 34 Å 20 Å A B 一段DNA多核苷酸链,显示连接相邻核苷酸的5'-3'磷酸二酯键。
(B) The double-helix model of DNA, as proposed by Watson and Crick.(B) 沃森和克里克提出的DNA双螺旋模型。
The horizontal “rungs” represent the paired bases.横向的“梯级”代表配对的碱基。
The helix is said to be right-handed because the strand going from lower left to upper right crosses over the opposite strand.该螺旋被称为右旋,因为从左下到右上的链跨越了对侧链。
The detailed portion of the figure illustrates the two complementary strands of DNA, showing the AT and GC base pairs.图中的细节部分展示了DNA的两条互补链,显示了AT和GC碱基对。
Note that the orientation of the two strands is antiparallel. )注意两条链的方向是反向平行的。)
8/42
Introduction to the Human Genome 9 DNA of the 46 human chromosomes in the nucleus plus the mitochondrial chromosome.
Ch2 — Segment 8
Introduction to the Human Genome 9 DNA of the 46 human chromosomes in the nucleus plus the mitochondrial chromosome.人类基因组导论 9 细胞核中46条人类染色体的DNA加上线粒体染色体。
Each human chromosome consists of a single, continuous DNA double helix; that is, each chromosome is one long, double-stranded DNA molecule, and the nuclear genome consists therefore of 46 linear DNA molecules, totaling more than 6 billion nucleotide pairs .每条人类染色体由一条连续不断的DNA双螺旋组成;也就是说,每条染色体是一个长的双链DNA分子,因此核基因组由46条线性DNA分子组成,总计超过60亿个核苷酸对。
Chromosomes are not naked DNA double helices, however.然而,染色体并非裸露的DNA双螺旋。
Within each cell, the genome is packaged as chromatin, in which genomic DNA is complexed with several classes of specialized proteins.在每个细胞中,基因组以染色质形式包装,其中基因组DNA与几类特殊蛋白质复合。
Except during cell division, chromatin is distributed throughout the nucleus and is relatively homogeneous in appearance under the microscope.除细胞分裂期外,染色质分布在细胞核各处,在显微镜下外观相对均一。
When a cell divides, however, its genome condenses to appear as microscopically visible chromosomes.然而,当细胞分裂时,其基因组浓缩,呈现为显微镜下可见的染色体。
Chromosomes are thus visible as discrete structures only in dividing cells, although they retain their integrity between cell divisions.因此,染色体仅在分裂细胞中可见为离散结构,尽管它们在细胞分裂之间保持完整性。
The DNA molecule of a chromosome exists in chromatin as a complex with a family of basic chromosomal proteins called histones.染色体的DNA分子在染色质中与一组称为组蛋白的碱性染色体蛋白形成复合物。
This fundamental unit interacts with a heterogeneous group of nonhistone proteins, which are involved in establishing a proper spatial and functional environment to ensure normal chromosome behavior and appropriate gene expression.这一基本单位与一组异质的非组蛋白相互作用,这些蛋白参与建立适当的空间和功能环境,以确保正常的染色体行为和适当的基因表达。
Five major types of histones play a critical role in the proper packaging of chromatin.五种主要类型的组蛋白在染色质的正确包装中起关键作用。
Two copies each of the four core histones H2A, H2B, H3, and H4 constitute an octamer, around which a segment of DNA double helix winds, like thread around a spool .四种核心组蛋白H2A、H2B、H3和H4各两个拷贝构成一个八聚体,一段DNA双螺旋缠绕在其周围,如同线绕在线轴上。
Approximately 140 base pairs (bp) of DNA are associated with each histone core, making just under two turns around the octamer.大约140个碱基对(bp)的DNA与每个组蛋白核心相关联,围绕八聚体绕行略少于两圈。
After a short (20- to 60-bp) “spacer” segment of DNA, the next core DNA complex forms, and so on, giving chromatin the appearance of beads on a string.在短(20至60 bp)的DNA“间隔”片段之后,下一个核心DNA复合物形成,如此重复,使染色质呈现串珠状外观。
Each complex of DNA with core histones is called a nucleosome , which is the basic structural unit of chromatin, and each of the 46 human chromosomes contains several hundred thousand to well over 1 million nucleosomes.每个DNA与核心组蛋白的复合物称为核小体,它是染色质的基本结构单位,每条人类染色体包含数十万至超过一百万个核小体。
A fifth histone, H1, appears to bind to DNA at the edge of each nucleosome, in the internucleosomal spacer region. 5' 5' 3' 5' 3' 5' 3' 3' Double helix Nucleosome fiber (“beads on a string”) Solenoid Each loop contains ~100–200 kb of DNA Histone octamer ~140 bp of DNA 2 nm ~10 nm ~30 nm Portion of an interphase chromosome Interphase nucleus Cell in early interphase第五种组蛋白H1似乎结合在每条核小体边缘的DNA上,位于核小体间的间隔区。
9/42
The amount of DNA associated with a core nucleosome, together with the spacer region, is approximately 200 bp.
Ch2 — Segment 9
The amount of DNA associated with a core nucleosome, together with the spacer region, is approximately 200 bp.核心核小体相关的DNA量,连同间隔区,约为200 bp。
In addition to the major histone types, a number of specialized histones can substitute for H3 or H2A and confer specific characteristics on the genomic DNA at that location.除主要组蛋白类型外,一些特化组蛋白可替代H3或H2A,并为该位置的基因组DNA赋予特定特征。
Histones can also be modified by chemical changes, and these modifications can change the properties of nucleosomes that contain them.组蛋白也可通过化学变化进行修饰,这些修饰可改变含有它们的核小体的性质。
As discussed further in Chapter 3, the pattern of major and specialized histone types and their modifications can vary from cell type to cell type and is thought to specify how DNA is packaged and how accessible it is to regulatory molecules that determine gene expression or other genome functions.正如第3章进一步讨论的,主要和特化组蛋白类型及其修饰的模式可因细胞类型而异,并被认为决定了DNA的包装方式以及调控分子(这些调控分子决定基因表达或其他基因组功能)对DNA的可及性。
During the cell cycle, as we will see later in this chapter, chromosomes pass through orderly stages of condensation and decondensation.在细胞周期中,正如我们将在本章后面看到的,染色体经历有序的凝缩和解凝缩阶段。
However, even when chromosomes are in their most decondensed state, in a stage of the cell cycle called interphase, DNA packaged in chromatin is substantially more condensed than it would be as a native, protein-free, double helix.然而,即使染色体处于其最解凝缩的状态(即细胞周期中称为间期的阶段),包装在染色质中的DNA也远比其天然的无蛋白质双螺旋形式更为凝缩。
Further, the long strings of nucleosomes are themselves compacted into a secondary helical structure, a cylindrical solenoid fiber (from the Greek solenoeides, “pipe shaped”) that appears to be the fundamental unit of chromatin organization .此外,长串的核小体自身压缩成二级螺旋结构,即圆柱状的螺线管纤维(源自希腊语solenoeides,“管状”),该结构似乎是染色质组织的基本单位。
The solenoids are themselves packed into loops or domains attached at intervals of approximately 100,000 bp (equivalent to 100 kilobase pairs [kb] because 1 kb = 1000 bp) to a protein scaffold within the nucleus.螺线管本身又包装成环或结构域,这些环或结构域以大约100,000 bp(相当于100千碱基对[kb],因为1 kb = 1000 bp)的间隔附着于细胞核内的蛋白质支架上。
It has been speculated that these loops are the functional units of the genome and that the attachment points of each loop are specified along the chromosomal DNA.据推测,这些环是基因组的功能单位,每个环的附着点沿染色体DNA被特定指定。
As we shall see, one level of control of gene expression depends on how DNA and genes are packaged into chromosomes and on their association with chromatin proteins in the packaging process.正如我们将要看到的,基因表达的一个控制层面取决于DNA和基因如何包装到染色体中,以及它们在包装过程中与染色质蛋白的关联。
The enormous amount of genomic DNA packaged into a chromosome can be appreciated when chromosomes are treated to release the DNA from the underlying protein scaffold .当处理染色体以将DNA从下方的蛋白质支架上释放出来时,可以体会到包装到一条染色体中的基因组DNA数量之巨大。
When DNA is released in this manner, long loops of DNA can be visualized, and the residual scaffolding can be seen to reproduce the outline of a typical chromosome.当以这种方式释放DNA时,可以观察到长的DNA环,并可看到残留的支架再现典型染色体的轮廓。
The Mitochondrial Chromosome As mentioned earlier, a small but important subset of genes encoded in the human genome resides in the cytoplasm in the mitochondria .线粒体染色体 如前所述,人类基因组中编码的一小部分但重要的基因子集存在于细胞质内的线粒体中。
Mitochondrial genes exhibit exclusively maternal inheritance (see Chapter 7).线粒体基因表现出严格的母系遗传(见第7章)。
Human cells can have hundreds to thousands of mitochondria, each containing a number of copies of a small circular molecule, the mitochondrial chromosome.人类细胞可拥有数百至数千个线粒体,每个线粒体含有多个拷贝的一个小的环状分子——线粒体染色体。
The mitochondrial DNA molecule is only 16 kb in length (just a tiny fraction of the length of even the smallest nuclear chromosome) and encodes only 37 genes.线粒体DNA分子长度仅为16 kb(即使是最小核染色体长度的一小部分),仅编码37个基因。
The products of these genes function in mitochondria, although the vast majority of proteins within the mitochondria are, in fact, the products of nuclear genes.这些基因的产物在线粒体中发挥功能,尽管线粒体内绝大多数蛋白质实际上是核基因的产物。
Pathogenic variants in mitochondrial genes have been demonstrated in several maternally inherited as well as sporadic disorders (Case 33) (see Chapters 7 and 13).线粒体基因的致病性变异已在多种母系遗传性疾病以及散发性疾病中被证实(病例33)(见第7章和第13章)。
The Human Genome Sequence With a general understanding of the structure and clinical importance of chromosomes and the genes they carry, scientists turned attention to the identification of specific genes and their location in the human genome.人类基因组序列 在对染色体及其携带基因的结构和临床重要性有了基本了解后,科学家将注意力转向识别特定基因及其在人类基因组中的位置。
From this broad effort emerged the Human Genome Project, an international consortium of hundreds of laboratories around the world, formed to map and determine the sequence of the 3. 3 billion bp of DNA located among the 24 types of human chromosomes.从这一广泛努力中产生了人类基因组计划,这是一个由全球数百个实验室组成的国际联盟,旨在绘制并测定分布于24种人类染色体中的33亿碱基对DNA的序列。
Over the course of 15 years, powered by major developments in DNA-sequencing technology, large sequencing centers collaborated to assemble sequences of each chromosome.在15年的时间里,得益于DNA测序技术的重大发展,大型测序中心合作组装了每条染色体的序列。
The genomes actually being sequenced came from several different individuals, and the consensus sequence that resulted at the conclusion of the Human Genome Project was reported in 2003 as a “reference” sequence assembly, to be used as a basis for later comparison with sequences of individual genomes.实际被测序的基因组来自几个不同的个体,人类基因组计划结束时得出的共有序列于2003年被报告为“参考”序列组装,用作后来与个体基因组序列进行比较的基础。
This reference sequence is maintained in publicly accessible databases to facilitate scientific discovery and its translation into useful advances for medicine.该参考序列保存于可公开访问的数据库中,以促进科学发现及其转化为医学上的有用进展。
Genome sequences are typically presented in a 5′ to 3′ direction on just one of the two strands of the double helix because—owing to the complementary nature of DNA structure described earlier—if one knows the sequence of one strand, one can infer the sequence of the other strand .基因组序列通常仅以双螺旋两条链中一条链的5′到3′方向呈现,因为——由于先前描述的DNA结构的互补性——如果知道一条链的序列,就可以推断出另一条链的序列。
Organization of the Human Genome Chromosomes are not just a random collection of different types of genes and other DNA sequences.人类基因组组织 染色体不仅仅是不同类型基因和其他DNA序列的随机集合。
Regions of the genome with similar characteristics tend to be clustered together, and the functional organization of the genome reflects its structural organization and sequence.具有相似特征的基因组区域倾向于聚集在一起,基因组的功能组织反映了其结构组织和序列。
Some chromosome regions, or even whole chromosomes, are high in gene content (“gene rich”), whereas others are low (“gene poor”) .一些染色体区域甚至整条染色体基因含量高(“基因丰富”),而其他区域则低(“基因贫乏”)。
The clinical consequences of abnormalities of genome structure reflect the specific nature of the genes and sequences involved.基因组结构异常的临床后果反映了所涉及基因和序列的特定性质。
Thus abnormalities of gene-rich chromosomes or chromosomal regions tend to be much more severe clinically than similar-sized defects involving gene-poor parts of the genome.因此,涉及基因丰富染色体或染色体区域的异常,其临床严重程度往往远大于涉及基因组中基因贫乏部分的相似大小的缺陷。
As a result of knowledge gained from the Human Genome Project, it is apparent that the organization of DNA in the human genome is both more varied and more complex than was once appreciated.由于从人类基因组计划中获得的知识,人类基因组中DNA的组织显然比过去所认为的更加多样和复杂。
Of the billions of base pairs of DNA in any genome, fewer than 1. 5% actually encodes proteins.在任何基因组的数十亿碱基对DNA中,实际上只有不到1.5%编码蛋白质。
This portion of the genome is referred to as the exome.基因组的这一部分称为外显子组。
Regulatory elements调控元件。
10/42
Introduction to the Human Genome 11 Double Helix Reference Sequence Individual 1 Individual 2 Individual 3 Individual 4 …
Ch2 — Segment 10
Introduction to the Human Genome 11 Double Helix Reference Sequence Individual 1 Individual 2 Individual 3 Individual 4 Individual 5- T T T T -................................................人类基因组导论 11 双螺旋参考序列 个体1 个体2 个体3 个体4 个体5- T T T T -................................................
A G CA T G C A AA 5´ 3´ 3´ 5´ By convention, sequences are presented from one strand of DNA only, because the sequence of the complementary strand can be inferred from the double-stranded nature of DNA (shown above the reference sequence).A G CA T G C A AA 5´ 3´ 3´ 5´ 按照惯例,序列仅从DNA的一条链表示,因为互补链的序列可以从DNA的双链性质推断(如上方的参考序列所示)。
The sequence of DNA from a group of individuals is similar but not identical to the reference, with single nucleotide changes in some individuals and a small deletion of two bases in another. 1500 1000 500 0 Number of Genes 0 50 100 150 200 250 Gene-rich chromosomes Gene-poor chromosomes Genome Average: ~6. 7 genes/Mb Chromosome Size (Mb) 1 2 3 6 5 4 7 9 X 8 10 15 13 18 21 Y 11 19 17 12 14 16 20 22 Dotted diagonal line corresponds to the average density of genes in the genome, approximately 6. 7 protein-coding genes per megabase (Mb).一组个体的DNA序列与参考序列相似但不完全相同,部分个体存在单核苷酸改变,另一个个体则有两个碱基的小缺失。1500 1000 500 0 基因数量 0 50 100 150 200 250 基因富集染色体 基因贫乏染色体 基因组平均值:约6.7基因/Mb 染色体大小(Mb) 1 2 3 6 5 4 7 9 X 8 10 15 13 18 21 Y 11 19 17 12 14 16 20 22 虚线对角线代表基因组中基因的平均密度,约每兆碱基(Mb)6.7个蛋白质编码基因。
Chromosomes that are relatively gene rich are above the diagonal and trend to the upper left.相对基因富集的染色体位于对角线上方,并趋向于左上角。
Chromosomes that are relatively gene poor are below the diagonal and trend to the lower right.相对基因贫乏的染色体位于对角线下方,并趋向于右下角。
Available from v 37.)可从v37获取。
11/42
that influence or determine patterns of gene expression during development or in tissues were believed to account for on…
Ch2 — Segment 11
that influence or determine patterns of gene expression during development or in tissues were believed to account for only approximately 5% of additional sequence, although more recent analyses of chromatin characteristics suggest that a much higher proportion of the genome may provide signals that are relevant to genome functions.在发育过程中或在组织中影响或决定基因表达模式的序列,曾被认为仅占额外序列的大约5%,尽管近期对染色质特征的分析表明,基因组中更高比例的序列可能提供与基因组功能相关的信号。
Only approximately half of the total linear length of the genome consists of so-called singlecopy or unique DNA, that is, DNA whose linear order of specific nucleotides is represented only once (or at most a few times) around the entire genome.基因组总线性长度中只有大约一半由所谓的单拷贝或独有DNA组成,即其特定核苷酸线性顺序在整个基因组中仅出现一次(或至多几次)的DNA。
This concept may appear surprising to some, given that there are only four different nucleotides in DNA.鉴于DNA中仅有四种不同的核苷酸,这一概念可能令一些人感到惊讶。
But consider even a tiny stretch of the genome that is only 10 bases long; with four types of bases, there are over 1 million possible sequences.但试想基因组中仅10个碱基长的一小段;有四种碱基类型,可能序列超过100万种。
Although the order of bases in the genome is not entirely random, any particular 16-base sequence would be predicted by chance alone to appear only once in any given genome.尽管基因组中碱基顺序并非完全随机,任何特定的16碱基序列仅凭偶然性预测在任一给定基因组中只会出现一次。
The rest of the genome consists of several classes of repetitive DNA and includes DNA whose nucleotide sequence is repeated, either identically or with some variation, hundreds to millions of times in the genome.基因组的其余部分由几类重复DNA组成,包括其核苷酸序列在基因组中相同或略有变异地重复数百至数百万次的DNA。
Whereas most (but not all) of the estimated 20,000 protein-coding genes in the genome (see box earlier in this chapter) are represented in single-copy DNA, sequences in the repetitive DNA fraction contribute to maintaining chromosome structure and are an important source of variation between different individuals; some of this variation can predispose to pathologic events in the genome, as we will see in Chapters 5 and 6.尽管基因组中估计的20,000个蛋白质编码基因(见本章前面的框)大多数(但非全部)存在于单拷贝DNA中,但重复DNA组分中的序列有助于维持染色体结构,并且是不同个体间变异的重要来源;其中一些变异可能易导致基因组中的病理事件,我们将在第5章和第6章中看到。
Single-Copy DNA Sequences Although single-copy DNA makes up at least half of the DNA in the genome, much of its function remains a mystery because, as mentioned, sequences actually encoding proteins (i. e., the coding portion of genes) constitute only a small proportion of all the single-copy DNA.单拷贝DNA序列 尽管单拷贝DNA占基因组DNA至少一半,但其大部分功能仍是个谜,因为如前所述,实际编码蛋白质的序列(即基因的编码部分)仅占所有单拷贝DNA的一小部分。
Most single-copy DNA is found in short stretches (several kilobase pairs or less), interspersed with members of various repetitive DNA families.大多数单拷贝DNA以短片段(几千碱基对或更短)形式存在,散布于各种重复DNA家族成员之间。
The organization of genes in single-copy DNA is addressed in depth in Chapter 3.单拷贝DNA中基因的组织结构将在第3章深入探讨。
Repetitive DNA Sequences Several different categories of repetitive DNA are recognized.重复DNA序列 可识别出几种不同类别的重复DNA。
A useful distinguishing feature is whether the repeated sequences (“repeats”) are clustered in one or a few locations or whether they are interspersed with single-copy sequences along the chromosome.一个有用的区分特征在于重复序列(“重复”)是聚集在一个或几个位置,还是沿着染色体与单拷贝序列相互散布。
Clustered repeated sequences constitute an estimated 10% to 15% of the genome and consist of arrays of various short repeats organized in tandem in a head-to-tail fashion.聚集的重复序列估计占基因组的10%至15%,由各种短重复序列以头尾相连的串联方式排列成的阵列组成。
The different types of such tandem repeats are collectively called satellite DNAs, so named because many of the original tandem repeat families could be separated by biochemical methods from the bulk of the genome as distinct (“satellite”) fractions of DNA.不同类型的此类串联重复统称为卫星DNA,之所以如此命名,是因为许多原始的串联重复家族可通过生物化学方法与基因组主体分离,成为独特的(“卫星”)DNA组分。
Tandem repeat families vary with regard to their location in the genome and the nature of sequences that make up the array.串联重复家族在其基因组中的位置以及构成阵列的序列性质方面有所不同。
In general, such arrays can stretch several million base pairs or more in length and constitute up to several percent of the DNA content of an individual human chromosome.一般而言,此类阵列可延伸至数百万碱基对或更长,占单个人类染色体DNA含量的百分之几。
Some tandem repeat sequences are important as tools that are useful in clinical cytogenetic analysis (see Chapter 5).某些串联重复序列在临床细胞遗传学分析中作为有用工具具有重要意义(见第5章)。
Long arrays of repeats based on repetitions (with some variation) of a short sequence such as a pentanucleotide are found in large genetically inert regions on chromosomes 1, 9, and 16 and make up more than half of the Y chromosome.基于短序列(如五核苷酸)重复(略有变异)的长重复阵列,存在于1、9和16号染色体上大的遗传惰性区域,并构成Y染色体的一半以上。
Other tandem repeat families are based on somewhat longer basic repeats.其他串联重复家族基于稍长的基本重复单元。
For example, the α-satellite family of DNA is composed of tandem arrays of an approximately 171-bp unit, found at the centromere of each human chromosome, which is critical for attachment of chromosomes to microtubules of the spindle apparatus during cell division.例如,α-卫星DNA家族由约171 bp单元的串联阵列组成,存在于每条人类染色体的着丝粒处,对细胞分裂期间染色体与纺锤体微管的附着至关重要。
In addition to tandem repeat DNAs, another major class of repetitive DNA in the genome consists of related sequences that are dispersed throughout the genome rather than clustered in one or a few locations.除串联重复DNA外,基因组中另一大类重复DNA由相关序列组成,这些序列散布于整个基因组,而非聚集在一个或几个位置。
Although many DNA families meet this general description, two in particular warrant discussion because together they make up a significant proportion of the genome and because they have been implicated in genetic conditions.尽管许多DNA家族符合这一总体描述,但有两个家族尤其值得讨论,因为它们共同构成了基因组的重要部分,并且与遗传疾病相关。
Among the best-studied dispersed repetitive elements are those belonging to the so-called Alu family.研究最充分的散布重复元件之一是所谓的Alu家族成员。
The members of this family are approximately 300 bp in length and are related to each other although not identical in DNA sequence.该家族成员长度约为300 bp,彼此在DNA序列上相关但不完全相同。
In total, there are more than 1 million Alu family members in the genome, making up at least 10% of human DNA.总计,基因组中有超过100万个Alu家族成员,至少占人类DNA的10%。
A second major dispersed repetitive DNA family is called the long interspersed nuclear element (LINE, sometimes called L1) family.第二大主要的散布重复DNA家族称为长散在核元件(LINE,有时称为L1)家族。
LINEs are up to 6 kb in length and are found in approximately 850,000 copies per genome, accounting for nearly 20% of the genome.LINE长度可达6 kb,每个基因组中约含850,000个拷贝,占基因组近20%。
Both of these families are plentiful in some regions of the genome but relatively sparse in others—regions rich in GC content tend to be enriched in Alu elements but depleted of LINE sequences, whereas the opposite is true of more AT-rich regions of the genome.这两个家族在基因组的某些区域丰富,而在其他区域相对稀少——富含GC的区域往往富含Alu元件但缺少LINE序列,而富含AT的区域则相反。
Repetitive DNA and Disease.重复DNA与疾病。
Both Alu and LINE sequences have been implicated as the cause of mutations in hereditary disease.Alu和LINE序列均已被证实是遗传病中突变的原因。
At least a few copies of the LINE and Alu families generate copies of themselves that can integrate elsewhere in the genome, occasionally causing insertional inactivation of a medically important gene.至少有少数LINE和Alu家族成员可产生自身拷贝并整合到基因组的其他位置,偶尔导致具有医学重要性的基因发生插入失活。
The frequency of such events causing genetic disease in humans is unknown, but they may account for as many as 1 in 500 mutations.此类事件引发人类遗传病的频率尚不清楚,但它们可能占突变的1/500之多。
In addition, aberrant recombination events between different LINE repeats or Alu repeats can also be a cause of variants in some genetic diseases.此外,不同LINE重复或Alu重复之间的异常重组事件也可能成为某些遗传病变异的原因。
12/42
Introduction to the Human Genome 13 An important additional type of repetitive DNA found in many different locations aro…
Ch2 — Segment 12
Introduction to the Human Genome 13 An important additional type of repetitive DNA found in many different locations around the genome includes sequences that are duplicated, often with extraordinarily high sequence conservation.人类基因组导论13 一种重要的额外类型的重复DNA存在于基因组中许多不同位置,包括重复序列,这些序列通常具有极高的序列保守性。
Duplications involving substantial segments of a chromosome, called segmental duplications, can span hundreds of kilobase pairs and account for at least 5% of the genome.涉及染色体大片段的重复,称为节段性重复,可跨越数百千碱基对,至少占基因组的5%。
When the duplicated regions contain genes, genomic rearrangements involving the duplicated sequences can result in the deletion of the region (and the genes) between the copies and thus give rise to disease (see Chapters 5 and 6).当重复区域含有基因时,涉及重复序列的基因组重排可导致拷贝间区域(及基因)的缺失,从而引发疾病(见第5章和第6章)。
VARIATION IN THE HUMAN GENOME With completion of the reference human genome sequence, much attention has turned to the discovery and cataloging of variation in sequence among different individuals (including both healthy individuals and those with various diseases) and among different populations around the globe.人类基因组的变异 随着参考人类基因组序列的完成,大量注意力转向发现和编录不同个体(包括健康个体和患有各种疾病的个体)以及全球不同人群之间的序列变异。
As we will explore in much more detail in Chapter 4, there are many tens of millions of common sequence variants that are seen at significant frequency in one or more populations; any given individual carries millions of these sequence variants.正如我们将在第4章中更详细探讨的那样,存在数千万个常见序列变异,这些变异在一个或多个人群中以显著频率出现;任何特定个体都携带数百万个此类序列变异。
In addition, there are countless very rare variants, many of which probably exist in only a single or a few individuals.此外,还有无数极为罕见的变异,其中许多可能仅存在于单个或少数个体中。
In fact, given the number of individuals in our species, essentially each and every base pair in the human genome is expected to vary in someone somewhere around the globe.事实上,鉴于我们物种的个体数量,人类基因组中的每个碱基对预计都会在全球某处的某个个体中存在变异。
It is for this reason that the original human genome sequence is considered a reference sequence for our species, and not one that is actually identical to any individual’s genome.正是由于这个原因,最初的人类基因组序列被视为我们物种的参考序列,而不是实际与任何个体基因组相同的序列。
Early estimates were that any two randomly selected individuals would have sequences that are 99. 9% identical or, put another way, that an individual genome would carry two different versions (alleles) of the human genome sequence at some 3 to 5 million positions, with different bases (e. g., a T or a G) at the maternally and paternally inherited copies of that particular sequence position .早期估计认为,任意两个随机选择的个体的序列有99.9%的相同性,或者换句话说,一个个体会在约300万至500万个位置上携带人类基因组序列的两个不同版本(等位基因),在这些特定序列位置的母系和父系遗传拷贝上具有不同的碱基(例如T或G)。
Although many of these allelic differences involve simply one nucleotide, much of the variation consists of insertions or deletions of (usually) short sequence stretches, variation in the number of copies of repeated elements (including genes), or inversions in the order of sequences at a particular position (locus) in the genome (see Chapter 4).尽管许多这些等位基因差异仅仅涉及一个核苷酸,但大部分变异包括(通常)短序列片段的插入或缺失、重复元件(包括基因)拷贝数的变异,或基因组中特定位置(位点)序列顺序的倒位(见第4章)。
These copy number variations account for at least 12 million bp of difference between any two unrelated individuals.这些拷贝数变异占任何两个无亲缘关系个体之间至少1200万碱基对的差异。
The total amount of the genome involved in such variation is now known to be substantially more than originally estimated and approaches 2% between any two randomly selected individuals.目前已知参与此类变异的基因组总量远高于最初估计,在任意两个随机选择的个体之间接近2%。
As will be addressed in future chapters, any and all of these types of variation can influence biologic function and thus must be accounted for in any attempt to understand the contribution of genetics and genomics to human health.正如将在后续章节中论述的那样,这些类型的任何和所有变异都可能影响生物学功能,因此在任何试图理解遗传学和基因组学对人类健康的贡献时,都必须考虑这些变异。
TRANSMISSION OF THE GENOME The chromosomal basis of heredity lies in the copying of the genome and its transmission from a cell to its progeny during typical cell division and from one generation to the next during reproduction, when single copies of the genome from each parent come together in a new embryo.基因组的传递 遗传的染色体基础在于基因组的复制及其在典型细胞分裂过程中从细胞传递到子代,以及在繁殖过程中从一代传递到下一代,此时来自每个亲本的单个基因组拷贝在一个新胚胎中结合。
To achieve these related but distinct forms of genome inheritance, there are two kinds of cell division, mitosis and meiosis.为了实现这些相关但不同的基因组遗传形式,存在两种细胞分裂:有丝分裂和减数分裂。
Mitosis is ordinary somatic cell division by which the body grows, differentiates, and effects tissue regeneration.有丝分裂是普通的体细胞分裂,身体通过其生长、分化并实现组织再生。
Mitotic division normally results in two daughter cells, each with chromosomes and genes identical to those of the parent cell.有丝分裂通常产生两个子细胞,每个子细胞的染色体和基因与亲代细胞相同。
There may be dozens or even hundreds of successive mitoses in a lineage of somatic cells.在体细胞谱系中可能发生数十次甚至数百次连续的有丝分裂。
In contrast, meiosis occurs only in cells of the germline.相比之下,减数分裂仅发生在生殖细胞系中。
Meiosis results in the formation of reproductive cells (gametes), each of which has only 23 chromosomes—one of each kind of autosome and either an X or a Y.减数分裂导致生殖细胞(配子)的形成,每个配子仅有23条染色体——每种常染色体一条,以及一条X或Y染色体。
Thus, whereas somatic cells have the diploid (diploos, “double”) or the 2n chromosome complement (i. e., 46 chromosomes), gametes have the haploid (haploos, “single”) or the n complement (i. e., 23 chromosomes).因此,体细胞具有二倍体(diploos,“双倍”)或2n染色体组成(即46条染色体),而配子具有单倍体(haploos,“单倍”)或n染色体组成(即23条染色体)。
Abnormalities of chromosome number or structure, which are usually clinically significant, can arise either in somatic cells or in cells of the germline by errors in cell division.染色体数目或结构的异常(通常具有临床意义)可由细胞分裂错误发生在体细胞或生殖细胞系细胞中。
The Cell Cycle A human being begins life as a fertilized ovum (zygote), a diploid cell from which all the cells of the body (estimated to be ~100 trillion in number) are derived by a series of dozens or even hundreds of mitoses.细胞周期 人类生命始于受精卵(合子),这是一个二倍体细胞,身体的所有细胞(数量估计约100万亿)通过一系列数十次甚至数百次有丝分裂由此产生。
Mitosis is obviously crucial for growth and differentiation, but it takes up only a small part of the life cycle of a cell.有丝分裂显然对生长和分化至关重要,但它只占细胞生命周期的一小部分。
The period between two successive mitoses is called interphase, the state in which most of the life of a cell is spent.两次连续有丝分裂之间的时期称为间期,这是细胞度过大部分生命的状态。
Immediately after mitosis, the cell enters a phase, called G1, in which there is no DNA synthesis .有丝分裂后立即进入称为G1期的阶段,在此阶段没有DNA合成。
Some cells pass through this stage in hours; others spend a long time, days or years, in G1.有些细胞在数小时内通过此阶段;其他细胞则在G1期停留
In fact, some cell types, such as neurons and red blood cells, do not divide at all once they are fully differentiated; rather, they are permanently arrested in a distinct phase known as G0 (“G zero”).[TL:missing]
Other cells, such as liver cells, may enter G0 but, after organ damage, return to G1 and continue through the cell cycle.[TL:missing]
The cell cycle is governed by a series of checkpoints that determine the timing of each step in mitosis.[TL:missing]
In addition, checkpoints monitor and control the accuracy of DNA synthesis as well as the assembly and attachment of an elaborate network of microtubules that facilitate chromosome movement.[TL:missing]
If damage to the genome is detected, these mitotic checkpoints halt cell cycle progression until repairs are made or, if the[TL:missing]
13/42
damage is excessive, until the cell is instructed to die by programmed cell death (a process called apoptosis).
Ch2 — Segment 13
damage is excessive, until the cell is instructed to die by programmed cell death (a process called apoptosis).损伤过度,直至细胞被指令通过程序性细胞死亡(一种称为凋亡的过程)而死亡。
During G1, each cell contains one diploid copy of the genome.在G1期,每个细胞含有一份二倍体基因组拷贝。
As the process of cell division begins, the cell enters S phase, the stage of programmed DNA synthesis, ultimately leading to the precise replication of each chromosome’s DNA.随着细胞分裂过程的开始,细胞进入S期,即程序性DNA合成阶段,最终导致每条染色体DNA的精确复制。
During this stage, each chromosome, which in G1 has been a single DNA molecule, is duplicated and consists of two sister chromatids , each of which contains an identical copy of the original linear DNA double helix.在此阶段,每条在G1期时为单条DNA分子的染色体被复制,并由两条姐妹染色单体组成,每条姐妹染色单体包含原始线性DNA双螺旋的相同拷贝。
The two sister chromatids are held together physically at the centromere, a region of DNA that associates with a number of specific proteins to form the kinetochore.两条姐妹染色单体在着丝粒处物理连接,着丝粒是一段与多种特定蛋白质结合形成动粒的DNA区域。
This complex structure serves to attach each chromosome to the microtubules of the mitotic spindle and to govern chromosome movement during mitosis.这一复杂结构用于将每条染色体附着于有丝分裂纺锤体的微管上,并调控有丝分裂期间的染色体运动。
DNA synthesis during S phase is not synchronous throughout all chromosomes or even within a single chromosome; rather, along each chromosome, it begins at hundreds to thousands of sites, called origins of DNA replication.S期的DNA合成并非在所有染色体上同步进行,甚至在同一染色体内部也不同步;相反,沿每条染色体,合成起始于数百至数千个称为DNA复制起始点的位点。
Individual chromosome segments have their own characteristic time of replication during the 6- to 8-hour S phase.在持续6至8小时的S期中,个别染色体片段有其特征性的复制时间。
The ends of each chromosome (or chromatid) are marked by telomeres, which consist of specialized repetitive DNA sequences that ensure the integrity of the chromosome during cell division.每条染色体(或染色单体)的末端由端粒标记,端粒由特化的重复DNA序列组成,确保细胞分裂过程中染色体的完整性。
Correct maintenance of the ends of chromosomes requires a special enzyme called telomerase, which ensures that the very ends of each chromosome are replicated.染色体末端的正确维持需要一种称为端粒酶的特殊酶,该酶确保每条染色体的最末端得到复制。
The essential nature of these structural elements of chromosomes and their role in ensuring genome integrity is illustrated by a range of clinical conditions that result from defects in elements of the telomere or kinetochore or cell cycle machinery or from inaccurate replication of even small portions of the genome (see 3).染色体这些结构元件的基本性质及其在确保基因组完整性中的作用,可通过一系列临床病症加以说明,这些病症源于端粒、动粒或细胞周期机制元件的缺陷,或基因组即使小部分的不准确复制(见第3节)。
Some of these conditions will be presented in greater detail in subsequent chapters.其中某些病症将在后续章节中更详细地介绍。
By the end of S phase, the DNA content of the cell has doubled, and each cell now contains two copies of the diploid genome.到S期结束时,细胞的DNA含量已翻倍,每个细胞现含有两份二倍体基因组拷贝。
After S phase, the cell enters a brief stage called G2.S期之后,细胞进入一个称为G2期的短暂阶段。
Throughout the whole cell cycle, the cell gradually enlarges, eventually doubling its total mass before the next mitosis.在整个细胞周期中,细胞逐渐增大,最终在下一次有丝分裂前使其总质量加倍。
G2 is ended by mitosis, which begins when individual chromosomes begin to condense and become visible under the microscope as thin, extended threads, a process that is considered in greater detail in the following section.G2期由有丝分裂终止,有丝分裂始于每条染色体开始凝缩并在显微镜下显现为细长丝状物,这一过程将在下一节中更详细地讨论。
The G1, S, and G2 phases together constitute interphase.G1期、S期和G2期共同构成间期。
In typical dividing human cells, the three phases take a total of 16 to 24 hours, whereas mitosis lasts only 1 to 2 hours .在典型的分裂人类细胞中,这三个阶段总共需要16至24小时,而有丝分裂仅持续1至2小时。
There is great variation, however, in the length of the cell cycle, which ranges from a few hours in rapidly dividing cells, such as those of the dermis of the skin or the intestinal mucosa, to months in other cell types.然而,细胞周期的长度存在很大差异,从快速分裂细胞(如皮肤真皮或肠黏膜细胞)的几小时到其他细胞类型的数月不等。
Mitosis During the mitotic phase of the cell cycle, an elaborate apparatus ensures that each of the two daughter cells receives a complete set of genetic information.有丝分裂 在细胞周期的有丝分裂阶段,一个精细的装置确保两个子细胞各获得一套完整的遗传信息。
This result is achieved by a mechanism that distributes one chromatid of each chromosome to each daughter cell .这一结果通过一种机制实现,该机制将每条染色体的一个染色单体分配到每个子细胞。
The process of distributing a copy of each chromosome to each daughter cell is called chromosome segregation.将每条染色体的一个拷贝分配到每个子细胞的过程称为染色体分离。
The importance of this process for normal Sister chromatids G1 (10–12 hr) S (6–8 hr) G2 (2–4 hr) M Telomere Telomere Centromere The telomeres, the centromere, and sister chromatids are indicated..该过程对正常姐妹染色单体G1(10–12小时)S(6–8小时)G2(2–4小时)M端粒 端粒 着丝粒 的重要性:端粒、着丝粒和姐妹染色单体已标明。
CLINICAL CONSEQUENCES OF ABNORMALITIES AND VARIATION IN CHROMOSOME STRUCTURE AND MECHANICS Medically relevant conditions arising from abnormal structure or function of chromosomal elements during cell division include the following: A broad spectrum of congenital abnormalities in children with inherited defects in genes encoding key components of the mitotic spindle checkpoint at the kinetochore A range of birth defects and developmental disorders due to anomalous segregation of chromosomes with multiple or missing centromeres (see Chapter 6) A variety of cancers associated with overreplication (amplification) or altered timing of replication of specific regions of the genome in S phase (see Chapter 16) Roberts syndrome of short stature, limb shortening, and microcephaly in children with abnormalities of a gene required for proper sister chromatid alignment and cohesion in S phase Premature ovarian failure as a major cause of female infertility due to deleterious variants in a meiosis-specific gene required for correct sister chromatid cohesion The so-called telomere syndromes, a number of degenerative disorders presenting from childhood to adulthood in patients with abnormal telomere shortening due to defects in components of telomerase (see Case 49) At the other end of the spectrum, common gene variants that correlate with the number of copies of the repeats at telomeres and with life expectancy and longevity染色体结构及力学异常与变异的临床后果 由细胞分裂过程中染色体元件结构或功能异常引起的医学相关病症包括以下内容:儿童中因编码动粒处有丝分裂纺锤体检查点关键成分的基因存在遗传缺陷而导致的广泛先天异常;因具有多个或缺失着丝粒的染色体异常分离而导致的一系列出生缺陷和发育障碍(见第6章);与S期特定基因组区域过度复制(扩增)或复制时间改变相关的多种癌症(见第16章);表现为身材矮小、肢体缩短和小头畸形的罗伯茨综合征,患儿存在S期姐妹染色单体正确排列和黏连所需基因的异常;因减数分裂特异性基因(负责正确姐妹染色单体黏连)的有害变异导致卵巢早衰,这是女性不孕的主要原因;所谓的端粒综合征,即一系列从儿童期至成年期发病的退行性疾病,患者因端粒酶组分缺陷而出现异常端粒缩短(见案例49);另一方面,常见的基因变异与端粒重复序列拷贝数及预期寿命、长寿相关联。
14/42
Introduction to the Human Genome 15 cell growth is illustrated by the observation that many tumors are invariably charac…
Ch2 — Segment 14
Introduction to the Human Genome 15 cell growth is illustrated by the observation that many tumors are invariably characterized by a state of genetic imbalance resulting from mitotic errors in the distribution of chromosomes to daughter cells.人类基因组导论15 细胞生长通过以下观察得以说明:许多肿瘤无一例外地以遗传失衡状态为特征,这种失衡源于染色体向子细胞分配过程中的有丝分裂错误。
The process of mitosis is continuous, but five stages are distinguished: prophase, prometaphase, metaphase, anaphase, and telophase.有丝分裂过程是连续的,但可区分为五个阶段:前期、前中期、中期、后期和末期。
This stage is marked by gradual condensation of the chromosomes, formation of the mitotic spindle, and formation of a pair of centrosomes, from which microtubules radiate and eventually take up positions at the poles of the cell.此阶段的特征为染色体的逐渐凝缩、有丝分裂纺锤体的形成以及一对中心体的形成,微管从中心体辐射发出并最终定位于细胞两极。
Prometaphase.前中期。
Here, the nuclear membrane dissolves, allowing the chromosomes to disperse within the cell and to attach, by their kinetochores, to microtubules of the mitotic spindle.在此阶段,核膜溶解,使染色体分散于细胞内,并通过其动粒附着于有丝分裂纺锤体的微管上。
Metaphase.中期。
At this stage, the chromosomes are maximally condensed and line up at the equatorial plane of the cell.在此阶段,染色体最大限度地凝缩,并排列于细胞的赤道板上。
The chromosomes separate at the centromere, and the sister chromatids of each chromosome now become independent daughter chromosomes, which move to opposite poles of the cell.染色体在着丝粒处分离,每条染色体的姐妹染色单体随即成为独立的子染色体,并移向细胞的两极。
Telophase.末期。
Now, the chromosomes begin to decondense from their highly contracted state, and a nuclear membrane begins to re-form around each of the two daughter nuclei, which resume their interphase appearance.此时,染色体开始从其高度凝缩状态解凝缩,并在两个子核周围重新形成核膜,子核恢复其间期形态。
To complete the process of cell division, the cytoplasm cleaves by a process known as cytokinesis.为完成细胞分裂过程,胞质通过称为胞质分裂的过程发生分裂。
There is an important difference between a cell entering mitosis and one that has just completed the process.进入有丝分裂的细胞与刚刚完成有丝分裂的细胞之间存在一个重要区别。
A cell in G2 has a fully replicated genome (i. e., a 4n complement of DNA), and each chromosome consists of a pair of sister chromatids.处于G2期的细胞拥有完全复制的基因组(即4n DNA含量),每条染色体由一对姐妹染色单体组成。
In contrast, after mitosis, the chromosomes of each daughter cell have only one copy of the genome.相反,在有丝分裂之后,每个子细胞的染色体仅含一份基因组拷贝。
This copy will not be duplicated until a daughter cell in its turn reaches the S phase of the next cell cycle .这一拷贝将不会被复制,直到该子细胞依次进入下一个细胞周期的S期。
The entire process of mitosis thus ensures the orderly duplication and distribution of the genome through successive cell divisions.因此,有丝分裂的整个过程确保了基因组在连续细胞分裂中的有序复制和分配。
The Human Karyotype The condensed chromosomes of a dividing human cell are most readily analyzed at metaphase or prometaphase.人类核型 正在分裂的人类细胞中的凝缩染色体在中期或前中期最易于分析。
At these stages, the chromosomes are visible under the microscope as a so-called chromosome spread; each chromosome consists of its sister chromatids, although in most chromosome preparations, the two chromatids are held together so tightly that they are rarely visible as separate entities.在这些阶段,染色体在显微镜下可见,形成所谓的染色体铺展;每条染色体由其姐妹染色单体组成,尽管在大多数染色体标本制备中,两条染色单体紧密贴合,很少能作为独立实体被观察到。
As stated earlier, there are 24 different types of human chromosomes, each of which can be distinguished cytologically by a combination of overall length, location of the centromere, and sequence content, the latter reflected by various staining methods.如前所述,人类染色体有24种不同类型,每种均可通过总长度、着丝粒位置以及序列内容(后者通过多种染色方法反映)的组合在细胞学上加以区分。
The centromere is apparent as a primary constriction, a narrowing or pinching-in of the sister chromatids due to formation of the kinetochore.着丝粒明显表现为一个初级缢痕,即由于动粒形成而导致的姐妹染色单体变窄或内缩。
This is a recognizable cytogenetic Cells in G1 S phase Interphase Prophase Prometaphase Metaphase Anaphase Telophase Cytokinesis Decondensed chromatin Cell in G2 Onset of mitosis Centrosomes Microtubules Only two chromosome pairs are shown.这是一个可识别的细胞遗传学现象:处于G1期的细胞、S期、间期、前期、前中期、中期、后期、末期、胞质分裂、解凝缩染色质、处于G2期的细胞、有丝分裂起始、中心体、微管;仅显示两对染色体。
For details, see text.详情请参阅正文。
15/42
landmark, dividing the chromosome into two arms, a short arm designated p (for petit) and a long arm designated q. metho…
Ch2 — Segment 15
landmark, dividing the chromosome into two arms, a short arm designated p (for petit) and a long arm designated q. method (also see Chapter 5).标志点,将染色体分为两臂,短臂称为p(源自petit),长臂称为q。方法(另见第5章)。
Each chromosome pair stains in a characteristic pattern of alternating light and dark bands (G bands) that correlates roughly with features of the underlying DNA sequence, such as base composition (i. e., the percentage of base pairs that are GC or AT) and the distribution of repetitive DNA elements.每对染色体以明暗交替的带型(G带)染色,该带型与底层DNA序列的特征大致相关,例如碱基组成(即GC或AT碱基对的百分比)以及重复DNA元件的分布。
With such banding techniques, all of the chromosomes can be individually distinguished, and the nature of many structural or numerical abnormalities can be determined, as we examine in greater detail in Chapters 5 and 6.借助这种显带技术,所有染色体均可被单独区分,并且许多结构或数量异常的性质得以确定,正如我们在第5章和第6章中更详细探讨的那样。
Although experts can often analyze metaphase chromosomes directly under the microscope, a common procedure is to cut out the chromosomes from a digital image or photomicrograph and arrange them in pairs in a standard classification .尽管专家通常可以在显微镜下直接分析中期染色体,但常用的操作是从数字图像或显微照片中剪切出染色体,并按标准分类将其配对排列。
The completed picture is called a karyotype.完成的图像被称为核型。
The word karyotype is also used to refer to the standard chromosome set of an individual (“a normal male karyotype”) or of a species (“the human karyotype”) and, as a verb, to the process of preparing such a standard figure (“to karyotype”).核型一词也用于指代个体(“正常男性核型”)或物种(“人类核型”)的标准染色体组,作为动词时指制备此类标准图像的过程(“进行核型分析”)。
Unlike the chromosomes seen in stained preparations under the microscope or in photographs, the chromosomes of living cells are fluid and dynamic structures.与在显微镜下或照片中看到的染色标本中的染色体不同,活细胞的染色体是流动且动态的结构。
During mitosis, the chromatin of each interphase chromosome condenses substantially .在有丝分裂过程中,每个间期染色体的染色质显著凝集。
When maximally condensed at metaphase, DNA in chromosomes is approximately 1/10,000 of its fully extended state.当在中期达到最大凝集时,染色体中的DNA约为其完全伸展状态的1/10,000。
When chromosomes are prepared to reveal bands (as in Figs. 2. 10 and 2. 11), as many as 1000 or more bands can be recognized in stained preparations of all the chromosomes.当染色体经处理以显示条带(如图2.10和2.11所示)时,在所有染色体的染色标本中可识别多达1000条或更多的条带。
Each cytogenetic band therefore contains as many as 50 or more genes, although the density of genes in the genome, as mentioned previously, is variable.因此,每条细胞遗传学条带包含多达50个或更多的基因,尽管如前所述,基因组中基因的密度是变化的。
Meiosis Meiosis, the process by which diploid cells give rise to haploid gametes, involves a type of cell division that is unique to germ cells.减数分裂 减数分裂是二倍体细胞产生单倍体配子的过程,涉及一种生殖细胞特有的细胞分裂类型。
In contrast to mitosis, meiosis consists of one round of DNA replication followed by two rounds of chromosome segregation and cell division (see meiosis I and meiosis II in .与有丝分裂不同,减数分裂包括一轮DNA复制,随后进行两轮染色体分离和细胞分裂(见减数分裂I和减数分裂II在...中)。
As outlined here and illustrated in occurs.如这里概述并在...中图示,发生。
In this process, as shown for one pair of chromosomes in[TL:missing]
16/42
The chromosomes are at the prometaphase stage of mitosis and are arranged in a standard classification, numbered 1 to 22…
Ch2 — Segment 16
The chromosomes are at the prometaphase stage of mitosis and are arranged in a standard classification, numbered 1 to 22 in order of length, with the X and Y chromosomes shown separately.染色体处于有丝分裂的前中期阶段,并按标准分类排列,按长度顺序编号为1至22,X和Y染色体单独显示。
Courtesy Stuart Schwartz, University Hospitals of Cleveland, Ohio.图片由俄亥俄州克利夫兰大学医院的Stuart Schwartz提供。
Interphase nucleus Decondensed chromatin Metaphase Decondensation as cell returns to interphase Condensation as mitosis begins Prophase chromosomes align themselves on the equatorial plane with their centromeres oriented toward different poles .间期核、去浓缩染色质、中期、细胞返回间期时的去浓缩、有丝分裂开始时的浓缩、前期染色体排列在赤道面上,其着丝粒朝向不同的极。
Anaphase of meiosis I again differs substantially from the corresponding stage of mitosis.减数分裂I的后期再次与有丝分裂的相应阶段有显著差异。
Here, it is the two members of each bivalent that move apart, not the sister chromatids (contrast .在这里,是每个二价体的两个成员分开,而不是姐妹染色单体(对比。
The homologous centromeres (with their attached sister chromatids) are drawn to opposite poles of the cell, a process termed disjunction.同源着丝粒(及其附着的姐妹染色单体)被拉向细胞的两极,这一过程称为分离。
Thus the chromosome number is halved, and each cellular product of meiosis I has the haploid chromosome number.因此,染色体数目减半,减数分裂I的每个细胞产物具有单倍体染色体数目。
The 23 pairs of homologous chromosomes assort independently of one another, and as a result the original paternal and maternal chromosome sets are sorted into random combinations.23对同源染色体彼此独立分配,因此原始父本和母本染色体组被随机组合。
The possible number of combinations of the 23 chromosome Chromosome replication Meiosis I Meiosis II Four haploid gametes Interphase Interphase Prophase I Metaphase I Anaphase I Metaphase II Meiosis I Meiosis II Anaphase II Gametes A single chromosome pair and a single crossover are shown, leading to formation of four distinct gametes.23条染色体的可能组合数:染色体复制、减数分裂I、减数分裂II、四个单倍体配子、间期、间期、前期I、中期I、后期I、中期II、减数分裂I、减数分裂II、后期II、配子。图中显示一对染色体和一个交叉,导致形成四个不同的配子。
The chromosomes replicate during interphase and begin to condense as the cell enters prophase of meiosis I.染色体在间期复制,并在细胞进入减数分裂I前期时开始浓缩。
In meiosis I, the chromosomes synapse and recombine.在减数分裂I中,染色体发生联会并重组。
A crossover is visible as the homologues align at metaphase I, with the centromeres oriented toward opposite poles.当同源染色体在中期I对齐时,可见交叉,着丝粒朝向相反的两极。
In anaphase I, the exchange of DNA between the homologues is apparent as the chromosomes are pulled to opposite poles.在后期I,当染色体被拉向两极时,同源染色体之间的DNA交换变得明显。
After completion of meiosis I and cytokinesis, meiosis II proceeds with a mitosis-like division.减数分裂I和胞质分裂完成后,减数分裂II以类似有丝分裂的分裂方式进行。
The sister kinetochores separate and move to opposite poles in anaphase II, yielding four haploid products.姐妹动粒在后期II分离并移向两极,产生四个单倍体产物。
17/42
Introduction to the Human Genome 19 pairs that can be present in the gametes is 223 (>8 million).
Ch2 — Segment 17
Introduction to the Human Genome 19 pairs that can be present in the gametes is 223 (>8 million).人类基因组导论:配子中可能存在的23对染色体组合数为2^23(>800万)。
Owing to the process of crossing over, however, the variation in the genetic material that is transmitted from parent to child is actually much greater than this (see 4).然而,由于交叉过程,从父母传递到子女的遗传物质变异实际上比这大得多(见第4点)。
As a result, each chromatid typically contains segments derived from each member of the original parental chromosome pair, as illustrated schematically in .因此,每条染色单体通常包含来自原始亲本染色体对中每个成员的片段,如图所示。
After telophase of meiosis I, the two haploid daughter cells enter meiotic interphase.在减数第一次分裂末期之后,两个单倍体子细胞进入减数分裂间期。
In contrast to mitosis, this interphase is brief, and meiosis II begins.与有丝分裂相反,此间期短暂,随后开始减数第二次分裂。
The notable point that distinguishes meiotic and mitotic interphase is that there is no S phase (i. e., no DNA synthesis and duplication of the genome) between the first and second meiotic divisions.区别减数分裂间期与有丝分裂间期的显著特点是:在第一次和第二次减数分裂之间没有S期(即没有DNA合成和基因组复制)。
Meiosis II is similar to an ordinary mitosis, except that the chromosome number is 23 instead of 46; the chromatids of each of the 23 chromosomes separate, and one chromatid of each chromosome passes to each daughter cell .减数第二次分裂与普通有丝分裂相似,区别在于染色体数目为23而非46;23条染色体各自的染色单体分离,每条染色体的一条染色单体进入每个子细胞。
However, as mentioned earlier, because of crossing over in meiosis I, the chromosomes of the resulting gametes are not identical .然而,如前所述,由于减数第一次分裂中的交叉,产生的配子染色体并不相同。
Grandpaternal DNA sequences Grandmaternal DNA sequences Paternal chromosomes Paternal chromosome inherited by Child 1 Paternal chromosome inherited by Child 2 In this example, representing the inheritance of sequences on a typical large chromosome, an individual has distinctive homologues, one containing sequences inherited from his father (blue) and one containing homologous sequences from his mother (purple).祖父DNA序列 祖母DNA序列 父本染色体 子1继承的父本染色体 子2继承的父本染色体 在此示例中,代表典型大染色体上序列的遗传,个体具有独特的同源染色体,一条包含来自其父亲(蓝色)的序列,另一条包含来自其母亲(紫色)的同源序列。
After meiosis in spermatogenesis, he transmits a single complete copy of that chromosome to his two offspring.在精子发生的减数分裂后,他将该染色体的一条完整拷贝传递给他的两个后代。
However, as a result of crossing over (arrows), the copy he transmits to each child consists of alternating segments of the two grandparental sequences.然而,由于交叉(箭头所示),他传递给每个孩子的拷贝由两个祖辈序列的交替片段组成。
Child 1 inherits a copy after two crossovers, whereas child 2 inherits a copy with three crossovers..子1继承经过两次交叉的拷贝,而子2继承经过三次交叉的拷贝。
GENETIC CONSEQUENCES AND MEDICAL RELEVANCE OF HOMOLOGOUS RECOMBINATION The take-home lesson of this portion of the chapter is a simple one: the genetic content of each gamete is unique because of random assortment of the parental chromosomes to shuffle the combination of sequence variants ­between chromosomes and because of homologous recombination to shuffle the combination of sequence variants within each and every chromosome.同源重组的遗传后果及医学意义 本章此部分的要点很简单:每个配子的遗传内容都是独特的,这是因为亲本染色体的随机分配可打乱染色体间序列变异的组合,同时同源重组可打乱每条染色体内部序列变异的组合。
This has significant consequences for patterns of genomic variation among and between different populations around the globe and for diagnosis and counseling of many common conditions with complex patterns of inheritance (see Chapters 9 and 10).这对全球不同
The amounts and patterns of meiotic recombination are determined by sequence variants in specific genes and at specific “hot spots” and differ between individuals, between the sexes, between families, and between populations (see Chapter 10).[TL:missing]
Because recombination involves the physical intertwining of the two homologues until the appropriate point during meiosis I, it is also critical for ensuring proper chromosome segregation during meiosis.[TL:missing]
Failure to recombine properly can lead to chromosome missegregation (nondisjunction) in meiosis I and is a frequent cause of pregnancy loss and of chromosome abnormalities such as Down syndrome (see Chapters 5 and 6).[TL:missing]
Major ongoing efforts to identify genes and their variants responsible for various medical conditions rely on tracking the inheritance of millions of sequence differences within families or the sharing of variants within groups of even unrelated individuals affected with a particular condition.[TL:missing]
The utility of this approach, which has uncovered thousands of gene-disease associations to date, depends on patterns of homologous recombination in meiosis (see Chapters 6, 11, and 12).[TL:missing]
Although homologous recombination is generally precise, areas of repetitive DNA in the genome and genes of variable copy number in the population are prone to occasional unequal crossing over during meiosis, leading to variations in clinically relevant traits such as drug response, to common disorders such as the thalassemias or autism, or to abnormalities of sexual differentiation (see Chapters 6, 11, and 12).[TL:missing]
Although homologous recombination is a normal and essential part of meiosis, it also occurs, albeit more rarely, in somatic cells.[TL:missing]
Anomalies in somatic recombination are one of the causes of genome instability in cancer (see Chapter 16).[TL:missing]
18/42
HUMAN GAMETOGENESIS AND FERTILIZATION The cells in the germline that undergo meiosis, primary spermatocytes or primary o…
Ch2 — Segment 18
HUMAN GAMETOGENESIS AND FERTILIZATION The cells in the germline that undergo meiosis, primary spermatocytes or primary oocytes, are derived from the zygote by a long series of mitoses before the onset of meiosis.人类配子发生与受精 生殖系中经历减数分裂的细胞,即初级精母细胞或初级卵母细胞,由受精卵在减数分裂开始前经过一系列有丝分裂衍生而来。
Male and female gametes have different histories, marked by different patterns of gene expression that reflect their developmental origin as an XY or XX embryo.雄性和雌性配子有不同的历史,表现为不同的基因表达模式,这些模式反映了它们作为XY或XX胚胎的发育起源。
The human primordial germ cells are recognizable by the fourth week of development outside the embryo proper, in the endoderm of the yolk sac.人类原始生殖细胞在发育第四周时可在胚胎本体外的卵黄囊内胚层中被识别。
From there, they migrate during the sixth week to the genital ridges and associate with somatic cells to form the primitive gonads, which soon differentiate into testes or ovaries, depending on the cells’ sex chromosome constitution (XY or XX), as we examine in greater detail in Chapter 6.从那里,它们在第六周迁移至生殖嵴,并与体细胞结合形成原始性腺,后者根据细胞的性染色体组成(XY或XX)很快分化成睾丸或卵巢,我们将在第6章中更详细地探讨。
Both spermatogenesis and oogenesis require meiosis but have important differences in detail and timing that may have clinical and genetic consequences for the offspring.精子发生和卵子发生都需要减数分裂,但在细节和时机上存在重要差异,这些差异可能对后代表现出临床和遗传学后果。
Female meiosis is initiated once, early during fetal life, in a limited number of cells.女性减数分裂在胎儿生命早期仅在有限数量的细胞中启动一次。
In contrast, male meiosis is initiated continuously in many cells from a dividing cell population throughout the adult life of a male.相比之下,男性减数分裂在男性整个成年期由分裂细胞群体中的许多细胞持续启动。
In the female, successive stages of meiosis take place over several decades—in the fetal ovary before the female in question is even born, in the oocyte near the time of ovulation in the sexually mature female, and after fertilization of the egg that can become that female’s offspring.在女性中,减数分裂的相继阶段在数十年间发生——在该女性出生前的胎儿卵巢中,在性成熟女性排卵时间附近的卵母细胞中,以及在成为该女性后代的卵子受精后。
Although postfertilization stages can be studied in vitro, access to the earlier stages is limited.尽管受精后阶段可以在体外研究,但早期阶段的获取受到限制。
Testicular material for the study of male meiosis is less difficult to obtain, inasmuch as testicular biopsy is included in the assessment of many men attending infertility clinics.用于研究男性减数分裂的睾丸材料相对容易获取,因为睾丸活检是许多就诊于不孕门诊的男性评估中的一部分。
Much remains to be learned about the cytogenetic, biochemical, and molecular mechanisms involved in normal meiosis and about the causes and consequences of meiotic irregularities.关于正常减数分裂中的细胞遗传学、生物化学和分子机制,以及减数分裂异常的成因和后果,仍有许多有待了解。
Spermatogenesis The stages of spermatogenesis are shown in are formed only after sexual maturity is reached.精子发生 精子发生的阶段显示于,仅在性成熟后才形成。
The last cell type in the developmental sequence is the primary spermatocyte, a diploid germ cell that undergoes meiosis I to form two haploid secondary spermatocytes.发育序列中的最后一种细胞类型是初级精母细胞,这是一种二倍体生殖细胞,经过减数分裂I形成两个单倍体次级精母细胞。
Secondary spermatocytes rapidly enter meiosis II, each forming two spermatids, which differentiate without further division into sperm.次级精母细胞迅速进入减数分裂II,每个形成两个精细胞,精细胞无需进一步分裂即可分化为精子。
In humans, the entire process takes approximately 64 days.在人类中,整个过程大约需要64天。
The enormous number of sperm produced, typically approximately 200 million per ejaculate and an estimated 1012 in a lifetime, requires several hundred successive mitoses.产生的大量精子,通常每次射精约2亿个,一生中估计达10^12个,需要进行数百次连续的有丝分裂。
Testis Spermatogonium 46,XY Primary spermatocyte 46,XY Secondary spermatocytes 23,X 23,X 23,X 23,Y 23,Y 23,X 23,X 23,Y 23,Y 23,Y Spermatids Meiosis I Meiosis II The sequence of events begins at puberty and takes approximately 64 days to be completed.睾丸 精原细胞 46,XY 初级精母细胞 46,XY 次级精母细胞 23,X 23,X 23,X 23,Y 23,Y 23,X 23,X 23,Y 23,Y 23,Y 精细胞 减数分裂I 减数分裂II 事件顺序始于青春期,大约需要64天完成。
The chromosome number (46 or j 23) and the sex chromosome constitution (X or Y) of each cell are shown. )显示每个细胞的染色体数目(46或23)和性染色体组成(X或Y)。
19/42
Introduction to the Human Genome 21 As discussed earlier, normal meiosis requires pairing of homologous chromosomes foll…
Ch2 — Segment 19
Introduction to the Human Genome 21 As discussed earlier, normal meiosis requires pairing of homologous chromosomes followed by recombination.人类基因组导论21 如前所述,正常减数分裂需要同源染色体配对,随后发生重组。
The autosomes and the X chromosomes in females present no unusual difficulties in this regard; but what of the X and Y chromosomes during spermatogenesis?女性的常染色体和X染色体在这方面没有异常困难;但精子发生过程中X和Y染色体又如何呢?
Although the X and Y chromosomes are different and are not homologues in a strict sense, they do have relatively short identical segments at the ends of their respective short arms (Xp and Yp) and long arms (Xq and Yq) (see Chapter 6).尽管X和Y染色体不同,严格意义上并非同源染色体,但它们在各自短臂(Xp和Yp)和长臂(Xq和Yq)的末端确实有相对较短的相同片段(参见第6章)。
Pairing and crossing over occurs in both regions during meiosis I.在减数分裂I期间,这两个区域均发生配对和交叉。
These homologous segments are called pseudoautosomal to reflect their autosomelike pairing and recombination behavior, despite being on different sex chromosomes.这些同源片段被称为假常染色体区,以反映它们类似常染色体的配对和重组行为,尽管位于不同的性染色体上。
Oogenesis Whereas spermatogenesis is initiated only at the time of puberty, oogenesis begins during a female’s development as a fetus .卵子发生 精子发生仅在青春期开始,而卵子发生始于女性胎儿发育期间。
The ova develop from oogonia, cells in the ovarian cortex that have descended from the primordial germ cells by a series of approximately 20 mitoses.卵子由卵原细胞发育而来,卵原细胞是卵巢皮质中的细胞,这些细胞通过约20次有丝分裂从原始生殖细胞衍生而来。
Each oogonium is the central cell in a developing follicle.每个卵原细胞是发育中卵泡的中央细胞。
By approximately the third month of fetal development, the oogonia of the embryo have begun to develop into primary oocytes, most of which have already entered prophase of meiosis I.大约在胎儿发育第三个月时,胚胎的卵原细胞已开始发育为初级卵母细胞,其中大多数已进入减数分裂I的前期。
The process of oogenesis is not synchronized, and both early and late stages coexist in the fetal ovary.卵子发生过程并非同步进行,早期和晚期阶段在胎儿卵巢中并存。
Although there are several million oocytes at the time of birth, most of these degenerate; the others remain arrested in prophase I for decades.尽管出生时有数百万个卵母细胞,但大部分会退化;其余的则停滞在前期I长达数十年。
Only approximately 400 eventually mature and are ovulated as part of a woman’s menstrual cycle.最终只有约400个卵母细胞成熟并作为女性月经周期的一部分被排出。
After a woman reaches sexual maturity, individual follicles begin to grow and mature, and a few (on average one per month) are ovulated.女性达到性成熟后,单个卵泡开始生长和成熟,少数(平均每月一个)被排出。
Just before ovulation, the oocyte rapidly completes meiosis I, dividing in such a way that one cell becomes the secondary oocyte (an egg or ovum), containing most of the cytoplasm with its organelles; the other cell becomes the first polar body .就在排卵前,卵母细胞迅速完成减数分裂I,以这样一种方式分裂:一个细胞成为次级卵母细胞(卵子),含有大部分细胞质及其细胞器;另一个细胞成为第一极体。
Meiosis II begins promptly and proceeds to the metaphase stage during ovulation, where it halts again, only to be completed if fertilization occurs.减数分裂II立即开始,并在排卵期间进行到中期阶段,在此再次停滞,只有在受精发生时才能完成。
Fertilization Fertilization of the egg usually takes place in the fallopian tube within a day or so of ovulation.受精 卵子的受精通常在排卵后一天左右在输卵管内发生。
Although many sperm may be present, the penetration of a single sperm into the ovum sets up a series of biochemical events that usually prevent the entry of other sperm.尽管可能有许多精子存在,但单个精子进入卵子会引发一系列生化事件,通常阻止其他精子进入。
Fertilization is followed by the completion of meiosis II, with the formation of the second polar body .受精后,减数分裂II完成,形成第二极体。
The chromosomes of the now-fertilized egg and sperm form pronuclei, each surrounded by its own Ovary Primary oocyte in follicle Suspended in prophase I until sexual maturity Secondary oocyte Meiotic spindle 1st polar body Ovulation Fertilization 2nd polar body Mature ovum Sperm Meiosis I Meiosis II The primary oocytes are formed prenatally and remain suspended in prophase of meiosis I for years until the onset of puberty.受精卵和精子的染色体形成原核,每个原核被其自身的卵巢、卵泡中的初级卵母细胞、停滞在减数分裂Ⅰ前期直至性成熟、次级卵母细胞、减数分裂纺锤体、第一极体、排卵、受精、第二极体、成熟卵子、精子、减数分裂Ⅰ、减数分裂Ⅱ所包围;初级卵母细胞在出生前形成,并停滞在减数分裂Ⅰ前期多年直至青春期开始。
An oocyte completes meiosis I as its follicle matures, resulting in a secondary oocyte and the first polar body.卵母细胞在其卵泡成熟时完成减数分裂I,产生一个次级卵母细胞和第一极体。
After ovulation, each oocyte continues to metaphase of meiosis II.排卵后,每个卵母细胞继续进行至减数分裂II的中期。
Meiosis II is completed only if fertilization occurs, resulting in a fertilized mature ovum and the second polar body.减数分裂II仅在受精发生时完成,产生一个受精的成熟卵子和第二极体。
20/42
nuclear membrane.
Ch2 — Segment 20
nuclear membrane.核膜。
It is only upon replication of the parental genomes after fertilization that the two haploid genomes become one diploid genome within a shared nucleus.只有在受精后亲本基因组复制时,两个单倍体基因组才在一个共享的细胞核内成为一个二倍体基因组。
The diploid zygote divides by mitosis to form two diploid daughter cells, the first in the series of cell divisions that initiate the process of embryonic development (see Chapter 15).二倍体合子通过有丝分裂形成两个二倍体子细胞,这是启动胚胎发育过程的一系列细胞分裂中的第一次(见第15章)。
Although development begins at the time of conception, with the formation of the zygote, in clinical medicine the stage and duration of pregnancy are usually measured as the “menstrual age,” dating from the beginning of the mother’s last menstrual period, typically approximately 14 days before conception.尽管发育始于受孕时刻,即合子形成之时,但在临床医学中,妊娠的阶段和持续时间通常以“月经龄”来衡量,从母亲末次月经开始算起,通常大约在受孕前14天。
MEDICAL RELEVANCE OF MITOSIS AND MEIOSIS The biologic significance of mitosis and meiosis lies in ensuring the constancy of chromosome number—and thus the integrity of the genome—from one cell to its progeny and from one generation to the next.有丝分裂和减数分裂的医学意义 有丝分裂和减数分裂的生物学意义在于确保染色体数目的恒定性——从而保证基因组的完整性——从一个细胞到其子代,以及从一代到下一代。
The medical relevance of these processes lies in errors of one or the other mechanism of cell division, leading to the formation of an individual or of a cell lineage with an abnormal number of chromosomes and thus an abnormal dosage of genomic material.这些过程的医学意义在于细胞分裂的某种机制出现错误,导致个体或细胞谱系形成异常数目的染色体,从而产生异常剂量的基因组物质。
As we see in detail in Chapter 5, meiotic nondisjunction, particularly in oogenesis, is the most common mutational mechanism in our species, responsible for chromosomally abnormal fetuses in at least several percent of all recognized pregnancies.正如我们在第5章中详细看到的,减数分裂不分离,尤其是在卵子发生中,是人类最常见的突变机制,导致在所有已确认的妊娠中至少有百分之几的染色体异常胎儿。
Among pregnancies that survive to term, chromosome abnormalities are a leading cause of developmental defects, failure to thrive in the newborn period, and intellectual disability.在存活至足月的妊娠中,染色体异常是发育缺陷、新生儿期生长迟缓和智力残疾的主要原因。
Mitotic nondisjunction in somatic cells also contributes to genetic disease.体细胞中的有丝分裂不分离也会导致遗传病。
Nondisjunction soon after fertilization, either in the developing embryo or in extraembryonic tissues like the placenta, leads to chromosomal mosaicism that can underlie some medical conditions, such as a proportion of patients with Down syndrome.受精后不久发生的不分离,无论是在发育中的胚胎还是胎盘等胚外组织中,都会导致染色体嵌合体,这可能是某些医学状况的基础,例如一部分唐氏综合征患者。
Further, abnormal chromosome segregation in rapidly dividing tissues, such as in cells of the colon, is frequently a step in the development of chromosomally abnormal tumors, and thus evaluation of chromosome and genome balance is an important diagnostic and prognostic test in many cancers.此外,在快速分裂的组织(如结肠细胞)中染色体异常分离通常是染色体异常肿瘤发展的一个步骤,因此染色体和基因组平衡的评估是许多癌症中重要的诊断和预后检测。
GENERAL REFERENCES Gates, AJ, et al: A wealth of discovery built on the Human Genome Project - by the numbers, Nature, 590, 212–215, 2021.一般参考文献 Gates, AJ, 等: 基于人类基因组计划的丰富发现——数字视角, Nature, 590, 212–215, 2021.
Green ED, et al: Mapping genomic loci implicates genes and synaptic biology in schizophrenia, Nature 604:502–508, 2022.Green ED, 等: 基因组位点作图提示基因和突触生物学在精神分裂症中的作用, Nature 604:502–508, 2022.
Miga KH, et al: Telomere-to-telomere assembly of a complete human X chromosome, Nature, 585, 79-84, 2020 Moore KL, Presaud TVN, Torchia MG: The developing human: ­clinically oriented embryology, ed 9, Philadelphia, 2013, WB Saunders.Miga KH, 等: 完整人类X染色体的端粒到端粒组装, Nature, 585, 79-84, 2020 Moore KL, Presaud TVN, Torchia MG: 发育中的人类:临床导向的胚胎学,第9版,费城,2013,WB Saunders.
REFERENCES FOR SPECIFIC TOPICS Deininger P: Alu elements: know the SINES, Genome Biol 12:236, 2011.特定主题参考文献 Deininger P: Alu元件:认识SINEs, Genome Biol 12:236, 2011.
Frazer KA: Decoding the human genome, Genome Res 22:1599– 1601, 2012.Frazer KA: 解码人类基因组, Genome Res 22:1599–1601, 2012.
International Human Genome Sequencing Consortium: Initial sequencing and analysis of the human genome, Nature 409:860– 921, 2001.国际人类基因组测序联盟: 人类基因组的初始测序与分析, Nature 409:860–921, 2001.
International Human Genome Sequencing Consortium: Finishing the euchromatic sequence of the human genome, Nature 431:931–945, 2004.国际人类基因组测序联盟: 完成人类基因组常染色质序列, Nature 431:931–945, 2004.
Nurk S, Koren S, Rhie A, et al: The complete sequence of a human genome, Science 376:44–53, 2022. abj 6987.Nurk S, Koren S, Rhie A, 等: 人类基因组的完整序列, Science 376:44–53, 2022. abj 6987.
Epub 2022 Mar 31.电子版2022年3月31日。
PMID: 35357919.PMID: 35357919.
Venter J, Adams M, Myers E, et al: The sequence of the human genome, Science 291:1304–1351, 2001.Venter J, Adams M, Myers E, 等: 人类基因组序列, Science 291:1304–1351, 2001.
PROBLEMS 1.问题 1.
At a certain locus, a person has two alleles, A and a. a.在某一基因座上,一个人有两个等位基因A和a。a.
What alleles will be present in this person’s gametes? b.这个人的配子中会出现哪些等位基因?b.
When do A and a segregate (1) if there is no crossing over between the locus and the centromere of the chromosome?当该基因座与染色体着丝粒之间没有交换时,A和a何时分离?(1)
(2) if there is a single crossover between the locus and the centromere?(2) 如果该基因座与着丝粒之间有一次交换呢?
What is the main cause of numerical chromosome abnormalities in humans?人类染色体数目异常的主要原因是什么?
Disregarding crossing over, which increases the amount of genetic variability, estimate the probability that all your chromosomes have come to you from your father’s mother and your mother’s mother.忽略会增加遗传变异量的交换,估算你所有染色体都来自你父亲的母亲和你母亲的母亲的概率。
Would you be male or female?你会是男性还是女性?
A chromosome entering meiosis is composed of two sister chromatids, each of which is a single DNA molecule. a.进入减数分裂的染色体由两条姐妹染色单体组成,每条单体是一个单链DNA分子。a.
In our species, at the end of meiosis I, how many chromosomes are there per cell?在我们人类中,减数分裂I结束时,每个细胞有多少条染色体?
How many chromatids? b.有多少条染色单体?b.
At the end of meiosis II, how many chromosomes are there per cell?减数分裂II结束时,每个细胞有多少条染色体?
How many chromatids? c.有多少条染色单体?c.
When is the diploid chromosome number restored?二倍体染色体数目何时恢复?
When is the two-chromatid structure of a typical ­metaphase chromosome restored?典型中期染色体的两条染色单体结构何时恢复?

The Human Genome — Gene Structure and Function

21/42
The Human Genome Gene Structure and Function Since the discovery of the structure of DNA and the development of technolo…
Ch3 — Segment 21
The Human Genome Gene Structure and Function Since the discovery of the structure of DNA and the development of technologies in molecular biology, remarkable progress has been made in our understanding of the structure and function of genes and chromosomes.人类基因组基因的结构与功能:自DNA结构发现及分子生物学技术发展以来,我们对基因和染色体结构与功能的理解取得了显著进展。
The development of even newer methods to globally study the entire genome provides additional tools for a distinctive new approach to medical genetics.开发更新颖的全局性研究整个基因组的方法,为医学遗传学独特的新方法提供了额外工具。
In this chapter we present an overview of gene structure and function and the aspects of molecular genetics required for an understanding of the principles underlying genomic medicine.在本章中,我们概述了基因结构与功能,以及理解基因组医学原理所需的分子遗传学方面的知识。
We provide additional material online.我们在网上提供额外资料。
The increased knowledge of genes and of their organization in the genome has had an enormous impact on medicine and on our perception of human physiology.对基因及其在基因组中组织的知识的增加,对医学及我们对人类生理学的认知产生了巨大影响。
As 1980 Nobel laureate Paul Berg stated presciently at the dawn of this new era: Just as our present knowledge and practice of medicine relies on a sophisticated knowledge of human anatomy, physiology, and biochemistry, so will dealing with disease in the future demand a detailed understanding of the molecular anatomy, physiology, and biochemistry of the human genome.… We shall need a more detailed knowledge of how human genes are organized and how they function and are regulated.正如1980年诺贝尔奖得主保罗·伯格在这个新时代开启时富有先见之明地指出:正如我们目前的医学知识和实践依赖于对人体解剖学、生理学和生物化学的深入了解,未来处理疾病也需要对人类基因组的分子解剖学、生理学和生物化学有详细理解……我们将需要更详细地了解人类基因如何组织、如何运作以及如何被调控。
We shall also have to have physicians who are as conversant with the molecular anatomy and physiology of chromosomes and genes as the cardiac surgeon is with the structure and workings of the heart.我们还需要有医生,他们对染色体和基因的分子解剖学与生理学的熟悉程度,如同心脏外科医生对心脏的结构与运作的熟悉程度。
INFORMATION CONTENT OF THE HUMAN GENOME How does the approximately 3-billion-letter digital code of the human genome guide the intricacies of human anatomy, physiology, and biochemistry to which Berg referred?人类基因组的信息内容:人类基因组大约30亿字母的数字代码如何指导伯格所提及的人类解剖学、生理学和生物化学的复杂性?
The answer lies in the enormous amplification and integration of information content that occurs as one moves from genes in the genome to their products in the cell and to the observable expression of that genetic information as cellular, morphologic, clinical, or biochemical traits—that is, the phenotype of the individual.答案在于信息内容的巨大放大与整合,这一过程发生在从基因组中的基因到其在细胞中的产物,再到该遗传信息作为细胞、形态、临床或生化特征(即个体的表型)的可观察表达的过程中。
This hierarchic expansion of information from the genome to phenotype includes a wide range of structural and regulatory RNA products, as well as protein products that orchestrate the many functions of cells, tissues, organs, and the entire organism, in addition to their interactions with the environment.这种从基因组到表型的信息层次扩展,包括了广泛的结构性和调控性RNA产物,以及协调细胞、组织、器官和整个生物体众多功能的蛋白质产物,此外还包括它们与环境的相互作用。
Even with the essentially complete sequence of the human genome in hand, we still do not know the precise number of genes in the genome.即使手中掌握了人类基因组的基本完整序列,我们仍然不知道基因组中基因的确切数量。
Our traditional definition of genes has also expanded.我们对基因的传统定义也已扩展。
Current estimates are that the genome contains ~20,000 protein-coding genes (see Box in Chapter 2), but this figure only begins to hint at the levels of complexity that emerge from the decoding of this digital information .目前估计基因组包含约20,000个蛋白质编码基因(见第2章专栏),但这个数字仅仅开始暗示从这一数字信息解码中涌现出的复杂性层次。
As introduced briefly in Chapter 2, the product of protein-coding genes is a protein whose structure ultimately determines its particular function(s) in the cell.正如第2章简要介绍的那样,蛋白质编码基因的产物是一种蛋白质,其结构最终决定了其在细胞中的特定功能。
But if there were a simple one-to-one correspondence between genes and proteins, we could have at most ~20,000 different proteins.但如果基因与蛋白质之间存在简单的一一对应关系,我们最多只能有约20,000种不同的蛋白质。
This number is insufficient to account for the vast array of functions that occur in human cells over the life span.这个数字不足以解释人类细胞在整个生命周期中发生的众多功能。
The answer to this dilemma is found in two features of gene structure and function.这一困境的答案在于基因结构与功能的两个特征。
First, many genes are capable of generating multiple different products, not just one .首先,许多基因能够产生多种不同的产物,而不仅仅是一种。
This process, discussed later in this chapter, is accomplished through the use of alternative coding segments in genes and through the subsequent biochemical modification of the encoded protein; these two features of complex genomes result in a substantial amplification of information content.这一过程将在本章后面讨论,通过使用基因中的可变编码片段以及后续对编码蛋白质的生化修饰来实现;复杂基因组的这两个特征导致信息内容的大幅放大。
Indeed, it has been estimated that in this way, these 20,000 human genes can encode many hundreds of thousands of different proteins, collectively referred to as the proteome.事实上,据估计,通过这种方式,这20,000个人类基因可以编码成千上万种不同的蛋白质,统称为蛋白质组。
Second, individual proteins do not function by themselves.其次,单个蛋白质并不单独发挥作用。
They form networks, often involving many different proteins and regulatory RNAs that respond in a coordinated and integrated fashion to many different genetic, developmental, or environmental signals.它们形成网络,通常涉及许多不同的蛋白质和调控RNA,这些网络以协调和整合的方式对多种不同的遗传、发育或环境信号作出响应。
The combinatorial nature of protein networks results in an even greater diversity of possible cellular functions.蛋白质网络的组合性质导致了可能的细胞功能具有更大的多样性。
Genes are located throughout the genome but tend to cluster in particular regions on particular chromosomes基因分布于整个基因组,但倾向于聚集在特定染色体的特定区域。
22/42
and to be relatively sparse in other regions or on other chromosomes.
Ch3 — Segment 22
and to be relatively sparse in other regions or on other chromosomes.并且在其他区域或其他染色体上相对稀疏。
For example, chromosome 11, an ~135 million-bp (megabase pairs [Mb]) chromosome, is relatively gene-rich with ~1300 protein-coding genes .例如,11号染色体,一条约1.35亿碱基对(兆碱基对[Mb])的染色体,相对基因密集,拥有约1300个蛋白质编码基因。
These genes are not distributed randomly along the chromosome, and their localization is particularly enriched in two chromosomal regions with gene density as high as one gene every 10 kb .这些基因并非沿染色体随机分布,其定位特别富集于两个染色体区域,基因密度高达每10 kb一个基因。
Some of the genes belong to families of related genes, as we will describe more fully later in this chapter.其中一些基因属于相关基因家族,我们将在本章后面更全面地描述。
Other regions are gene-poor, and there are several so-called gene deserts of 1 million bp or more without any identified protein-coding genes.其他区域基因贫乏,存在若干所谓的基因荒漠,长度达100万碱基对或更多,未鉴定出任何蛋白质编码基因。
There are two caveats here: first, the process of gene identification and genome annotation remains an ongoing process despite the apparent robustness of recent estimates.这里有两个注意事项:首先,尽管近期估计看似可靠,基因鉴定和基因组注释仍是一个持续的过程。
It is virtually certain that there are some genes, including clinically relevant genes, that are currently undetected or that display characteristics that we do not currently recognize as being associated with genes.几乎可以肯定,存在一些基因,包括临床相关基因,目前未被检测到,或表现出我们目前尚未识别为与基因相关的特征。
Second, as mentioned in Chapter 2, many genes are not protein coding; their products are functional RNA molecules (noncoding RNAs [ncRNAs]) that play a variety of roles in the cell, many of which are only just being uncovered.其次,如第2章所述,许多基因并非编码蛋白质;其产物是功能性RNA分子(非编码RNA[ncRNAs]),在细胞中发挥多种作用,其中许多作用才刚刚被发现。
For genes located on the autosomes, there are two copies of each gene, one on the chromosome inherited from the mother and one on the chromosome inherited from the father.对于位于常染色体上的基因,每个基因有两个拷贝,一个位于来自母亲的染色体上,另一个位于来自父亲的染色体上。
For most autosomal genes, both copies are expressed and generate a product.对于大多数常染色体基因,两个拷贝均表达并产生产物。
There are, however, a growing number of genes in the genome that are exceptions to this general rule and are expressed at characteristically different levels from the two copies, including some that, at the extreme, are expressed from only one of the two homologues.然而,基因组中越来越多的基因是该普遍规则的例外,它们从两个拷贝中以特征性不同的水平表达,包括一些极端情况,仅从两个同源染色体中的一个表达。
These examples of allelic imbalance are discussed in greater detail later in this chapter, as well as in Chapters 6 and 7.这些等位基因不平衡的例子将在本章后面以及第6章和第7章中更详细地讨论。
In addition, many genes are present in variable numbers at a particular location on a chromosome.此外,许多基因在染色体特定位置上以可变拷贝数存在。
One example is the variability in the copy number of the genes for amylase, an enzyme important in starch digestion; AMY1 exists in two to eight copies per chromosome and is expressed in the salivary glands.一个例子是淀粉酶(一种在淀粉消化中重要的酶)基因拷贝数的变异性;AMY1在每条染色体上存在2至8个拷贝,并在唾液腺中表达。
THE CENTRAL DOGMA: DNA → RNA → PROTEIN How does the genome specify the functional complexity and diversity evident in and noncoding RNA (ncRNA) genes (red).中心法则:DNA → RNA → 蛋白质。基因组如何指定在和非编码RNA(ncRNA)基因(红色)中明显的功能复杂性和多样性?
Many genes in the genome use alternative coding information to generate multiple different products.基因组中的许多基因利用替代编码信息产生多种不同的产物。
Both small and large ncRNAs participate in gene regulation.小分子和大分子ncRNA均参与基因调控。
Many proteins participate in multigene networks that respond to cellular signals in a coordinated and combinatorial manner, thus further expanding the range of cellular functions that underlie organismal phenotypes.许多蛋白质参与多基因网络,以协调和组合的方式响应细胞信号,从而进一步扩展构成生物表型基础的细胞功能范围。
23/42
The Human Genome 25 functions, takes place in the cytoplasm.
Ch3 — Segment 23
The Human Genome 25 functions, takes place in the cytoplasm.人类基因组25功能在细胞质中发生。
This compartmentalization reflects the fact that the human organism is a eukaryote.这种区室化反映了人类有机体是真核生物的事实。
This means that human cells have a nucleus containing the genome, which is separated by a nuclear membrane from the cytoplasm.这意味着人类细胞具有包含基因组的细胞核,该细胞核通过核膜与细胞质分隔。
In contrast, in prokaryotes like the intestinal bacterium Escherichia coli, DNA is not enclosed within a nucleus.相比之下,在像肠道细菌大肠杆菌这样的原核生物中,DNA不包被在细胞核内。
Because of the compartmentalization of eukaryotic cells, information transfer from the nucleus to the cytoplasm is a complex process that has been a focus of much attention among molecular and cellular biologists.由于真核细胞的区室化,从细胞核到细胞质的信息传递是一个复杂的过程,一直是分子和细胞生物学家关注的焦点。
The molecular link between these two related types of information—the DNA code of genes and the amino acid code of protein—is ribonucleic acid (RNA).这两种相关信息——基因的DNA密码和蛋白质的氨基酸密码——之间的分子连接是核糖核酸(RNA)。
The chemical structure of RNA is similar to that of DNA, except that each nucleotide in RNA has a ribose sugar component instead of a deoxyribose; in addition, uracil (U) replaces thymine as one of the pyrimidine bases of RNA .RNA的化学结构与DNA相似,不同之处在于RNA中的每个核苷酸含有核糖糖基而非脱氧核糖;此外,尿嘧啶(U)取代胸腺嘧啶作为RNA的嘧啶碱基之一。
An additional difference between RNA and DNA is that RNA in most organisms exists as a single-stranded molecule, whereas DNA, as we saw in Chapter 2, exists as a double helix.RNA与DNA的另一个区别是,在大多数生物体中RNA以单链分子存在,而DNA如我们在第2章所见,以双螺旋存在。
The informational relationships among DNA, RNA, and protein are intertwined: genomic DNA directs the synthesis and sequence of RNA, RNA directs the synthesis and sequence of polypeptides, and specific proteins are involved in the synthesis and metabolism of DNA and RNA.DNA、RNA和蛋白质之间的信息关系相互交织:基因组DNA指导RNA的合成和序列,RNA指导多肽的合成和序列,而特定蛋白质参与DNA和RNA的合成与代谢。
This flow of information is referred to as the central dogma of molecular biology.这种信息流被称为分子生物学的中心法则。
Genetic information is stored in the DNA of the genome by means of a code (the genetic code, discussed later) in which the sequence of adjacent bases ultimately determines the sequence of amino acids in the encoded polypeptide.遗传信息通过一种密码(遗传密码,稍后讨论)存储于基因组的DNA中,其中相邻碱基的序列最终决定编码多肽中的氨基酸序列。
First, RNA is synthesized from the DNA template through the process of transcription.首先,通过转录过程从DNA模板合成RNA。
The RNA, carrying the coded information in a form called messenger RNA (mRNA), is then transported from the nucleus to the cytoplasm, where the RNA sequence is decoded, or translated, to determine the sequence of amino acids in the protein being synthesized.携带编码信息的RNA(称为信使RNA,mRNA)随后从细胞核运输到细胞质,在那里RNA序列被解码或翻译,以确定正在合成的蛋白质中的氨基酸序列。
The process of translation occurs on ribosomes, which are cytoplasmic organelles with binding sites for all of the interacting molecules, including the mRNA, involved in protein synthesis.翻译过程发生在核糖体上,核糖体是细胞质细胞器,具有与蛋白质合成相关的所有相互作用分子(包括mRNA)的结合位点。
Ribosomes are themselves made up of many different structural proteins in association with specialized types of RNA known as ribosomal RNA (rRNA).核糖体本身由许多不同的结构蛋白与称为核糖体RNA(rRNA)的特殊类型RNA结合而成。
Translation involves yet a third type of RNA, transfer RNA (tRNA), which provides the molecular link between the code contained in the base sequence of each mRNA and the amino acid sequence of the protein encoded by that mRNA.......翻译还涉及第三种RNA,即转运RNA(tRNA),它提供了每个mRNA碱基序列中包含的密码与该mRNA所编码蛋白质的氨基酸序列之间的分子连接……
Chromosome 11 No. genes 5. 2 5. 15 5. 25 5. 3 5. 35 OR genes OR genes 0 20 40 60 kb Direction of transcription β ε Aγ Gγ δ β-like globin genes Mb A B C The distribution of genes is indicated along the chromosome and is high in two regions of the chromosome and low in other regions.染色体11 基因编号5.2 5.15 5.25 5.3 5.35 OR基因 OR基因 0 20 40 60 kb 转录方向 β ε Aγ Gγ δ β样珠蛋白基因 Mb A B C 基因分布沿染色体指示,在染色体的两个区域高而在其他区域低。
(B) An expanded region from 5. 15 to 5. 35 Mb (measured from the short-arm telomere), which contains 10 known protein-coding genes, five belonging to the olfactory receptor (OR) gene family and five belonging to the globin gene family.(B) 从5.15到5.35 Mb(从短臂端粒测量)的扩展区域,包含10个已知的蛋白质编码基因,其中5个属于嗅觉受体(OR)基因家族,5个属于珠蛋白基因家族。
(C) The five β-like globin genes expanded further.(C) 五个β样珠蛋白基因进一步扩展。
(Data from European Bioinformatics Institute and Wellcome Trust Sanger Institute: Ensembl release 70, January 2013.(数据来自欧洲生物信息学研究所和威康信托桑格研究所:Ensembl第70版,2013年1月。
Available from _ _ O O CH C CH C H HN O O O O O P CH2 C OH H H H OH C C H C 5' 3' N Uracil (U) Phosphate Ribose Base Note that the sugar ribose replaces the sugar deoxyribose of DNA.可从 _ _ O O CH C CH C H HN O O O O O P CH2 C OH H H H OH C C H C 5' 3' N 尿嘧啶(U) 磷酸 核糖 碱基 注意核糖取代了DNA的脱氧核糖。
Compare with比较。
24/42
Because of the interdependent flow of information represented by the central dogma, one can begin discussion of the mole…
Ch3 — Segment 24
Because of the interdependent flow of information represented by the central dogma, one can begin discussion of the molecular genetics of gene expression at any of its three informational levels: DNA, RNA, or protein.由于中心法则所表征的信息相互依赖流动,人们可以从DNA、RNA或蛋白质这三个信息层面中的任意一个开始讨论基因表达的分子遗传学。
We begin by examining the structure of genes in the genome as a foundation for discussion of the genetic code, transcription, and translation.我们首先检查基因组中基因的结构,以此作为讨论遗传密码、转录和翻译的基础。
GENE ORGANIZATION AND STRUCTURE In its simplest form, a protein-coding gene can be visualized as a segment of a DNA molecule containing the code for the amino acid sequence of a polypeptide chain and the regulatory sequences necessary for its expression.基因的组织与结构 在其最简单的形式中,编码蛋白质的基因可被视为一段DNA分子,包含编码多肽链氨基酸序列的密码以及其表达所必需的调控序列。
This description, however, is inadequate for genes in the human genome (and indeed in most eukaryotic genomes) because few genes exist as continuous coding sequences.然而,这一描述对于人类基因组(实际上对于大多数真核生物基因组)中的基因并不充分,因为很少有基因是连续编码序列存在的。
Rather, in the majority of genes, the coding sequences are interrupted by one or more noncoding regions .相反,在大多数基因中,编码序列被一个或多个非编码区域所中断。
These intervening sequences, called introns, are initially transcribed into RNA in the nucleus but are not present in the mature mRNA in the cytoplasm because they are removed (“spliced out”) by a process we will discuss later.这些插入序列称为内含子,最初在细胞核中被转录成RNA,但不在细胞质中的成熟mRNA中出现,因为它们通过我们稍后将讨论的过程被移除(“剪接掉”)。
Thus information from the intronic sequences is not normally represented in the final protein product.因此,来自内含子序列的信息通常不会在最终蛋白质产物中体现。
Introns alternate with exons, the Exons Direction of transcription CAT TATA CG-rich 3′ 5′ “Upstream” Start of transcription “Downstream” Termination codon Introns (intervening sequences) Initiator codon Promoter 5′ untranslated region 3′ untranslated region Polyadenylation signal b-Globin BRCA1 MYH7 0 0. 5 1. 0 1. 5 2. 0 kb 0 20 40 60 80 kb 0 5 10 15 20 kb A B Individual labeled features are discussed in the text.内含子与外显子交替排列,外显子转录方向 CAT TATA CG丰富区 3′ 5′ “上游” 转录起始点 “下游” 终止密码子 内含子(插入序列) 起始密码子 启动子 5′非翻译区 3′非翻译区 多聚腺苷酸化信号 β-珠蛋白 BRCA1 MYH7 0 0.5 1.0 1.5 2.0 kb 0 20 40 60 80 kb 0 5 10 15 20 kb A B 各标记特征在正文中讨论。
(B) Examples of three medically important human genes.(B) 三个医学上重要的人类基因示例。
Different deleterious variants in the β-globin gene, with three exons, cause a variety of important disorders of hemoglobin (Case 25).β-珠蛋白基因(含有三个外显子)中不同的有害变异导致多种重要的血红蛋白疾病(病例25)。
Mutations in the BRCA1 gene (24 exons) are responsible for many cases of inherited breast or breast and ovarian cancer (Case 7).BRCA1基因(24个外显子)的突变导致许多遗传性乳腺癌或乳腺卵巢癌病例(病例7)。
Mutations in the β-myosin heavy chain (MYH7) gene (40 exons) lead to inherited hypertrophic cardiomyopathy.β-肌球蛋白重链(MYH7)基因(40个外显子)的突变导致遗传性肥厚型心肌病。
25/42
The Human Genome 27 segments of genes that ultimately determine the amino acid sequence of the protein.
Ch3 — Segment 25
The Human Genome 27 segments of genes that ultimately determine the amino acid sequence of the protein.人类基因组中,最终决定蛋白质氨基酸序列的基因有27个片段。
In addition, the collection of coding exons in any particular gene is flanked by additional sequences that are transcribed but untranslated, called the 5′ and 3′ untranslated regions .此外,任何特定基因中编码外显子的集合两侧均存在被转录但不翻译的额外序列,称为5′和3′非翻译区。
Although a few genes in the human genome have no introns, most genes contain at least one, with nine exons spanning ~25 kb found in an average gene.尽管人类基因组中有少数基因不含内含子,但大多数基因至少含有一个内含子,平均每个基因含有约25 kb的9个外显子。
In many genes, the cumulative length of the introns makes up a far greater proportion of a gene’s total length than do the exons.在许多基因中,内含子的累积长度在基因总长度中所占比例远大于外显子。
Whereas some genes are only a few kilobase pairs in length, others stretch on for hundreds of kilobase pairs.有些基因仅长几千碱基对,而另一些则延伸至数百千碱基对。
Also, a few genes are exceptionally large (e. g., the CTNAP2 gene on chromosome 7 and the dystrophin gene on the X chromosome [pathogenic variants that lead to Duchenne/Becker muscular dystrophy (Case 14)] span >2 Mb, of which, remarkably, <1% consists of coding exons).此外,少数基因异常庞大(例如,7号染色体上的CTNAP2基因和X染色体上的肌营养不良蛋白基因[导致杜氏/贝氏肌营养不良(病例14)的致病性变异]长度超过2 Mb,值得注意的是,其中编码外显子占比不足1%)。
The KCNIP4 potassium channel gene has a single intron that is over 1 Mb in size.KCNIP4钾通道基因含有一个单一内含子,其大小超过1 Mb。
Structural Features of a Typical Human Gene A range of features characterize human genes .典型人类基因的结构特征 人类基因具有一系列特征。
In Chapters 1 and 2, we briefly defined gene in general terms.在第1章和第2章中,我们曾对基因进行了概括性定义。
At this point, we can provide a molecular definition of a gene as a sequence of DNA that specifies production of a functional product, be it a polypeptide or a functional RNA molecule.至此,我们可以给出基因的分子定义:一段DNA序列,其指定功能产物的生成,无论是多肽还是功能性RNA分子。
A gene includes not only the actual coding sequences but also adjacent nucleotide sequences required for the proper expression of the gene—that is, for the production of normal mRNA or other RNA molecules in the correct amount, in the correct place, and at the correct time during development or during the cell cycle.基因不仅包括实际编码序列,还包括基因正确表达所需的相邻核苷酸序列——即在发育或细胞周期过程中,在正确时间、正确位置以正确量产生正常mRNA或其他RNA分子所必需的序列。
The adjacent nucleotide sequences provide the molecular start and stop signals for the synthesis of mRNA transcribed from the gene.相邻核苷酸序列提供了从该基因转录的mRNA合成的分子起始和终止信号。
Because the primary RNA transcript is synthesized in a 5′ to 3′ direction, the transcriptional start is referred to as the 5′ end of the transcribed portion of a gene .由于初级RNA转录物以5′到3′方向合成,转录起始点被称为基因转录部分的5′端。
By convention, the genomic DNA that precedes the transcriptional start site in the 5′ direction is referred to as the upstream sequence, whereas DNA sequence located in the 3′ direction past the end of a gene is the downstream sequence.按惯例,在5′方向上位于转录起始位点之前的基因组DNA称为上游序列,而位于基因末端3′方向之后的DNA序列称为下游序列。
At the 5′ end of each gene lies a promoter region that includes sequences responsible for the proper initiation of transcription.每个基因的5′端都有一个启动子区域,包含负责正确启动转录的序列。
Within this region are several DNA elements whose sequence is often conserved among many different genes; this conservation, together with functional studies of gene expression, indicates that these particular sequences play an important role in gene regulation.该区域内存在若干DNA元件,其序列在许多不同基因中常呈保守性;这种保守性,结合基因表达的功能研究,表明这些特定序列在基因调控中起重要作用。
Importantly, only a subset of genes in the genome is expressed in any given tissue or at any given time during development.重要的是,在任一给定组织或发育过程中的任一给定时间,基因组中仅有一小部分基因被表达。
Several different types of promoter are found in the human genome, with different regulatory properties that specify the patterns as well as the levels of expression of a particular gene in different tissues and cell types, both during development and throughout the life span.人类基因组中存在几种不同类型的启动子,它们具有不同的调控特性,这些特性指定了特定基因在不同组织和细胞类型中,在发育过程中及整个生命周期内的表达模式和表达水平。
Some of these properties are encoded in the genome, whereas others are specified by features of chromatin associated with those sequences, as discussed later in this chapter.其中一些特性编码于基因组中,而另一些则由与这些序列相关的染色质特征所规定,本章后续将讨论。
Both promoters and other regulatory elements (located either 5′ or 3′ of a gene or in its introns) can be sites of variation causing genetic disease that can interfere with the normal expression of a gene.启动子及其他调控元件(位于基因的5′或3′端,或位于其内含子中)均可成为导致遗传病的变异位点,这些变异可能干扰基因的正常表达。
These regulatory elements, including enhancers, insulators, and locus control regions, are discussed more fully later in this chapter.这些调控元件,包括增强子、绝缘子和基因座控制区,将在本章后面更详细地讨论。
Some of these elements lie a significant distance away from the coding portion of a gene, thus reinforcing the concept that the genomic environment in which a gene resides is an important feature of its evolution and regulation.其中一些元件位于基因编码部分相当远的距离处,从而强化了一个概念:基因所处的基因组环境是其进化和调控的重要特征。
The 3′ untranslated region contains a signal for the addition of a sequence of adenosine residues (the so-called poly A tail) to the end of the mature RNA.3′非翻译区包含一个信号,用于在成熟RNA末端添加腺苷酸残基序列(即所谓的poly A尾)。
Although it is generally accepted that such closely neighboring regulatory sequences are part of what is called a gene, the precise dimensions of any particular gene will remain somewhat uncertain until the potential functions of more distant sequences are fully characterized.尽管普遍认为这些紧密相邻的调控序列属于所谓的基因的一部分,但任何特定基因的精确边界仍存在一定不确定性,直到更远序列的潜在功能被完全阐明。
Gene Families Many genes belong to gene families, which share closely related DNA sequences and encode polypeptides with closely related amino acid sequences.基因家族 许多基因属于基因家族,它们共享高度相似的DNA序列,并编码氨基酸序列高度相关的多肽。
Members of two such gene families are located within a small region on chromosome 11 and illustrate a number of features that characterize gene families in general.两个此类基因家族的成员位于11号染色体上的一个小区段内,体现了基因家族普遍具有的若干特征。
One small and medically important gene family is composed of genes that encode the protein chains found in hemoglobins.一个规模较小但具有医学重要性的基因家族由编码血红蛋白中蛋白质链的基因组成。
The β-globin gene cluster on chromosome 11 and the related α-globin gene cluster on chromosome 16 are believed to have arisen by duplication of a primitive precursor gene ~500 million years ago.11号染色体上的β-珠蛋白基因簇和16号染色体上的相关α-珠蛋白基因簇被认为源于约5亿年前一个原始前体基因的重复。
These two clusters contain multiple genes coding for closely related globin chains expressed at different developmental stages, from embryo to adult.这两个基因簇包含多个编码高度相关珠蛋白链的基因,这些基因在从胚胎到成人的不同发育阶段表达。
Each cluster is believed to have evolved by a series of sequential gene duplication events within the past 100 million years.每个基因簇被认为在过去1亿年内通过一系列连续的基因重复事件进化而来。
The exon-intron patterns of the functional globin genes have been remarkably conserved during evolution; each of the functional globin genes has two introns at similar locations (see the β-globin gene in , although the sequences contained within the introns have accumulated far more nucleotide base changes over time than have the coding sequences of each gene.功能性珠蛋白基因的外显子-内含子模式在进化过程中高度保守;每个功能性珠蛋白基因在相似位置有两个内含子(参见β-珠蛋白基因),尽管内含子内部的序列随时间积累的核苷酸碱基变化远多于每个基因的编码序列。
The control of expression of the various globin genes, in the normal state as well as in the many inherited disorders of hemoglobin, is considered in more detail both later in this chapter and in Chapter 12.各种珠蛋白基因的表达调控——包括正常状态及多种遗传性血红蛋白疾病中的情况——将在本章后续及第12章中更详细地讨论。
The second gene family shown in genes.所示的第二个基因家族位于基因中。
There are estimated to be as many as 1000 OR genes in the genome基因组中估计有多达1000个OR基因。
26/42
(390 putatively functional genes and 465 pseudogenes).
Ch3 — Segment 26
(390 putatively functional genes and 465 pseudogenes).(390个推定功能性基因和465个假基因)。
ORs are responsible for our acute sense of smell that can recognize and distinguish thousands of structurally diverse chemicals.OR基因负责我们敏锐的嗅觉,能识别和区分数千种结构各异的化学物质。
OR genes are found throughout the genome on nearly every chromosome, although more than half are found on chromosome 11, including a number of family members near the β-globin cluster.OR基因遍布基因组中几乎所有染色体,但半数以上位于11号染色体,包括β-珠蛋白基因簇附近的多个家族成员。
Pseudogenes Within both the β-globin and OR gene families are sequences that are related to the functional globin and OR genes but that do not produce any functional RNA or protein product.假基因 在β-珠蛋白和OR基因家族中,存在与功能性珠蛋白和OR基因相关但不产生任何功能性RNA或蛋白质产物的序列。
DNA sequences that closely resemble known genes but are nonfunctional are called pseudogenes, and there are ~20,000 pseudogenes related to many different genes and gene families located all around the genome.与已知基因高度相似但无功能的DNA序列称为假基因,基因组中分布有约20,000个与多种不同基因和基因家族相关的假基因。
Pseudogenes are of two general types, processed and nonprocessed.假基因分为两大类:加工型假基因和非加工型假基因。
Nonprocessed pseudogenes are thought to be byproducts of evolution, representing “dead” genes that were once functional but are now vestigial, having been inactivated by variants in critical coding or regulatory sequences.非加工型假基因被认为是进化的副产物,代表曾经有功能但现已退化的“死亡”基因,因关键编码或调控序列中的变异而失活。
In contrast to nonprocessed pseudogenes, processed pseudogenes are pseudogenes that have been formed, not by mutation, but by a process called retrotransposition, which involves transcription, generation of a DNA copy of the mRNA (a so-called cDNA) by reverse transcription, and finally integration of such DNA copies back into the genome at a location usually quite distant from the original gene.与非加工型假基因不同,加工型假基因并非通过突变形成,而是通过一种称为逆转录转座的过程形成,该过程涉及转录、通过逆转录生成mRNA的DNA拷贝(即所谓的cDNA),以及最终将这些DNA拷贝整合回基因组中通常远离原始基因的位置。
Because such pseudogenes are created by retrotransposition of a DNA copy of processed mRNA, they lack introns and are usually not on the same chromosome (or chromosomal region) as their progenitor gene.由于此类假基因由加工后mRNA的DNA拷贝经逆转录转座产生,因此它们缺乏内含子,且通常不与祖先基因位于同一染色体(或染色体区域)。
In many gene families there are as many or even more pseudogenes as there are functional gene members.在许多基因家族中,假基因的数量与功能性基因成员相当或更多。
Noncoding RNA Genes Many genes are protein coding and are transcribed into mRNAs that are ultimately translated into their respective proteins; their products comprise the enzymes, structural proteins, receptors, and regulatory proteins that are found in various human tissues and cell types.非编码RNA基因 许多基因为蛋白质编码基因,转录为mRNA后最终翻译成相应蛋白质;其产物包括各种人体组织和细胞类型中的酶、结构蛋白、受体和调控蛋白。
However, as introduced briefly in Chapter 2, there are additional genes whose functional product appears to be the RNA itself .然而,如第2章简要介绍,还存在一些功能产物似乎是RNA本身的额外基因。
These so-called noncoding RNAs (ncRNAs) have a range of functions in the cell, although many do not as yet have any identified function.这些所谓的非编码RNA(ncRNA)在细胞中具有一系列功能,尽管许多尚未明确其功能。
By current estimates, there are some 15,000 to 20,000 ncRNA genes in addition to the ~20,000 protein-coding genes that we introduced earlier.根据当前估计,除了之前介绍的大约20,000个蛋白质编码基因外,还有约15,000至20,000个ncRNA基因。
Thus the collection of ncRNAs represents approximately half of all identified human genes.因此,ncRNA集合约占所有已鉴定人类基因的一半。
Some of the types of ncRNA play largely generic roles in cellular infrastructure, including the tRNAs and rRNAs involved in translation of mRNAs on ribosomes, other RNAs involved in control of RNA splicing, and small nucleolar RNAs (snoRNAs) involved in modifying rRNAs.某些类型的ncRNA在细胞基础设施中主要发挥通用作用,包括参与核糖体上mRNA翻译的tRNA和rRNA、参与RNA剪接调控的其他RNA,以及参与rRNA修饰的小核仁RNA(snoRNA)。
Additional ncRNAs can be quite long (thus sometimes called long ncRNAs [lncRNAs]) and play roles in gene regulation, gene silencing, and human disease, as we explore in more detail later in this chapter and in Case Report 35.其他ncRNA可以相当长(因此有时称为长链非编码RNA [lncRNA]),并在基因调控、基因沉默和人类疾病中发挥作用,我们将在本章后续部分及病例报告35中更详细地探讨。
A particular class of small RNAs of growing importance are the micro RNAs (miRNAs), ncRNAs of only ~22 bases in length that suppress translation of target genes by binding to their respective mRNAs and regulating protein production from the target transcript(s).一类日益重要的小RNA是微小RNA(miRNA),其长度仅约22个碱基,通过结合目标mRNA并调控目标转录本的蛋白质产生,从而抑制靶基因的翻译。
Well over 1000 miRNA genes have been identified in the human genome; some are evolutionarily conserved, whereas others appear to be of quite recent origin.人类基因组中已鉴定出远超1000个miRNA基因;其中一些在进化上保守,而另一些则似乎起源较新。
Some miRNAs have been shown to down-regulate hundreds of mRNAs each, with different combinations of target RNAs in different tissues; combined, the miRNAs are thus predicted to control the activity of as many as 30% of all protein-coding genes in the genome.已显示某些miRNA可下调数百个mRNA,且在不同组织中具有不同的靶RNA组合;因此,预测miRNA共同控制基因组中多达30%的蛋白质编码基因的活性。
Although this is a fast-moving area of genome biology, pathogenic variants in several ncRNA genes have already been implicated in human diseases, including cancer, developmental disorders, and various diseases of both early and adult onset (see 1)..尽管这是基因组生物学中一个快速发展的领域,但多个ncRNA基因的致病性变异已涉及人类疾病,包括癌症、发育障碍以及各种早发和晚发疾病(见1)。
NONCODING RNAS AND DISEASE The importance of various types of ncRNAs for medicine is underscored by their roles in a range of human diseases, from early developmental syndromes to adult-onset disorders.非编码RNA与疾病 各类ncRNA在从早期发育综合征到晚发疾病等一系列人类疾病中的作用,凸显了它们对医学的重要性。
Deletion of a cluster of miRNA genes on chromosome 13 leads to a form of Feingold syndrome, a developmental syndrome of skeletal and growth defects, including microcephaly, short stature, and digital anomalies.13号染色体上miRNA基因簇的缺失导致一种费因戈尔德综合征,这是一种以骨骼和生长缺陷为特征的发育综合征,包括小头畸形、身材矮小和手指畸形。
Pathogenic variants in the miRNA gene MIR96, in the region of the gene critical for the specificity of recognition of its target mRNA(s), can result in progressive hearing loss in adults.miRNA基因MIR96中对其靶mRNA识别特异性至关重要的区域内的致病性变异,可导致成人进行性听力丧失。
Aberrant levels of certain classes of miRNAs have been reported in a wide variety of cancers, central nervous system disorders, and cardiovascular disease (see Chapter 15).已在多种癌症、中枢神经系统疾病和心血管疾病中报告了某些类别miRNA的异常水平(见第15章)。
Deletion of clusters of snoRNA genes on chromosome 15 results in Prader-Willi syndrome, a disorder characterized by obesity, hypogonadism, and cognitive impairment (see Case 38).15号染色体上snoRNA基因簇的缺失导致普拉德-威利综合征,这是一种以肥胖、性腺功能减退和认知障碍为特征的疾病(见病例38)。
Abnormal expression of a specific lncRNA on chromosome 12 has been reported in patients with a pregnancy-associated disease called HELLP (hemolysis, elevated liver enzymes, and low platelets) syndrome.已在患有妊娠相关疾病HELLP综合征(溶血、肝酶升高和血小板减少)的患者中报告了12号染色体上特定lncRNA的异常表达。
Deletion, abnormal expression, and/or structural abnormalities in different lncRNAs with roles in long-range regulation of gene expression and genome function underlie a variety of disorders involving telomere length maintenance, monoallelic expression of genes in specific regions of the genome, and X chromosome dosage (see Chapter 6).不同lncRNA的缺失、异常表达和/或结构异常(这些lncRNA在基因表达的远程调控和基因组功能中发挥作用)是多种疾病的基础,包括端粒长度维持、基因组特定区域基因的单等位表达以及X染色体剂量补偿(见第6章)。
27/42
The Human Genome 29 FUNDAMENTALS OF GENE EXPRESSION For genes that encode proteins, the flow of information from gene to…
Ch3 — Segment 27
The Human Genome 29 FUNDAMENTALS OF GENE EXPRESSION For genes that encode proteins, the flow of information from gene to polypeptide involves several steps .人类基因组29 基因表达基础 对于编码蛋白质的基因,从基因到多肽的信息流动涉及多个步骤。
Initiation of transcription of a gene is under the influence of promoters and other regulatory elements, as well as specific proteins known as transcription factors, which interact with specific sequences within these regions and determine the spatial and temporal pattern of expression of a gene.基因转录的起始受启动子和其他调控元件以及称为转录因子的特定蛋白质的影响,这些蛋白质与这些区域内的特定序列相互作用,并决定基因表达的空间和时间模式。
Transcription of a gene is initiated at the transcriptional start site on chromosomal DNA at the beginning of a 5′ transcribed but untranslated region (called the 5′ UTR), just upstream from the coding sequences, and continues along the chromosome for anywhere from several hundred base pairs to more than a million base pairs, through both introns and exons and past the end of the coding sequences.基因的转录起始于染色体DNA上转录起始位点,该位点位于5′转录但非翻译区(称为5′ UTR)的起始处,恰好在编码序列上游,并沿着染色体延续,长度从几百个碱基对到超过一百万个碱基对不等,穿过内含子和外显子,并越过编码序列的末端。
After modification at both the 5′ and 3′ ends of the primary RNA transcript, the portions corresponding to introns are removed, and the segments corresponding to exons are spliced together, a process called RNA splicing.在初级RNA转录本的5′和3′端进行修饰后,对应于内含子的部分被移除,对应于外显子的片段被拼接在一起,这个过程称为RNA剪接。
After splicing, the resulting mRNA (containing a central segment that is now colinear with the coding portions of the gene) is transported from the nucleus to the cytoplasm, where the mRNA is finally translated into the amino acid sequence of the encoded polypeptide.剪接后,产生的mRNA(含有一个现在与基因编码部分共线性的中心片段)从细胞核运输到细胞质,在那里mRNA最终被翻译成所编码多肽的氨基酸序列。
Each of the steps in this complex pathway is subject to error, and DNA variations that interfere with the individual steps have been implicated in a number of inherited disorders (see Chapters 12 and 13).这一复杂途径中的每一步都可能出错,干扰各个步骤的DNA变异已与多种遗传性疾病相关(参见第12章和第13章)。
Transcription Transcription of protein-coding genes by RNA polymerase II (one of several classes of RNA polymerases) is initiated at the transcriptional start site, the point in the 5′ UTR that corresponds to the 5′ end of the final RNA product (see Figs. 3. 4 and 3. 5).转录 由RNA聚合酶II(几种RNA聚合酶中的一种)对蛋白质编码基因的转录起始于转录起始位点,该位点是5′ UTR中对应于最终RNA产物5′端的点(参见图3.4和图3.5)。
Synthesis of the primary RNA transcript proceeds in a 5′ to 3′ direction, whereas the strand of the gene that is transcribed and that serves as the template for RNA synthesis is read in a 3′ to 5′ direction with respect to the direction of the deoxyribose phosphodiester backbone .初级RNA转录本的合成以5′到3′方向进行,而被转录并作为RNA合成模板的基因链相对于脱氧核糖磷酸二酯骨架的方向则以3′到5′方向被读取。
Because the RNA synthesized corresponds both in polarity and in base sequence (substituting U for T) to the 5′ to 3′ strand of DNA, this 5′ to 3′ strand of nontranscribed DNA is sometimes called the coding, or sense, DNA strand.由于合成的RNA在极性和碱基序列(以U代替T)上与DNA的5′到3′链相对应,这条未转录的5′到3′DNA链有时被称为编码链或有义链。
The 3′ to 5′ strand of DNA that is used as a template for transcription is then referred to as the noncoding, or antisense, strand.用作转录模板的3′到5′DNA链随后被称为非编码链或反义链。
Transcription Transcribed strand Completed polypeptide Growing polypeptide chain Cytoplasm Nucleus Exons: RNA 3' 3' 3' 3' 3' 3' 5' 5' 5' 5' 5' 5' 1 3 2 3' 5' 4.转录 转录链 完成的多肽 生长的多肽链 细胞质 细胞核 外显子: RNA 3' 3' 3' 3' 3' 3' 5' 5' 5' 5' 5' 5' 1 3 2 3' 5' 4.
Translation 5.翻译 5.
Protein assembly 3.蛋白质组装 3.
Transport Nontranscribed strand 1.运输 未转录链 1.
Transcription 2.转录 2.
RNA processing and splicing Ribosomes poly A addition CAP AAAA AAAA AAAA Within the exons, purple indicates the coding sequences.RNA加工和剪接 核糖体 poly A添加 帽 AAAA AAAA AAAA 在外显子内,紫色表示编码序列。
Steps include transcription, RNA processing and splicing, RNA transport from the nucleus to the cytoplasm, and translation.步骤包括转录、RNA加工和剪接、RNA从细胞核到细胞质的运输以及翻译。
28/42
continues through both intronic and exonic portions of the gene, beyond the position on the chromosome that eventually c…
Ch3 — Segment 28
continues through both intronic and exonic portions of the gene, beyond the position on the chromosome that eventually corresponds to the 3′ end of the mature mRNA.继续穿过基因的内含子和外显子区域,超出染色体上最终对应于成熟mRNA 3′端的位置。
Whether transcription ends at a predetermined 3′ termination point is unknown.转录是否在预定的3′终止点结束尚不清楚。
The primary RNA transcript is processed by addition of a chemical cap structure to the 5′ end of the RNA and cleavage of the 3′ end at a specific point downstream from the end of the coding information.初级RNA转录物通过在其5′端添加化学帽状结构以及在编码信息末端下游的特定点切割3′端进行加工。
This cleavage is followed by addition of a poly A tail to the 3′ end of the RNA; the poly A tail appears to increase the stability of the resulting polyadenylated RNA.此切割之后,在RNA的3′端添加多聚腺苷酸尾;多聚腺苷酸尾似乎增加了所得多聚腺苷酸化RNA的稳定性。
The location of the polyadenylation point is specified in part by the sequence AAUAAA (or a variant of this), usually found in the 3′ untranslated portion of the RNA transcript.多聚腺苷酸化位点的位置部分由序列AAUAAA(或其变体)指定,该序列通常位于RNA转录物的3′非翻译区。
All of these posttranscriptional modifications take place in the nucleus, as does the process of RNA splicing.所有这些转录后修饰均在细胞核内进行,RNA剪接过程也是如此。
The fully processed RNA, now called mRNA, is then transported to the cytoplasm, where translation takes place .经过完全加工的RNA(现称为mRNA)随后被转运至细胞质,在那里进行翻译。
Translation and the Genetic Code In the cytoplasm, mRNA is translated into protein by the action of a variety of short RNA adaptor molecules, the tRNAs, each specific for a particular amino acid.翻译与遗传密码 在细胞质中,mRNA通过多种短RNA适配分子(即tRNA)的作用被翻译成蛋白质,每种tRNA特异于一种特定氨基酸。
These remarkable molecules, each only 70 to 100 nucleotides long, have the job of bringing the correct amino acids into position along the mRNA template, to be added to the growing polypeptide chain.这些非凡的分子每个仅长70至100个核苷酸,其任务是将正确的氨基酸带到mRNA模板上的相应位置,以便添加到正在延伸的多肽链中。
Protein synthesis occurs on ribosomes, macromolecular complexes made up of rRNA (encoded by the 18 S and 28 S rRNA genes), and several dozen ribosomal proteins .蛋白质合成发生在核糖体上,核糖体是由rRNA(由18S和28S rRNA基因编码)和数十种核糖体蛋白组成的大分子复合物。
The key to translation is a code that relates specific amino acids to combinations of three adjacent bases along the mRNA.翻译的关键是一种将特定氨基酸与mRNA上三个相邻碱基的组合联系起来的密码。
Each set of three bases constitutes a codon, specific for a particular amino acid ( In theory, almost infinite variations are possible in the arrangement of the bases along a polynucleotide chain.每组三个碱基构成一个密码子,特异于一种特定氨基酸(理论上,多核苷酸链上的碱基排列方式几乎有无穷多种可能。
At any one position, there are four possibilities (A, T, C, or G); thus, for three bases, there are 43, or 64, possible triplet combinations.在任何位置,有四种可能性(A、T、C或G);因此,对于三个碱基,有4的3次方即64种可能的三联体组合。
These 64 codons constitute the genetic code.这64个密码子构成了遗传密码。
Because there are only 20 amino acids and 64 possible codons, most amino acids are specified by more than one codon; hence the code is said to be degenerate.由于仅有20种氨基酸和64种可能的密码子,大多数氨基酸由不止一个密码子指定;因此该密码被称为简并密码。
For instance, the base in the third position of the triplet can often be either purine (A or G) or either pyrimidine (T or C) or, in some cases, any one of the four bases, without altering the coded message (see Leucine and arginine are each specified by six codons.例如,三联体第三位的碱基常常可以是嘌呤(A或G)或嘧啶(T或C),在某些情况下可以是四种碱基中的任何一种,而不改变编码信息(参见亮氨酸和精氨酸各由六个密码子指定。
Only methionine and tryptophan are each specified by a single, unique codon.只有甲硫氨酸和色氨酸各由一个独特的密码子指定。
Three of the codons are called stop (or nonsense) codons because they designate termination of translation of the mRNA at that point.其中三个密码子被称为终止(或无义)密码子,因为它们指示mRNA翻译在该点终止。
Codons are shown in terms of mRNA, which are complementary to the corresponding DNA codons.密码子以mRNA表示,它们与相应的DNA密码子互补。
29/42
The Human Genome 31 Translation of a processed mRNA is always initiated at a codon specifying methionine.
Ch3 — Segment 29
The Human Genome 31 Translation of a processed mRNA is always initiated at a codon specifying methionine.人类基因组31 经过加工的mRNA的翻译始终从指定甲硫氨酸的密码子开始。
Methionine is therefore the first encoded (amino-terminal) amino acid of each polypeptide chain, although it is usually removed before protein synthesis is completed.因此甲硫氨酸是每条多肽链中第一个编码的(氨基末端)氨基酸,尽管它通常在蛋白质合成完成前被移除。
The codon for methionine (the initiator codon, AUG) establishes the reading frame of the mRNA; each subsequent codon is read in turn to predict the amino acid sequence of the protein.甲硫氨酸的密码子(起始密码子AUG)确立了mRNA的阅读框;随后依次读取每个后续密码子以预测蛋白质的氨基酸序列。
The molecular links between codons and amino acids are the specific tRNA molecules.密码子与氨基酸之间的分子联系是特定的tRNA分子。
A particular site on each tRNA forms a three-base anticodon that is complementary to a specific codon on the mRNA.每个tRNA上的特定位点形成一个三碱基反密码子,与mRNA上的特定密码子互补。
Bonding between the codon and anticodon brings the appropriate amino acid into the next position on the ribosome for attachment, by formation of a peptide bond, to the carboxyl end of the growing polypeptide chain.密码子与反密码子之间的结合将相应的氨基酸带入核糖体上的下一个位置,通过形成肽键连接到生长中多肽链的羧基末端。
The ribosome then slides along the mRNA exactly three bases, bringing the next codon into line for recognition by another tRNA with the next amino acid.随后核糖体沿mRNA精确滑动三个碱基,使下一个密码子进入位置,以便被携带下一个氨基酸的另一个tRNA识别。
Thus proteins are synthesized from the amino terminus to the carboxyl terminus, which corresponds to translation of the mRNA in a 5′ to 3′ direction.因此蛋白质从氨基端向羧基端合成,这对应mRNA从5′到3′方向的翻译。
As mentioned earlier, translation ends when a stop codon (UGA, UAA, or UAG) is encountered in the same reading frame as the initiator codon.如前所述,当在与起始密码子相同的阅读框内遇到终止密码子(UGA、UAA或UAG)时,翻译结束。
(Stop codons in either of the other unused reading frames are not read, and therefore have no effect on translation.) The completed polypeptide is then released from the ribosome, which becomes available to begin synthesis of another protein.(在另外两个未使用的阅读框中的终止密码子不被读取,因此对翻译无影响。)完成的多肽随后从核糖体释放,核糖体可开始合成另一个蛋白质。
Transcription of the Mitochondrial Genome The previous sections described fundamentals of gene expression for genes contained in the nuclear genome.线粒体基因组的转录 前面章节描述了核基因组中基因表达的基本原理。
The mitochondrial genome has its own transcription and protein-synthesis system.线粒体基因组有其自身的转录和蛋白质合成系统。
A specialized RNA polymerase, encoded in the nuclear genome, is used to transcribe the 16-kb mitochondrial genome, which contains two related promoter sequences, one for each strand of the circular genome.一种由核基因组编码的特化RNA聚合酶用于转录16-kb的线粒体基因组,该基因组包含两个相关的启动子序列,分别对应环状基因组的两条链。
Each strand is transcribed in its entirety, and the mitochondrial transcripts are then processed to generate the various individual mitochondrial mRNAs, tRNAs, and rRNAs.每条链被完整转录,然后线粒体转录产物经过加工生成各种独立的线粒体mRNA、tRNA和rRNA。
GENE EXPRESSION IN ACTION The flow of information outlined in the preceding sections can best be appreciated by reference to a particular well-studied gene, the β-globin gene.基因表达的实际运作 前面章节概述的信息流可通过参考一个研究透彻的基因——β-珠蛋白基因——得到最佳理解。
The β-globin chain is a 146–amino acid polypeptide, encoded by a gene that occupies ~1. 6 kb on the short arm of chromosome 11.β-珠蛋白链是一个146个氨基酸的多肽,由位于11号染色体短臂上约1.6 kb的基因编码。
The gene has three exons and two introns .该基因有三个外显子和两个内含子。
The β-globin gene, as well as the other genes in the β-globin cluster , is transcribed in a centromere-to-telomere direction.β-珠蛋白基因以及β-珠蛋白簇中的其他基因,以着丝粒到端粒的方向转录。
The orientation, however, is different for different genes in the genome and depends on which strand of the chromosomal double helix is the coding strand for a particular gene.然而,基因组中不同基因的方向不同,取决于染色体双螺旋的哪条链是特定基因的编码链。
DNA sequences required for accurate initiation of transcription of the β-globin gene are located in the promoter within ~200 bp upstream from the transcription start site (see 2).β-珠蛋白基因精确起始转录所需的DNA序列位于转录起始位点上游约200 bp内的启动子中(见2)。
The double-stranded DNA sequence of this region of the β-globin gene, the corresponding RNA sequence, and the translated sequence of the first 10 amino acids are depicted in .β-珠蛋白基因该区域的双链DNA序列、相应的RNA序列以及前10个氨基酸的翻译序列如图所示。
Because of this correspondence, the 5′ to 3′ DNA strand of a gene (i. e., the strand that is not transcribed) is the strand generally reported in the scientific literature or in databases.由于这种对应关系,基因的5′到3′ DNA链(即不被转录的链)通常是科学文献或数据库中报告的链。
In accordance with this convention, the complete sequence of ~2. 0 kb of chromosome 11 that includes the β-globin gene is shown in Within these 2. 0 kb are contained most, but not all, of the sequence elements required to encode and regulate the expression of this gene.按照这一惯例,包含β-珠蛋白基因的11号染色体约2.0 kb的完整序列如图所示。这2.0 kb内包含了编码和调控该基因表达所需的大部分(但非全部)序列元件。
Indicated in .如图所示。
The polypeptide chain that is the primary translation product folds on itself and forms intramolecular bonds to create a specific 3D structure that is determined by the amino acid sequence itself.作为初级翻译产物的多肽链自身折叠并形成分子内键,从而产生由氨基酸序列本身决定的特定三维结构。
Two or more polypeptide chains, products of the same gene or of different genes, may combine to form a single multiprotein complex.两条或多条多肽链(来自同一基因或不同基因的产物)可能结合形成单个多蛋白复合物。
For example, two α-globin chains and two β-globin chains associate noncovalently to form a tetrameric hemoglobin molecule (see Chapter 12).例如,两条α-珠蛋白链和两条β-珠蛋白链通过非共价结合形成四聚体血红蛋白分子(见第12章)。
The protein products may also be modified chemically by, for example, addition of methyl groups, phosphates, or carbohydrates at specific sites.蛋白质产物也可能发生化学修饰,例如在特定位点添加甲基、磷酸基团或碳水化合物。
These modifications can have significant influence on the function or abundance of the modified protein.这些修饰可能对修饰后蛋白质的功能或丰度产生显著影响。
Other modifications may involve cleavage of the protein, either to remove specific amino-terminal sequences after they have functioned to direct a protein to its correct location within the cell (e. g., proteins that function within mitochondria) or to split the molecule into smaller polypeptide chains.其他修饰可能涉及蛋白质的切割,要么在氨基末端序列发挥引导蛋白质到达细胞内正确位置(例如在线粒体内起作用的蛋白质)的功能后将其移除,要么将分子分割成更小的多肽链。
For example, the two chains that make up mature insulin, one 21 and the other 30 amino acids long, are originally part of an 82–amino acid primary translation product called proinsulin.例如,组成成熟胰岛素的两条链(一条含21个氨基酸,另一条含30个氨基酸)最初是名为胰岛素原的82个氨基酸初级翻译产物的一部分。
The SSH gene involved in holoprosencephaly (see Case Report 23) encodes a protein that also goes through a series of processing steps before being secreted from the cell.参与前脑无裂畸形(见病例报告23)的SSH基因编码一种蛋白质,该蛋白质在从细胞分泌前也需经过一系列加工步骤。
30/42
known to be mutated in various inherited defects of the β-globin gene (see Chapter 12).
Ch3 — Segment 30
known to be mutated in various inherited defects of the β-globin gene (see Chapter 12).已知在β-珠蛋白基因的各种遗传缺陷中发生突变(见第12章)。
Initiation of Transcription The β-globin promoter, like many other gene promoters, consists of a series of relatively short functional elements that interact with specific regulatory proteins (generically called transcription factors) that control transcription, including, in the case of the globin genes, those proteins that restrict expression of these genes to erythroid cells, the cells in which hemoglobin is produced.转录起始 β-珠蛋白启动子与许多其他基因启动子一样,由一系列相对较短的功能元件组成,这些元件与特异性调控蛋白(统称为转录因子)相互作用,从而控制转录;就珠蛋白基因而言,这些调控蛋白包括将基因表达限制在红细胞(产生血红蛋白的细胞)中的那些蛋白质。
There are well over a thousand sequence-specific, DNA-binding transcription factors in the genome, some of which are ubiquitous in their expression, whereas others are cell type or tissue specific.基因组中存在超过一千种序列特异性的DNA结合转录因子,其中一些表达普遍存在,而另一些则具有细胞类型或组织特异性。
One important promoter sequence found in many but not all genes is the TATA box, a conserved region rich in adenines and thymines that is ~25 to 30 bp upstream of the start site of transcription (see Figs. 3. 4 and 3. 7).在许多但并非所有基因中发现的一个重要的启动子序列是TATA盒,这是一个富含腺嘌呤和胸腺嘧啶的保守区域,位于转录起始位点上游约25至30 bp处(见图3.4和图3.7)。
The TATA box appears to be important for determining the position of the start of transcription, which in the β-globin gene is ~50 bp upstream from the translation initiation site .TATA盒对于确定转录起始位置似乎很重要,在β-珠蛋白基因中,该位置距离翻译起始位点约50 bp上游。
Thus in this gene, there are ~50 bp of sequence at the 5′ end that are transcribed but are not translated; in other genes, the 5′ UTR can be much longer and can even be interrupted by one or more introns.因此,在该基因中,5′端约有50 bp的序列被转录但不被翻译;在其他基因中,5′ UTR可能更长,甚至可能被一个或多个内含子打断。
A second conserved region, the so-called CAT box (actually CCAAT), is a few dozen base pairs farther upstream .第二个保守区域,即所谓的CAT盒(实际上是CCAAT),位于更上游几十个碱基对处。
Both experimentally induced and naturally occurring variants in either of these sequence elements, as well as in other regulatory sequences even farther upstream, lead to a sharp reduction in the level of transcription, thereby demonstrating the importance of these elements for normal gene expression.这些序列元件中以及更上游的其他调控序列中的实验诱导变异和自然发生变异都会导致转录水平急剧下降,从而证明了这些元件对正常基因表达的重要性。
Many variants in these regulatory elements have been identified in individuals with the hemoglobin disorder β-thalassemia (see Chapter 12).在患有血红蛋白疾病β-地中海贫血的个体中已发现这些调控元件的许多变异(见第12章)。
Not all gene promoters contain the two specific elements just described.并非所有基因启动子都包含刚才描述的两种特异性元件。
Importantly, genes that are constitutively expressed in most or all tissues (so-called housekeeping genes) often lack the CAT and TATA boxes, which are more typical of tissue-specific genes.重要的是,在大多数或所有组织中组成型表达的基因(即所谓的管家基因)通常缺乏CAT盒和TATA盒,而这些元件更典型于组织特异性基因。
Promoters of many housekeeping genes contain a high proportion of cytosines and guanines in relation to the surrounding DNA (see the promoter of the BRCA1 breast cancer gene in .许多管家基因的启动子相对于周围DNA含有高比例的胞嘧啶和鸟嘌呤(参见BRCA1乳腺癌基因的启动子)。
Such CG-rich promoters are often located in regions of the genome called Cp G islands, so named because of the unusually high concentration of the dinucleotide 5′-Cp G-3′ (the p representing the phosphate group between adjacent bases) that stands out from the more general AT-rich genomic landscape.此类富含CG的启动子通常位于称为CpG岛的基因组区域,之所以如此命名是因为二核苷酸5′-CpG-3′(p表示相邻碱基之间的磷酸基团)的浓度异常高,从而在普遍富含AT的基因组背景中显得突出。
Some of the CG-rich sequence elements found in these promoters are thought to serve as binding sites for specific transcription factors.这些启动子中发现的一些富含CG的序列元件被认为是特定转录因子的结合位点。
Cp G islands are also important because they are targets for DNA methylation.CpG岛也很重要,因为它们是DNA甲基化的靶标。
Extensive DNA methylation at Cp G islands is usually associated with repression of gene transcription, as we will discuss later in the context of chromatin and its role in the control of gene expression (see Chapter 8).CpG岛上的广泛DNA甲基化通常与基因转录抑制相关,我们将在后面讨论染色质及其在基因表达调控中的作用时详述(见第8章)。
Transcription by RNA polymerase II (RNA pol II) is subject to regulation at multiple levels, including binding to the promoter, initiation of transcription, unwinding of the DNA double helix to expose the template strand, and elongation as RNA pol II moves along the DNA.RNA聚合酶II(RNA pol II)的转录在多个水平受到调控,包括与启动子的结合、转录起始、解开DNA双螺旋以暴露模板链,以及RNA pol II沿DNA移动时的延伸。
Although some silenced genes are devoid of RNA pol II binding altogether, consistent with their inability to be transcribed in a given cell type, others have RNA pol II poised bidirectionally at the transcriptional start site, perhaps as a means of fine-tuning transcription in response to particular cellular signals.尽管一些沉默基因完全缺乏RNA pol II的结合,这与它们在特定细胞类型中无法转录一致,但另一些基因则在转录起始位点处有RNA pol II双向待命,这可能是作为一种响应特定细胞信号微调转录的方式。
In addition to the sequences that constitute a promoter itself are other sequence elements that can markedly alter the efficiency of transcription.除了构成启动子本身的序列之外,还有其他序列元件可以显著改变转录效率。
The best characterized of these activating sequences are called enhancers.这些激活序列中表征最清楚的是增强子。
Enhancers are sequence elements that can act at a distance from a gene to stimulate transcription.增强子是能够远离基因发挥作用以刺激转录的序列元件。
Enhancers can be located several or even hundreds of kilobases away from a gene, and in the case of the Sonic hedgehog (SHH) gene there can be many, with some being 1 million bp away, acting in different tissues.增强子可以位于距离基因几kb甚至数百kb处,以Sonic hedgehog(SHH)基因为例,其增强子可能很多,有些甚至位于100万bp之外,并在不同组织中发挥作用。
Unlike promoters, enhancers are both position and orientation independent and can be located either 5′ or 3′ of the transcription start site.与启动子不同,增强子具有位置和方向独立性,可以位于转录起始位点的5′或3′端。
Specific enhancer elements function only in certain cell types and thus appear to be involved in establishing the tissue specificity or level of..................特定的增强子元件仅在某些细胞类型中起作用,因此似乎参与建立组织特异性或表达水平..................
Val His Leu Thr Pro Glu Glu Lys Ser Ala DNA mRNA Start Transcription Translation Reading frame β-globin Transcription of the 3′ to 5′ (lower) strand begins at the indicated start site to produce β-globin messenger RNA (mRNA).Val His Leu Thr Pro Glu Glu Lys Ser Ala DNA mRNA 起始 转录 翻译 阅读框 β-珠蛋白 从3′到5′(下方)链的转录在所示起始位点开始,产生β-珠蛋白信使RNA(mRNA)。
The translational reading frame is determined by the AUG initiator codon (); subsequent codons specifying amino acids are indicated in blue.翻译阅读框由AUG起始密码子决定;后续指定氨基酸的密码子以蓝色表示。
The other two potential frames are not used.其他两个潜在阅读框未被使用。
31/42
The Human Genome 33 expression of many genes, in concert with one or more transcription factors.
Ch3 — Segment 31
The Human Genome 33 expression of many genes, in concert with one or more transcription factors.人类基因组33许多基因的表达,与一个或多个转录因子协同作用。
In the case of the β-globin gene, several tissue-specific enhancers are present both within the gene itself and in its flanking regions.就β-珠蛋白基因而言,多个组织特异性增强子存在于基因本身及其侧翼区域。
The interaction of enhancers with specific regulatory proteins leads to increased levels of transcription.增强子与特定调节蛋白的相互作用导致转录水平升高。
Normal expression of the β-globin gene during development also requires more distant sequences called the locus control region (LCR), located upstream of the ε-globin gene , which is required for establishing the proper chromatin context needed for appropriate high-level expression.β-珠蛋白基因在发育过程中的正常表达还需要更远的序列,称为位点控制区(LCR),位于ε-珠蛋白基因上游,该区域对于建立适当高水平表达所需的正确染色质环境是必需的。
As expected, variants that disrupt or delete either enhancer or LCR sequences interfere with or prevent β-globin gene expression (see Chapter 12).正如预期,破坏或删除增强子或LCR序列的变异体会干扰或阻止β-珠蛋白基因的表达(参见第12章)。
RNA Splicing The primary RNA transcript of the β-globin gene contains two introns, ~100 and 850 bp in length, that need *** Exon 1 Exon 2 Exon 3 The sequence of the 5′ to 3′ strand of the gene is shown.RNA剪接 β-珠蛋白基因的初级RNA转录本包含两个内含子,长度分别约为100和850 bp,需要*** 外显子1 外显子2 外显子3 显示了基因5′至3′链的序列。
Tan areas with capital letters represent exonic sequences corresponding to mature mRNA.带有大写字母的棕褐色区域代表对应于成熟mRNA的外显子序列。
Lowercase letters indicate introns and flanking sequences.小写字母表示内含子和侧翼序列。
The CAT and TATA box sequences in the 5′ flanking region are indicated in brown.5′侧翼区域中的CAT和TATA盒序列以棕色标示。
The GT and AG dinucleotides important for RNA splicing at the intron-exon junctions and the AATAAA signal important for addition of a poly A tail are also highlighted.对内含子-外显子连接处RNA剪接重要的GT和AG二核苷酸,以及对添加多聚腺苷酸尾重要的AATAAA信号也以高亮显示。
The ATG initiator codon (AUG in mRNA) and the TAA stop codon (UAA in mRNA) are shown in red letters.ATG起始密码子(mRNA中为AUG)和TAA终止密码子(mRNA中为UAA)以红色字母显示。
The amino acid sequence of β-globin is shown above the coding sequence; the three-letter abbreviations in (Original data from Lawn RM, Efstratiadis A, O'Connell C, et al: The nucleotide sequence of the human β-globin gene.β-珠蛋白的氨基酸序列显示在编码序列上方;括号中的三字母缩写(原始数据来自Lawn RM, Efstratiadis A, O'Connell C, 等:人类β-珠蛋白基因的核苷酸序列。
Cell 21:647–651, 1980.)《细胞》21:647–651, 1980。)
32/42
to be removed and the remaining RNA segments joined together to form the mature mRNA.
Ch3 — Segment 32
to be removed and the remaining RNA segments joined together to form the mature mRNA.待移除部分,剩余RNA片段连接形成成熟mRNA。
The process of RNA splicing, described generally earlier, is typically an exact and highly efficient one; 95% of β-globin transcripts are thought to be accurately spliced to yield functional globin mRNA.如前所述,RNA剪接过程通常精确且高效;据认为95%的β-珠蛋白转录本被准确剪接,产生功能性珠蛋白mRNA。
The splicing reactions are guided by specific sequences in the primary RNA transcript at both 5′ and 3′ ends of introns.剪接反应由初级RNA转录本中内含子5′端和3′端的特定序列引导。
The 5′ sequence consists of nine nucleotides, of which two (the dinucleotide GT [GU in the RNA transcript] located in the intron immediately adjacent to the splice site) are virtually invariant among splice sites in different genes .5′端序列由九个核苷酸组成,其中两个(位于剪接位点紧邻内含子中的二核苷酸GT [RNA转录本中为GU])在不同基因的剪接位点间几乎不变。
The 3′ sequence consists of approximately a dozen nucleotides, of which, again, two—the AG located immediately 5′ to the intron-exon boundary—are obligatory for normal splicing.3′端序列由约十二个核苷酸组成,其中同样有两个——位于内含子-外显子边界紧邻5′端的AG——对正常剪接是必需的。
The splice sites themselves are unrelated to the reading frame of the particular mRNA.剪接位点本身与特定mRNA的阅读框无关。
In some instances, as in the case of intron 1 of the β-globin gene, the intron actually splits a specific codon .在某些情况下,如β-珠蛋白基因的内含子1,内含子实际上分裂了一个特定密码子。
The medical significance of RNA splicing is illustrated by the fact that variants within the conserved sequences at the intron-exon boundaries commonly impair RNA splicing, with a concomitant reduction in the amount of normal, mature β-globin mRNA; alterations in the GT or AG dinucleotides mentioned earlier invariably eliminate normal splicing of the intron containing the variant.RNA剪接的医学意义体现在:内含子-外显子边界保守序列内的变异通常会损害RNA剪接,同时减少正常成熟β-珠蛋白mRNA的量;前述GT或AG二核苷酸的改变必然消除含该变异的内含子的正常剪接。
Representative splice site variants identified in patients with β-thalassemia are discussed in detail in Chapter 12.在β-地中海贫血患者中鉴定出的代表性剪接位点变异将在第12章详细讨论。
Alternative Splicing As just discussed, when introns are removed from the primary RNA transcript by RNA splicing, the remaining exons are spliced together to generate the final, mature mRNA.选择性剪接 如上所述,当内含子通过RNA剪接从初级RNA转录本中移除后,剩余外显子连接在一起生成最终成熟mRNA。
However, for most genes, the primary transcript can follow multiple alternative splicing pathways, leading to the synthesis of multiple related but different mRNAs, each of which can be subsequently translated to generate different protein products .然而,对于大多数基因,初级转录本可遵循多条选择性剪接途径,导致合成多种相关但不同的mRNA,每一种随后可被翻译生成不同蛋白质产物。
Some of these alternative events are tissue or cell type specific, and, to the extent that such events are determined by primary sequence, they are subject to allelic variation between different individuals.其中一些选择性事件具有组织或细胞类型特异性,并且,在此类事件由初级序列决定的情况下,它们在不同个体间存在等位基因变异。
Nearly all human genes undergo alternative splicing to some degree, and it has been estimated that there are an average of two or three alternative transcripts per gene in the human genome, thus greatly expanding the information content of the human genome beyond the ~20,000 protein-coding genes.几乎所有人类基因都在一定程度上经历选择性剪接,据估计人类基因组中每个基因平均有两到三种选择性转录本,从而将人类基因组的信息含量大幅扩展至约20,000个蛋白编码基因之上。
The regulation of alternative splicing appears to play a particularly impressive role during brain development, where it may contribute to generating the high levels of functional diversity needed in the nervous system.选择性剪接的调控在大脑发育中似乎扮演着尤为突出的角色,可能有助于产生神经系统所需的高水平功能多样性。
The reason for this may be because genes expressed in the brain tend to be larger in size and have more exons than those expressed in other tissues.其原因可能是大脑中表达的基因通常比其他组织中表达的基因更大且拥有更多外显子。
Consistent with this, susceptibility to a number of neurodevelopmental conditions has been associated with shifts or disruption of alternative splicing patterns and other spontaneous, rare germline, and even somatic events.与此一致,多种神经发育条件的易感性与选择性剪接模式的转变或破坏以及其他自发、罕见胚系甚至体细胞事件相关联。
Polyadenylation The mature β-globin mRNA contains ~130 bp of 3′ untranslated material (the 3′ UTR) between the stop codon and the location of the poly A tail .多聚腺苷酸化 成熟的β-珠蛋白mRNA在终止密码子和多聚A尾位置之间含有约130 bp的3′非翻译区(3′ UTR)。
As in other genes, cleavage of the 3′ end of the mRNA and addition of the poly A tail is controlled, at least in part, by an AAUAAA sequence ~20 bp before the polyadenylation site.与其他基因一样,mRNA 3′端的切割和多聚A尾的添加至少部分由多聚腺苷酸化位点前约20 bp处的AAUAAA序列控制。
Pathogenic variants in this polyadenylation signal in patients with β-thalassemia document the importance of this signal for proper 3′ cleavage and polyadenylation (see Chapter 12).β-地中海贫血患者中该多聚腺苷酸化信号的致病变异证实了该信号对正确3′切割和多聚腺苷酸化的重要性(见第12章)。
The 3′ UTR of some genes can be up to several kb in length.某些基因的3′ UTR长度可达数kb。
Other genes have a number of alternative polyadenylation sites, selection among which may influence the stability of the resulting mRNA and thus the steady-state level of each mRNA.其他基因拥有多个选择性多聚腺苷酸化位点,对其中位点的选择可能影响所得mRNA的稳定性,进而影响每种mRNA的稳态水平。
RNA Editing and RNA-DNA Sequence Differences Recent findings suggest that the conceptual principle underlying the central dogma—that RNA and protein sequences reflect the underlying genomic sequence—may not always hold true.RNA编辑与RNA-DNA序列差异 近期发现表明,中心法则的基本概念——RNA和蛋白质序列反映潜在基因组序列——可能并非始终成立。
RNA editing to change the nucleotide sequence of the mRNA has been demonstrated in a number of organisms, including humans.已在包括人类在内的多种生物中证明RNA编辑可改变mRNA的核苷酸序列。
This process involves deamination of adenosine at particular sites, converting an A in the DNA sequence to an inosine in the resulting RNA; this is then read by the translational machinery as a G, leading to changes in gene expression and protein function, especially in the nervous system.该过程涉及特定位点腺苷的脱氨基作用,将DNA序列中的A转变为所得RNA中的肌苷;随后翻译机器将其读作G,导致基因表达和蛋白质功能的改变,尤其在神经系统中。
More widespread RNA-DNA differences involving other bases (with corresponding changes in the encoded amino acid sequence) have also been reported, at levels that vary among individuals.涉及其他碱基(并导致编码氨基酸序列相应改变)的更广泛RNA-DNA差异也已被报道,其水平在不同个体间存在差异。
Although the mechanism(s) and clinical relevance of these events remain controversial, they illustrate the existence of a range of processes capable of increasing transcript and proteome diversity.尽管这些事件的机制和临床意义仍存争议,但它们表明存在一系列能增加转录本和蛋白质组多样性的过程。
EPIGENETIC AND EPIGENOMIC ASPECTS OF GENE EXPRESSION Given the range of functions and fates that different cells in any organism must adopt over its lifetime, it is apparent that not all genes in the genome can be actively expressed in every cell at all times.基因表达的表观遗传与表观基因组方面 鉴于任何生物体中不同细胞在其生命周期中必须承担的功能和命运范围,显然基因组中并非所有基因在每个细胞中都随时活跃表达。
As important as completion of the Human Genome Project has been for contributing to our understanding of human biology and disease, identifying the genomic sequences and features that direct developmental, spatial, and temporal aspects of gene expression remains a formidable challenge.尽管人类基因组计划的完成对我们理解人类生物学和疾病至关重要,但识别指导基因表达发育、空间和时间方面的基因组序列和特征仍是一项艰巨挑战。
Several decades of work in molecular biology have defined critical regulatory elements for many individual genes, as we saw in the previous section, and正如我们在前一节所见,数十年的分子生物学工作已为许多单个基因定义了关键调控元件。
33/42
The Human Genome 35 more recent attention has been directed toward performing such studies on a genome-wide scale.
Ch3 — Segment 33
The Human Genome 35 more recent attention has been directed toward performing such studies on a genome-wide scale.人类基因组35 近期的注意力已转向在全基因组规模上进行此类研究。
In Chapter 2 we introduced general aspects of chromatin that package the genome and its genes in all cells.在第2章中,我们介绍了染色质的一般特性,这些特性在所有细胞中包装基因组及其基因。
Here, we explore the specific characteristics of chromatin that are associated with active or repressed genes as a step toward identifying the regulatory code for expression of the human genome.在此,我们探讨与活跃或抑制基因相关的染色质特定特征,作为迈向识别人类基因组表达调控代码的一步。
Such studies focus on reversible changes in the chromatin landscape as determinants of gene function rather than on changes to the genome sequence itself and are thus called epigenetic or, when considered in the context of the entire genome, epigenomic (Greek epi-, over or upon).此类研究关注染色质景观中的可逆变化作为基因功能的决定因素,而非基因组序列本身的变化,因此被称为表观遗传的,或当放在整个基因组背景下考虑时,称为表观基因组(希腊语epi-,意为在…之上或之上)。
The field of epigenetics is growing rapidly and is the study of heritable changes in cellular function or gene expression that can be transmitted from cell to cell (and even generation to generation), as a result of chromatinbased molecular signals .表观遗传学领域正在迅速发展,它研究由染色质分子信号导致的细胞功能或基因表达的可遗传变化,这些变化可以在细胞间(甚至代际间)传递。
Complex epigenetic states can be established, maintained, and transmitted by a variety of mechanisms: modifications to the DNA, such as DNA methylation; numerous histone modifications that alter chromatin packaging or access; and substitution of specialized histone variants that mark chromatin associated with particular sequences or Chromosome Expressed gene Nucleosomes Modifications on histone tails Histone variants to mark specific regions DNA methylation Repressed gene Me CG GC Me Me CG GC Me Me CG GC Me Me CG GC Me Not to scale. regions in the genome.复杂的表观遗传状态可以通过多种机制建立、维持和传递:对DNA的修饰(如DNA甲基化);大量改变染色质包装或可及性的组蛋白修饰;以及替换专门的组蛋白变体,这些变体标记与特定序列或染色体表达的基因、核小体、组蛋白尾部修饰、组蛋白变体以标记特定区域、DNA甲基化、抑制基因Me CG GC Me Me CG GC Me Me CG GC Me Me CG GC Me(未按比例缩放)基因组中的区域相关的染色质。
These chromatin changes can be highly dynamic and transient, capable of responding rapidly and sensitively to changing needs in the cell, or they can be long lasting, capable of being transmitted through multiple cell divisions or even to subsequent generations.这些染色质变化可以是高度动态和瞬时性的,能够快速灵敏地响应细胞中变化的需求,也可以是持久的,能够通过多次细胞分裂甚至传递给后代。
In either instance, the key concept is that epigenetic mechanisms do not alter the underlying DNA sequence, and this distinguishes them from genetic mechanisms, which are sequence based.无论哪种情况,关键概念是表观遗传机制不改变潜在的DNA序列,这使其与基于序列的遗传机制区分开来。
Together, the epigenetic marks and the DNA sequence make up the set of signals that guide the genome to express its genes at the right time, in the right place, and in the right amounts.共同地,表观遗传标记和DNA序列构成了一组信号,引导基因组在正确的时间、正确的位置以正确的量表达其基因。
These topics are covered in detail in Chapter 8.这些主题在第8章中详细讨论。
DNA Methylation DNA methylation involves the modification of cytosine bases by methylation of the carbon at the fifth position in the pyrimidine ring .DNA甲基化 DNA甲基化涉及通过甲基化嘧啶环中第五位碳来修饰胞嘧啶碱基。
Extensive DNA methylation is a mark of repressed genes and is a widespread mechanism associated with the establishment of specific programs of gene expression during cell differentiation and development.广泛的DNA甲基化是抑制基因的标志,并且是一种广泛存在的机制,与细胞分化和发育过程中特定基因表达程序的建立相关。
Typically, DNA methylation occurs on the C of Cp G dinucleotides通常,DNA甲基化发生在CpG二核苷酸的C上。
34/42
and inhibits gene expression by recruitment of specific methyl-Cp G–binding proteins that, in turn, recruit chromatin-mo…
Ch3 — Segment 34
and inhibits gene expression by recruitment of specific methyl-Cp G–binding proteins that, in turn, recruit chromatin-modifying enzymes to silence transcription.[TL:failed]
The presence of 5-methylcytosine (5-m C) is considered to be a stable epigenetic mark that can be faithfully transmitted through cell division; however, altered methylation states are frequently observed in cancer, with hypomethylation of large genomic segments or with regional hypermethylation (particularly at Cp G islands) in others (see Chapter 16).[TL:failed]
Extensive demethylation occurs during germ cell development and in the early stages of embryonic development, consistent with the need to reset the chromatin environment and restore totipotency or pluripotency of the zygote and of various stem cell populations.[TL:failed]
Although the details are still incompletely understood, these reprogramming steps appear to involve the enzymatic conversion of 5-m C to 5-hydroxymethylcytosine (5-hm C) , as a likely intermediate in the demethylation of DNA.[TL:failed]
Overall, 5-m C levels are stable across adult tissues (~5% of all cytosines), whereas 5-hm C levels are much lower and much more variable (0. 1–1% of all cytosines).[TL:failed]
Interestingly, although 5-hm C is widespread in the genome, its highest levels are found in known regulatory regions, suggesting a possible role in the regulation of specific promoters and enhancers.[TL:failed]
Histone Modifications A second class of epigenetic signals consists of an extensive inventory of modifications to any of the core histone types: H2A, H2B, H3, and H4 (see Chapter 2).[TL:failed]
Such modifications include histone methylation, phosphorylation, acetylation, and others at specific amino acid residues, mostly located on the N-terminal tails of histones that extend out from the core nucleosome itself .[TL:failed]
These epigenetic modifications are believed to influence gene expression by affecting chromatin compaction or accessibility and by signaling protein complexes that—depending on the nature of the signal—activate or silence gene expression at that site.[TL:failed]
There are dozens of modified sites that can be experimentally queried genome-wide by using antibodies that recognize specifically modified sites—for example, histone H3 methylated at lysine position 9 (H3K9 methylation, using the one-letter abbreviation K for lysine; see The former is a repressive mark associated with silent regions of the genome, whereas the latter is a mark for activating regulatory regions.[TL:failed]
Histone Variants The histone modifications just discussed involve modification of the core histones themselves, which are all encoded by multigene clusters in a few locations in the genome.[TL:failed]
In contrast, the many dozens of histone variants are products of entirely different genes located elsewhere in the genome, and their amino acid sequences are distinct from, although related to, those of the canonical histones.[TL:failed]
Different histone variants are associated with different functions, and they replace—all or in part—the related member of the core histones found in typical nucleosomes to generate specialized chromatin structures .[TL:failed]
Some variants mark specific regions or loci in the genome with highly specialized functions (e. g., the CENP-A histone is a histone H3-related variant that is found exclusively at functional centromeres in the genome and contributes to essential features of centromeric chromatin that mark the location of kinetochores along the chromosome fiber).[TL:failed]
Other variants are more transient and mark regions of the genome with particular attributes (e. g., H2A.[TL:failed]
X is a histone H2A variant involved in the response to DNA damage to mark regions of the genome that require DNA repair).[TL:failed]
Chromatin Architecture In contrast to the impression one gets from viewing the genome as a linear string of sequence , the genome adopts a highly ordered and dynamic arrangement within the space of the nucleus, correlated with and likely guided by the epigenetic and epigenomic signals just discussed.[TL:failed]
This three-dimensional (3D) landscape is highly predictive of the map of all expressed sequences in any given cell type (the transcriptome) and reflects dynamic changes in chromatin architecture at different levels .[TL:failed]
First, large chromosomal domains (up to millions of base pairs in size) can exhibit coordinated patterns of gene expression at the chromosome level, involving dynamic interactions between different intrachromosomal and interchromosomal points of contact within the nucleus.[TL:failed]
At a finer level, technical advances to map and sequence points of contact around the genome in the context of 3D space have pointed to ordered loops of chromatin that position and orient genes precisely, exposing or blocking critical regulatory regions for access by RNA pol II, transcription factors, 5-Methylcytosine (5-m C) 5-Hydroxymethylcytosine (5-hm C) N 1 2 4 5 6 N 1 2 3 4 5 6 3 NH2 O CH C C C H N CH3 NH2 O CH C C C H N CH2 OH Compare to the structure of cytosine in[TL:failed]
35/42
The Human Genome 37 and other regulators (see Chapter 2).
Ch3 — Segment 35
The Human Genome 37 and other regulators (see Chapter 2).人类基因组37及其他调控因子(见第2章)。
Lastly, specific and dynamic patterns of nucleosome positioning differ among cell types and tissues in the face of changing environmental and developmental cues .最后,核小体定位的特异性和动态模式在不同细胞类型和组织中,面对不断变化的环境和发育信号而有所不同。
The biophysical, epigenomic, and/or genomic properties that facilitate or specify the orderly and dynamic packaging of each chromosome during each cell cycle, without reducing the genome to a disordered tangle within the nucleus, remain a marvel of landscape engineering.促进或规定每个细胞周期中每条染色体有序且动态包装的生物物理、表观基因组和/或基因组特性,同时避免基因组在细胞核内陷入无序缠结,仍是一项景观工程的奇迹。
Topologically Associating Domains Philipp Maass Chromosomes are organized in the 3D space of the nucleus to fulfill gene regulation.拓扑关联结构域 Philipp Maass 染色体在细胞核的三维空间中组织以实现基因调控。
Within this 3D genome organization, chromosomes are spatially segregated in A- and B-type genomic compartments that represent active (euchromatin = open chromatin) and inactive (heterochromatin = repressive chromatin) domains, respectively.在这种三维基因组组织中,染色体在空间上被分隔为A型和B型基因组区室,分别代表活跃(常染色质=开放染色质)和不活跃(异染色质=抑制性染色质)结构域。
Euchromatin is typically associated with higher gene density and early replication, while repressive chromatin—which tends to be transcriptionally silent—is densely packed and protects chromosome integrity.常染色质通常与较高的基因密度和早期复制相关,而抑制性染色质——往往在转录上沉默——密集包装并保护染色体完整性。
Genomic compartmentalization can comprise multiple subcompartments with specific histone marks, which contribute to the general organization of the genome in the nucleus by establishing subnuclear domains with various functions.基因组区室化可包含多个具有特定组蛋白修饰的亚区室,这些亚区室通过建立具有多种功能的核内结构域,有助于细胞核内基因组的整体组织。
For example, the nucleolus is the largest subnuclear compartment and forms at the site of hundreds of ribosomal genes of the five acrocentric chromosomes (13, 14, 15, 21, and 22) for the preassembly of the ribosomal subunits.例如,核仁是最大的核内区室,在五条近端着丝粒染色体(13、14、15、21和22)的数百个核糖体基因位点形成,用于核糖体亚基的预组装。
Splicing or nuclear speckles are nuclear domains enriched for splicing machinery; Cajal bodies relate to mRNA processing, and promyelocytic leukemia (PML) bodies are involved in cell cycle processes and DNA repair.剪接体或核斑点是富含剪接机制的核结构域;卡哈尔体与mRNA加工相关,早幼粒细胞白血病(PML)体参与细胞周期过程和DNA修复。
Collectively, these 3D organizational hubs in the nuclei regulate gene expression and posttranscriptional processing of RNAs, and facilitate tissue-specific gene regulation.总的来说,细胞核中的这些三维组织枢纽调节基因表达和RNA的转录后加工,并促进组织特异性基因调控。
Individual chromosome territories Long-range regulatory element Active gene expression Promoter Chromatin loops and intrachromosomal and interchromosomal interactions Nucleosome positioning allows access to exposed DNA elements A B C D Within interphase nuclei, each chromosome occupies a particular territory, represented by the different colors.单个染色体领土 长程调控元件 活跃基因表达 启动子 染色质环和染色体内及染色体间相互作用 核小体定位允许接触暴露的DNA元件 A B C D 在间期细胞核内,每条染色体占据一个特定的领土,以不同颜色表示。
(B) Chromatin is organized into large subchromosomal domains within each territory, with loops that bring certain sequences and genes into proximity with each other, with detectable intrachromosomal and interchromosomal interactions.(B) 染色质在每个领土内组织成大的亚染色体结构域,环状结构使某些序列和基因彼此靠近,并可检测到染色体内和染色体间相互作用。
(C) Loops bring long-range regulatory elements (e. g., enhancers or locus-control regions) into association with promoters, leading to active transcription and gene expression.(C) 环状结构将长程调控元件(例如增强子或基因座控制区)与启动子相关联,导致活跃转录和基因表达。
(D) Positioning of nucleosomes along the chromatin fiber provides access to specific DNA sequences for binding by transcription factors and other regulatory proteins.(D) 核小体沿染色质纤维的定位提供了特定DNA序列以供转录因子和其他调控蛋白结合的途径。
36/42
On the molecular level genomic compartments are subdivided into clusters of genomic interactions, termed topologically a…
Ch3 — Segment 36
On the molecular level genomic compartments are subdivided into clusters of genomic interactions, termed topologically associating domains (TADs).在分子水平上,基因组区室被细分为基因组相互作用的簇,称为拓扑关联结构域(TADs)。
They range from several hundred kilobase- to megabase-long genomic regions that interact with themselves.它们的大小从几百千碱基到兆碱基长的基因组区域不等,这些区域内部相互作用。
The organization of TADs between cell types and species is consistent.TADs在不同细胞类型和物种间的组织模式是一致的。
TADs are separated from one another by regions with less frequent interactions, termed TAD boundaries.TADs之间由相互作用频率较低的区域分隔,这些区域称为TAD边界。
A model called loop-extrusion can explain the formation of TADs as organizational structures at the subchromosomal level.一种称为环挤压的模型可以解释TADs作为亚染色体水平组织结构形成的过程。
Loop-extrusion suggests bringing otherwise distal gene-regulatory elements (i. e., enhancers and silencers) into 3D proximity of target genes to regulate their expression.环挤压模型提出,原本远端的基因调控元件(即增强子和沉默子)被带入靶基因的三维邻近区域,以调节其表达。
Loop extrusion forms the majority of TADs by two cohesin/condensin molecules sliding toward each other while extruding the DNA from in between, until two convergent CTCF (CCCTCbinding factor) sites are recognized.环挤压通过两个粘连蛋白/凝缩蛋白分子向彼此滑动,同时将中间的DNA挤出,直至识别到两个聚合的CTCF(CCCTC结合因子)位点,从而形成大多数TADs。
CTCF is a highly conserved zinc finger protein, considered the master regulator of the genome because it acts as a transcriptional repressor to regulate the communication between gene-regulatory elements and genes.CTCF是一种高度保守的锌指蛋白,被认为是基因组的主调控因子,因为它作为转录抑制因子,调节基因调控元件与基因之间的通讯。
TAD boundaries often show enrichment of CTCF, cohesion, and mediator proteins, which contribute to the 3D topology of the genome by establishing chromatin loops and the TAD structure.TAD边界通常富含CTCF、粘连蛋白和中介蛋白复合物,这些蛋白通过建立染色质环和TAD结构,帮助形成基因组的三维拓扑结构。
Weaker TAD boundaries may allow for interactions between different TADs (inter-TAD) to regulate genes at larger genomic distances; however, this occurs less frequently than interactions within the same TAD (intra-TAD).较弱的TAD边界可能允许不同TAD之间的相互作用(跨TAD),从而在更大的基因组距离上调节基因;然而,这种情况的发生频率低于同一TAD内的相互作用(TAD内)。
Most gene-regulatory processes occur by intra-TAD chromatin loops; an enhancer with bound transcription factors reaches spatial proximity with the target gene’s promoter site and its core transcriptional machinery (i. e., polymerase II), just upstream of its transcription start site.大多数基因调控过程通过TAD内的染色质环发生;结合了转录因子的增强子与靶基因启动子位点及其核心转录机制(即RNA聚合酶II)在转录起始位点上游达到空间邻近。
The analysis of intra-TAD interactions in different cell types and single cell studies showed high variability, indicating that tissue-specific gene expression can be partially explained by specific cells and by genomic contacts within TADs.对不同细胞类型和单细胞研究中TAD内相互作用的分析显示出高度变异性,表明组织特异性基因表达可部分由特定细胞及TAD内的基因组接触解释。
The reorganization of the TAD architecture by chromosomal rearrangements can alter gene expression and may cause clinically apparent disease phenotypes (see Chapter 6).染色体重排导致TAD结构的重组可能改变基因表达,并可能引起临床可识别的疾病表型(见第6章)。
Disrupting higher-order chromatin features and gene-regulatory elements may affect transcriptional programs and development.破坏高级染色质特征和基因调控元件可能影响转录程序与发育。
For example, genomic deletions can lead to TADs fusing; duplications may form neo-TADs; inversions reshuffle TADs; and translocations could alter interchromosomal contacts between nonhomologous chromosomes .例如,基因组缺失可导致TAD融合;重复可能形成新TAD;倒位可重排TAD;易位可能改变非同源染色体之间的染色体间接触。
GENE EXPRESSION AS THE INTEGRATION OF GENOMIC AND EPIGENOMIC SIGNALS The gene expression program of a cell encompasses the specific subset of the ~20,000 protein-coding genes in the genome that are actively transcribed and translated into their respective functional products, the subset of the estimated 20,000 to 25,000 ncRNA genes that are transcribed, the amount of product produced, and the particular sequence (alleles) of those products.基因表达作为基因组与表观基因组信号的整合 细胞的基因表达程序包括:基因组中约20,000个蛋白质编码基因中被主动转录并翻译为其相应功能产物的特定子集;估计20,000至25,000个非编码RNA基因中被转录的子集;产物的生成量;以及这些产物的特定序列(等位基因)。
The gene expression profile of any particular cell or cell type in a given individual at a given time (whether in the context of the cell cycle, early development, or one’s entire life span) and under a given set of circumstances (as influenced by environment, lifestyle, or disease) is thus the integrated sum of several different but interrelated effects, including the following: The primary sequence of genes, their allelic variants, and their encoded products Regulatory sequences and their epigenetic positioning in chromatin Interactions with the thousands of transcriptional factors, ncRNAs, and other proteins involved in the control of transcription, splicing, translation, and posttranslational modification Organization of the genome into subchromosomal domains Programmed interactions between different parts of the genome The genome is organized in the three-dimensional architecture of the nucleus, with A and B compartments representing open chromatin and heterochromatin, respectively.因此,在给定个体、给定时间(无论处于细胞周期、早期发育还是整个生命周期)以及给定环境条件(受环境、生活方式或疾病影响)下,任何特定细胞或细胞类型的基因表达谱是多种不同但相互关联效应的综合总和,包括:基因的一级序列、其等位基因变异及其编码产物;调控序列及其在染色质中的表观遗传定位;与数千种转录因子、非编码RNA及其他涉及转录、剪接、翻译和翻译后修饰调控的蛋白质的相互作用;基因组在亚染色体结构域中的组织;基因组不同部分之间的程序化相互作用。基因组在细胞核三维结构中组织,其中A区室和B区室分别代表开放染色质和异染色质。
Further functional nuclear subdomains are the nucleolus, splicing speckles, and promyelocytic leukemia (PML) bodies.其他功能性核亚结构域包括核仁、剪接斑点及早幼粒细胞白血病(PML)小体。
The next organization level of chromatin involves TADs that are formed by the loop-extrusion model, with CCCTC-binding factor (CTCF) and cohesion to facilitate spatial proximity of gene-regulatory elements (enhancers), and gene promoters to regulate gene expression.染色质的下一级组织涉及通过环挤压模型形成的TADs,依赖CCCTC结合因子(CTCF)和粘连蛋白促进基因调控元件(增强子)与基因启动子的空间邻近,以调节基因表达。
(Courtesy Philipp Maass.)(由Philipp Maass提供。)
37/42
The Human Genome 39 Dynamic 3D chromatin packaging in the nucleus All of these orchestrate in an efficient, hierarchic, …
Ch3 — Segment 37
The Human Genome 39 Dynamic 3D chromatin packaging in the nucleus All of these orchestrate in an efficient, hierarchic, and highly programmed fashion.人类基因组39 细胞核中的动态三维染色质包装 所有这些都以高效、层级化和高度程序化的方式协调运作。
Disruption of any—due to genetic variation, to epigenetic changes, and/or to disease-related processes—would be expected to alter the overall cellular program and its functional output (see 3).由于遗传变异、表观遗传改变和/或疾病相关过程导致的任何破坏,都将预期改变整体细胞程序及其功能输出(参见3)。
Although an average cell might contain ~300,000 copies of mRNA in total, the abundance of specific mRNAs can differ over many orders of magnitude; among genes that are active, most are expressed at low levels (estimated to be <10 copies of that gene’s mRNA per cell), whereas others are expressed at much higher levels (several hundred to a few thousand copies of that mRNA per cell).尽管一个平均细胞可能总共含有约30万份mRNA拷贝,但特定mRNA的丰度可能相差多个数量级;在活跃的基因中,大多数表达水平较低(估计每个细胞中该基因的mRNA拷贝数少于10份),而其他基因的表达水平则高得多(每个细胞中该mRNA的拷贝数为几百到几千份)。
Only in highly specialized cell types are particular genes expressed at very high levels (many tens of thousands of copies) that account for a significant proportion of all mRNA in those cells.仅在高度特化的细胞类型中,特定基因以非常高的水平表达(数万份拷贝),这类表达占这些细胞中所有mRNA的显著比例。
Now consider an expressed gene with a sequence variant that allows one to distinguish between the RNA products (whether mRNA or ncRNA) transcribed from each of two alleles, one allele with a T that is transcribed to yield RNA with an A and the other allele with a C that is transcribed to yield RNA with a G .现在考虑一个表达基因,其序列变异使得能够区分从两个等位基因各自转录的RNA产物(无论是mRNA还是ncRNA),其中一个等位基因含有T,转录产生含有A的RNA,另一个等位基因含有C,转录产生含有G的RNA。
By sequencing individual RNA molecules and comparing the number of sequences generated that contain an A or G at that position, one can infer the ratio of transcripts from the two alleles in that sample.通过对单个RNA分子进行测序,并比较在该位置含有A或G的序列数量,可以推断该样本中来自两个等位基因的转录本比例。
Although most genes show essentially equivalent levels of biallelic expression, recent analyses of this type have demonstrated widespread unequal allelic expression for 5% to 20% of autosomal genes in the genome ( For most of these genes, the extent of imbalance is twofold or less, although up to tenfold differences have been observed for some genes.尽管大多数基因显示出基本相等的双等位基因表达水平,但最近此类分析已证明,基因组中5%至20%的常染色体基因存在广泛的不平等等位基因表达(对于这些基因中的大多数,不平衡程度为两倍或更低,尽管某些基因观察到高达十倍的差异)。
This allelic imbalance may reflect interactions between genome sequence and gene regulation; for example, sequence changes can alter the relative binding of various transcription factors or other transcriptional regulators to the two alleles or the extent of DNA methylation observed at the two alleles (see Monoallelic Gene Expression Some genes, however, show a much more complete form of allelic imbalance, resulting in monoallelic gene expression .这种等位基因不平衡可能反映了基因组序列与基因调控之间的相互作用;例如,序列变化可以改变各种转录因子或其他转录调控因子对两个等位基因的相对结合,或者改变两个等位基因上观察到的DNA甲基化程度(参见单等位基因表达。然而,一些基因表现出更完全形式的等位基因不平衡,导致单等位基因表达)。
Several different mechanisms have been shown to account for allelic imbalance of this type for particular subsets of genes in the genome: DNA rearrangement including copy number variation, random monoallelic expression, parent-of-origin imprinting, and, for genes on the X chromosome in females, X chromosome inactivation.多种不同机制已被证明可解释基因组中特定基因子集的此类等位基因不平衡:DNA重排(包括拷贝数变异)、随机单等位基因表达、亲本印记,以及对于女性X染色体上的基因,X染色体失活。
Their distinguishing characteristics are summarized in Somatic Rearrangement A highly specialized form of monoallelic gene expression is observed in the genes encoding immunoglobulins and T-cell receptors, expressed in B cells and T cells, respectively, as part of the immune response.它们的区别特征总结于体细胞重排。一种高度特化的单等位基因表达形式见于编码免疫球蛋白和T细胞受体的基因中,这些基因分别在B细胞和T细胞中表达,作为免疫反应的一部分。
Antibodies are encoded in the germline by a relatively small number of genes that, during B-cell development, undergo a unique process of somatic rearrangement that involves the cutting and pasting of DNA sequences in lymphocyte precursor cells (but not in any other cell lineages) 3 THE EPIGENETIC LANDSCAPE OF THE GENOME AND MEDICINE Different chromosomes and chromosomal regions occupy characteristic territories within the nucleus.抗体在生殖系中由相对较少数量的基因编码,在B细胞发育过程中,这些基因经历一个独特的体细胞重排过程,涉及淋巴细胞前体细胞中DNA序列的剪切和粘贴(但在任何其他细胞谱系中不发生)3 基因组与医学的表观遗传景观 不同的染色体和染色体区域在细胞核内占据特征性区域。
The probability of physical proximity influences the ­incidence of specific chromosome abnormalities (see Chapters 5 and 6).物理邻近的概率影响特定染色体异常的发生率(参见第5章和第6章)。
The genome is organized into megabase-sized domains with locally shared characteristics of base pair composition (i. e., GC rich or AT rich), gene density, timing of replication in the S phase, and presence of particular histone modifications (see Chapter 5).基因组被组织成兆碱基大小的结构域,这些结构域具有局部共享的碱基对组成特征(即GC富集或AT富集)、基因密度、S期复制时间以及特定组蛋白修饰的存在(参见第5章)。
Modules of coexpressed genes correspond to distinct anatomic or developmental stages in, for example, the human brain or the hematopoietic lineage.共表达基因模块对应于不同的解剖或发育阶段,例如在人脑或造血谱系中。
Such coexpression networks are revealed by shared regulatory networks and epigenetic signals, by clustering within genomic domains, and by overlapping patterns of altered gene expression in various disease states.这种共表达网络通过共享的调控网络和表观遗传信号、在基因组结构域内的聚类以及不同疾病状态下基因表达改变的重叠模式来揭示。
Although monozygotic twins share virtually identical genomes, they can be quite discordant for certain traits, including susceptibility to common diseases.尽管同卵双胞胎共享几乎相同的基因组,但他们在某些性状上可能相当不一致,包括对常见疾病的易感性。
Significant changes in DNA methylation occur during the lifetime of such twins, implicating epigenetic regulation of gene expression as a source of diversity.在这类双胞胎的一生中,DNA甲基化发生显著变化,暗示基因表达的表观遗传调控是多样性的来源。
ALLELIC IMBALANCE IN GENE EXPRESSION It was once assumed that genes present in two copies in the genome would be expressed from both homologues at comparable levels.基因表达中的等位基因不平衡 曾经假设基因组中两份拷贝的基因会从两个同源染色体以相当水平表达。
However, it has become increasingly evident that there can be extensive imbalance between alleles, reflecting both the amount of sequence variation in the genome and the interplay between genome sequence and epigenetic patterns that were just discussed.然而,越来越明显的是,等位基因之间可能存在广泛的不平衡,这既反映了基因组中序列变异的总量,也反映了刚刚讨论的基因组序列与表观遗传模式之间的相互作用。
In Chapter 2, we introduced the general finding that any individual genome carries two different alleles at a minimum of 3 to 5 million positions around the genome, thus distinguishing by sequence the maternally and paternally inherited copies of that sequence position .在第2章中,我们介绍了一般发现:任何个体基因组在基因组中至少300万到500万个位置上携带两个不同的等位基因,从而通过序列区分该序列位置来自母本和父本遗传的拷贝。
Here we explore ways in which those sequence differences reveal allelic imbalance in gene expression, both at autosomal loci and at X chromosome loci in females.在此,我们探讨这些序列差异揭示基因表达中等位基因不平衡的方式,包括在常染色体位点和女性X染色体位点上。
By determining the sequences of all the RNA products—the transcriptome—in a population of cells, one can quantify the relative level of transcription of all the genes (both protein coding and noncoding) that are transcriptionally active in those cells.通过确定一群细胞中所有RNA产物(转录组)的序列,可以量化这些细胞中转录活跃的所有基因(包括蛋白质编码和非编码)的相对转录水平。
Consider, for example, the collection of protein-coding genes.例如,考虑蛋白质编码基因的集合。
38/42
Different underlying mechanisms for allelic imbalance are compared in SNP, Single nucleotide polymorphism. to rearrange …
Ch3 — Segment 38
Different underlying mechanisms for allelic imbalance are compared in SNP, Single nucleotide polymorphism. to rearrange genes in somatic cells to generate enormous antibody diversity.在单核苷酸多态性(SNP)中比较了等位基因不平衡的不同潜在机制,以重排体细胞中的基因,从而产生巨大的抗体多样性。
The highly orchestrated DNA rearrangements occur across many hundreds of kilobases but involve only one of the two alleles, which is chosen randomly in any given B cell (see Thus expression of mature mRNAs for the immunoglobulin heavy or light chain subunits is exclusively monoallelic.高度协调的DNA重排发生在数百千碱基的范围内,但仅涉及两个等位基因中的一个,该等位基因在任何给定的B细胞中被随机选择(见因此,免疫球蛋白重链或轻链亚基的成熟mRNA的表达是严格单等位基因的)。
This mechanism of somatic rearrangement and random monoallelic gene expression is also observed at the T-cell receptor genes in the T-cell lineage.这种体细胞重排和随机单等位基因表达的机制也在T细胞谱系的T细胞受体基因中观察到。
However, such behavior is unique to these gene families and cell lineages; the rest of the genome, even DNA segments bearing genomic repeats, remains surprisingly stable throughout development and differentiation.然而,这种行为仅限于这些基因家族和细胞谱系;基因组的其余部分,甚至带有基因组重复序列的DNA片段,在整个发育和分化过程中仍然保持惊人的稳定性。
Random Monoallelic Expression In contrast to this highly specialized form of DNA rearrangement, monoallelic expression typically results from随机单等位基因表达 与这种高度特化的DNA重排形式相反,单等位基因表达通常源于
39/42
The Human Genome 41 differential epigenetic regulation of the two alleles.
Ch3 — Segment 39
The Human Genome 41 differential epigenetic regulation of the two alleles.人类基因组41号差异表观遗传调控两个等位基因。
One well-studied example of random monoallelic expression involves the OR gene family described earlier .随机单等位基因表达的一个经典例子涉及前面描述的OR基因家族。
In this case, only a single allele of one OR gene is expressed in each olfactory sensory neuron; the many hundred other copies of the OR family remain repressed in that cell.在这种情况下,每个嗅觉感觉神经元中仅表达一个OR基因的一个等位基因;该细胞中OR家族的数百个其他拷贝保持抑制状态。
Other genes with chemosensory or immune system functions also show random monoallelic expression, suggesting that this mechanism may be a general one for increasing the diversity of responses for cells that interact with the outside world.其他具有化学感觉或免疫系统功能的基因也表现出随机单等位基因表达,表明这种机制可能是一种增加与外部世界相互作用的细胞反应多样性的通用机制。
However, this mechanism is apparently not restricted to the immune and sensory systems because a substantial subset of all human genes (5–10% in different cell types) has been shown to undergo random allelic silencing; these genes are broadly distributed on all autosomes, have a wide range of functions, and vary in terms of the cell types and tissues in which monoallelic expression is observed.然而,这种机制显然不限于免疫和感觉系统,因为所有人类基因中相当一部分(在不同细胞类型中占5-10%)已被证明经历随机等位基因沉默;这些基因广泛分布在所有常染色体上,具有多种功能,并且在观察到单等位基因表达的细胞类型和组织方面各不相同。
Parent-of-Origin Imprinting For the examples just described, the choice of which allele is expressed is not dependent on parental origin; either the maternal or paternal copy can be expressed in different cells and their clonal descendants.亲本起源印记 对于刚才描述的例子,表达哪个等位基因的选择不依赖于亲本起源;母本或父本拷贝可以在不同细胞及其克隆后代中表达。
This distinguishes random forms of monoallelic expression from genomic imprinting, in which the choice of the allele to be expressed is nonrandom and is determined solely by parental origin.这将随机形式的单等位基因表达与基因组印记区分开来,在基因组印记中,要表达的等位基因的选择是非随机的,完全由亲本起源决定。
Imprinting is a process involving the introduction of epigenetic marks in the germline of one parent, but not the other, at specific locations in the genome.印记是一个过程,涉及在一个亲本(而非另一个亲本)的生殖系中,在基因组的特定位置引入表观遗传标记。
These lead to monoallelic expression of a gene or, in some cases, of multiple genes within the imprinted region.这些标记导致一个基因的单等位基因表达,或者在某些情况下,导致印记区域内多个基因的单等位基因表达。
Imprinting takes place during gametogenesis, before fertilization, and marks certain genes as having come from the mother or father .印记发生在配子发生期间、受精之前,并标记某些基因来自母亲或父亲。
After conception, the parent-of-origin imprint is maintained in some or all of the somatic tissues of the embryo and silences gene expression on allele(s) within the imprinted region; whereas some imprinted genes show monoallelic expression throughout the embryo, others show tissue-specific imprinting, especially in the placenta, with biallelic expression in other tissues.受孕后,亲本起源印记在胚胎的部分或全部体组织中得以维持,并沉默印记区域内等位基因上的基因表达;而有些印记基因在整个胚胎中表现出单等位基因表达,其他基因则表现出组织特异性印记,尤其是在胎盘中,而在其他组织中为双等位基因表达。
The imprinted state persists postnatally into adulthood through hundreds of cell divisions so that only the maternal or paternal copy of the gene is expressed.印记状态通过数百次细胞分裂持续到出生后直至成年,因此只有母本或父本基因拷贝得以表达。
Yet, imprinting must be reversible: a paternally derived allele, when it is inherited by a female, must be converted in her germline so that she can then pass it on with a maternal imprint to her offspring.然而,印记必须是可逆的:一个父本来源的等位基因,当被女性继承时,必须在其生殖系中转换,以便她随后能以母本印记将其传递给后代。
Likewise, an imprinted maternally derived allele, when it is inherited by a male, must be converted in his germline so that he can pass it on as a paternally imprinted allele to his offspring .同样,一个印记的母本来源等位基因,当被男性继承时,必须在其生殖系中转换,以便他能以父本印记等位基因将其传递给后代。
Control over this conversion process appears to be governed by specific DNA elements called imprinting control regions or imprinting centers that are located within imprinted regions throughout the genome; although their mechanism of action is not fully known, many appear to involve ncRNAs that initiate the epigenetic change in chromatin, which then spreads outward along the chromosome over the imprinted region.对这一转换过程的控制似乎由特定的DNA元件调控,这些元件称为印记控制区或印记中心,位于整个基因组的印记区域内;尽管它们的作用机制尚未完全明了,但许多似乎涉及ncRNA,这些ncRNA启动染色质中的表观遗传改变,然后沿着染色体向外扩散至印记区域。
Notably, although the imprinted region can encompass more than a single gene, this form of monoallelic expression is confined to a delimited genomic segment, typically a few hundred kilobase pairs to a few megabases in overall size; this distinguishes genomic imprinting both from the more general form of random monoallelic expression described earlier (which appears to involve individual genes under locusspecific control) and from X chromosome inactivation, described in the next section (which involves genes along the entire chromosome).值得注意的是,虽然印记区域可以包含不止一个基因,但这种形式的单等位基因表达仅限于一个限定的基因组片段,总大小通常为几百千碱基对到几兆碱基;这使基因组印记既区别于前述更普遍的随机单等位基因表达形式(似乎涉及受基因座特异性控制的单个基因),也区别于下一节描述的X染色体失活(涉及整条染色体上的基因)。
To date, ~100 imprinted genes have been identified on many different autosomes.迄今为止,已在许多不同的常染色体上鉴定出约100个印记基因。
The involvement of these genes in various chromosomal disorders is described more fully in Chapter 6.这些基因在各种染色体疾病中的参与将在第6章中更全面地描述。
For clinical conditions due to a single imprinted gene, such as Prader-Willi syndrome (Case 38) and Beckwith-Wiedemann syndrome (Case 6), the effect of genomic imprinting on inheritance patterns in pedigrees is discussed in Chapter 8.对于由单个印记基因导致的临床病症,如普拉德-威利综合征(病例38)和贝克威斯-威德曼综合征(病例6),基因组印记对家系遗传模式的影响将在第8章中讨论。
X Chromosome Inactivation The chromosomal basis for sex determination, introduced in Chapter 2 and discussed in more detail in Chapter 6, results in a dosage difference between typical males and females with respect to genes on the X chromosome.X染色体失活 性别决定的染色体基础在第2章中介绍,并在第6章中更详细讨论,导致典型男性和女性之间在X染色体基因方面的剂量差异。
Here we discuss the chromosomal and molecular mechanisms of X chromosome inactivation, the most extensive example of random monoallelic expression in the genome and a mechanism of dosage compensation that results in the epigenetic silencing of most genes on one of the two X chromosomes in females.这里我们讨论X染色体失活的染色体和分子机制,这是基因组中最广泛的随机单等位基因表达实例,也是一种剂量补偿机制,导致女性两条X染色体中一条上的大多数基因发生表观遗传沉默。
In normal female cells, the choice of which X chromosome is to be inactivated is a random one that is then maintained in each clonal lineage.在正常女性细胞中,选择哪条X染色体失活是随机的,然后每个克隆谱系中维持这一选择。
Thus females are mosaic with respect to X-linked gene expression; some cells express alleles on the paternally inherited X but not the maternally inherited X, whereas other cells do the opposite .因此,女性在X连锁基因表达方面呈嵌合状态;一些细胞表达父本来源X染色体上的等位基因而不表达母本来源X染色体上的等位基因,而其他细胞则相反。
This mosaic pattern of gene expression distinguishes most X-linked genes from imprinted genes, whose expression, as we just noted, is determined strictly by parental origin.这种基因表达的嵌合模式将大多数X连锁基因与印记基因区分开来,正如我们刚才所述,印记基因的表达严格由亲本起源决定。
Although the inactive X chromosome was first identified cytologically by the presence of a heterochromatic mass (called the Barr body) in interphase cells, many epigenetic features at the molecular level distinguish the active and inactive X chromosomes, including DNA methylation, histone modifications, and a specific histone variant, macro H2A, that is particularly enriched in chromatin on the inactive X.尽管失活的X染色体最早是通过间期细胞中异染色质团块(称为巴氏小体)的细胞学存在而识别的,但分子水平上的许多表观遗传特征区分了活性X染色体和失活X染色体,包括DNA甲基化、组蛋白修饰以及一种特定的组蛋白变体macro H2A,后者在失活X染色质的染色质中特别富集。
As well as providing insights into the mechanisms of X inactivation, these features can be useful diagnostically for identifying inactive X chromosomes in clinical material, as we will see in Chapter 6.除了提供对X失活机制的见解外,这些特征在诊断上可用于识别临床样本中的失活X染色体,正如我们将在第6章中看到的。
40/42
Oogenesis Spermatogenesis Sperm Imprint erasure Establishment of imprint Oocyte Fertilization Embryo Embryo Within a hyp…
Ch3 — Segment 40
Oogenesis Spermatogenesis Sperm Imprint erasure Establishment of imprint Oocyte Fertilization Embryo Embryo Within a hypothetical imprinted region on a pair of homologous autosomes, paternally imprinted genes are indicated in blue, whereas a maternally imprinted gene is indicated in red.卵子发生 精子发生 精子 印记擦除 印记建立 卵母细胞 受精 胚胎 胚胎 在一对同源常染色体上的一个假设印记区域内,父本印记基因以蓝色表示,而母本印记基因以红色表示。
After fertilization, both male and female embryos have one copy of the chromosome carrying a paternal imprint and one copy carrying a maternal imprint.受精后,雄性和雌性胚胎均各携带一条带有父本印记的染色体拷贝和一条带有母本印记的染色体拷贝。
During oogenesis (top) and spermatogenesis (bottom), the imprints are erased by removal of epigenetic marks, and new imprints determined by the sex of the parent are established within the imprinted region.在卵子发生(上方)和精子发生(下方)过程中,印记通过去除表观遗传标记而被擦除,并且由亲本性别决定的新印记在印记区域内建立。
Gametes thus carry a monoallelic imprint appropriate to the parent of origin, whereas somatic cells in both sexes carry one chromosome of each imprinted type.因此,配子携带与亲本来源相符的单等位印记,而两性的体细胞各携带一条每种印记类型的染色体。
Although X inactivation is clearly a chromosomal phenomenon, not all genes on the X chromosome show monoallelic expression in female cells.尽管X失明显然是一种染色体现象,但并非X染色体上的所有基因在女性细胞中都表现出单等位表达。
Extensive analysis of expression of nearly all X-linked genes has demonstrated that at least 15% of the genes show biallelic expression and are expressed from both active and inactive X chromosomes, at least to some extent; a proportion of these show significantly higher levels of mRNA production in female cells compared to male cells and are interesting candidates for a role in explaining sexually dimorphic traits.对几乎所有X连锁基因表达的广泛分析表明,至少有15%的基因表现出双等位表达,并且在一定程度上从活性和失活的X染色体上均有表达;其中一部分基因在女性细胞中的mRNA产量显著高于男性细胞,且是解释性别二态性特征的有趣候选基因。
A special subset of genes is located in the pseudoautosomal segments, which are essentially identical on the X and Y chromosomes and undergo recombination during一个特殊的基因子集位于假常染色体区段,这些区段在X和Y染色体上基本一致,并在期间发生重组。
41/42
The Human Genome 43 X X pat mat mat Barr body or Xi Expresses maternal alleles X inactivation Clonal maintenance Clonal …
Ch3 — Segment 41
The Human Genome 43 X X pat mat mat Barr body or Xi Expresses maternal alleles X inactivation Clonal maintenance Clonal maintenance Xi X inactivation Expresses paternal alleles pat X inactivation center Shortly after conception of a female embryo, both the paternally and maternally inherited X chromosomes (pat and mat, respectively) are active.人类基因组 43 X X 父源 母源 母源 巴氏小体或Xi 表达母源等位基因 X染色体失活 克隆维持 克隆维持 Xi X染色体失活 表达父源等位基因 父源 X失活中心 在女性胚胎受孕后不久,父源性和母源性遗传的X染色体(分别为pat和mat)均为活跃状态。
Within the first week of embryogenesis, one or the other X is chosen at random to become the future inactive X, through a series of events involving the X inactivation center (black box).在胚胎发生的第一周内,通过涉及X失活中心(黑色方框)的一系列事件,随机选择其中一条X染色体成为未来的失活X染色体。
That X then becomes the inactive X (Xi, indicated by the shading) in that cell and its progeny and forms the Barr body in interphase nuclei.该X染色体随后在该细胞及其子代中变为失活X染色体(Xi,以阴影表示),并在间期细胞核中形成巴氏小体。
The resulting female embryo is thus a clonal mosaic of two epigenetically determined cell types: one expresses alleles from the maternal X (pink cells), whereas the other expresses alleles from the paternal X (blue cells).由此形成的女性胚胎因此是两种表观遗传决定细胞类型的克隆嵌合体:一种表达来自母源X染色体的等位基因(粉色细胞),而另一种表达来自父源X染色体的等位基因(蓝色细胞)。
The ratio of the two cell types is determined randomly but varies among normal females and among females who are carriers of X-linked disease alleles (see Chapters 6 and 7). spermatogenesis (see Chapter 2).两种细胞类型的比例是随机确定的,但在正常女性以及X连锁疾病等位基因携带者女性中有所不同(见第6章和第7章)。精子发生(见第2章)。
These genes have two copies in both females (two X-linked copies) and males (one X-linked and one Y-linked copy) and thus do not undergo X inactivation; as expected, these genes show balanced biallelic expression, as one sees for most autosomal genes.这些基因在女性(两个X连锁拷贝)和男性(一个X连锁拷贝和一个Y连锁拷贝)中均有两个拷贝,因此不经历X染色体失活;正如预期,这些基因表现出平衡的双等位基因表达,与大多数常染色体基因所见一致。
The X Inactivation Center and the XIST Gene.X失活中心与XIST基因。
X inactivation occurs very early in female embryonic development, and determination of which X will be designated the inactive X in any given cell in the embryo is a random choice under the control of a complex locus called the X inactivation center.X染色体失活发生在女性胚胎发育的极早期,在胚胎任何特定细胞中决定哪条X染色体被指定为失活X染色体是一个随机选择,受控于一个称为X失活中心的复杂位点。
This region contains a unique ncRNA gene, XIST, that appears to be a key master regulatory locus for X inactivation.该区域包含一个独特的非编码RNA基因XIST,它似乎是X失活的关键主调控位点。
XIST (an acronym for inactive X [Xi]–specific transcripts) has the novel feature that it is expressed only from the allele on the inactive X; it is transcriptionally silent on the active X in both male and female cells.XIST(失活X染色体特异性转录本的缩写)具有一个新颖特征:它仅从失活X染色体上的等位基因表达;在男性和女性细胞的活跃X染色体上,它处于转录沉默状态。
Although the exact mode of action of XIST is unknown, X inactivation cannot occur in its absence.尽管XIST的确切作用机制尚不清楚,但在其缺失的情况下X失活无法发生。
The product of XIST is a long ncRNA that stays in the nucleus in close association with the inactive X chromosome.XIST的产物是一种长链非编码RNA,在细胞核内与失活X染色体紧密关联。
Additional aspects and consequences of X chromosome inactivation will be discussed in Chapter 6, in the context of individuals with structurally abnormal X chromosomes or an abnormal number of X chromosomes, and in Chapter 7, in the case of females carrying deleterious mutant alleles for X-linked disease.X染色体失活的其他方面和后果将在第6章(涉及结构异常X染色体或X染色体数目异常的个体)和第7章(携带X连锁疾病有害突变等位基因的女性)中讨论。
VARIATION IN GENE EXPRESSION AND ITS RELEVANCE TO MEDICINE The regulated expression of genes in the human genome involves a set of complex interrelationships among different levels of control, including proper gene dosage (controlled by mechanisms of chromosome replication and segregation), gene structure, chromatin packaging and epigenetic regulation, transcription, RNA splicing, and, for protein-coding loci, mRNA stability, translation, protein processing, and protein degradation.基因表达变异及其与医学的相关性 人类基因组中基因的调控表达涉及不同控制层级之间的复杂相互关系,包括适当的基因剂量(由染色体复制和分离机制控制)、基因结构、染色质包装与表观遗传调控、转录、RNA剪接,以及对于蛋白质编码位点,mRNA稳定性、翻译、蛋白质加工和蛋白质降解。
For some genes, fluctuations in the level of functional gene product, due either to inherited variation in the structure of a particular gene or to changes induced by nongenetic factors such as diet or the environment, are of relatively little importance.对于某些基因,功能性基因产物水平的波动(无论是由于特定基因结构的遗传变异,还是由饮食或环境等非遗传因素引起的变化)相对不那么重要。
For other genes, even relatively minor changes in the level of expression can have dire clinical consequences, reflecting the importance of those gene products in particular biologic pathways.对于其他基因,即使是表达水平相对微小的变化也可能导致严重的临床后果,这反映了这些基因产物在特定生物学通路中的重要性。
The nature of inherited variation in the structure and function of chromosomes and the genes they contain, combined with the influence of this variation on the expression of specific traits, is the very essence of medical genetics and is dealt with in subsequent chapters.染色体及其所含基因的结构与功能的遗传变异性质,连同这些变异对特定性状表达的影响,正是医学遗传学的核心所在,并将在后续章节中讨论。
42/42
GENERAL REFERENCES Brown TA: Genomes, ed 3, New York, 2007, Garland Science.
Ch3 — Segment 42
GENERAL REFERENCES Brown TA: Genomes, ed 3, New York, 2007, Garland Science.一般参考文献 Brown TA: 基因组学,第3版,纽约,2007年,Garland Science。
Lodish H, Berk A, Kaiser CA, et al: Molecular cell biology, ed 9, New York, 2021, WH Freeman.Lodish H, Berk A, Kaiser CA, 等: 分子细胞生物学,第9版,纽约,2021年,WH Freeman。
Strachan T, Read A: Human molecular genetics, ed 5, New York, 2018, Garland Science.Strachan T, Read A: 人类分子遗传学,第5版,纽约,2018年,Garland Science。
REFERENCES FOR SPECIFIC TOPICS Bartolomei MS, Ferguson-Smith AC: Mammalian genomic imprinting, Cold Spring Harbor Perspect Biol 3:1002592, 2011.特定主题参考文献 Bartolomei MS, Ferguson-Smith AC: 哺乳动物基因组印记,冷泉港展望生物学 3:1002592, 2011。
Beck CR, Garcia-Perez JL, Badge RM, et al: LINE-1 elements in structural variation and disease, Annu Rev Genomics Hum Genet 12:187–215, 2011.Beck CR, Garcia-Perez JL, Badge RM, 等: 结构变异与疾病中的LINE-1元件,年度综述:基因组学与人类遗传学 12:187–215, 2011。
Berg P: Dissections and reconstructions of genes and chromosomes (Nobel Prize lecture), Science 213:296–303, 1981.Berg P: 基因与染色体的解剖与重构(诺贝尔奖演讲),科学 213:296–303, 1981。
Chess A: Mechanisms and consequences of widespread random monoallelic expression, Nat Rev Genet 13:421–428, 2012.Chess A: 广泛随机单等位基因表达的机制与后果,自然综述·遗传学 13:421–428, 2012。
Dekker J: Gene regulation in the third dimension, Science 319:1793– 1794, 2008.[TL:missing]
Djebali S, Davis CA, Merkel A, et al: Landscape of transcription in human cells, Nature 489:101–108, 2012.[TL:missing]
ENCODE Project Consortium: An integrated encyclopedia of DNA elements in the human genome, Nature 489:57–74, 2012.[TL:missing]
Gerstein MB, Bruce C, Rozowsky JS, et al: What is a gene, postENCODE?[TL:missing]
Genome Res 17:669–681, 2007.[TL:missing]
Guil S, Esteller M: Cis-acting noncoding RNAs: friends and foes, Nat Struct Mol Biol 19:1068–1074, 2012.[TL:missing]
Heyn H, Esteller M: DNA methylation profiling in the clinic: applications and challenges, Nature Rev Genet 13:679–692, 2012.[TL:missing]
Hubner MR, Spector DL: Chromatin dynamics, Annu Rev Biophys 39:471–489, 2010.[TL:missing]
Li M, Wang IX, Li Y, et al: Widespread RNA and DNA sequence differences in the human transcriptome, Science 333:53–58, 2011.[TL:missing]
Nagano T, Fraser P: No-nonsense functions for long noncoding RNAs, Cell 145:178–181, 2011.[TL:missing]
Willard HF: The human genome: a window on human genetics, biology and medicine.[TL:missing]
In Ginsburg GS, Willard HF, editors: Genomic and personalized medicine, ed 2, New York, 2013, Elsevier.[TL:missing]
Zhou VW, Goren A, Bernstein BE: Charting histone modifications and the functional organization of mammalian genomes, Nat Rev Genet 12:7–18, 2012.[TL:missing]
PROBLEMS 1.[TL:missing]
The following amino acid sequence represents part of a protein.[TL:missing]
The wild-type sequence and four variant forms are shown.[TL:missing]
By consulting Which strand is the strand that RNA polymerase “reads”?[TL:missing]
What would be the sequence of the resulting mRNA?[TL:missing]
What kind of variation is each altered protein most likely to represent?[TL:missing]
Normal -lys-arg-his-his-tyr-leuMutant 1 -lys-arg-his-his-cys-leuMutant 2 -lys-arg-ile-ile-ileMutant 3 -lys-glu-thr-ser-leu-serMutant 4 -asn-tyr-leu2.[TL:missing]
The following items are related to each other in a hierarchic fashion: chromosome, base pair, nucleosome, kilobase pair, intron, gene, exon, chromatin, codon, nucleotide, promoter.[TL:missing]
What are these relationships?[TL:missing]
Describe how a variant in each of the following might alter or interfere with gene function and thus cause human disease: promoter, initiator codon, splice sites at intronexon junctions, one base-pair deletion in the coding sequence, stop codon.[TL:missing]
Most of the human genome consists of sequences that are not transcribed and do not directly encode gene products.[TL:missing]
For each of the following, consider ways in which these genome elements might contribute to human disease: introns, Alu or LINE repetitive sequences, locus control regions, pseudogenes.[TL:missing]
Contrast the mechanisms and consequences of RNA splicing and somatic rearrangement.[TL:missing]
Contrast the mechanisms and consequences of genomic imprinting and X chromosome inactivation.[TL:missing]