Part 6

PART 6: Clinical Genetics and Genetic Counseling

← Back to Genetics Contents
Ch11 — Clinical Genetic Evaluation (22) Ch12 — Genetic Counseling and Risk Assessment (24)

Clinical Genetic Evaluation

1/46
Identifying the Genetic Basis for Human Disease This chapter provides an overview of how geneticists study families and …
Ch11 — Segment 1
Identifying the Genetic Basis for Human Disease This chapter provides an overview of how geneticists study families and populations to identify genetic contributions to disease.识别人类疾病的遗传基础 本章概述了遗传学家如何研究家族和人群以识别疾病中的遗传因素。
Whether a disease is inherited in a recognizable mendelian pattern, as illustrated in Chapter 7, or occurs at a higher frequency in relatives of affected individuals, as explored in Chapter 9, it is specific genomic variants that either cause disease directly or influence the susceptibility to disease.无论是如第7章所示以可识别的孟德尔模式遗传的疾病,还是如第9章所探讨的在患者亲属中发病率较高的疾病,都是由特定的基因组变异直接导致疾病或影响疾病易感性。
Genome research has provided geneticists with a catalogue of all known human genes, knowledge of their location and structure, and an ever-growing list of tens of millions of variants in DNA sequence found among individuals in different populations.基因组研究为遗传学家提供了所有已知人类基因的目录、关于其位置和结构的知识,以及在不同人群个体中发现的数以千万计的DNA序列变异不断增长的列表。
As we saw in previous chapters, some of these variants are common, others are rare, and still others are ultrarare or even private to families or individuals.正如我们在前几章所见,其中一些变异是常见的,其他是罕见的,还有一些是超罕见甚至为家族或个人所特有。
Whereas some variants clearly have functional consequences associated with disease risk, others are certainly neutral.尽管某些变异显然具有与疾病风险相关的功能后果,但其他变异肯定是中性的。
For most, their significance for human health and disease is unknown.对于大多数变异而言,其对人类健康和疾病的意义尚不清楚。
In Chapter 4, we dealt with the effect of mutation, which alters one or more genes or loci to generate variant alleles and polymorphism.在第4章中,我们讨论了突变的影响,突变会改变一个或多个基因或位点,从而产生变异等位基因和多态性。
In Chapters 7 and 9, we examined the role of genetic factors in the pathogenesis of various mendelian or complex disorders.在第7章和第9章中,我们探讨了遗传因素在各种孟德尔疾病或复杂疾病发病机制中的作用。
In this ­chapter, we discuss how geneticists go about discovering the particular genes implicated in disease and the variants they contain that underlie or contribute to human diseases, focusing on three approaches: The first approach, linkage analysis, is family based.在本章中,我们讨论遗传学家如何着手发现与疾病相关的特定基因及其所含的、构成人类疾病基础或促发疾病的变异,重点介绍三种方法:第一种方法是连锁分析,基于家系。
Linkage analysis takes explicit advantage of family pedigrees to follow the inheritance of a disease among family members and to test for consistent, repeated coinheritance of the disease with a particular genomic region or even with a specific variant(s), whenever the disease is passed on in a family.连锁分析明确利用家系谱来追踪疾病在家庭成员中的遗传情况,并在疾病在家系中传递时,检验该疾病是否与特定基因组区域甚至特定变异持续、重复地共遗传。
The second approach, genome-wide association analysis, is population based.第二种方法是全基因组关联分析,基于人群。
Association analysis takes advantage of the entire history of a population to look for increased or decreased frequency of a particular allele or set of alleles in a cohort of affected individuals compared with a control set of unaffected individuals from that same population.关联分析利用整个人群的历史,比较来自同一人群的患者队列与未患病对照组中特定等位基因或一组等位基因频率的升高或降低。
It is particularly useful for complex and multifactorial diseases that do not show a mendelian inheritance pattern.该方法对于不表现孟德尔遗传模式的复杂多因素疾病特别有用。
The third approach involves direct genome-wide sequencing of affected individuals and their parents and/or other individuals in the family or population.第三种方法涉及对患者及其父母和/或家系或人群中其他个体进行直接的全基因组测序。
Genome-wide sequencing refers to sequencing the entire genome or sequencing the coding portion of the genome, the exome.全基因组测序是指对整个基因组进行测序,或对基因组的编码部分即外显子组进行测序。
This approach has been broadly adopted by geneticists, mainly due to the advancement of nextgeneration sequencing technologies that have reduced the cost of DNA sequencing a millionfold from the original reference genome sequenced for the Human Genome Project.这种方法已被遗传学家广泛采用,主要得益于下一代测序技术的进步,这些技术使DNA测序成本比人类基因组计划最初测序参考基因组时降低了一百万倍。
Genome-wide sequencing is particularly useful for rare mendelian disorders in which linkage analysis is not possible because there are not enough families to do such analysis, or because the disorder is a genetic lethal that always results from new mutations and is never inherited.全基因组测序对于罕见的孟德尔疾病特别有用,在这些疾病中连锁分析无法进行,因为缺乏足够的家系来进行此类分析,或者因为该疾病是一种遗传致死性疾病,总是由新生突变引起且从不遗传。
Although this approach allows an unbiased analysis of genes, examination of the resulting billions (or in the case of the exome, tens of millions) of bases of DNA requires a robust filtering strategy for identification of disease alleles.尽管这种方法允许对基因进行无偏倚分析,但检查所产生的数十亿(或对于外显子组而言是数千万)个DNA碱基需要一种强大的过滤策略来识别疾病等位基因。
Using a strategy to pare down variants coupled with the emergence of matchmaking programs has been extremely successful in gene discovery for hundreds of rare genetic disorders.使用精简变异策略,加上配对程序的出现,已经在数百种罕见遗传病的基因发现中取得了极大成功。
Use of linkage, association, and genome-wide sequencing to identify disease-associated genes has had an enormous impact on our understanding of the pathogenesis and pathophysiology of many diseases.利用连锁分析、关联分析和全基因组测序来识别疾病相关基因,对我们理解许多疾病的发病机制和病理生理学产生了巨大影响。
In time, knowledge of the genetic contributions to disease will also suggest new methods of prevention, management, and treatment.随着时间的推移,对疾病遗传因素的认识也将提出预防、管理和治疗的新方法。
GENETIC BASIS FOR LINKAGE ANALYSIS AND ASSOCIATION A fundamental feature of human biology is that each generation reproduces by combining haploid gametes containing 23 chromosomes that resulted from independent assortment and recombination of homologous chromosomes (see Chapter 2).连锁分析与关联的遗传基础 人类生物学的一个基本特征是,每一代通过含有23条染色体的单倍体配子进行繁殖,这些配子来自同源染色体的独立分配和重组(见第2章)。
To understand fully the concepts underlying genetic linkage analysis and tests for association, it is necessary to review briefly the为了充分理解遗传连锁分析和关联检验背后的概念,有必要简要回顾一下
2/46
behavior of chromosomes and genes during meiosis, as they are passed from one generation to the next.
Ch11 — Segment 2
behavior of chromosomes and genes during meiosis, as they are passed from one generation to the next.染色体和基因在减数分裂过程中传递至下一代的行为。
Some of this information repeats the classic material on gametogenesis presented in Chapter 2, illustrating it with new information that has become available as a result of the Human Genome Project and its applications to the study of human variation.其中部分信息重复了第二章介绍的配子发生经典内容,并以人类基因组计划及其应用于人类变异研究的新信息加以说明。
Independent Assortment and Homologous Recombination in Meiosis During meiosis I, homologous chromosomes line up in pairs along the meiotic spindle.减数分裂中的自由组合与同源重组 在减数第一次分裂期间,同源染色体沿减数分裂纺锤体成对排列。
The paternal and maternal homologues exchange homologous segments by crossing over and creating new chromosomes that are a patchwork-like effect consisting of alternating portions of the grandparental chromosomes .父本和母本同源染色体通过交叉互换交换同源片段,产生新的染色体,这些染色体呈现镶嵌效应,由祖父母染色体的交替部分构成。
In the family illustrated in .在所示家系中。
The creation of such patchwork chromosomes emphasizes the notion of human genetic individuality: each chromosome inherited by a child from a parent is never exactly the same as either of the two copies of that chromosome in the parent.这种镶嵌染色体的产生强调了人类遗传个体性的概念:孩子从父母遗传的每条染色体,从来不会与父母体内该染色体的两个拷贝中的任何一个完全相同。
Homologous chromosomes differ substantially at the DNA sequence level.同源染色体在DNA序列水平上存在显著差异。
As discussed in Chapter 4, these differences at the same position (locus) on a pair of homologous chromosomes are alleles.如第4章所述,一对同源染色体上同一位置(基因座)的这些差异即为等位基因。
Alleles that are common (generally considered to be those carried by 1% or more of the population) constitute a polymorphic locus.常见的等位基因(通常指群体中1%或以上个体携带的等位基因)构成多态性基因座。
Allelic variants on homologous chromosomes allow geneticists to trace each segment of a chromosome inherited by a particular child to determine if and where recombination events have occurred along the homologous chromosomes.同源染色体上的等位基因变异使遗传学家能够追踪特定孩子遗传的每条染色体片段,以确定沿同源染色体是否发生重组事件及其发生位置。
There are now hundreds of millions of genetic markers across diverse populations available to serve as genetic markers for this purpose.目前在不同人群中已有数亿个遗传标记可用作此目的的遗传标记。
Alleles at Loci on Different Chromosomes Assort Independently Assume there are two polymorphic loci, 1 and 2, on different chromosomes, with alleles A and a at locus 1 and alleles B and b at locus 2 .不同染色体上基因座的等位基因自由组合 假设存在两个多态性基因座1和2,位于不同染色体上,基因座1有等位基因A和a,基因座2有等位基因B和b。
Suppose an I II III Because of crossing over in meiosis, the copy of the chromosome the boy (generation III) inherited from his mother is a mosaic of segments of all four of his grandparents’ copies of that chromosome.假设一个 I II III 由于减数分裂中的交叉互换,男孩(III代)从母亲遗传的染色体副本是其祖父母该染色体全部四个副本片段的嵌合体。
The blank chromosome represents the chromosome inherited from the boy’s father.空白染色体代表男孩从父亲遗传的染色体。
A Locus 1→ b←Locus 2 B b B B a A A A a a b B A a Meiosis: homologous chromosomes line up randomly in one of two orientations or Gametes Parental combinations (AB and ab) Nonparental combinations (Ab and a B) b b a B Assume that alleles A and B were inherited from one parent, a and b from the other.一个 基因座1→ b←基因座2 B b B B a A A A a a b B A a 减数分裂:同源染色体随机以两种方向之一排列 配子 亲本组合(AB和ab) 非亲本组合(Ab和aB) b b a B 假设等位基因A和B遗传自一方亲本,a和b遗传自另一方亲本。
The two chromosomes can line up on the metaphase plate in meiosis I in one of two equally likely combinations, resulting in independent assortment of the alleles on these two chromosomes.在减数第一次分裂中,这两条染色体可以以两种概率均等的组合之一排列在中期板上,导致这两条染色体上的等位基因自由组合。
3/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 205 individual’s genotype at these loci is Aa and Bb; that is, she is he…
Ch11 — Segment 3
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 205 individual’s genotype at these loci is Aa and Bb; that is, she is heterozygous at both loci, with alleles A and B inherited from her father and alleles a and b inherited from her mother.识别人类疾病的遗传基础:205个体的这些位点基因型为Aa和Bb;即她在两个位点上均为杂合子,等位基因A和B遗传自父亲,等位基因a和b遗传自母亲。
The two different chromosomes will line up on the metaphase plate at meiosis I in one of two combinations with equal likelihood.两条不同的染色体在减数分裂I的中期板上以两种组合之一排列,每种可能性均等。
After recombination and chromosomal segregation are complete, there will be four possible combinations of alleles in a gamete: AB, ab, Ab, and a B.重组和染色体分离完成后,配子中将有四种可能的等位基因组合:AB、ab、Ab和aB。
Each combination is as likely to occur as any other, a phenomenon known as independent assortment.每种组合发生的可能性与其他组合相同,这一现象称为自由组合。
Because AB gametes ­contain only her paternally derived alleles, and ab gametes only her maternally derived alleles, these gametes are designated parental.由于AB配子仅包含她源自父亲的等位基因,ab配子仅包含她源自母亲的等位基因,这些配子被指定为亲本型。
In contrast, Ab or a B gametes, each containing one paternally derived allele and one maternally derived allele, are termed nonparental gametes.相比之下,Ab或aB配子各包含一个源自父亲和一个源自母亲的等位基因,被称为非亲本型配子。
On average, half (50%) of gametes will be parental (AB or ab) and 50% nonparental (Ab or a B).平均而言,一半(50%)的配子为亲本型(AB或ab),50%为非亲本型(Ab或aB)。
Alleles at Loci on the Same Chromosome Assort Independently If At Least One Crossover Between Them Always Occurs Now suppose that an individual is heterozygous at two loci, 1 and 2, with alleles A and B paternally derived and a and b maternally derived, but the loci are on the same chromosome .如果同一条染色体上的两个位点之间总是至少发生一次交换,则这些位点上的等位基因会独立分配;现在假设一个个体的两个位点1和2均为杂合,等位基因A和B源自父亲,a和b源自母亲,但这两个位点位于同一条染色体上。
Genes that reside on the same chromosome are said to be syntenic (literally, “on the same thread”), regardless of how close together or how far apart they lie on that chromosome.位于同一条染色体上的基因被称为同线基因(字面意思为“在同一条线上”),无论它们在染色体上的距离是近还是远。
How will these alleles behave during meiosis?这些等位基因在减数分裂期间会如何表现?
We know that between one and four crossovers occur between homologous chromosomes during meiosis I when there are two chromatids per homologous chromosome.我们知道,当每条同源染色体有两条染色单体时,减数分裂I期间同源染色体之间会发生一到四次交换。
If no crossing over occurs within the segment of the chromatids between the loci 1 and 2 (and ignoring whatever happens in segments outside the interval between these loci), then the chromosomes we see in the gametes will be AB and ab, which are the same as the original parental chromosomes; a parental chromosome is therefore a nonrecombinant chromosome.如果位点1和2之间的染色单体片段内没有发生交换(并忽略这些位点区间外片段中发生的任何事件),那么我们在配子中看到的染色体将是AB和ab,与原始亲本染色体相同;因此亲本染色体是非重组染色体。
If crossing over occurs at least once in the segment between the loci, the resulting chromatids may be either nonrecombinant or Ab and a B, which are not the same as the parental chromosomes; such a nonparental chromosome is therefore a recombinant chromosome (shown in .如果在位点之间的片段中至少发生一次交换,则产生的染色单体可能为非重组型或Ab和aB,它们与亲本染色体不同;这样的非亲本染色体因此是重组染色体(如图所示)。
One, two, or more recombinations occurring between two loci at the four-chromatid stage result in gametes that are 50% nonrecombinant (parental) and 50% recombinant (nonparental), which is precisely the same proportions one sees with independent assortment of alleles at loci on different chromosomes.在四染色单体阶段,两个位点之间发生一次、两次或更多次重组,产生的配子中50%为非重组型(亲本型),50%为重组型(非亲本型),这与不同染色体上位点等位基因自由组合时所见的比例完全相同。
Thus if two syntenic loci are sufficiently far apart on the same chromosome to ensure at least one crossover between them in every meiosis, then the ratio of recombinant to A A A a a A a A a a b b b b A A A a a A a A a a b b A A A a a a b b A A A a a a b b A A A a a a b b b b A a A a b b A a A a A a A a b b b b b 8 NR 8 R NR:R=1:1 2 AB 2 ab 0 Ab 0 a B 0 AB 0 ab 2 Ab 2 a B 1 AB 1 ab 1 Ab 1 a B 1 AB 1 ab 1 Ab 1 a B NR NR R R NR NR R R NR NR R R NR NR R R 2 NR 2 R NR:R=1:1 B B B B B b B B B b B B B b B B B b B B B B B B B B B B B Crossovers result in new combinations of maternally and paternally derived alleles on the recombinant chromosomes present in gametes, shown on the right.因此,如果两个同线位点在同一条染色体上相距足够远以确保每次减数分裂中它们之间至少发生一次交换,则重组型与A A A a a A a A a a b b b b A A A a a A a A a a b b A A A a a a b b A A A a a a b b A A A a a a b b b b A a A a b b A a A a A a A a b b b b b 8 NR 8 R NR:R=1:1 2 AB 2 ab 0 Ab 0 a B 0 AB 0 ab 2 Ab 2 a B 1 AB 1 ab 1 Ab 1 a B 1 AB 1 ab 1 Ab 1 a B NR NR R R NR NR R R NR NR R R NR NR R R 2 NR 2 R NR:R=1:1 B B B B B b B B B b B B B b B B B b B B B B B B B B B B B的比率,交换导致配子中重组染色体上源自母亲和父亲的等位基因产生新组合,如右图所示。
If no crossing over occurs in the interval between loci 1 and 2, only parental (nonrecombinant) allele combinations, AB and ab, occur in the offspring.如果位点1和2之间的区间内没有发生交换,则后代中只会出现亲本型(非重组)等位基因组合AB和ab。
If one or two crossovers occur in the interval between the loci, half the gametes will contain a nonrecombinant combination of alleles and half the recombinant combination.如果位点之间的区间内发生一次或两次交换,则一半配子将包含非重组等位基因组合,另一半包含重组组合。
The same is true if more than two crossovers occur between the loci (not illustrated here).如果位点之间发生超过两次交换,情况也是如此(此处未图示)。
NR, Nonrecombinant; R, recombinant. nonrecombinant genotypes will be, on average, 1: 1— just as if the loci were on separate chromosomes and assorting independently.NR,非重组;R,重组;非重组基因型平均比例为1:1,就像位点位于不同染色体上独立分配一样。
4/46
Recombination Frequency and Map Distance Frequency of Recombination as a Measure of Distance Between Loci Suppose now th…
Ch11 — Segment 4
Recombination Frequency and Map Distance Frequency of Recombination as a Measure of Distance Between Loci Suppose now that two loci are on the same chromosome but are either far apart, very close together, or somewhere in between .重组频率与图谱距离 重组频率作为基因座间距离的度量 现在假设两个基因座位于同一条染色体上,但它们要么相距很远,要么非常接近,要么介于两者之间。
As we just saw, when the loci are far apart , at least one crossover will occur in the segment of the chromosome between loci 1 and 2, and there will be gametes of both the nonrecombinant genotypes (AB and ab) and recombinant genotypes (Ab and a B), in equal proportions (on average) in the offspring.正如我们刚刚看到的,当基因座相距很远时,在基因座1和2之间的染色体片段中至少会发生一次交叉,并且在后代中会以(平均)相等的比例同时存在非重组基因型(AB和ab)和重组基因型(Ab和aB)的配子。
On the other hand, if two loci are so close together on the same chromosome that crossovers never occur between them, there will be no recombination; the nonrecombinant genotypes (parental chromosomes AB and ab in are transmitted together all of the time, and the frequency of the recombinant genotypes Ab and a B will be 0.另一方面,如果两个基因座在同一条染色体上如此接近以至于它们之间从不发生交叉,则不会有重组;非重组基因型(亲本染色体AB和ab)始终一起传递,重组基因型Ab和aB的频率为0。
In between these two extremes is the situation in which two loci are far enough apart that one recombination between the loci occurs in some meioses but not in others .在这两个极端之间的情况是,两个基因座之间的距离足够远,以至于在一些减数分裂中发生一次重组,而在其他减数分裂中则不发生。
In this situation we observe nonrecombinant combinations of alleles in the offspring when no crossover occurred and recombinant combinations when a recombination has occurred.在这种情况下,当没有发生交叉时,我们在后代中观察到等位基因的非重组组合;当发生重组时,观察到重组组合。
The frequency of recombinant chromosomes at the two loci will fall between 0 and 50%.这两个基因座的重组染色体频率将介于0%和50%之间。
The crucial point is that the closer together two loci are, the smaller the recombination frequency and the fewer recombinant genotypes in the offspring.关键点在于,两个基因座越接近,重组频率就越小,后代中的重组基因型就越少。
Detecting Recombination Events Requires Heterozygosity and Knowledge of Phase Detecting the recombination events between loci requires that (1) a parent be heterozygous (informative) at both loci and (2) we know which allele at locus 1 is on the same chromosome as which allele at locus 2.检测重组事件需要杂合性和相位信息 检测基因座间的重组事件要求:(1)亲本在两个基因座均为杂合(信息性),并且(2)我们知道基因座1的哪个等位基因与基因座2的哪个等位基因位于同一条染色体上。
In an individual who is heterozygous at two syntenic loci, one with alleles A and a, the other B and b, phase refers to which allele at the first locus is on the same chromosome with which allele at the second locus .在一个在两位点(一个具有等位基因A和a,另一个具有等位基因B和b)均为杂合的个体中,相位指的是第一个基因座的哪个等位基因与第二个基因座的哪个等位基因位于同一条染色体上。
The set of alleles on the same homologue (A and B, or a and b) are said to be in cis (or in coupling) and form what is referred to as a haplotype.位于同一同源染色体上的等位基因组合(A和B,或a和b)被称为顺式(或相引),并形成所谓的单倍型。
In contrast, alleles on the different homologues (A and b, or a and B) are in trans (or in repulsion) .相比之下,位于不同同源染色体上的等位基因(A和b,或a和B)处于反式(或相斥)状态。
A Locus 1→ Locus 2→ a B b A B a b A b a B Nonrecombinant = Recombinant (AB + ab) = (Ab + a B) A Locus 1→ Locus 2→ a B b A B a b Nonrecombinant only (AB + ab) only A Locus 1→ Locus 2→ a B b A B a b A B a B Nonrecombinant > Recombinant (AB + ab) > (Ab + a B) A B C The loci are far apart and at least one crossover between them is likely to occur in every meiosis.A 基因座1→基因座2→ a B b A B a b A b a B 非重组=重组 (AB+ab)=(Ab+aB) A 基因座1→基因座2→ a B b A B a b 仅非重组 (AB+ab) A 基因座1→基因座2→ a B b A B a b A B a B 非重组>重组 (AB+ab)>(Ab+aB) A B C 基因座相距很远,每次减数分裂中它们之间至少可能发生一次交叉。
(B) The loci are so close together that crossing over between them is not observed, regardless of the presence of crossovers elsewhere on the chromosome.(B) 基因座非常接近,以至于它们之间观察不到交叉,无论染色体其他部位是否存在交叉。
(C) The loci are close together on the same chromosome but far enough apart that crossing over occurs in the interval between the two loci only in some meioses but not in most others.(C) 基因座在同一条染色体上相距很近,但距离足够远,以至于只在某些减数分裂中两个基因座之间的区间发生交叉,而在大多数其他减数分裂中不发生。
In coupling (cis): Aand Baandb In repulsion (trans): aand BAandb B A b a在相引(顺式)中:A和B,a和b。在相斥(反式)中:a和B,A和b。B A b a
5/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 207 , a degenerative disease of the retina that causes progressive blind…
Ch11 — Segment 5
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 207 , a degenerative disease of the retina that causes progressive blindness in association with abnormal retinal pigmentation.鉴定人类疾病的遗传基础 207 是一种视网膜退行性疾病,伴有异常视网膜色素沉着,导致进行性失明。
As shown, individual I-1 is heterozygous at both marker locus 1 (with alleles A and a) and marker locus 2 (with alleles B and b), as well as heterozygous for the disorder (D is the disease allele, d is the normal allele).如图所示,个体 I-1 在标记基因座1(等位基因A和a)和标记基因座2(等位基因B和b)均为杂合,并且对该疾病也为杂合(D为疾病等位基因,d为正常等位基因)。
The alleles A-D-B form one haplotype, and a-d-b the other.等位基因A-D-B构成一个单倍型,a-d-b构成另一个单倍型。
Because we know her spouse is homozygous at all three loci and can only pass on the a, b, and d alleles, we can easily determine which alleles the children received from their mother and thus trace the inheritance of her RP-causing allele or her normal allele at that locus, as well as the alleles at both marker loci in her children.由于我们知道她的配偶在所有三个基因座上均为纯合,且只能传递a、b和d等位基因,因此我们可以轻松确定子代从母亲那里获得了哪些等位基因,从而追踪她在该基因座上致RP等位基因或正常等位基因的遗传,以及她子代中两个标记基因座的等位基因。
Close inspection of had been homozygous bb at locus 2, then all children would have inherited a maternal b allele, regardless of whether they received a mutant D or normal d allele at the RP9 locus.仔细检查若在基因座2为bb纯合,则所有子代都会遗传母亲的b等位基因,无论他们在RP9基因座上是获得突变型D还是正常型d等位基因。
Because she is not informative at locus 2 in this scenario, it would be impossible to determine whether recombination had occurred.因为在此情景下她在基因座2上无信息性,所以不可能确定是否发生了重组。
Similarly, if the information provided for the family in is the Greek letter theta, θ, where θ varies from 0 (no recombination at all) to 0. 5 (independent assortment).类似地,如果为该家庭提供的信息是希腊字母θ,θ的取值范围从0(无重组)到0.5(独立分配)。
If two loci are so close together that θ = 0 between them (as in , they are said to be completely linked; if they are so far apart that θ = 0. 5 (as in , they are assorting independently and are unlinked.如果两个基因座非常接近,使得θ=0(如...),则称它们完全连锁;如果它们相距甚远,使得θ=0.5(如...),则它们独立分配且不连锁。
In between these two extremes are various degrees of linkage.在这两个极端之间存在着不同程度的连锁。
Genetic Maps and Physical Maps The map distance between two loci is a theoretical concept that is based on actual data—the extent of observed recombination, θ, between the loci.遗传图谱与物理图谱 两个基因座之间的图距是一个基于实际数据(即观察到的重组程度θ)的理论概念。
Map distance is measured in units called centimorgans (c M), defined as the genetic length over which, on average, one crossover occurs in 1% of meioses.图距以名为厘摩(cM)的单位测量,定义为平均有1%减数分裂发生一次交换的遗传长度。
(The centimorgan is 1100 of a “morgan,” named after Thomas Hunt Morgan, who first observed genetic recombination in the fruit fly, Drosophila.) Therefore a recombination fraction of 1% (i. e., θ = 0. 01) translates approximately into a map distance of 1 c M.(厘摩是“摩根”的1/100,以托马斯·亨特·摩根命名,他首次在果蝇中观察到遗传重组。)因此,重组率1%(即θ=0.01)大致相当于1 cM的图距。
As we discussed before in this chapter, the recombination frequency between two loci increases proportionately with the distance between two loci only up to a point, because once markers are far enough apart that at least one recombination will always occur, the observed recombination frequency will equal 50% (θ = 0. 5), no matter how physically far apart the two loci are.正如我们本章前面所讨论的,两个基因座之间的重组频率仅在一定范围内与它们之间的距离成正比,因为一旦标记相距足够远以至于总是发生至少一次重组,观察到的重组频率将等于50%(θ=0.5),无论这两个基因座在物理上相距多远。
To accurately measure true genetic map distance between two widely spaced loci, therefore, one has to use markers spaced at short genetic distances (≤1 c M) in the interval between these two loci, and then add up the values of θ between the intervening markers; the values of θ between pairs of closely neighboring markers will be good approximations of the genetic distances between them.因此,要精确测量两个相距较远的基因座之间的真实遗传图距,必须使用这两个基因座之间间隔中短遗传距离(≤1 cM)的标记,然后累加中间各标记之间的θ值;相邻近标记对之间的θ值将是它们之间遗传距离的良好近似。
Using this approach, the genetic length of an entire human genome has been measured and, 1 I II 2 1 2 3 4 5 6 7 8 Locus 2 RP9 Locus 1 B D A b d a b d A b d A b D A b d A B D a B D A b d a b d a b d A b d A Only the mother’s contribution to the children’s genotypes is shown.采用这一方法,整个人类基因组的遗传长度已被测量,并且 1 I II 2 1 2 3 4 5 6 7 8 基因座2 RP9 基因座1 B D A b d a b d A b d A b D A b d A B D a B D A b d a b d a b d A b d A 仅显示母亲对子代基因型的贡献。
The mother (I-1) is affected with this dominant disease and is heterozygous at the RP9 locus (Dd) as well as at loci 1 and 2.母亲(I-1)患有这种显性遗传病,并且在RP9基因座(Dd)以及基因座1和2上均为杂合。
She carries the A and B alleles on the same chromosome as the mutant RP9 allele (D).她在携带突变型RP9等位基因(D)的同一染色体上携带A和B等位基因。
The unaffected father is homozygous normal (dd) at the RP9 locus as well as at the two marker loci (AA and bb); his contributions to his offspring are not considered further.未患病的父亲在RP9基因座以及两个标记基因座(AA和bb)上均为正常纯合(dd);他对其后代的贡献不再考虑。
Two of the three affected offspring have inherited the B allele at locus 2 from their mother, whereas individual II-3 inherited the b allele.三名患病子代中的两人从母亲那里继承了基因座2的B等位基因,而个体II-3继承了b等位基因。
The five unaffected offspring have also inherited the b allele.五名未患病子代也继承了b等位基因。
Thus, seven of eight offspring are nonrecombinant between the RP9 locus and locus 2.因此,八名子代中的七人在RP9基因座与基因座2之间是非重组的。
However, individuals II-2, II-4, II-6, and II-8 are recombinant for RP9 and locus 1, indicating that meiotic crossover has occurred between these two loci.然而,个体II-2、II-4、II-6和II-8在RP9与基因座1之间是重组的,表明这两个基因座之间发生了减数分裂交换。
6/46
interestingly, found to differ between the sexes.
Ch11 — Segment 6
interestingly, found to differ between the sexes.有趣的是,发现在两性之间存在差异。
When measured in female meiosis, genetic length of the human genome is ~60% greater (≈4596 c M) than when it is measured in male meiosis (2868 c M), and this sex difference is consistent and uniform across each autosome.在女性减数分裂中测量时,人类基因组的遗传长度比在男性减数分裂中测量时长约60%(≈4596 cM),且此性别差异在每条常染色体上一致且均匀。
The sex-averaged genetic length of the entire haploid human genome, which is estimated to contain ~3. 3 billion base pairs of DNA, or ≈3300 Mb (see Chapter 2), is 3790 c M, for an average of ~1. 15 c M/Mb.整个单倍体人类基因组的性别平均遗传长度为3790 cM,该基因组估计含有约33亿个碱基对的DNA,即约3300 Mb(见第2章),平均约为1.15 cM/Mb。
Pairwise measurements of recombination between genetic markers separated by 1 Mb or more gives a fairly constant ratio of genetic distance to physical distance of ~1 c M/Mb.对相距1 Mb或以上的遗传标记进行成对重组测量,得到的遗传距离与物理距离之比相当恒定,约为1 cM/Mb。
However, when recombination is measured at much higher resolution, such as between markers spaced less than 100 kb apart, recombination per unit length becomes nonuniform and can range over four orders of magnitude (0. 01–100 c M/ Mb).然而,当以更高分辨率测量重组时,例如在间距小于100 kb的标记之间,单位长度的重组变得不均匀,其范围可达四个数量级(0.01–100 cM/Mb)。
When viewed on the scale of a few tens of kilobase pairs of DNA, the apparent linear relationship between physical distance in base pairs and recombination between polymorphic markers located millions of base pairs of DNA apart is, in fact, the result of an averaging of so-called hot spots of recombination interspersed among regions of little or no recombination.当从几十千碱基对DNA的尺度观察时,碱基对物理距离与相隔数百万碱基对的DNA多态标记之间表观的线性关系,实际上是穿插在极少或没有重组区域之间的所谓重组热点的平均结果。
Hot spots occupy only ~6% of sequence in the genome and yet account for ~60% of all the meiotic recombination in the human genome.热点仅占基因组序列的约6%,却贡献了人类基因组中约60%的减数分裂重组。
The impact of this nonuniformity of recombination at high resolution is discussed next, as we address the phenomenon of linkage disequilibrium.接下来在讨论连锁不平衡现象时,将探讨高分辨率下这种重组不均匀性的影响。
Linkage Disequilibrium It is generally the case that the alleles at two loci will not show any preferred phase in the population if the loci are linked, but at a distance of 0. 1 to 1 c M or more.连锁不平衡 通常情况下,如果两个位点连锁,但相距0.1至1 cM或更远,它们在群体中不会表现出任何优势相。
For example, suppose loci 1 and 2 are 1 c M apart.例如,假设位点1和位点2相距1 cM。
Suppose further that allele A is present on 50% of the chromosomes in a population and allele a on the other 50%, whereas at locus 2, a disease susceptibility allele S is present on 10% of chromosomes and the protective allele s is on 90% .进一步假设,在群体中,等位基因A存在于50%的染色体上,等位基因a存在于另外50%的染色体上,而在位点2,疾病易感等位基因S存在于10%的染色体上,保护性等位基因s存在于90%的染色体上。
Because the frequency of the A-S haplotype (freq(A-S)) is simply the product of the frequencies of the two alleles—freq(A) × freq(S) = 0. 5 × 0. 1 = 0. 05—the alleles are said to be in linkage equilibrium .由于A-S单倍型的频率(freq(A-S))仅为两个等位基因频率的乘积——freq(A)×freq(S)=0.5×0.1=0.05——因此称这些等位基因处于连锁平衡。
That is, the frequencies of the four possible haplotypes, A-S, A-s, a-S, and a-s follow directly from the allele frequencies of A, a, S, and s.也就是说,四种可能单倍型A-S、A-s、a-S和a-s的频率直接由等位基因A、a、S和s的频率得出。
However, as we examine haplotypes involving loci that are very close together, we find that knowing the allele frequencies for these loci individually does not allow us to predict the four haplotype frequencies.然而,当我们检查涉及非常接近的位点的单倍型时,我们发现,单独知道这些位点的等位基因频率并不能预测四种单倍型频率。
The frequency of any one of the haplotypes, freq(A-S) for example, may not be equal to the product of the frequencies of the individual alleles that make up that haplotype; in this situation, freq(A-S) ≠ freq(A) × freq(S), and the alleles are thus said to be in linkage disequilibrium (LD).任何一种单倍型的频率,例如freq(A-S),可能不等于构成该单倍型的各个等位基因的频率的乘积;在这种情况下,freq(A-S)≠freq(A)×freq(S),因此称这些等位基因处于连锁不平衡(LD)。
The deviation (“delta”) between the expected and actual haplotype frequencies is called D and is given by: D freq(-) freq( - freq - freq( - A S a s A s a s) ()) D ≠ 0 is equivalent to saying the alleles are in LD, whereas D = 0 means the alleles are in linkage equilibrium.期望单倍型频率与实际单倍型频率之间的偏差(“delta”)称为D,由下式给出:D = freq(A-S)freq(a-s) - freq(A-s)freq(a-S)。D≠0相当于说等位基因处于LD,而D=0意味着等位基因处于连锁平衡。
Examples of LD are illustrated in = 0. 1 freq(s) = 0. 9 Haplotype A-S freq(A-S) = 0. 05 Haplotype A-s freq(A-s) = 0. 45 Haplotype a-S freq(a-S) = 0. 05 freq(A) = 0. 5 freq(a) = 0. 5 Haplotype a-s freq(a-s) = 0. 45 Allele frequencies at locus 2 Allele frequencies at locus 1 freq(S) = 0. 1 freq(s) = 0. 9 Haplotype A-S freq(A-S)=0 Haplotype A-s freq(A-s) = 0. 5 Haplotype a-S freq(a-S) = 0. 1 freq(A) = 0. 5 freq(a) = 0. 5 Haplotype a-s freq(a-s) = 0. 4 Allele frequencies at locus 2 Allele frequencies at locus 1 freq(S) = 0. 1 freq(s) = 0. 9 Haplotype A-S freq(A-S) = 0. 01 Haplotype A-s freq(A-s) = 0. 49 Haplotype a-S freq(a-S) = 0. 09 freq(A) = 0. 5 freq(a) = 0. 5 Haplotype a-s freq(a-s) = 0. 41 Allele frequencies at locus 2 Linkage equilibrium: Haplotype frequencies are as expected from allele frequencies Linkage disequilibrium: Haplotype frequencies diverge from what is expected from allele frequencies Partial linkage disequilibrium: Haplotype frequencies are rarer than expected from allele frequencies A B C Under linkage equilibrium, haplotype frequencies are as expected from the product of the relevant allele frequencies.LD的示例图示如下:freq(S)=0.1, freq(s)=0.9;单倍型A-S freq(A-S)=0.05,单倍型A-s freq(A-s)=0.45,单倍型a-S freq(a-S)=0.05,freq(A)=0.5, freq(a)=0.5,单倍型a-s freq(a-s)=0.45;位点2的等位基因频率:freq(S)=0.1, freq(s)=0.9;位点1的等位基因频率:freq(A)=0.5, freq(a)=0.5;单倍型A-S freq(A-S)=0,单倍型A-s freq(A-s)=0.5,单倍型a-S freq(a-S)=0.1,单倍型a-s freq(a-s)=0.4;位点2的等位基因频率:freq(S)=0.1, freq(s)=0.9;位点1的等位基因频率:freq(A)=0.5, freq(a)=0.5;单倍型A-S freq(A-S)=0.01,单倍型A-s freq(A-s)=0.49,单倍型a-S freq(a-S)=0.09,单倍型a-s freq(a-s)=0.41;连锁平衡:单倍型频率与根据等位基因频率预期的相符;连锁不平衡:单倍型频率偏离根据等位基因频率预期的结果;部分连锁不平衡:单倍型频率比根据等位基因频率预期的更罕见;在连锁平衡下,单倍型频率与根据相关等位基因频率的乘积预期的结果相符。
(B) Loci 1 and 2 are located very close to one another, and alleles at these loci show strong linkage disequilibrium.(B)位点1和位点2彼此非常接近,这些位点上的等位基因表现出强烈的连锁不平衡。
Haplotype A-S is absent and a-s is less frequent (0. 4 instead of 0. 45) compared to what is expected from allele frequencies.与根据等位基因频率预期的结果相比,单倍型A-S缺失,a-s频率较低(0.4而非0.45)。
(C) Alleles at loci 1 and 2 show partial linkage disequilibrium.(C)位点1和位点2上的等位基因表现出部分连锁不平衡。
Haplotypes, A-S and as are underrepresented compared to what is expected from allele frequencies.与根据等位基因频率预期的结果相比,单倍型A-S和a-s的代表性不足。
Note that the allele frequencies for A and a at locus 1 and for S and s at locus 2 are the same in all three tables; it is the way the alleles are distributed in haplotypes, shown in the central four cells of the table, that differs.注意,在所有三个表格中,位点1的等位基因A和a的频率以及位点2的等位基因S和s的频率均相同;不同的是等位基因在单倍型中的分布方式,如表格中央四个单元格所示。
7/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 209 haplotype is present on only 1% of chromosomes in the population .
Ch11 — Segment 7
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 209 haplotype is present on only 1% of chromosomes in the population .识别人类疾病的遗传基础 209 单倍型在人群中仅存在于 1% 的染色体上。
The A-S haplotype has a frequency much below what one would expect on the basis of the frequencies of alleles A and S in the population as a whole, and D &lt; 0, whereas the haplotype a-S has a frequency much greater than expected and D &gt; 0.A-S单倍型的频率远低于基于整个群体中等位基因A和S频率所预期的值,且D < 0,而a-S单倍型的频率远高于预期,且D > 0。
In other words, chromosomes carrying the susceptibility allele S are enriched for allele a at the expense of allele A, compared with chromosomes that carry the protective allele s.换言之,与携带保护性等位基因s的染色体相比,携带易感性等位基因S的染色体上等位基因a富集而等位基因A减少。
Note, however, that the individual allele frequencies are unchanged; it is only how they are distributed into haplotypes that differ, and this is what determines if there is LD.但请注意,各等位基因频率本身并未改变;改变的只是它们在单倍型中的分布方式,而这正是决定是否存在连锁不平衡的因素。
Linkage Disequilibrium Has Both Biologic and Historical Causes What causes LD?连锁不平衡具有生物学和历史双重原因 什么导致连锁不平衡?
When a pathogenic allele first enters the population (by mutation or by immigration of a founder who carries the altered allele), the particular set of alleles at polymorphic loci linked to the disease locus constitutes a disease-associated haplotype .当致病等位基因首次进入群体时(通过突变或携带改变等位基因的奠基者迁入),与疾病位点连锁的多态性位点上的一组特定等位基因构成了疾病相关单倍型。
The degree to which this original disease-associated haplotype will persist over time depends in part on the probability that recombination removes the diseaseassociated allele from the original haplotype and onto chromosomes with different sets of alleles at these linked loci.这种原始疾病相关单倍型能够持续存在的时间部分取决于重组将疾病相关等位基因从原始单倍型移除并转移到这些连锁位点上具有不同等位基因组合的染色体上的概率。
The speed with which recombination will move the pathogenic allele onto a new haplotype depends on a number of factors: The number of generations (and therefore the number of opportunities for recombination) since the mutation first appeared.重组将致病等位基因转移到新单倍型上的速度取决于多个因素:自突变首次出现以来的世代数(因此也是重组机会的数量)。
The frequency of recombination per generation between the loci.位点之间每代重组发生的频率。
The smaller the value of θ, the greater is the chance that the disease-associated haplotype will persist intact.θ值越小,疾病相关单倍型保持完整的可能性就越大。
Processes of natural selection for or against particular haplotypes.针对特定单倍型的自然选择过程,无论有利还是不利。
If a haplotype combination undergoes either positive selection (and is preferentially passed on) or experiences negative selection (and is less readily passed on), it will be either over- or under-represented in that population.如果一种单倍型组合经历正向选择(因而优先传递)或负向选择(因而较难传递),那么它在群体中的出现频率将相应偏高或偏低。
Measuring Linkage Disequilibrium Although conceptually valuable, the discrepancy, D, between the expected and observed frequencies of haplotypes is not a good way to quantify LD because it varies not only with degree of LD but also with the allele frequencies themselves.测量连锁不平衡 尽管在概念上很有价值,但单倍型预期频率与观察频率之间的差值D并不是量化连锁不平衡的良好方式,因为它不仅随连锁不平衡程度变化,还受等位基因频率本身影响。
To quantify varying degrees of LD, therefore, geneticists often use a measure derived from D, referred to as D′ (see 1).因此,为量化不同程度的连锁不平衡,遗传学家常使用一种源自D的测量指标,称为D′(见1)。
D′ is designed to vary from 0, indicating linkage equilibrium, to a maximum of ±1, indicating very strong LD.D′的设计取值范围从0(表示连锁平衡)到最大±1(表示非常强的连锁不平衡)。
LD is a result, not only of genetic distance, but of the amount of time A B Fragmentation of original chromosome by recombination as population expands through multiple generations Pathogenic variant located within region of linkage disequilibrium Pathogenic variant on founder chromosome Over many generations, the only alleles that remain in coupling phase with the variant are those at loci so close to the disease associated locus that recombination between the loci is very rare.连锁不平衡不仅是遗传距离的结果,也是时间量的结果 A B 随着群体通过多代扩张,重组使原始染色体片段化 致病变异位于连锁不平衡区域内 致病变异位于奠基者染色体上 经过许多世代,唯一与变异保持相偶相的等位基因是那些位于与疾病相关位点极近、以至于位点间重组非常罕见的位点上的等位基因。
These alleles are in linkage disequilibrium with the pathogenic variant and constitute a disease-associated haplotype.这些等位基因与致病变异处于连锁不平衡状态,并构成疾病相关单倍型。
(B) Affected individuals in the current generation (arrows) carry the pathogenic variant (X) in linkage disequilibrium with the disease-associated haplotype (individuals in blue).(B) 当前世代中的受累个体(箭头所示)携带与疾病相关单倍型处于连锁不平衡的致病变异(X)(蓝色个体)。
Depending on the age of the pathogenic variant and other population genetic factors, a disease-associated haplotype ordinarily spans a region of DNA of a few kb to a few hundred kb. )取决于致病变异的年龄及其他群体遗传学因素,疾病相关单倍型通常跨越从几kb到几百kb的DNA区域。 )
8/46
during which recombination had a chance to occur and the possible effects of selection for or against particular haploty…
Ch11 — Segment 8
during which recombination had a chance to occur and the possible effects of selection for or against particular haplotypes.在此期间,重组有机会发生,并且可能产生针对特定单倍型的选择或反选择效应。
Different populations, therefore, living in different environments and with different histories can have different values of D′ between the same two alleles at the same locus in the genome. shorter haplotypes more rapidly than average, resulting in linkage equilibrium between SNPs on one side and the other side of the hot spot.因此,生活在不同环境且具有不同历史的人群,在基因组同一基因座上相同两个等位基因之间的 D′ 值可能存在差异;较短的单倍型比平均更快地被打破,导致热点一侧与另一侧的 SNP 之间形成连锁平衡。
The correlation is by no means exact, and many apparent boundaries between LD blocks are not located over evident recombination hot spots.这种相关性绝非精确,许多 LD 区块之间的明显边界并不位于明显的重组热点之上。
This lack of perfect correlation should not be surprising, given what we have already surmised about LD: it is affected not only by how likely a recombination event is (i. e., where the hot spots are) but also by the age of the population, the frequency of the haplotypes originally present in the founding members of that population, and whether there has been either positive or negative selection for particular haplotypes.考虑到我们对 LD 已有的推断——它不仅受重组事件发生概率(即热点位置)的影响,还受群体年龄、该群体奠基成员原始存在的单倍型频率,以及特定单倍型是否经历过正向或负向选择的影响——这种不完全相关并不令人意外。
STRATEGIES FOR DISCOVERY OF DISEASEASSOCIATED GENES In clinical medicine, a disease state is defined by a collection of phenotypic findings seen in a patient or group of patients.疾病相关基因发现策略 在临床医学中,疾病状态由患者或患者群体中观察到的表型发现集合来定义。
Designating such a disease as “genetic”—inferring the existence of a gene whose alteration is responsible for or contributes to the disease—comes from detailed genetic analysis, applying the principles outlined in Chapters 7 and 9.将此类疾病指定为“遗传性”——推断存在一个基因,其改变导致或促成该疾病——需要依据第7章和第9章概述的原则进行详细的遗传分析。
However, surmising the existence of a gene or genes in such a way does not tell us which of the ~20,000 coding and ~18,000 noncoding genes in the genome is involved, what its function might be, or how it causes or contributes to the disease.然而,以这种方式推断一个或多个基因的存在,并不能告诉我们基因组中约20000个编码基因和约18000个非编码基因中哪一个参与其中、其功能可能是什么,或者它如何导致或促成疾病。
Strategies for discovery of genes associated with human disease have evolved over the years, from gene mapping to genome-wide sequencing.与人类疾病相关基因的发现策略多年来不断发展,从基因定位到全基因组测序。
Combining these two approaches has provided an effective strategy.将这两种方法结合起来提供了一种有效的策略。
Mapping of such genes has historically been a critical and necessary first step in identifying the gene(s) in which certain variants are responsible for causing or increasing susceptibility to disease.历史上,此类基因的定位是识别基因(其中特定变异导致疾病或增加易感性)的关键且必要的第一步。
Mapping the gene focuses attention on a region of the genome in which to carry out a systematic analysis of all the genes in that region, to identify variation that contributes to the disease.基因定位将注意力聚焦于基因组的一个区域,以便对该区域内所有基因进行系统分析,从而识别导致疾病的变异。
The marked fall in cost of DNA sequencing over the last decade has made it feasible to take a genomewide sequencing approach to gene discovery.过去十年 DNA 测序成本的显著下降,使得采用全基因组测序方法进行基因发现变得可行。
Sequencing the genomes (or just the coding portion of the genome, the exome) of cohorts with similar phenotypes, followed by systematic filtering, has proven a powerful approach for gene discovery.对具有相似表型的队列进行基因组测序(或仅测序基因组的编码部分,即外显子组),随后进行系统性过滤,已被证明是一种强大的基因发现方法。
This is especially true for disorders with a new (de novo) dominant mechanism that may be intractable to gene mapping.对于可能难以通过基因定位解决的、具有新发(de novo)显性机制疾病尤其如此。
Incorporation of gene mapping information into the filtering strategy of genomewide sequencing has also proven an effective approach to narrow in on loci that are associated with disease.将基因定位信息整合到全基因组测序的过滤策略中,也被证明是缩小与疾病相关位点范围的有效方法。
Regardless of strategy, identification of the gene that harbors the DNA variants responsible for either causing a mendelian disorder or increasing susceptibility to a genetically complex disease, allows the full spectrum of variation in that gene to be studied.无论采用何种策略,识别携带导致孟德尔遗传病或增加遗传复杂性疾病易感性的 DNA 变异的基因,可以研究该基因变异的全谱。
We can determine the degree of allelic heterogeneity, the penetrance of 1 MEASURING LINKAGE DISEQUILIBRIUM D′ = D/F where D freq(-)freq(-) freq(-)freq(-) A S a s A s a S and F is a correction factor that helps account for the allele frequencies.我们可以确定等位基因异质性的程度、外显率,以及测量连锁不平衡的 D′ = D/F,其中 D = freq(A)freq(s) - freq(a)freq(S),F 是一个有助于校正等位基因频率的校正因子。
The value of F depends on whether D itself is a positive or negative number.F 的值取决于 D 本身是正数还是负数。
F = the smaller of freq(A) × freq(s) or freq(a) × freq(S) if D&gt;0 F = the smaller of freq(A) × freq(S) or freq(a) × freq(s) if D&lt;0 Clusters of Alleles Form Blocks Defined by Linkage Disequilibrium Analysis of pairwise measurements of D′ for neighboring variants, particularly common single nucleotide variants (SNVs), across the genome reveals a complex genetic architecture for LD.若 D>0,F = freq(A)×freq(s) 和 freq(a)×freq(S) 中较小者;若 D<0,F = freq(A)×freq(S) 和 freq(a)×freq(s) 中较小者。等位基因簇形成由连锁不平衡定义的区块 对基因组中相邻变异(尤其是常见单核苷酸变异,SNV)的 D′ 成对测量进行分析,揭示了 LD 的复杂遗传结构。
Contiguous SNVs can be grouped into clusters of varying size, in which the SNVs in any one cluster show high levels of LD with each other but not with SNPs outside that cluster .连续 SNV 可归为大小不一的簇,其中任一簇内的 SNV 彼此之间显示出高 LD,但与簇外的 SNP 则不然。
For example, the nine polymorphic loci in cluster 1 , each consisting of two alleles, have the potential to generate 29 = 512 different haplotypes; yet, only five haplotypes constitute 98% of all haplotypes seen.例如,簇1中的九个多态位点(每个由两个等位基因组成)理论上可产生 2⁹ = 512 种不同的单倍型;然而,仅五种单倍型就构成了所见单倍型的 98%。
The absolute values of |D′| between SNVs within the cluster are well above 0. 8.簇内 SNV 之间的 |D′| 绝对值远高于 0.8。
Clusters of loci with alleles in high LD across segments of only a few kilobase pairs to a few dozen kilobase pairs are termed LD blocks.在仅有数千碱基对到数万碱基对片段上,等位基因呈高 LD 的位点簇称为 LD 区块。
The size of an LD block encompassing alleles at a particular set of polymorphic loci is not identical in all populations.覆盖特定多态位点等位基因的 LD 区块大小在所有人种中并不相同。
African populations have smaller blocks, averaging 7. 3 kb per block across the genome, compared with 16. 3 kb in Europeans; Chinese and Japanese block sizes are comparable to each other and are intermediate, averaging 13. 2 kb.非洲人群的区块较小,基因组平均每区块 7.3 kb,而欧洲人为 16.3 kb;中国人和日本人的区块大小彼此相近且居中,平均为 13.2 kb。
This difference in block size is almost certainly the result of the smaller number of generations since the founding of the non-African populations compared with populations in Africa, thereby limiting the time in which there has been opportunity for recombination to break up regions of LD.这种区块大小的差异几乎可以肯定是因为非非洲人群自奠基以来的世代数少于非洲人群,从而限制了重组有机会破坏 LD 区域的时间。
Is there a biologic basis for LD blocks, or are they simply genetic phenomena reflecting human (and genome) history?LD 区块是否存在生物学基础,或者它们仅仅是反映人类(及基因组)历史的遗传现象?
It appears that biology contributes to LD block structure in that the boundaries between LD blocks often coincide with meiotic recombination hot spots, discussed earlier .生物学似乎对 LD 区块结构有贡献,因为 LD 区块之间的边界常常与减数分裂重组热点重合,这一点前面已讨论过。
Such recombination hot spots would break up any haplotypes spanning them into two此类重组热点会将跨越它们的任何单倍型分裂成两部分。
9/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 21140% 30% 11% 9% 8% 98% Haplotype frequency 40% 60% 100% Allele frequen…
Ch11 — Segment 9
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 21140% 30% 11% 9% 8% 98% Haplotype frequency 40% 60% 100% Allele frequency 1 2 3 4 5 6 7 8 9 C T C G G T T T C A A T 42% 31% 26% 99% Haplotype frequency G A 10 CLUSTER 1 CLUSTER 2 1 SNP# 2 3 4 5 6 7 8 9 10 11 14 13 12 50 kb 100 kb 150 kb 11 12 13 14 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 0. 9 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 1. 0 0. 8 1. 0 1. 0 1 2 3 4 5 6 7 8 9 10 11 12 13 1. 0 14 1. 0 0 50 100 100 50 1. 0 c M/Mb Mb 89 kb 14 kb 42 kb 1. 0 0. 8 1. 0 1. 0 0. 6 1. 0 1. 0 0. 9 0. 9 1. 0 A B C In cluster 1, containing SNPs 1 through 9, five of the 29 = 512 theoretically possible haplotypes are responsible for 98% of all the haplotypes in the population, reflecting substantial linkage disequilibrium (LD) among these SNP loci.在包含SNP1至9的簇1中,理论上可能的2^9=512种单倍型中有五种构成了群体中所有单倍型的98%,反映了这些SNP位点间存在显著的连锁不平衡(LD)。
Similarly, in cluster 2, only three of the 24 = 16 theoretically possible haplotypes involving SNPs 11 to 14 represent 99% of all the haplotypes found.类似地,在簇2中,涉及SNP11至14的理论上可能的2^4=16种单倍型中仅有三种代表了所发现全部单倍型的99%。
In contrast, alleles at SNP 10 are found in linkage equilibrium with the SNPs in cluster 1 and cluster 2.相反,SNP10的等位基因被发现与簇1和簇2中的SNP处于连锁平衡状态。
(B) A schematic diagram in which each red box contains the pairwise measurement of the degree of LD between two SNPs (e. g., the arrow points to the box, outlined in black, containing the value of D′ for SNPs 2 and 7).(B) 示意图,其中每个红色方框包含两个SNP之间LD程度的成对测量值(例如,箭头指向黑色轮廓的方框,内含SNP2和SNP7的D′值)。
The higher the degree of LD, the darker the color in the box, with maximum D′ values of 1. 0 occurring when there is complete LD.LD程度越高,方框颜色越深,当存在完全LD时,D′最大值达到1.0。
Two LD blocks are detectable, the first containing SNPs 1 through 9, and the second SNPs 11 through 14.可检测到两个LD区块,第一个包含SNP1至9,第二个包含SNP11至14。
Between blocks, the 14-kb region containing SNP 10 shows no LD with neighboring SNPs 9 or 11 or with any of the other SNP loci.在区块之间,包含SNP10的14kb区域与邻近的SNP9或SNP11或任何其他SNP位点均未显示LD。
(C) A graph of the ratio of map distance to physical distance (c M/Mb), showing that a recombination hot spot is present in the region between SNP 10 and cluster 2, with values of recombination that are 50- to 60-fold above the average of ~1. 15 c M/Mb for the genome. )(C) 图谱距离与物理距离之比(cM/Mb)的图表,显示在SNP10与簇2之间的区域存在一个重组热点,其重组值比基因组平均约1.15 cM/Mb高50至60倍。
10/46
different alleles, whether there is a correlation between certain alleles and various aspects of the phenotype (genotype…
Ch11 — Segment 10
different alleles, whether there is a correlation between certain alleles and various aspects of the phenotype (genotype-phenotype correlation), and the frequency of disease-causing or predisposing variants in various populations.不同的等位基因,某些等位基因与表型各个方面之间是否存在相关性(基因型-表型相关性),以及在不同人群中致病性或易感性变异的频率。
Those with the same or similar disorders can also be examined to determine locus heterogeneity.患有相同或相似疾病的个体也可以进行检查,以确定位点异质性。
Once the gene and variants in that gene are identified in affected individuals, highly specific methods of diagnosis— including prenatal diagnosis and carrier screening (see Chapter 18)—can be offered to patients and their families.一旦在受累个体中鉴定出基因及该基因内的变异,就可以向患者及其家属提供高度特异的诊断方法,包括产前诊断和携带者筛查(见第18章)。
The variants associated with disease can then be modeled in other organisms, which allows us to use powerful genetic, biochemical, and physiologic tools to better understand the disease pathogenesis.与疾病相关的变异随后可以在其他生物体中建模,这使我们能够利用强大的遗传学、生物化学和生理学工具来更好地理解疾病的发病机制。
Finally, armed with an understanding of gene function and how disease-alleles affect that function, we can begin to develop specific therapies to prevent or ameliorate the disorder (see Chapter 14).最后,在掌握了基因功能以及疾病等位基因如何影响该功能的知识后,我们可以开始开发特异性疗法来预防或改善该疾病(见第14章)。
Indeed, much of the material in the next few chapters about the etiology, pathogenesis, mechanism, and treatment of various diseases begins with identification of the genes involved.事实上,接下来几章中关于各种疾病的病因学、发病机制、机制和治疗的许多内容,都始于对相关基因的鉴定。
Here, we examine the major approaches used to discover these genes, as outlined at the beginning of this chapter.在此,我们探讨用于发现这些基因的主要方法,正如本章开头所概述的那样。
MAPPING HUMAN DISEASE GENES BY LINKAGE ANALYSIS Determining Whether Two Loci Are Linked Linkage analysis is a method of mapping genes that uses studies of recombination in families to determine whether two genes show linkage when passed from one generation to the next.通过连锁分析定位人类疾病基因 确定两个位点是否连锁 连锁分析是一种基因定位方法,它利用家系中重组的研究来确定两个基因在代际传递时是否表现出连锁。
We use information from the known or suspected mendelian inheritance pattern (dominant, recessive, X-linked) to determine which family members have inherited a recombinant or a nonrecombinant chromosome.我们利用已知或疑似的孟德尔遗传模式(显性、隐性、X连锁)的信息,来确定哪些家庭成员遗传了重组染色体或非重组染色体。
To decide whether two loci are linked and, if so, how close or far apart they are, we rely on two pieces of information.为了判断两个位点是否连锁,以及如果连锁的话,它们之间的距离远近,我们依赖于两方面的信息。
First, using the family data in hand, we estimate θ, the recombination frequency between the two loci.首先,利用手头的家系数据,我们估算θ,即两个位点之间的重组频率。
Next, we ascertain whether θ is statistically significantly different from 0. 5, which is the fraction expected for unlinked loci.接下来,我们确定θ是否在统计学上显著不同于0.5,后者是不连锁位点的预期比例。
Estimating θ and, at the same time, determining the statistical significance of any deviation of θ from 0. 5, relies on a statistical tool called the likelihood ratio (as discussed later in the chapter).估算θ,同时确定θ与0.5的任何偏差的统计学显著性,依赖于一个称为似然比的统计工具(如本章后面所述)。
Linkage analysis begins with a set of actual family data with N individuals. .连锁分析从一组包含N个个体的实际家系数据开始。
The number of chromosomes that do not show recombination is therefore N − r.因此,不显示重组的染色体数目为N − r。
With each meiosis, the recombination fraction, θ, is the unknown probability that a recombination will occur between the two loci; the probability that no recombination occurs is, therefore, 1 − θ.在每次减数分裂中,重组分数θ是两个位点之间发生重组的未知概率;因此,不发生重组的概率为1 − θ。
Because each meiosis is an independent event, one multiplies the probability of a recombination, θ, or of no recombination, (1 − θ), for each chromosome.由于每次减数分裂是独立事件,因此将每条染色体重组概率θ或不重组概率(1 − θ)相乘。
The formula for the likelihood (probability) of observing this number of recombinant and nonrecombinant chromosomes when θ is unknown is {N!/r!(N − r)!}θr (1 − θ)(N−r).当θ未知时,观察到这一数量的重组和非重组染色体的似然(概率)公式为{N!/r!(N − r)!}θr (1 − θ)(N−r)。
(The factorial term, N!/r!(N − r)!, accounts for all the possible birth orders in which the recombinant and nonrecombinant children can appear in the pedigree.) Calculate a second likelihood based on the null hypothesis that the two loci are unlinked, i. e., that θ = 0. 50.(阶乘项N!/r!(N − r)!考虑了重组和非重组儿童在系谱中可能出现的所有出生顺序。)基于两个位点不连锁(即θ = 0.50)的原假设,计算第二个似然值。
The ratio of the likelihood of the family data supporting linkage with unknown θ to the likelihood that the loci are unlinked is the odds in favor of linkage and is given by: Likelihood of the data if loci are linked at distance Likelihood of t he data if loci are unlinked( 0. 5) {N!/r!(N r)!} (1) { r N r () N!/r!(N r)!} r N r   12 12 () Fortunately, the factorial terms are always the same in the numerator and denominator of the likelihood ratio, and therefore they cancel each other out and can be ignored.支持连锁(θ未知)的家系数据似然值与位点不连锁(θ=0.5)的似然值之比即为支持连锁的优势比,其公式为:似然比值 = [N!/(r!(N-r)!)] θ^r (1-θ)^(N-r) / [N!/(r!(N-r)!)] (0.5)^N。幸运的是,阶乘项在似然比的分子和分母中总是相同的,因此它们相互抵消,可以忽略不计。
If θ = 0. 5, the numerator and denominator are the same and the odds equal 1.如果θ = 0.5,分子和分母相同,优势比等于1。
Statistical theory tells us that when the value of the likelihood ratio for all values of θ between 0 and 0. 5 are calculated, the value of θ with the greatest likelihood ratio is, in fact, the best estimate of the recombination fraction and is referred to as θmax.统计学理论告诉我们,当计算θ在0到0.5之间所有值的似然比值时,具有最大似然比的θ值实际上是重组分数的最佳估计值,被称为θmax。
By convention, the computed likelihood ratio for different values of θ is usually expressed as the log 10 and is called the LOD score (Z), where LOD stands for logarithm of the odds.按照惯例,针对不同θ值计算出的似然比通常以log10表示,称为LOD评分(Z),其中LOD代表优势比的对数。
The use of logarithms allows likelihood ratios calculated from different families to be combined by simple addition instead of having to multiply them together.使用对数允许将来自不同家系的似然比通过简单相加进行合并,而不必相乘。
How is LOD score analysis actually carried out in families with mendelian disorders?在患有孟德尔遗传病的家系中,LOD评分分析实际上是如何进行的?
(See the accompanying box.) Return to the family shown in that the disease allele D is in coupling with allele A at locus 1 and allele B at locus 2.(见随附框。)回到所示的家系,其中疾病等位基因D与位点1的等位基因A和位点2的等位基因B呈相引。
Given this phase, one can see that there has been recombination between RP and locus 2 in only one of her eight children: her daughter II-3.鉴于这种相,可以看到在她的八个孩子中,只有她的女儿II-3在RP和位点2之间发生了重组。
The alleles at the disease locus, however, show no tendency to follow the然而,疾病位点的等位基因没有表现出遵循的趋势。
11/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 213 alleles at locus 1 or alleles at any of the other hundreds of marker…
Ch11 — Segment 11
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 213 alleles at locus 1 or alleles at any of the other hundreds of marker loci tested on the other autosomes.识别人类疾病的遗传基础 213 个等位基因位于位点1,或位于其他常染色体上测试的数百个标记位点中的任何一个。
Thus, although the RP locus involved in this family could, in principle, have mapped anywhere in the human genome the linkage data suggest that the responsible RP locus lies in the region of chromosome 7 near marker locus 2.因此,尽管原则上该家系中涉及的RP位点可能定位在人类基因组的任何位置,但连锁数据表明,责任RP位点位于7号染色体上标记位点2附近的区域。
To provide a quantitative assessment of this suspicion, suppose we let θ be the “true” recombination fraction between RP and locus 2—the fraction we would see if we had unlimited offspring to test.为了对这一怀疑进行定量评估,假设θ是RP与位点2之间的“真实”重组分数——即如果我们有无限多的后代可供测试时会观察到的分数。
The likelihood ratio for this family is ()(1) ( ( 1 7 1 7 12 12)) and reaches a maximum LOD score of Zmax = 1. 1 at θmax = 0. 125.该家系的似然比为 ()(1) ( (1 7 1 7 12 12)),并在θmax = 0.125 时达到最大LOD得分 Zmax = 1.1。
The value of θ that maximizes the likelihood ratio, θmax, may be the best estimate for θ given the data, but how good an estimate is it?使似然比最大化的θ值,即θmax,可能是给定数据下对θ的最佳估计,但这个估计有多好?
This is reflected in the magnitude of the LOD score.这反映在LOD得分的大小上。
By convention, a LOD score of +3 or greater (equivalent to greater than 1000: 1 odds in favor of linkage) is considered firm evidence that two loci are linked; that is, that θmax is statistically significantly different from 0. 5.按照惯例,LOD得分为+3或更高(相当于支持连锁的几率大于1000:1)被认为是两个位点连锁的有力证据;即θmax在统计学上显著不同于0.5。
In our RP example, 78 of the offspring are nonrecombinant and 18 are recombinant.在我们的RP例子中,78个子代为非重组型,18个子代为重组型。
The θmax = 0. 125, but the LOD score is only 1. 1: enough to raise a suspicion of linkage but insufficient to prove linkage because Zmax falls far short of 3.θmax = 0.125,但LOD得分仅为1.1:足以引起对连锁的怀疑,但不足以证明连锁,因为Zmax远低于3。
Combining LOD Score Information Across Families Just as each meiosis in a family that produces a nonrecombinant or recombinant offspring is an independent event, so too are the meioses that occur in different families.跨家系合并LOD得分信息 正如一个家系中产生非重组或重组子代的每一次减数分裂是一个独立事件,不同家系中发生的减数分裂也是独立事件。
We can, therefore, combine the likelihoods in the numerators and denominators of each family’s likelihood odds ratio.因此,我们可以合并每个家系似然比中分子和分母的似然值。
Suppose two additional families with RP were studied; one showed no recombination between locus 2 and RP in four children and the other showed no recombination in five children.假设又研究了两个患有RP的家系;一个家系在四个孩子中显示位点2与RP之间无重组,另一个家系在五个孩子中显示无重组。
The individual LOD scores can be generated for each family and added together ( Because the maximum LOD score, Zmax, exceeds 3 at θmax = ≈0. 06, the RP gene in this group of families is linked to locus 2 at a recombination distance of ≈0. 06.可以为每个家系生成单独的LOD得分并相加(由于在θmax ≈ 0.06时最大LOD得分Zmax超过3,因此这组家系中的RP基因与位点2连锁,重组距离约为0.06。
Because the genomic location of marker locus 2 is known to be at 7p14, the RP in this family can be mapped to the 7p14 region.由于已知标记位点2的基因组位置在7p14,该家系中的RP可定位到7p14区域。
The RP9 gene is a likely candidate – one of the identified loci for a form of autosomal dominant RP.RP9基因是一个可能的候选基因——它是常染色体显性RP的一种已知位点之一。
If, however, some of the families being used for the study were to have RP due to pathogenic variants at a different locus, the LOD scores between families would diverge, with some showing a trend to being positive at small values of θ and others showing strongly negative LOD scores at these values.但是,如果用于研究的某些家系中RP是由不同位点的致病变异引起的,那么这些家系之间的LOD得分会发散,有些在θ值较小时显示正向趋势,而有些则在相同θ值下显示强烈负向的LOD得分。
Thus, in linkage analysis involving more than one family, unsuspected locus heterogeneity can obscure what may be real evidence for linkage in a subset of families.因此,在涉及多个家系的连锁分析中,未被识别的位点异质性可能会掩盖在某亚组家系中可能存在的真正连锁证据。
Phase-Known and Phase-Unknown Pedigrees In the RP example just discussed, we assumed that we knew the phase of marker alleles on chromosome 7 in the affected mother in that family.相位已知与相位未知的家系图 在刚刚讨论的RP例子中,我们假设知道该家系中患病母亲在7号染色体上标记等位基因的相位。
Let us now look at the implications of knowing phase in more detail.现在让我们更详细地探讨知晓相位的意义。
Consider the three-generation family with autosomal dominant neurofibromatosis, type 1 (NF1) (Case 34) in and a marker locus (A/a), but (as shown in we have no genotype information on her parents.考虑一个三代家系,患有常染色体显性神经纤维瘤病1型(NF1)(病例34),以及一个标记位点(A/a),但(如图所示)我们对其父母没有基因型信息。
The two affected children received the A alleles along with the D disease allele, and the one unaffected child received the a allele along with the normal d allele.两个患病孩子接受了A等位基因和D疾病等位基因,一个未患病孩子接受了a等位基因和正常d等位基因。
Without knowing the phase of these alleles in the mother, either all three offspring are recombinants or all three are nonrecombinants.在不知道母亲这些等位基因相位的情况下,要么三个子代全是重组型,要么三个全是非重组型。
Because both possibilities are equally likely in the absence of any other information, we consider the phase on her two chromosomes to be D-a and d-A half of the time and 00 0. 01 0. 05 0. 06 0. 07 0. 10 0. 125 0. 20 0. 30 0. 40 Family 1 — 0. 38 0. 95 1. 00 1. 03 1. 09 1. 1 1. 03 0. 80 0. 46 Family 2 1. 2 1. 19 1. 11 1. 10 1. 08 1. 02 0. 97 0. 82 0. 58 0. 32 Family 3 1. 5 1. 48 1. 39 1. 37 1. 35 1. 28 1. 22 1. 02 0. 73 0. 39 Total — 3. 05 3. 45 3. 47 3. 46 3. 39 3. 29 2. 87 2. 11 1. 17 Individual Zmax for each family is shown in bold.由于在缺乏其他信息时这两种可能性同等可能,我们认为她的两条染色体上的相位一半时间是D-a和d-A,以及00 0.01 0.05 0.06 0.07 0.10 0.125 0.20 0.30 0.40 家系1 — 0.38 0.95 1.00 1.03 1.09 1.1 1.03 0.80 0.46 家系2 1.2 1.19 1.11 1.10 1.08 1.02 0.97 0.82 0.58 0.32 家系3 1.5 1.48 1.39 1.37 1.35 1.28 1.22 1.02 0.73 0.39 总计 — 3.05 3.45 3.47 3.46 3.39 3.29 2.87 2.11 1.17 每个家系的个体Zmax以粗体显示。
The overall Zmax = 3. 47 at θmax = 0. 06..在θmax = 0.06时,整体Zmax = 3.47。
LINKAGE ANALYSIS OF MENDELIAN DISEASES Linkage analysis is used when there is a particular mode of inheritance (autosomal dominant, autosomal recessive, or X-linked) that explains the inheritance pattern.孟德尔疾病的连锁分析 当存在特定遗传模式(常染色体显性、常染色体隐性或X连锁)能够解释遗传模式时,使用连锁分析。
LOD score analysis allows mapping of genes with variants for phenotypes that follow mendelian inheritance.LOD得分分析允许对具有遵循孟德尔遗传表型的变异基因进行定位。
The LOD score gives both: a best estimate of the recombination frequency, θmax, between a marker locus and the disease locus; and an assessment of the strength of evidence for linkage at that value of θmax.LOD得分同时提供:标记位点与疾病位点之间重组频率θmax的最佳估计;以及在该θmax值下支持连锁的证据强度评估。
For the LOD score, Z, values above 3 are considered strong evidence.对于LOD得分Z,高于3的值被认为是强有力的证据。
Linkage at a particular θmax of a given gene locus to a marker with known physical location implies proximity between the two.某个基因位点在特定θmax下与已知物理位置的标记连锁,意味着两者之间距离较近。
The smaller the θmax the closer the locusof-interest is to the linked marker locus.θmax越小,目标位点与连锁标记位点越接近。
12/46
D-A and d-a the other half (which assumes the alleles in these haplotypes are in linkage equilibrium).
Ch11 — Segment 12
D-A and d-a the other half (which assumes the alleles in these haplotypes are in linkage equilibrium).D-A 和 d-a 是另一半(假设这些单倍型中的等位基因处于连锁平衡)。
To calculate the overall likelihood of this pedigree, we then add the likelihood calculated assuming one phase in the mother to that assuming the other phase.为了计算该系谱的整体似然度,我们将假设母亲中一种相位的似然度与假设另一种相位的似然度相加。
The overall likelihood = 12 0 12 () ()() 1 1 3 3 0 ; the likelihood ratio for this pedigree, then, is: 12 12 18 () ()() () 1 1 3 0 3 0 giving a maximum LOD score of Zmax = 0. 602 at θmax = 0.总体似然度 = 12 0 12 () ()() 1 1 3 3 0;那么该系谱的似然比是:12 12 18 () ()() () 1 1 3 0 3 0,在 θmax = 0 时得到最大 LOD 分数 Zmax = 0.602。
If, however, additional genotype information in the maternal grandfather, I-1, becomes available (as in , the phase can now be determined to be D-A (i. e., the NF1 allele D was in coupling with the A in individual II-2).然而,如果可获得外祖父 I-1 的额外基因型信息(如图所示),那么现在可以确定相位为 D-A(即,NF1 等位基因 D 与个体 II-2 中的 A 呈相引)。
In light of this new information, the three children can now be scored definitively as nonrecombinants, and we no longer have to consider the possibility of the opposite phase.根据这一新信息,三个孩子现在可以被明确判定为非重组体,我们不再需要考虑相反相位的可能性。
The numerator of the likelihood ratio now becomes (1 − θ)3(θ0) and the maximum LOD score, Zmax = 0. 903 at θmax = 0.似然比的分子现在变为 (1 − θ)^3 (θ^0),最大 LOD 分数在 θmax = 0 时为 Zmax = 0.903。
Thus knowing the phase increases the power of the data available to test for linkage.因此,了解相位增加了可用于检验连锁的数据的效力。
Gene Finding in a Common Mendelian Disorder by Linkage Mapping: Cystic Fibrosis The application of linkage mapping to medical genetics using the approaches outlined has met with many successes.通过连锁作图发现常见孟德尔遗传病基因:囊性纤维化,使用所概述的方法将连锁作图应用于医学遗传学已取得许多成功。
Here we describe one such historical example, using linkage analysis and LD to narrow down the location of the gene responsible for the common autosomal recessive disease, cystic fibrosis (CF) (Case 12).这里我们描述一个这样的历史案例,使用连锁分析和 LD 来缩小常见常染色体隐性遗传病——囊性纤维化(CF)(病例 12)的致病基因的位置。
Because of its relatively high frequency, particularly in populations of European descent, and (at the time) the little understanding of its pathogenesis, CF represented a prime candidate for identifying the gene responsible by first using linkage to find the gene’s location.由于其相对较高的频率,尤其是在欧洲裔人群中,以及(当时)对其发病机制了解甚少,CF 成为通过首先使用连锁寻找基因位置来鉴定致病基因的主要候选对象。
DNA samples from nearly 50 multiplex CF families were analyzed for linkage between CF and hundreds of DNA markers throughout Dd AA dd AA Dd Aa dd Aa Dd AA 1 2 1 2 1 2 3 I II III NR R NR R NR R Phase in II-2 If D-A/d-a: If D-a/d-A: Dd AA dd AA Dd Aa dd Aa Dd AA 1 2 dd aa Dd Aa 2 1 1 2 3 I II III NR NR NR Phase in II-2 D-A/d-a: A B Phase of the disease allele D and marker alleles A and a in individual II-2 is unknown.来自近 50 个多发性 CF 家庭的 DNA 样本被用于分析 CF 与遍布整个基因组的数百个 DNA 标记之间的连锁:Dd AA dd AA Dd Aa dd Aa Dd AA 1 2 1 2 1 2 3 I II III NR R NR R NR R II-2 中的相位 如果 D-A/d-a: 如果 D-a/d-A: Dd AA dd AA Dd Aa dd Aa Dd AA 1 2 dd aa Dd Aa 2 1 1 2 3 I II III NR NR NR II-2 中的相位 D-A/d-a: A B 个体 II-2 中疾病等位基因 D 和标记等位基因 A 与 a 的相位未知。
(B) Availability of genotype information for generation I allows a determination that the disease allele D and marker allele A are in coupling in individual II-2.(B)获得第一代的基因型信息可以确定个体 II-2 中疾病等位基因 D 和标记等位基因 A 呈相引。
NR, Nonrecombinant; R, recombinant. the genome.NR,非重组体;R,重组体。基因组。
Eventually, linkage of CF to markers on the long arm of chromosome 7 was identified.最终,确定了 CF 与 7 号染色体长臂上的标记之间的连锁。
Linkage to additional DNA markers in 7q31-q 32 narrowed the location of the CF gene to a ~500 kb region of chromosome 7.与 7q31-q32 中其他 DNA 标记的连锁将 CF 基因的位置缩小到 7 号染色体上约 500 kb 的区域。
At this point, however, an important feature of CF genetics emerged: although the closest linked markers were still some distance from the CF gene, it became clear that there was significant LD between the disease locus and a particular haplotype of nearby markers.然而,此时 CF 遗传学的一个重要特征出现了:尽管最近连锁的标记仍与 CF 基因有一定距离,但明显的是,疾病位点与附近标记的特定单倍型之间存在显著的 LD。
Regions with the greatest degree of LD were analyzed for gene sequences, leading to the isolation of the gene responsible in 1989.对具有最大 LD 程度的区域进行了基因序列分析,从而在 1989 年分离出了致病基因。
As described in detail in Chapter 13, this gene, which was named the CF transmembrane conductance regulator (CFTR), showed an interesting spectrum of variants.如第 13 章详细所述,这个被命名为 CF 跨膜传导调节因子(CFTR)的基因展现出有趣的变异谱。
A 3 bp deletion (then called ΔF508) that removed a phenylalanine at position 508 in the protein was found in ~70% of all variant CF alleles in northern European populations, but never among normal alleles at this locus.一个 3 bp 的缺失(当时称为 ΔF508)去除了蛋白质中第 508 位的苯丙氨酸,在北欧人群的所有变异 CF 等位基因中约占 70%,但在此位点的正常等位基因中从未发现。
Although subsequent studies have demonstrated many hundreds of variant CFTR alleles worldwide, it was the high frequency of the ΔF508 variant in the families used to map the CF locus and the LD between it and alleles at polymorphic marker loci nearby, that proved so helpful in the ultimate identification of the CFTR gene.尽管后续研究已在全球范围内证明了数百种 CFTR 变异等位基因,但正是用于定位 CF 位点的家庭中 ΔF508 变异的高频率以及其与附近多态性标记位点等位基因之间的 LD,对最终鉴定 CFTR 基因如此有帮助。
Mapping of the CF locus and cloning of the CFTR gene made possible a wide range of research advances and clinical applications, from basic pathophysiology to molecular diagnosis for genetic counseling, prenatal diagnosis, animal models, and finally effective treatments for the disorder (see Chapters 13 and 14).CF 位点的定位和 CFTR 基因的克隆使得广泛的研究进展和临床应用成为可能,从基础病理生理学到用于遗传咨询的分子诊断、产前诊断、动物模型,以及最终针对该疾病的有效治疗(见第 13 章和第 14 章)。
Mapping Human Disease Genes by Association Designing an Association Study An entirely different approach to identification of the genetic contribution to disease relies on finding particular alleles that are associated with the disease in a sample from the population.通过关联作图发现人类疾病基因:设计一项关联研究,一种完全不同的鉴定疾病遗传贡献的方法依赖于在来自人群的样本中发现与疾病相关的特定等位基因。
In contrast to linkage analysis, this approach does not depend upon there being a mendelian inheritance pattern and is, therefore, better与连锁分析不同,这种方法不依赖于存在孟德尔遗传模式,因此更好。
13/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 215 Alternatively, if the association study was designed as a cross-sect…
Ch11 — Segment 13
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 215 Alternatively, if the association study was designed as a cross-sectional or cohort study, the strength of an association can be measured by the relative risk (RR).识别人类疾病的遗传基础 215 或者,如果关联研究设计为横断面研究或队列研究,则关联的强度可以通过相对危险度(RR)来衡量。
The RR is the ratio of the proportion of those with the disease who carry a particular marker allele ([a/(a + b)]) to the proportion of those without the disease who carry that marker ([c/(c + d)]).RR是患病者中携带特定标记等位基因的比例([a/(a+b)])与未患病者中携带该标记的比例([c/(c+d)])之比。
RR a (a b c (c d) ) Again, an RR that differs from 1 means there is an association of disease with the genetic marker, whereas RR = 1 means there is no association.RR a (a b c (c d) ) 同样,偏离1的RR意味着疾病与遗传标记存在关联,而RR=1意味着无关联。
(The RR introduced here should not be confused with Relative Risk Ratio (λr), (i. e., risk ratio in relatives) discussed in Chapter 9. λr is the prevalence of a particular disease phenotype in an affected individual’s relatives versus that in the general population.) For diseases that are rare (i. e., a &lt;&lt; b and c &lt;&lt; d), a case-control design with calculation of the OR is best because any random sample of a population is unlikely to contain sufficient numbers of affected individuals to be suitable for a cross-sectional or cohort study design.(此处介绍的RR不应与第9章讨论的相对危险度比(λr)混淆,λr是患病个体的亲属中特定疾病表型的患病率与一般人群中的患病率之比。)对于罕见疾病(即a<<b且c<<d),采用计算OR的病例对照设计最佳,因为任何随机人群样本都不太可能含有足够数量的患病个体以适用于横断面或队列研究设计。
Note, however, that when a disease is rare, and calculating an OR in a case-control study is the only practical approach, OR is a good approximation for RR.然而,需要注意的是,当疾病罕见且病例对照研究中计算OR是唯一可行的方法时,OR是RR的良好近似值。
(Examine the formula for RR and convince yourself that, when a &lt;&lt; b and c &lt;&lt; d, (a + b) ≈ b and (c + d) ≈ d; thus, RR ≈ OR.) The information obtained in an association study comes in two parts.(检查RR的公式并自行验证,当a<<b且c<<d时,(a+b)≈b且(c+d)≈d,因此RR≈OR。)关联研究获得的信息分为两部分。
The first is the magnitude of the association itself.第一部分是关联本身的强度。
The further the RR or OR diverges from 1, the greater is the effect of the genetic variant on the association.RR或OR偏离1的程度越大,遗传变异对关联的影响就越大。
However, an OR or RR for an association is a statistical measure and requires a test of statistical significance.然而,关联的OR或RR是一种统计度量,需要进行统计学显著性检验。
The significance of any association can be assessed with a chi-square test, asking whether if the frequencies of the marker classifications (a, b, c, and d in the two-by-two table) differ significantly from what would be expected if there were no association (i. e., if the OR or RR were equal to 1. 0).任何关联的显著性可通过卡方检验进行评估,检验标记分类(四格表中的a、b、c、d)的频率是否与无关联(即OR或RR等于1.0)时的预期频率存在显著差异。
A common way of expressing whether there is statistical significance to an estimate of OR or RR is to provide a 95% (or 99%) confidence interval.表示OR或RR估计值是否具有统计学显著性的常用方法是提供95%(或99%)置信区间。
The confidence interval is the range within which one would expect the OR or RR to fall 95% (or 99%) of the time by chance alone in a sample taken from the population.置信区间是指从人群中抽取样本时,仅凭随机因素,OR或RR有95%(或99%)的可能性落在此范围内。
If a confidence interval excludes the value 1. 0, then the OR or RR deviates significantly from what would be expected if there were no association with the marker locus being tested; thus, the null hypothesis (of no association) can be rejected at the corresponding significance level.如果置信区间排除1.0,则OR或RR显著偏离与待测标记位点无关联时的预期值;因此,可以在相应的显著性水平上拒绝零假设(无关联)。
(Later in this chapter we will explain why a level of 0. 05 or 0. 01 is inadequate for assessing statistical significance when multiple genomic marker loci are tested simultaneously for association.) suited for discovering the genetic contributions to disorders with complex inheritance (see Chapter 9).(本章后面将解释,当同时检测多个基因组标记位点的关联时,为何0.05或0.01的显著性水平不足以评估统计学显著性。)适用于发现复杂遗传疾病中的遗传贡献(见第9章)。
It is important to note that here, alleles can harbor either SNVs or CNVs associated with disease.重要的是要注意,此处的等位基因可能携带与疾病相关的SNV或CNV。
The increased or decreased frequency of a particular allele in affected individuals, compared to that in unaffected individuals, is known as a disease association.与未患病个体相比,患病个体中特定等位基因频率的增加或减少称为疾病关联。
There are two designs commonly used for association studies: Case-control studies.关联研究常用的两种设计:病例对照研究。
Individuals with the disease (cases) are selected in a population, along with a matching group without disease (controls).在人群中选择患病的个体(病例),以及匹配的无疾病组(对照)。
Genotypes of individuals in the two groups are determined and used to populate a two-by-two table (see table below).确定两组个体的基因型,并用于填充四格表(见下表)。
Cross-sectional or cohort studies.横断面或队列研究。
A random sample of the population is chosen and analyzed with respect to the disease phenotype in question and individually genotyped for selected markers.选择人群的随机样本,分析其相关疾病表型,并对选定标记进行个体基因分型。
A cross-sectional study involves a single time point, whereas a cohort study takes place over time.横断面研究涉及单个时间点,而队列研究则随时间推移进行。
The numbers of individuals with and without disease and with and without an allele (or genotype or haplotype) of interest are used to fill out the cells of a two-by-two table.患病与未患病个体以及携带与不携带目标等位基因(或基因型、单倍型)的个体数量用于填写四格表的单元格。
Odds Ratios and Relative Risks The two different types of association studies report the strength of association, using either the odds ratio (OR) or relative risk (RR).比值比与相对危险度 两种不同类型的关联研究使用比值比(OR)或相对危险度(RR)报告关联强度。
In a case-control study, the frequency of a particular marker (e. g., a human leukocyte antigen [HLA] haplotype or a particular SNV allele or haplotype) is compared between the selected affected and unaffected individuals.在病例对照研究中,比较所选患病与未患病个体中特定标记(如人类白细胞抗原[HLA]单倍型或特定SNV等位基因或单倍型)的频率。
Association between disease and genotype is then calculated by an odds ratio (OR).然后通过比值比(OR)计算疾病与基因型之间的关联。
Cases Controls Totals With genetic marker* a b a+b Without genetic marker c d c+d Totals a+c b+d *A genetic marker can be an allele, a genotype, or a haplotype.病例 对照 合计 有遗传标记* a b a+b 无遗传标记 c d c+d 合计 a+c b+d *遗传标记可以是等位基因、基因型或单倍型。
Using the two-by-two table, the odds of a marker carrier developing the disease is the ratio (a/b) of the number of marker carriers who develop the disease (a) to the number of marker carriers who do not develop the disease (b).使用四格表,标记携带者患病的比值是患病标记携带者数(a)与未患病标记携带者数(b)之比(a/b)。
Similarly, the odds of a noncarrier developing the disease is the ratio (c/d) of noncarriers who develop the disease (c) to the number of noncarriers who do not develop the disease (d).类似地,非携带者患病的比值是患病非携带者数(c)与未患病非携带者数(d)之比(c/d)。
The disease OR is then the ratio of these odds.疾病OR即为这些比值的比值。
OR a b c d ad bc = = An OR that differs from 1 means there is an association of the disease with the genetic marker, whereas OR = 1 means there is no association.OR a b c d ad bc = = 不等于1的OR意味着疾病与遗传标记存在关联,而OR=1意味着无关联。
14/46
To illustrate these approaches we first consider a case-control study of cerebral vein thrombosis (CVT), which we introd…
Ch11 — Segment 14
To illustrate these approaches we first consider a case-control study of cerebral vein thrombosis (CVT), which we introduced in Chapter 9.为了说明这些方法,我们首先考虑一个脑静脉血栓形成(CVT)的病例对照研究,该研究已在第9章中介绍。
In this study, suppose a group of 120 individuals with CVT and 120 matched controls were genotyped for the 20210 G&gt;A allele in the prothrombin gene (see Chapter 9).在该研究中,假设对120名CVT患者和120名匹配对照者进行了凝血酶原基因20210 G>A等位基因的基因分型(参见第9章)。
Cases With CVT Controls Without CVT Totals 20210 G&gt;A allele present 23 4 27 20210 G&gt;A allele absent 97 116 213 Total 120 120 240 CVT, Cerebral vein thrombosis.病例(有CVT) 对照(无CVT) 总计 携带20210 G>A等位基因 23 4 27 不携带20210 G>A等位基因 97 116 213 总计 120 120 240 CVT,脑静脉血栓形成。
Because this is a case-control study, we will calculate an odds ratio: OR = (23/4)/(97/116) = ≈6. 9 with 95% confidence limits of 2. 3 to 20. 6.由于这是一项病例对照研究,我们将计算比值比:OR = (23/4)/(97/116) ≈ 6.9,其95%置信区间为2.3至20.6。
The effect size of 6. 9 is substantial, and 95% confidence limits exclude 1. 0, thereby demonstrating a strong and statistically significant association between the 20210 G&gt;A allele and CVT.效应量6.9显著,且95%置信区间排除了1.0,从而证明了20210 G>A等位基因与CVT之间存在强烈且具有统计学显著性的关联。
Stated simply, individuals carrying the prothrombin 20210 G&gt;A allele have nearly seven times greater odds of having the disease than do those who do not carry this allele.简而言之,携带凝血酶原20210 G>A等位基因的个体患该疾病的几率比不携带该等位基因的个体高出近七倍。
To illustrate a longitudinal cohort study, calculating RR instead of OR—consider statin-induced myopathy, a rare but well-recognized adverse drug reaction that can develop in some individuals during statin therapy to lower cholesterol.为说明纵向队列研究中计算RR而非OR,我们考虑他汀类药物诱导的肌病——这是一种罕见但已充分认识的不良药物反应,可在部分个体接受他汀类药物降胆固醇治疗期间发生。
In one study, subjects enrolled in a cardiac protection study were randomized to receive 40 mg of the statin drug, simvastatin, or placebo.在一项研究中,参与心脏保护研究的受试者被随机分配接受40毫克他汀类药物辛伐他汀或安慰剂。
Over 16,600 participants exposed to the statin were genotyped for a variant (Val 174Ala) in the SLCO1B1 gene— which encodes a hepatic drug transporter, and were watched for development of the adverse drug response.超过16600名暴露于他汀类药物的受试者就其SLCO1B1基因(编码肝脏药物转运体)中的一个变异(Val174Ala)进行了基因分型,并观察其是否出现该不良药物反应。
Out of the entire genotyped group exposed to the statin, 21 developed myopathy.在暴露于他汀类药物的整个基因分型组中,有21人发生了肌病。
Examination of their genotypes showed that the RR for developing myopathy associated with the presence of the Val 174Ala allele was ~2. 6, with 95% confidence limits of 1. 3 to 5. 1.对其基因型的检查显示,携带Val174Ala等位基因者发生肌病的RR约为2.6,其95%置信区间为1.3至5.1。
Thus, there is a statistically significant association between the Val 174Ala allele and statin-induced myopathy.因此,Val174Ala等位基因与他汀类药物诱导的肌病之间存在统计学显著关联。
Those carrying this allele are at moderately increased risk for developing this adverse drug reaction, relative to those who do not carry this allele.与不携带该等位基因者相比,携带该等位基因者发生此不良药物反应的风险中度增加。
One common misconception concerning an association study is that the more significant the p-value, the stronger is the association.关于关联研究的一个常见误解是,p值越显著,关联就越强。
In fact, a significant p-value for an association does not provide information concerning the magnitude of the effect of an associated allele on disease susceptibility.事实上,关联的显著性p值并不能提供关于关联等位基因对疾病易感性影响程度的信息。
Significance is a statistical measure that describes how likely it is that the population sample used for the association study could have yielded an observed OR or RR that differs from 1. 0, simply by chance.显著性是一种统计度量,描述的是用于关联研究的人群样本仅仅由于随机性而得出与1.0不同的观察OR或RR的可能性有多大。
In contrast, the actual magnitude of the OR or RR—how far it diverges from 1. 0—is a measure of the impact a particular variant (or genotype or haplotype) on increasing or decreasing disease likelihood.相比之下,OR或RR的实际幅度——即其偏离1.0的程度——是衡量某个特定变异(或基因型或单倍型)对增加或减少疾病可能性影响程度的指标。
Genome-Wide Association Studies The Haplotype Map (Hap Map) Association studies for human disease genes were once limited to particular sets of variants in restricted sets of genes.全基因组关联研究 单倍型图谱(Hap Map) 人类疾病基因的关联研究曾局限于特定基因集合中的特定变异组。
These were chosen, either for convenience or because they were thought to be involved in a pathophysiologic pathway relevant to a disease, making them logical candidate genes for the disease under investigation.这些变异的选择,或是出于便利,或是因其被认为参与与疾病相关的病理生理通路,从而成为所研究疾病的合理候选基因。
Many such association studies were undertaken before the Human Genome Project era, using HLA or blood group loci, for example, because these were highly polymorphic and easily genotyped in case-control studies.在人类基因组计划时代之前,许多此类关联研究已开展,例如使用HLA或血型位点,因为这些位点高度多态性且易于在病例对照研究中进行基因分型。
Ideally, however, one would like to test systematically for an association between any disease of interest and every one of the tens of millions of rare and common alleles in the genome, in an unbiased fashion without preconception of what genes and genetic variants might be contributing to the disease.然而,理想情况下,人们希望以无偏倚的方式,不对哪些基因和遗传变异可能促成疾病做预先设想,系统性地检验任何感兴趣的疾病与基因组中数以千万计的稀有和常见等位基因之间的关联。
Association analyses on a genome scale are referred to as genome-wide association studies (GWAS).全基因组范围内的关联分析被称为全基因组关联研究(GWAS)。
Such an undertaking for all known variants is impractical for many reasons.由于诸多原因,对所有已知变异进行这种尝试是不切实际的。
It can, however, be approximated by genotyping cases and controls for a mere 300,000 to 1 million individual variants located throughout the genome, to search for association with the disease or trait in question.然而,可以通过对病例和对照仅对分布于全基因组的30万至100万个个体变异进行基因分型来近似实现,以搜索与所研究疾病或性状的关联。
The success of this approach depends on exploiting LD: as long as a variant responsible for altering disease susceptibility is in LD with one or more of the genotyped variants within an LD block, a positive association should be detectable between that disease and the alleles in the LD block.该方法的成功依赖于利用LD:只要改变疾病易感性的变异与LD区块内的一个或多个已分型变异处于LD状态,就应该能检测到该疾病与LD区块内等位基因之间的阳性关联。
Developing such a set of markers led to the launch of the Haplotype Mapping (Hap Map) Project, one of the biggest human genomics efforts to follow completion of the Human Genome Project.开发这样一套标记物促成了单倍型图谱(Hap Map)项目的启动,这是人类基因组计划完成后最大的基因组学努力之一。
The Hap Map Project began in four geographically distinct groups—a primarily European population, a West African population, a Han Chinese population, and a population from Japan—and included collecting and characterizing millions of SNP loci and developing methods to genotype them rapidly and inexpensively.Hap Map项目始于四个地理上不同的群体——主要欧洲人群、西非人群、中国汉族人群和日本人群——包括收集和表征数百万个SNP位点,并开发快速且低成本对其进行基因分型的方法。
Hap Map version 3 expanded coverage and diversity to include genotyping of 1. 6 million common SNP loci as well as common CNVs in more than 1000 reference individuals from 11 global populations.Hap Map第3版扩大了覆盖范围和多样性,包括对来自全球11个群体的1000多名参考个体的160万个常见SNP位点以及常见CNV进行基因分型。
Subsequently, whole genome sequencing has been applied to many populations in what is referred to as the 1000 Genomes Project, resulting in a massive expansion in the database of DNA variants available for GWAS among different populations around the globe.随后,全基因组测序已应用于被称为“千基因组计划”的众多人群,导致全球不同人群可用于GWAS的DNA变异数据库大幅扩展。
Gene Mapping by Genome-Wide Association Studies The purpose of the Hap Map was not just to gather basic information about the distribution of LD across通过全基因组关联研究进行基因定位 Hap Map的目的不仅是收集关于LD分布的基础信息。
15/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 217 the human genome.
Ch11 — Segment 15
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 217 the human genome.识别人类疾病的遗传基础 217 人类基因组。
Its primary purpose was to provide a powerful new tool for finding the genetic variants that contribute to human disease and other traits, by making possible an approximation to an idealized, fullscale, genome-wide association.其主要目的是通过实现近似于理想化的全基因组关联分析,为发现导致人类疾病及其他性状的遗传变异提供强有力的新工具。
The driving principle behind this approach is straightforward: detecting an association with alleles within an LD block pinpoints the genomic region within the block as likely to contain the disease-associated allele.该方法背后的驱动原理直截了当:检测与LD区块内等位基因的关联,可精确锁定该区块内可能包含疾病相关等位基因的基因组区域。
Consequently, although the approach does not typically pinpoint the actual variant responsible functionally for the disease association, this region will be the place to focus additional studies to find the allelic variant(s) directly involved in the disease process.因此,尽管该方法通常无法精确定位功能上导致疾病关联的实际变异,但该区域将成为进一步研究以寻找直接参与疾病过程的等位基因变异的重点区域。
Historically, detailed analysis of conditions associated with high-density variants in the class I and class II HLA regions has exemplified this approach (see 3).历史上,对与I类和II类HLA区域高密度变异相关的疾病进行详细分析,已例证了该方法(参见3)。
However, with the tens of ­millions of variants now available in different populations, this approach can be broadened to examine the genetic basis of virtually any complex disease or trait.然而,随着不同人群中现已可获得数千万个变异,该方法可拓展用于研究几乎任何复杂疾病或性状的遗传基础。
Indeed, to date, thousands of GWAS have uncovered an enormous number of naturally occurring variants associated with a variety of genetically common and complex multifactorial diseases.事实上,迄今为止,数千项全基因组关联研究已发现了大量自然存在的变异,这些变异与多种遗传上常见且复杂的多因素疾病相关。
These range from diabetes and inflammatory bowel disease to rheumatoid arthritis and neuropsychiatric disease, and include traits such as stature and pigmentation.这些疾病包括糖尿病、炎症性肠病、类风湿性关节炎和神经精神疾病,并涵盖身高和色素沉着等性状。
Research to uncover the underlying biologic basis for these associations will be ongoing for years to come.揭示这些关联背后生物学基础的研究将在未来多年持续进行。
Finding the Genes Contributing to a Complex Disease by Genome-Wide Association: Age-Related Macular Degeneration Genome-wide association has proven effective in identifying hundreds of genes and alleles associated with genetically complex disorders.通过全基因组关联寻找导致复杂疾病的基因:年龄相关性黄斑变性。全基因组关联已被证明能有效识别与遗传复杂疾病相关的数百个基因和等位基因。
The power of these approaches has increased enormously with the introduction of highly efficient and less expensive technologies for genome analysis.随着高效且廉价的基因组分析技术的引入,这些方法的能力大幅提升。
Here we describe an example of using GWAS to find multiple allelic variants in genes that increase susceptibility to age-related macular degeneration (AMD) (Case 3), a devastating disorder that robs older adults of their vision.这里我们描述了一个利用全基因组关联研究寻找基因中多个增加年龄相关性黄斑变性(AMD)易感性的等位基因变异的例子(病例3),这是一种剥夺老年人视力的破坏性疾病。
AMD is a progressive degenerative disease of the portion of the retina responsible for central vision.AMD是一种负责中心视力的视网膜区域进行性退行性疾病。
It causes blindness in 1. 75 million Americans older than 50 years.它在50岁以上的美国人中导致175万人失明。
The disease is characterized by the presence of drusen, which are clinically visible, discrete extracellular deposits of protein and lipids behind the retina in the region of the macula (Case 3).该疾病以玻璃膜疣为特征,玻璃膜疣是临床可见的、位于视网膜后方黄斑区域的离散性细胞外蛋白质和脂质沉积物(病例3)。
Although there is ample evidence for a genetic contribution to the disease, most individuals with AMD are not in families with a likely mendelian pattern of inheritance.尽管有充分证据表明遗传因素对该疾病有贡献,但大多数AMD患者并不属于具有典型孟德尔遗传模式的家族。
Environmental contributions are also important, as shown by the increased risk for AMD in cigarette smokers compared with nonsmokers.环境因素也很重要,吸烟者相比非吸烟者AMD风险增加即证明了这一点。
Initial case-control GWAS of AMD revealed association of two common SNP loci near the complement factor H (CFH) gene.最初的AMD病例-对照全基因组关联研究揭示了补体因子H(CFH)基因附近两个常见SNP位点的关联。
The most frequent at-risk haplotype containing these alleles was seen in 50% of cases versus only 29% of controls (OR = 2. 46; 95% confidence interval [CI], 1. 9–53. 11).包含这些等位基因的最常见风险单倍型在50%的病例中观察到,而对照组仅为29%(比值比=2.46;95%置信区间[CI],1.9–53.11)。
Homozygosity for this haplotype was found in 24. 2% of cases, compared to only 8. 3% of the controls (OR = 3. 51; 95% CI, 2. 13 − 5. 78).该单倍型的纯合性在24.2%的病例中发现,而对照组仅为8.3%(比值比=3.51;95% CI,2.13–5.78)。
A search through the SNPs within the LD block containing the AMD-associated haplotype revealed a nonsynonymous SNP in the CFH gene that substituted a histidine for tyrosine at position 402 of the CFH 3 HUMAN LEUKOCYTE ANTIGEN AND DISEASE ASSOCIATION Among more than 1000 genome-trait or genome-disease associations from around the genome, the region with the highest concentration of associations to different phenotypes is the human leukocyte antigen (HLA) region.对包含AMD相关单倍型的LD区块内的SNP进行搜索,发现CFH基因中存在一个非同义SNP,该SNP将CFH蛋白第402位的酪氨酸替换为组氨酸。
In addition to the association of specific alleles and haplotypes to type 1 diabetes discussed in Chapter 9, association of various HLA polymorphisms has been demonstrated for a wide range of conditions.除了第9章讨论的特定等位基因和单倍型与1型糖尿病的关联外,多种HLA多态性与广泛疾病的关联已被证实。
Most, but not all of these are autoimmune; that is, associated with an abnormal immune response apparently directed against one or more self-antigens.这些疾病中大多数(但非全部)是自身免疫性的;即与明显针对一种或多种自身抗原的异常免疫反应相关。
These associations are thought to be related to variation in the immune response resulting from polymorphism in immune response genes.这些关联被认为与免疫反应基因多态性导致的免疫反应变异有关。
The functional basis of most HLA-disease associations is unknown.大多数HLA-疾病关联的功能基础尚不清楚。
HLA molecules are integral to T-cell recognition of antigens.HLA分子是T细胞识别抗原所必需的。
Different HLA alleles are thought to result in structural variation in these cell surface molecules, leading to differences in capacity of the proteins to interact with antigen and the T-cell receptor in the initiation of an immune response.不同的HLA等位基因被认为导致这些细胞表面分子的结构变异,从而在启动免疫反应时,影响蛋白质与抗原及T细胞受体相互作用的能力。
This affects such critical processes as immunity against infections and self-tolerance to prevent autoimmunity.这影响诸如抗感染免疫和防止自身免疫的自身耐受等关键过程。
Ankylosing spondylitis, a chronic inflammatory disease of the spine and sacroiliac joints, is one example.强直性脊柱炎,一种脊柱和骶髂关节的慢性炎症性疾病,就是一个例子。
More than 95% of those with ankylosing spondylitis are HLA-B27 positive; the risk for developing ankylosing spondylitis is at least 150 times higher for people who have certain HLA-B27 alleles than for those who do not.超过95%的强直性脊柱炎患者为HLA-B27阳性;携带某些HLA-B27等位基因的人发生强直性脊柱炎的风险至少是不携带者的150倍。
These alleles lead to HLA-B27 heavy chain misfolding and inefficient antigen presentation.这些等位基因导致HLA-B27重链错误折叠和抗原呈递效率低下。
In other disorders, the association between a particular HLA allele or haplotype and a disease is not due to functional differences in immune response genes themselves.在其他疾病中,特定HLA等位基因或单倍型与疾病之间的关联并非由于免疫反应基因本身的功能差异。
Instead, the association is due to a particular allele being present at a very high frequency on chromosomes that also happen to contain disease-causing variants in another gene within the major histocompatibility complex region.相反,该关联是由于某一特定等位基因在染色体上高频出现,而这些染色体恰好也在主要组织相容性复合体区域内另一个基因上含有致病变异。
One example is hemochromatosis (Case 20), a common disorder of iron overload.一个例子是血色病(病例20),一种常见的铁过载疾病。
More than 80% of individuals with hemochromatosis are homozygous for a common variant, Cys 282Tyr, in the hemochromatosis gene (HFE) and have HLA-A*0301 alleles at their HLA-A locus.超过80%的血色病患者在血色病基因(HFE)中为常见变异Cys282Tyr的纯合子,并在其HLA-A位点携带HLA-A*0301等位基因。
The association is not the result of HLA-A*0301, however.然而,该关联并非HLA-A*0301所致。
HFE is involved with iron transport or metabolism in the intestine; HLA-A, as a class I immune response gene, has no effect on iron transport.HFE参与肠道铁转运或代谢;而HLA-A作为I类免疫反应基因,对铁转运无影响。
The association is due to proximity of the two loci and LD between the Cys 282Tyr HFE mutation and the A*0301 allele at HLA-A.该关联是由于两个位点邻近以及HFE基因Cys282Tyr突变与HLA-A位点A*0301等位基因之间的连锁不平衡所致。
16/46
protein (Tyr 402His).
Ch11 — Segment 16
protein (Tyr 402His).蛋白质(Tyr402His)。
The Tyr 402His alteration, which has an allele frequency of 26 to 29% in European and African populations, showed an even stronger association with AMD than did the two SNPs that showed an association in the original GWAS.Tyr402His变异在欧裔和非洲裔人群中的等位基因频率为26%至29%,与年龄相关性黄斑变性的关联性比原始全基因组关联研究中显示关联的两个单核苷酸多态性更强。
Given that drusen contain complement factors and that CFH is found in retinal tissues around drusen, it is believed that the Tyr 402His variant is less protective against the inflammation that is thought to be responsible for drusen formation and retinal damage.鉴于玻璃膜疣含有补体因子,且CFH存在于玻璃膜疣周围的视网膜组织中,一般认为Tyr402His变异对被认为导致玻璃膜疣形成和视网膜损伤的炎症的保护作用较弱。
Thus, Tyr 402His is likely to be the variant at the CFH locus responsible for increasing the risk for AMD.因此,Tyr402His很可能是CFH基因座上增加年龄相关性黄斑变性风险的变异。
More recent GWAS of AMD, using more than 7600 cases and more than 50,000 controls and millions of variants genome wide, have revealed that alleles at a minimum of 19 loci are associated with AMD, with genome-wide significance of P &lt; 5 × 10−8.更近期的年龄相关性黄斑变性全基因组关联研究使用了超过7600例病例、超过50000例对照以及全基因组数百万个变异,揭示了至少19个基因座的等位基因与年龄相关性黄斑变性相关,其全基因组显著性为P < 5 × 10⁻⁸。
A popular way to summarize GWAS in graphic form is to plot the −log 10 significance levels for each associated variant in a Manhattan plot (so named because it is thought to bear a somewhat fanciful similarity to the skyline of New York City) .以图形形式总结全基因组关联研究的一种常用方法是将每个相关变异的−log₁₀显著性水平绘制在曼哈顿图中(因其被认为与纽约市天际线有某种异想天开的相似性而得名)。
The ORs for AMD of these variants range from a high of 2. 76 for a gene of unknown function, ARMS2, and 2. 48 for CFH to 1. 1 for many other genes involved in multiple pathways, including the complement system, atherosclerosis, blood vessel formation, and others.这些变异对年龄相关性黄斑变性的比值比范围从功能未知基因ARMS2的高值2.76、CFH的2.48,到涉及补体系统、动脉粥样硬化、血管形成等多条通路的许多其他基因的1.1。
In this example of AMD, a complex disease, GWAS led to the identification of strongly associated common SNPs that in turn were in LD with a common coding SNP in the gene that appears to be the functional variant involved in the disease.在这个复杂疾病年龄相关性黄斑变性的例子中,全基因组关联研究导致鉴定出强关联的常见单核苷酸多态性,这些单核苷酸多态性反过来又与基因中一个常见编码单核苷酸多态性处于连锁不平衡,该编码单核苷酸多态性似乎是参与疾病的功能性变异。
This discovery in turn led to the identification of other SNPs in the complement cascade and elsewhere that can also predispose to or protect against the disease.这一发现进而导致鉴定出补体级联反应及其他通路中也可易感或保护该疾病的其他单核苷酸多态性。
Taken together, these results give important clues to the pathogenesis of AMD and suggest that the complement pathway might be a fruitful target for novel therapies.综合来看,这些结果为年龄相关性黄斑变性的发病机制提供了重要线索,并提示补体通路可能是新疗法的一个有前景的靶点。
Equally interesting is that GWAS revealed that a novel gene of unknown function, ARMS2, is also involved, thereby opening up an entirely new line of research into the pathogenesis of AMD.同样有趣的是,全基因组关联研究揭示了一个功能未知的新基因ARMS2也参与其中,从而为年龄相关性黄斑变性的发病机制开辟了一个全新的研究方向。
Gene Mapping by Analysis of Copy Number Variation Association studies have been successful in uncovering both common and rare risk alleles for common neuropsychiatric disorders.通过拷贝数变异分析进行基因定位:关联研究已成功揭示常见神经精神疾病的常见和罕见风险等位基因。
The Psychiatric Genomics Consortium (PGC) is a large international consortium that promotes global collaboration for the study of 11 psychiatric disorders: ADHD, Alzheimer disease, autism, bipolar disorder, eating disorders, major depressive disorder, obsessive-compulsive disorder/Tourette syndrome, posttraumatic stress disorder, schizophrenia, substance use disorders, and all other anxiety disorders.精神疾病基因组学联盟是一个大型国际联盟,旨在促进对11种精神疾病的全球合作研究:注意缺陷/多动障碍、阿尔茨海默病、孤独症、双相障碍、进食障碍、重度抑郁障碍、强迫症/抽动秽语综合征、创伤后应激障碍、精神分裂症、物质使用障碍及所有其他焦虑障碍。
For complex brain disorders such as these, it has become clear that extremely large numbers of cases and controls (i. e., &gt;10,000 samples) are necessary for markers to reach statistical significance; these numbers are well beyond what single study analysis can achieve.对于此类复杂脑部疾病,已明确需要极大数量的病例和对照(即>10,000份样本)才能使标记物达到统计学显著性;这一数量远超单一研究分析所能达到的范围。
The PGC GWAS group aims to conduct rigorous large-scale-analysis for 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 0 5 10 15 100 200 300 400 CFH CF1 C2-CFB ARMS2-HTRA1 B3GALTL RAD51B UPC CETP C3 APOE TIMP3 SLC16A8 VEGFA TNFRSF10A COL15A1-TGFBR1 FRK-COL10A1 IER3-DDR1 ADAMS9 COL8A1 –log 10P Chromosome Each blue dot represents the statistical significance [expressed as −log 10(P) plotted on the y-axis], confirming a previously known association; green dots are the statistical significance for novel associations.每个蓝点代表统计学显著性[以−log₁₀(P)表示,绘制在y轴上],确认了先前已知的关联;绿点代表新关联的统计学显著性。
The discontinuity in the y-axis is needed because some of the associations have extremely small P values &lt; 1 × 10−16.y轴出现间断是因为某些关联的P值极小,< 1 × 10⁻¹⁶。
(From Fritsche LG, Chen W, Schu M, et al: Seven new loci associated with age-related macular degeneration.(摘自Fritsche LG, Chen W, Schu M, 等:与年龄相关性黄斑变性相关的七个新基因座,《自然·遗传学》17:1783–1786, 2013。)
Nature Genet 17:1783–1786, 2013.)(此句缺失原文,根据编号规则,[17]应为空或待补充,但输入中无[17]内容,故跳过。实际输出时按照要求输出所有给定编号,此处[17]无对应英文,不输出。)
17/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 219 psychiatric disorders (i. e., to gather data from multiple studies a…
Ch11 — Segment 17
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 219 psychiatric disorders (i. e., to gather data from multiple studies and platforms into one large dataset and perform GWAS).识别人类疾病的遗传基础 219 精神障碍(即从多项研究和平台收集数据整合为一个大型数据集,并进行GWAS)。
This approach has proven effective in increasing the number of associated loci recognized for disorders (e. g., for schizophrenia, the number of associated loci increased from 22 using 20,000 subjects to 108 using 150,000 subjects).这种方法已被证明能有效增加疾病相关位点的识别数量(例如,精神分裂症的相关位点从使用20,000名受试者时的22个增加到使用150,000名受试者时的108个)。
One of the beneficial by-products of genotyping so many individuals, typically performed with SNP microarray, is the ability to interrogate the data for copy number variation (CNV).对如此多个体进行基因分型(通常使用SNP微阵列进行)的一个有益副产品是能够从数据中查询拷贝数变异(CNV)。
Rare CNVs are known to cause several neuropsychiatric conditions, some with high penetrance.已知罕见CNV会导致多种神经精神疾病,其中一些具有高外显率。
Since highly penetrant CNVs are individually rare and tend to have nonrecurrent breakpoints, large cohorts are needed to show statistical association and precisely map regions and genes.由于高外显率CNV各自罕见且往往具有非复发性断点,因此需要大型队列来显示统计关联并精确映射区域和基因。
For example, in .例如,在。
If a population is stratified into separate subpopulations (e. g., by ethnicity or religion) and members of one subpopulation rarely mate with members of other subpopulations, then a disease that happens to be more common in one subpopulation can appear (incorrectly) to be associated with any alleles that also happen to be more common in that subpopulation than in the population as a whole.如果一个群体被分层为不同的亚群(例如,按种族或宗教),且一个亚群的成员很少与其他亚群的成员交配,那么碰巧在该亚群中更常见的疾病可能(错误地)显示与任何也在该亚群中比整个群体更常见的等位基因相关联。
Factitious association due to population stratification can be minimized, however, by careful selection of matched controls.然而,通过仔细选择匹配的对照,可以最小化由群体分层导致的人为关联。
In particular, one form of quality control is to make sure the cases and controls have similar frequencies of alleles whose frequencies differ markedly between populations (ancestry informative markers as we discussed in Chapter 10).特别是,一种质量控制形式是确保病例和对照中那些在群体间频率差异显著的等位基因(如我们在第10章讨论的祖先信息标记)的频率相似。
If the frequencies seen in cases and controls are similar, then unsuspected or cryptic stratification is unlikely.如果病例和对照中观察到的频率相似,则不太可能存在未预料到的或隐秘的分层。
In addition to the problem of stratification producing false-positive associations, false-positive results in GWAS can arise if an inappropriately lax test for statistical significance is applied.除了分层产生假阳性关联的问题外,如果使用了不恰当的宽松统计显著性检验,GWAS中也可能出现假阳性结果。
This is because, as the number of alleles being tested for a disease association Position (hg 18) Genes Breakpoint association: –log(z P) 4 2 0 Case CNVs Control CNVs 50M NRXN1INM_004801 NRXN1INM_001135659 NRXN1INM_138735 51M Manhattan plot of copy number variation breakpoint associations at the NRXN1 locus in cases with schizophrenia versus controls.这是因为,随着测试疾病关联的等位基因数量增加 Position (hg 18) Genes Breakpoint association: –log(z P) 4 2 0 Case CNVs Control CNVs 50M NRXN1INM_004801 NRXN1INM_001135659 NRXN1INM_138735 51M 精神分裂症病例与对照中NRXN1位点拷贝数变异断点关联的曼哈顿图。
The three isoforms of NRXN1 are shown in pink.NRXN1的三种亚型以粉色显示。
Rare copy number deletions (red bars) and duplications (blue) are mapped in schizophrenia cases (n = 21,094) and population controls (n = 20,227).在精神分裂症病例(n = 21,094)和群体对照(n = 20,227)中映射了罕见拷贝数缺失(红色条)和重复(蓝色)。
18/46
increases, the risk of finding associations by chance alone also increases—a concept in statistics known as the problem …
Ch11 — Segment 18
increases, the risk of finding associations by chance alone also increases—a concept in statistics known as the problem of multiple hypothesis testing.增加时,仅凭偶然发现关联的风险也会增加——这一统计学概念被称为多重假设检验问题。
To understand why the cutoff for statistical significance must be much more stringent when multiple hypotheses are being tested, imagine flipping a coin 50 times and having it come up heads 40 times.要理解为何在检验多重假设时,统计学显著性的临界值必须严格得多,可以设想抛一枚硬币50次,结果有40次正面朝上。
Such a result has a probability of occurring only once in ~100,000 times.这样的结果发生的概率大约仅为十万分之一。
However, if the same experiment were repeated a million times, chances are greater than 99. 999% that at least one coin flip experiment out of the million performed will result in 40 or more heads!然而,如果重复同一实验一百万次,那么在一百万次抛硬币实验中,至少有一次出现40次或更多正面的几率超过99.999%!
Thus, even rare events that occur by chance alone in an experiment become frequent when the experiment is repeated over and over again.因此,即使在实验中仅凭偶然发生的罕见事件,当实验反复重复时也会变得频繁。
This is why, when testing for an association with hundreds of thousands to millions of variants across the genome, tens of thousands of variants could appear associated with P &lt; 0. 05 by chance alone.这就是为何当检验全基因组范围数十万到数百万个变异与疾病的关联时,可能有数万个变异仅凭偶然就显示P < 0.05的关联。
This makes a typical cutoff for statistical significance of P &lt; 0. 05 far too low to point to a true association.这使得通常的统计学显著性临界值P < 0.05过于宽松,不足以指向真正的关联。
Instead, a significance level of P &lt; 5 × 10−8 is considered to be more appropriate for GWAS that tests hundreds of thousands to millions of variants.相反,对于检验数十万到数百万个变异的GWAS,P < 5 × 10⁻⁸的显著性水平被认为更为合适。
Even with appropriately stringent cutoffs for genome-wide significance, however, false-positive results due to chance alone will still occur.然而,即使采用适当的全基因组显著性严格临界值,仍会出现仅凭偶然造成的假阳性结果。
To take this into account, a properly performed GWAS usually includes a replication study in a different, completely independent group of individuals to show that alleles near the same locus are associated.为考虑这一点,正确执行的GWAS通常会在另一个完全独立的人群中进行重复研究,以显示同一基因座附近的等位基因确实相关。
A caveat, however, is that alleles that show association may be different in different ancestral groups.但需要注意的是,显示关联的等位基因在不同祖先群体中可能不同。
Finally, it is important to emphasize that if an association is found between a disease and a marker allele that is part of a dense haplotype map, one cannot infer a functional role for that marker allele in increasing disease susceptibility.最后,必须强调的是,如果发现疾病与作为密集单倍型图谱一部分的标记等位基因存在关联,不能推断该标记等位基因在增加疾病易感性方面具有功能性作用。
Because of the nature of LD, all alleles in LD with an allele at a locus involved in the disease will show apparently positive association, whether or not they have any functional relevance in disease predisposition.由于连锁不平衡(LD)的性质,所有与疾病相关基因座上的等位基因处于LD的等位基因,无论它们在疾病易感性中是否具有任何功能相关性,都将显示明显阳性关联。
An association based on LD is still quite useful, however; for the marker alleles to appear associated, they likely sit within an LD block that also harbors the actual disease locus.然而,基于LD的关联仍然非常有用;因为标记等位基因要显示关联,它们很可能位于也包含真正疾病基因座的LD区块之内。
Importance of Associations Discovered with GWAS There is vigorous debate regarding the interpretation of GWAS results and their value as a tool for human genetic studies.通过GWAS发现的关联的重要性 关于GWAS结果的解读及其作为人类遗传学研究工具的价值,存在激烈的争论。
The debate arises primarily from a misunderstanding of what an OR or RR means.争论主要源于对OR或RR含义的误解。
Many properly executed GWAS yield significant associations, but of very modest effect size (similar to the OR of 1. 1 just mentioned for AMD).许多正确执行的GWAS产生了显著关联,但效应量非常适中(类似于上文提到的AMD的OR为1.1)。
In fact, significant associations of smaller and smaller effect size have become more common as larger and larger sample sizes are used.事实上,随着样本量越来越大,效应量越来越小的显著关联变得更为常见。
This has led to the suggestion that GWAS are of little value because the effect size of the association, as measured by OR or RR, is too small to implicate the gene and pathway identified by that variant in the pathogenesis of the disease.这导致有人提出GWAS价值不大,因为由OR或RR衡量的关联效应量太小,无法提示该变异所涉及的基因和通路在疾病发病机制中的作用。
This is faulty reasoning on two accounts.这种推理在两个方面是错误的。
First, ORs are a measure of the impact of a specific allele (e. g., the CFH Tyr 402His allele for AMD) on complex pathogenetic pathways, such as the alternative complement pathway, of which CFH is a component.首先,OR衡量的是特定等位基因(例如AMD的CFH Tyr402His等位基因)对复杂致病通路(例如CFH作为其组分的替代补体通路)的影响。
The subtlety of that impact is determined by how that allele perturbs the biologic function of the gene in which it is located, not by whether the gene harboring that allele might be important in disease pathogenesis.这种影响的微妙性取决于该等位基因如何干扰其所在基因的生物学功能,而非取决于携带该等位基因的基因是否在疾病发病机制中重要。
Studies, for example, of individuals with different autoimmune disorders, such as rheumatoid arthritis, systemic lupus erythematosus, and Crohn disease, reveal modest associations, but with some of the same variants.例如,对患有不同自身免疫性疾病(如类风湿关节炎、系统性红斑狼疮和克罗恩病)的个体的研究显示,存在适度的关联,但部分变异相同。
This suggests common pathways leading to these distinct but related diseases, an observation that may illuminate their pathogenesis.这表明这些不同但相关的疾病存在共同通路,这一观察可能有助于阐明其发病机制。
Second, even if the effect size of any one variant is small, GWAS demonstrate that many of these disorders are indeed extremely polygenic, even more so than previously suspected.其次,即使每个变异效应量很小,GWAS也表明许多这类疾病确实具有极高的多基因性,甚至比之前预想的还要高。
Thousands of variants, most of which individually contribute little to disease likelihood (ORs between 1. 01 and 1. 1) in aggregate, account for a substantial fraction of clustering of these diseases within certain families (see Chapter 9).数千个变异,其中大多数单独对疾病可能性贡献很小(OR在1.01到1.1之间),但总体上解释了这些疾病在特定家族中聚集的相当大一部分(见第9章)。
Indeed genotypes from all of the variants can be combined into a number called a polygenic risk score (PRS) (see Chapter 9).实际上,所有变异的基因型可以组合成一个称为多基因风险评分(PRS)的数字(见第9章)。
The PRS reflects a person’s inherited susceptibility to a disease, and the effect size can be as high as with some rare penetrant variants.PRS反映了一个人患某种疾病的遗传易感性,其效应量可与某些罕见的高外显率变异相媲美。
Most alleles found by GWAS may indeed have modest effect size, but there is a critical and perhaps most fundamental finding of GWAS: the genetic architecture of some of the most common complex diseases may involve hundreds to thousands of loci harboring variants of small effect in many genes and pathways.确实,GWAS发现的大多数等位基因效应量适中,但GWAS还有一个关键且可能是最基本的发现:一些最常见复杂疾病的遗传结构可能涉及数百到数千个基因座,这些基因座包含许多基因和通路中效应微小的变异。
These genes and pathways are important to our understanding of how complex diseases occur, even if each allele exerts but subtle effects on gene regulation or protein function and on disease susceptibility.这些基因和通路对于我们理解复杂疾病如何发生非常重要,即使每个等位基因仅对基因调控或蛋白质功能以及疾病易感性产生微妙影响。
It is also important to note the potential interplay between common and rare variants in disease susceptibility.同样重要的是要注意常见和罕见变异在疾病易感性中的潜在相互作用。
This has been studied in depth in schizophrenia, with the identification of both rare variants with high OR and common variants of low OR contributing to risk .在精神分裂症中对此进行了深入研究,发现了高OR的罕见变异和低OR的常见变异共同导致风险。
An overall risk profile that combines a PRS with rare variants will perhaps be come available for many disorders.结合PRS与罕见变异的总体风险谱未来或许可用于许多疾病。
Thus, GWAS remains an important human genetics research tool for dissecting the many contributions to complex disease, regardless of whether the individual variants associated substantially raise the likelihood of the disease in individuals who carry them (see Chapter 17).因此,GWAS仍然是剖析复杂疾病多种贡献的重要人类遗传学研究工具,无论单个变异是否显著提高了携带者患病的可能性(见第17章)。
There are more than 300,000 mapped associations in the genome (基因组中已定位超过300,000个关联(
19/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 221 anticipate many more genetic variants responsible for complex diseas…
Ch11 — Segment 19
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 221 anticipate many more genetic variants responsible for complex diseases to be identified by genome-wide association and that deep sequencing of such regions will uncover the variants or collections of variants functionally responsible for the disease.识别人类疾病的遗传基础 221 预期通过全基因组关联分析将发现更多复杂疾病相关的遗传变异,并且对这些区域的深度测序将揭示功能上导致疾病的变异或变异集合。
Such findings should provide powerful insights and potential therapeutic targets for many of the common diseases that cause so much morbidity and mortality.这些发现应为许多导致高发病率和死亡率的常见疾病提供强有力的见解和潜在的治疗靶点。
FINDING GENES RESPONSIBLE FOR DISEASE BY GENOME-WIDE SEQUENCING Thus far in this chapter we have focused on two approaches to map and then identify genes involved in disease: linkage analysis and GWAS.通过全基因组测序发现致病基因 到目前为止,本章我们专注于两种定位和识别疾病相关基因的方法:连锁分析和GWAS。
Now we turn to a third approach, involving direct genome-wide sequencing of affected individuals, along with their parents and/ or other family members, or a population cohort with the same clinical diagnosis.现在我们转向第三种方法,涉及对受影响的个体及其父母和/或其他家庭成员,或具有相同临床诊断的人群队列进行直接的全基因组测序。
Characteristics, strengths, and weaknesses of linkage, association, and genomewide sequencing methods for disease gene identification are summarized in 4.连锁分析、关联分析和全基因组测序方法在疾病基因鉴定中的特点、优势和劣势总结于4中。
The development of vastly improved and highthroughput methods of DNA sequencing has cut the cost of sequencing by six orders of magnitude from that spent for the Human Genome Project’s reference sequence.大幅改进的高通量DNA测序方法的发展,使得测序成本比人类基因组计划参考序列所花费的成本降低了六个数量级。
This has opened new possibilities for discovering the genes and variants responsible for diseases, particularly for rare mendelian disorders.这为发现致病基因和变异开辟了新的可能性,特别是对于罕见的孟德尔遗传病。
As introduced in Chapter 4, these new technologies make it possible to generate a whole genome sequence (GS) or, in what is a cost-effective compromise, sequence for the less than 2% of the genome containing the exons of genes, referred to as a whole exome sequence or exome sequencing (ES).如第4章所述,这些新技术使得生成全基因组序列(GS)成为可能,或者作为一种经济有效的折中方案,对基因组中不到2%的包含基因外显子的区域进行测序,称为全外显子组序列或外显子组测序(ES)。
Comparison of Exome and Whole Genome Sequencing Both exome and genome sequencing fall into the category of genome-wide sequencing (i. e., taking an unbiased approach to interrogating a genome).外显子组测序与全基因组测序的比较 外显子组测序和全基因组测序都属于全基因组测序的范畴(即采用无偏倚的方法来检测基因组)。
ES involves sequencing of the coding portion of the genome, where, during the library preparation step, gene exons are targeted or captured.ES涉及对基因组编码部分的测序,在文库制备步骤中,基因外显子被靶向或捕获。
There are several commercially available kits available for ES and all involve either a hybridization or PCR step to enrich for the exonic portion of the genome.有几种商业上可用的试剂盒可用于ES,所有这些试剂盒都涉及杂交或PCR步骤来富集基因组的外显子部分。
Further customization is often possible to boost and thus provide better coverage for certain variant types (e. g., mitochondrial variants) or clinically relevant regions (e. g., known pathogenic variants).通常还可以进一步定制,以提高并因此为某些变异类型(例如线粒体变异)或临床相关区域(例如已知致病性变异)提供更好的覆盖。
The application of ES has been instrumental in the discovery of genes responsible for rare mendelian disorders, given its cost effectiveness, allowing the sequencing of more samples compared to more costly GS.ES的应用在发现罕见孟德尔遗传病致病基因方面起到了关键作用,因为它具有成本效益,与更昂贵的GS相比,允许对更多样本进行测序。
However, there are both design and technical limitations in using ES, including the inability to analyze noncoding changes (e. g., deep intronic or regulatory regions) unless specifically targeted, the dropout of some coding sequence due to capture inefficiencies (high GC content exons), and limited ability to resolve more complex genetic mechanisms (e. g., structural rearrangements, repeat expansions).然而,使用ES存在设计和技术上的局限性,包括除非特别靶向,否则无法分析非编码变化(例如深度内含子或调控区域),由于捕获效率低下(高GC含量外显子)导致部分编码序列丢失,以及解决更复杂遗传机制(例如结构重排、重复扩增)的能力有限。
With the cost of sequencing continuing to drop it is becoming more feasible to use GS for gene discovery.随着测序成本的持续下降,使用GS进行基因发现变得更加可行。
The application of GS can address many of the technical limitations of ES (see 5).GS的应用可以解决ES的许多技术局限性(见5)。
First, since there is no capture process, GS is not prone to dropout of more complex or high GC regions and, thus, provides better coverage of the coding regions of the genomes.首先,由于没有捕获过程,GS不易丢失更复杂或高GC含量的区域,因此对基因组编码区域提供更好的覆盖。
Second, noncoding regions of the genome are 22q11. 2 del 20 1. 0 0. 0001 0. 01 Allele frequency in population Penetrance/OR 0. 1 0. 5 1q21 del 3q29 del CNVs with Low frequency with high OR SNPs with High frequency with low OR 16q11. 2 dup Penetrance in the form of an odds ratio (OR) is shown on the y-axis with allele frequency on the y-axis.其次,基因组的非编码区域是22q11.2缺失 20 1.0 0.0001 0.01 人群等位基因频率 外显率/OR 0.1 0.5 1q21缺失 3q29缺失 低频高OR的CNV 高频低OR的SNP 16q11.2重复 外显率以比值比(OR)形式显示在y轴上,等位基因频率在y轴上。
Both rare and common variation contribute to schizophrenia risk with rare CNVs acting with OR of 10 to 20.罕见变异和常见变异均对精神分裂症风险有贡献,其中罕见CNV的OR值为10至20。
Common SNPs associated with schizophrenia have OR of ~1 to 1. 5.与精神分裂症相关的常见SNP的OR值约为1至1.5。
(From Sullivan PF, Daly MJ, O’Donovan M.(来自Sullivan PF, Daly MJ, O'Donovan M.
Genetic architectures of psychiatric disorders: the emerging picture and its implications.精神疾病的遗传结构:新出现的图景及其意义。
Nat Rev Genet 13(8):537-551, 2012. org/10. 1038/nrg 3240.)Nat Rev Genet 13(8):537-551, 2012. org/10.1038/nrg3240.)
20/46
4 METHODS OF DISCOVERY: COMPARISON OF LINKAGE, ASSOCIATION METHODS, AND GENOME-WIDE SEQUENCING Linkage Association Genom…
Ch11 — Segment 20
4 METHODS OF DISCOVERY: COMPARISON OF LINKAGE, ASSOCIATION METHODS, AND GENOME-WIDE SEQUENCING Linkage Association Genome-Wide Sequencing Follows inheritance of a disease trait and regions of the genome from individual to individual in family pedigrees Looks for regions of the genome harboring disease alleles; uses polymorphic loci to mark which region an individual has inherited from which parent Uses hundreds to thousands of informative markers across the genome Not designed to find the specific variant responsible for or predisposing to the disease; can only demarcate where the variant can be found within (usually) one or a few megabases Relies on recombination events occurring in families during only a few generations to allow measurement of the genetic distance between a disease gene and markers on chromosomes Requires sampling of families, not just people affected by the disease Loses power when disease has complex inheritance with substantial lack of penetrance Most often used to map disease-causing variants with strong enough effects to cause a mendelian inheritance pattern Tests for altered frequency of particular alleles or haplotypes in individuals with a disease, compared with population controls Examines particular alleles or haplotypes for their contribution to the disease Uses anywhere from a few markers in targeted genes to hundreds of thousands of markers for genome-wide analyses Can occasionally pinpoint the variant functionally responsible for the disease; more often, defines a disease-containing haplotype over a 1- to 10-kb interval Relies on finding a set of alleles, including the disease gene, that remained together for many generations due to lack of recombination among the markers Can be carried out on case-control or cohort samples from populations Is sensitive to population stratification artifact, although this can be controlled by proper case-control designs or the use of family-based approaches Is the best approach for finding variants with small effect that contribute to complex traits Determines variation in the whole genome or coding sequence (exome) in unbiased approach in families or cohorts Requires robust filtering strategy to narrow down rare variants based on segregation and expected disease mode Filtering can be within a family or across individuals with the same clinical diagnosis Can be used in conjunction with linkage data to aid in narrowing down region of the genome with causative variant Designed to precisely identify causative variant that is functionally responsible for the disease Does not rely on linked variants and is particularly useful for finding de novo dominant variants that are intractable to linkage or association studies Can use either family- or cohort-based analysis Designed to identify rare disease-causing variants but can perform genome-wide association studies Confirmation of gene disease association greatly enhanced with submission to Gene Matcher sequenced, including deep intronic, regulatory regions and noncoding RNA, which may harbor pathogenic variation.四种发现方法:连锁分析、关联分析和全基因组测序的比较;连锁分析追踪疾病性状和基因组区域在家族系谱中从个体到个体的遗传,寻找携带疾病等位基因的基因组区域,利用多态性位点标记个体从父母继承的哪个区域,使用全基因组数百至数千个信息标记,并非旨在发现导致或易感疾病的特定变异,只能划定变异所在区域(通常在一到几兆碱基内),依赖家族中仅几代发生的重组事件来测量疾病基因与染色体上标记之间的遗传距离,需要采样家族而不仅仅是患病个体,当疾病具有复杂遗传且显著缺乏外显率时效力降低,最常用于定位效应足够强以致产生孟德尔遗传模式的致病性变异;关联分析检测患病个体与群体对照相比特定等位基因或单倍型频率的改变,检查特定等位基因或单倍型对疾病的贡献,使用从靶向基因的几个标记到全基因组分析的数十万个标记,偶尔能准确定位功能上负责疾病的变异,更常见的是在1-10 kb区间内定义包含疾病的单倍型,依赖找到一组等位基因(包括疾病基因)由于标记间缺乏重组而保持在一起许多世代,可在群体的病例-对照或队列样本中进行,对群体分层伪像敏感但可通过适当的病例-对照设计或基于家族的方法加以控制,是发现对复杂性状有贡献的小效应变异的最佳方法;全基因组测序以无偏方式在家族或队列中确定全基因组或编码序列(外显子组)的变异,需要强大的过滤策略根据分离和预期的疾病模式缩小罕见变异范围,过滤可在家族内或具有相同临床诊断的个体间进行,可与连锁数据结合使用以帮助缩小存在致病变异的基因组区域,旨在精确识别功能上负责疾病的致病变异,不依赖连锁变异尤其适用于发现连锁或关联研究难以处理的de novo显性变异,可使用基于家族或基于队列的分析,旨在发现罕见致病变异但也可进行全基因组关联研究,通过提交至Gene Matcher进行测序(包括深部内含子、调控区域和非编码RNA,可能含有致病性变异)可大大增强基因-疾病关联的确认。
Third, GS has the potential to detect nearly all classes of genetic variation, including those that are intractable to ES, such as complex structural variation, repeat expansions, and medically important genes with high homology (e. g., SMN1).第三,全基因组测序具有检测几乎所有类型遗传变异的潜力,包括那些外显子组测序难以处理的变异,例如复杂结构变异、重复扩增以及具有高度同源性的医学重要基因(如SMN1)。
A comparison of variant classes detectable by ES and GS is shown in 5.外显子组测序和全基因组测序可检测的变异类型比较见表5。
There are still many challenges in interpretation of the large amount of rare variation sequenced in a GS; however, advancements in algorithms and annotation allow us to take better advantage of GS and have led to finding causative variants in many cases whose ES results had been uninformative.全基因组测序中测序的大量罕见变异的解读仍面临许多挑战;然而,算法和注释的进展使我们能够更好地利用全基因组测序,并在许多外显子组测序结果无信息的情况下发现了致病性变异。
Filtering Genome-Wide Sequence Data to Find Potential Causative Variants It is now possible, given our current knowledge of the genome, to systematically filter sequence data from millions of variants to yield a handful of rare variants that are potentially functional.过滤全基因组测序数据以发现潜在致病性变异——基于当前对基因组的认识,现在可以从数百万个变异中系统过滤序列数据,从而得到少量可能具有功能的罕见变异。
For example, consider a family trio consisting of a child affected with a rare disorder and his parents.例如,考虑一个由患有罕见病的儿童及其父母组成的家庭三人组。
GS is performed for all three, yielding, typically, over 4 to 5 million differences relative to the human genome reference sequence (see Chapter 4).对三者均进行全基因组测序,通常会产生超过400万至500万个与人类基因组参考序列的差异(见第4章)。
Which of these variants is responsible for the disease?这些变异中哪一个是致病性的?
Extracting useful information from this massive amount of data relies on creating a variant filtering strategy, based on a variety of reasonable assumptions about which variants are more likely to be causative.从海量数据中提取有用信息依赖于创建一种变异过滤策略,该策略基于关于哪些变异更可能致病的多种合理假设。
There are many possible filtering strategies that will arrive at the same causative variant.存在许多可能的过滤策略,最终会得到相同的致病性变异。
Outcome can be influenced by the type of data generated (GS or ES), the potential inheritance pattern in the family, previous genetic testing, and whether you are looking for a known cause (e. g., diagnostic testing) versus gene discovery.结果可能受生成的数据类型(全基因组测序或外显子组测序)、家族的潜在遗传模式、既往遗传检测以及您是在寻找已知病因(如诊断检测)还是进行基因发现的影响。
Regardless, it is important to create a robust and systematic filtering scheme that can be used consistently to produce repeatable results in subsequent cases.无论如何,建立一个稳健且系统的过滤方案是重要的,该方案可一致用于后续病例以产生可重复的结果。
Most filtering strategies will rely on location of the variant, its predicted functional effect on the gene 5 COMPARISON OF VARIANT CLASSES DETECTED FROM WHOLE EXOME AND WHOLE GENOME SEQUENCING Category ES GS Small Variants Yes Yes Copy Number Variants Yes, limited sensitivity Yes Balanced Structural Variants No Yes Mitochondrial Variants Yes, requires boosting Yes Repeat Expansions No Yes, but limited Homologous Regions Some Some大多数过滤策略将依赖于变异的位置、其对基因的预测功能效应,以及表5:全外显子组和全基因组测序检测到的变异类型比较;类别包括:小变异(外显子组测序:是,全基因组测序:是)、拷贝数变异(外显子组测序:是但灵敏度有限,全基因组测序:是)、平衡结构变异(外显子组测序:否,全基因组测序:是)、线粒体变异(外显子组测序:是但需增强,全基因组测序:是)、重复扩增(外显子组测序:否,全基因组测序:是但有限)、同源区域(外显子组测序:部分,全基因组测序:部分)。
21/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 223 gene or in regulatory sequence located some distance from a gene, as…
Ch11 — Segment 21
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 223 gene or in regulatory sequence located some distance from a gene, as introduced in Chapter 3.识别人类疾病的遗传基础 223 个基因或位于距基因一定距离的调控序列中,如第三章所述。
However, these are currently more difficult to assess; as a simplifying assumption, it is reasonable to focus initially on protein coding genes.然而,这些目前更难以评估;作为简化假设,合理的方法是先聚焦于蛋白质编码基因。
If the filtering scheme does not yield interesting candidates, one can always go back and assess noncoding variants.如果过滤方案未产生有意义的候选,可随时回头评估非编码变异。
Splicing assessment algorithms are becoming better at predicting splice variants in deep intronic regions.剪接评估算法在预测深部内含子区域的剪接变异方面正变得更好。
Population frequency.人群频率。
Keep rare variants from step 1 by discarding common variants with allele frequencies greater than expected for a rare disorder.通过丢弃等位基因频率高于罕见病预期的常见变异,保留步骤1中的罕见变异。
For filtering, one would typically pick an allele frequency cutoff of between 0. 03 and 0. 05.过滤时,通常选择等位基因频率截断值在0.03至0.05之间。
Common variants are highly unlikely to be responsible for a rare disease whose population prevalence is much less than the q 2 predicted by Hardy-Weinberg equilibrium (see Chapter 10).常见变异极不可能导致罕见疾病,后者的人群患病率远小于哈迪-温伯格平衡预测的q²(见第十章)。
Deleterious nature of the variant.变异的有害性质。
Keep variants from step 2 that cause loss of function changes, including nonsense, frameshift, or those that alter highly conserved (canonical) splice sites.保留步骤2中导致功能丧失改变的变异,包括无义突变、移码突变或改变高度保守(经典)剪接位点的变异。
Keep nonsynomous variants that are predicted to be damaging.保留预测有害的非同义变异。
Discard synonymous or intronic changes that have no predicted effect on gene function.丢弃对基因功能无预测影响的同义或内含子改变。
Consistency with likely inheritance pattern.与可能的遗传模式一致性。
If the disorder is considered most likely to be autosomal recessive, keep any variants from step 3 that are found in both copies of a gene in an affected child.如果认为该疾病最可能为常染色体隐性遗传,则保留步骤3中在患病儿童的一个基因的两个拷贝中均发现的任何变异。
The child need not be homozygous for the same deleterious variant but could be a compound heterozygote for two different deleterious changes in the same gene (see Chapter 7).该儿童不必对同一有害变异纯合,但可能是同一基因中两个不同有害改变的复合杂合子(见第七章)。
If this hypothesized mode of inheritance is correct, then the parents should both be heterozygous for the variant(s).如果此假设的遗传模式正确,则父母双方均应对该变异杂合。
If the parents are consanguineous, the candidate genes and variants may be further filtered by requiring that the child be a true homozygote for the same variant derived from a single common ancestor (see Chapter 10).如果父母为近亲结婚,则可进一步筛选候选基因和变异,要求儿童是来自单个共同祖先的同一变异的真正纯合子(见第十章)。
If the disorder is severe and seems more likely to be due to a new mutation for a dominant trait, keep variants from step 3 that are de novo changes in the child and are not present in either parent.如果疾病严重且更可能由显性性状的新突变引起,则保留步骤3中儿童中新生且父母均不存在的变异。
Finally, if there is suspicion of an X-linked disorder and the child is male, one can focus on hemizygous variants that are inherited from a heterozygous mother.最后,如果怀疑为X连锁疾病且儿童为男性,则可聚焦于从杂合母亲遗传的半合子变异。
In the end, millions of variants can be filtered down to a handful occurring in a small number of genes.最终,数百万个变异可被过滤至仅剩少数基因中的少量变异。
Once the filtering reduces the number of genes and alleles to a manageable number, they can be further assessed for other characteristics.一旦过滤将基因和等位基因数量减少到可管理范围,即可进一步评估其他特征。
Do any of the genes have a known function or tissue expression pattern that would be expected of a potential disease gene?是否有任何基因具有已知功能或组织表达模式,符合潜在疾病基因的预期?
Is the gene involved in other disease phenotypes, or does it have a role in pathways with other genes in which variation can cause similar or different phenotypes?该基因是否参与其他疾病表型,或在与其它基因共同作用的通路中发挥作用,而其他基因的变异可导致相似或不同表型?
Is the gene under constraint such that loss of function variants are 4–5,000,000 variants 3,000,000,000 bp in genome Not located within or near an exon Too frequent in variant databases to cause disease Variants with no predicted functional consequence ~40,000 variants ~1500 variants ~200 variants Inconsistent with Inheritance model Clinical correlation and/or makes biological sense 5–10 variants 1–2 variants The initial enormous collection of variants is reduced into smaller and smaller bins by applying filters that remove variants unlikely to be causative, based on assuming that variants of interest are likely to be located near a gene, will disrupt its function, and are rare.该基因是否受约束,使得功能丧失变异为4–5,000,000个变异,基因组30亿碱基对,不在外显子内或附近,在变异数据库中过于常见而不致病,无预测功能后果的变异约40,000个变异,约1500个变异,约200个变异,与遗传模型不一致,临床关联和/或具有生物学意义,5–10个变异,1–2个变异;初始的大量变异通过应用过滤器被逐步缩小,这些过滤器移除不太可能致病的变异,基于假设感兴趣的变异可能位于基因附近、破坏其功能且罕见。
Each remaining candidate gene is then assessed for whether the variants are inherited in a manner that fits the most likely inheritance pattern of the disease, whether a variant occurs in a candidate gene that makes biologic sense given the phenotype in the affected child, and whether other affected individuals also have causative variants in that gene. product, population frequency, and inheritance pattern.然后评估每个剩余候选基因:变异是否以最符合疾病遗传模式的方式遗传,变异是否出现在候选基因中且从生物学角度与患病儿童的表型相符,以及其他患病个体是否也在该基因中存在致病性变异,涉及产物、人群频率和遗传模式。
Once these basic filters have been applied to narrow the number of variants, they can be further interrogated for clinical correlation, plausible biologic function, known expression, prior observation in other cases, and previous classifications.一旦应用这些基本过滤器缩小变异数量,即可进一步查询临床相关性、合理的生物学功能、已知表达、既往其他病例中的观察结果以及先前的分类。
One example of a filtering scheme that can be used to sort through these variants is shown in可用于筛选这些变异的一个过滤方案示例如图所示。
22/46
rarely observed?
Ch11 — Segment 22
rarely observed?很少观察到?
Finally, has the variant been observed and classified in others, or has pathogenic variation in the gene been observed in others with the disease?最后,该变异是否在其他个体中被观察到并分类,或者该基因的致病性变异是否在其他患病个体中被观察到?
Finding causative variation in one of these genes in other affected individuals would lend evidence that this was the responsible gene and variant in the original trio.在其他患病个体中发现这些基因之一的致病性变异将提供证据表明这是原始三人组中的致病基因和变异。
In some cases, one gene from the list in step 4 may rise to the top as a candidate because its involvement makes biologic or genetic sense, or it is known to be causative in other affected individuals.在某些情况下,步骤4列表中的一个基因可能成为首选候选基因,因为其参与具有生物学或遗传学意义,或者已知在其他患病个体中具有致病性。
In other cases, however, the gene responsible may turn out to be entirely unanticipated on biologic grounds, or may not be causative in other affected individuals because of locus heterogeneity (i. e., pathogenic variants in other as yet undiscovered genes can cause a similar disease).然而,在其他情况下,致病基因可能在生物学基础上完全出乎意料,或者由于位点异质性(即其他尚未发现的基因中的致病性变异可能导致类似疾病)而在其他患病个体中并不致病。
Such variant assessments require extensive use of public genomic databases and software tools.此类变异评估需要广泛使用公共基因组数据库和软件工具。
These include the human genome reference sequence, databases of allele frequencies (e. g., 1000 genomes and gnom AD), software that assesses how deleterious an amino acid substitution might be to gene function, collections of known diseasecausing variants (e. g., Clin Var or disease specific), and databases of functional networks and biologic pathways.这些包括人类基因组参考序列、等位基因频率数据库(例如1000基因组计划和gnomAD)、评估氨基酸置换对基因功能有害程度的软件、已知致病性变异集合(例如ClinVar或疾病特异性数据库)以及功能网络和生物学通路数据库。
The enormous expansion of this information over the past few years, fueled by genome-wide sequencing of millions of cases and controls, has played a crucial role in facilitating gene discovery and molecular diagnosis of rare mendelian disorders, as we discuss in the next section.过去几年中,在数百万病例和对照的全基因组测序推动下,这些信息的大规模扩展在促进罕见孟德尔疾病的基因发现和分子诊断方面发挥了关键作用,我们将在下一节讨论。
FILTERING STRATEGIES FOR IDENTIFICATION OF DISEASE-CAUSING GENES In the previous section we discussed filtering schemes for a single family to identify variants causing disease.识别致病基因的过滤策略 在上一节中,我们讨论了针对单个家系的过滤方案以识别致病变异。
But what if nothing is identified or there are interesting candidates that are lacking enough evidence to definitively link to disease?但如果未识别出任何变异,或者有候选变异但缺乏足够证据将其与疾病明确关联,该怎么办?
For rare mendelian disorders, analysis of a single affected individual is often inadequate for gene discovery, making other study designs and strategies necessary.对于罕见孟德尔疾病,单个患病个体的分析通常不足以发现基因,因此需要其他研究设计和策略。
This involves traditional mapping approaches as well as other common genetic strategies that have been adapted for genome-wide sequencing.这涉及传统定位方法以及其他已适用于全基因组测序的常见遗传策略。
Generally, sequencing of multiple affected individuals within a family or sequencing of unrelated individuals with the same clinical diagnosis is necessary.通常,需要对家系内多个患病个体进行测序,或对具有相同临床诊断的无关联个体进行测序。
Deciding which cases are the most informative to sequence is highly influenced by the suspected mode of inheritance and whether the variants are expected to be inherited or due to new (de novo) mutations. .决定哪些病例对测序最具信息量,很大程度上受疑似的遗传方式以及变异是遗传性还是新发(de novo)突变的影响。
In those families with expected or known consanguinity, causative variants are expected to be homozygous.在预期或已知存在近亲婚配的家系中,致病变异预计为纯合子。
Sequencing of sibling pairs or other affected relatives and prioritizing homozygous variants is an extremely effective strategy that has led to many new gene discoveries.对同胞对或其他患病亲属进行测序并优先考虑纯合变异,是一种极为有效的策略,已导致许多新基因的发现。
If possible, sequencing of more distantly related individuals will aid in further reducing the number of shared homozygous variants.如果可能,对更远缘亲属进行测序将有助于进一步减少共享纯合变异的数量。
Homozygosity mapping with SNP microarrays to predetermine chromosomal loci with overlapping stretches of homozygosity in affected individuals can be used to narrow regions.使用SNP微阵列进行纯合子定位,以预先确定患病个体中具有重叠纯合片段的染色体位点,可用于缩小区域。
However, with the cost of sequencing decreasing, current strategies typically only use the genome-wide data for homozygosity mapping, or to sequence more individuals and look for shared homozygous variants.然而,随着测序成本的降低,当前策略通常仅使用全基因组数据进行纯合子定位,或对更多个体进行测序并寻找共享纯合变异。
Finally, for those families showing X-linked recessive inheritance, one can look for shared variants on chromosome X in males and filter out autosome variants.最后,对于显示X连锁隐性遗传的家系,可以寻找男性患者中X染色体上的共享变异,并过滤掉常染色体变异。
However, one must A B Autosomal Recessive Autosomal Dominant X-linked C D E Strategies (A–E) are detailed in the text.然而,必须注意 A B 常染色体隐性 常染色体显性 X连锁 C D E 策略(A–E)在正文中有详细说明。
Filled symbols indicate affected individuals, and empty symbols are unaffected.实心符号表示患病个体,空心符号表示未患病个体。
Obligate carriers are denoted with a dot.强制携带者用圆点表示。
Cross lines depict individuals undergoing genome-wide sequencing, with circles below indicating variant sets and overlapping strategies for detection of causative variants.交叉线表示进行全基因组测序的个体,下方的圆圈表示变异集以及用于检测致病变异的重叠策略。
23/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 225 be cautious in this approach that there is unequivocal evidence that…
Ch11 — Segment 23
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 225 be cautious in this approach that there is unequivocal evidence that the transmission is X-linked and not potentially autosomal recessive.识别人类疾病的遗传基础 225 在采用该方法时应谨慎,即必须有明确证据表明遗传方式为X连锁而非潜在的常染色体隐性遗传。
The use of genome-wide sequencing for identifying genes causing inherited dominant disorders can be more challenging.使用全基因组测序来识别导致遗传性显性疾病的基因可能更具挑战性。
This strategy involves looking for shared variants in affected individuals and excluding variants present in unaffected individuals.该策略涉及在患病个体中寻找共享变异,并排除未患病个体中存在的变异。
In principle, finding the causative variant can be difficult owing to the number of rare variants that are private to a family.原则上,由于存在许多家族特有的罕见变异,找到致病变异可能很困难。
To effectively filter variants down to a manageable number to interpret, sequencing of many affected and unaffected individuals is necessary, which can be costly especially if using GS.为了有效将变异过滤到可解释的数目,需要对许多患病和未患病个体进行测序,这可能会很昂贵,尤其是使用基因组测序时。
This can be partially mitigated in large families by using the principles of linkage that we discussed earlier in the chapter.在大家族中,可以通过使用本章前面讨论的连锁原理来部分缓解这一问题。
Performing linkage to define a disease locus, then strategically sequencing a few distantly related individuals, can be a cost-effective approach to gene identification.进行连锁分析以确定疾病位点,然后策略性地对少数远亲个体进行测序,是一种经济有效的基因鉴定方法。
If there are no candidates in the linked region, it raises the possibility of variant classes that may not be detectable due to technical limitations of genome-wide sequencing or that the causative variant is noncoding and difficult to interpret without functional assays.如果连锁区域内没有候选基因,则可能涉及因全基因组测序技术限制而无法检测的变异类型,或者致病变异为非编码变异,在没有功能检测的情况下难以解读。
Genome-wide sequencing has been particularly successful for discovery of disorders caused by de novo dominant variants—a disease mode that is intractable to linkage studies.全基因组测序在发现由新生显性变异引起的疾病方面特别成功——这种疾病模式难以通过连锁研究处理。
These disorders tend to be severe genetic lethals, with variants not passed on to subsequent generations.这些疾病往往为严重的遗传致死性疾病,其变异不会传递给后代。
There are relatively few de novo events per exome (1–2) and genome (50–70); overlapping de novo events in two unrelated individuals with a similar phenotype can be sufficient to show disease causation, although overlap by chance is possible.每个外显子组(1–2个)和基因组(50–70个)的新生事件相对较少;两个具有相似表型的无关个体中出现重叠的新生事件足以表明疾病因果关系,尽管偶然重叠是可能的。
The de novo overlap strategy has been particularly useful in finding causes for disorders with high genetic and phenotypic heterogeneity, such as syndromic forms of autism spectrum disorder and intellectual disability.新生重叠策略在寻找具有高度遗传和表型异质性的疾病原因方面特别有用,例如自闭症谱系障碍和智力障碍的综合征型。
Pooling of large case cohorts is often necessary to find de novo variants in the same gene, given their individually rare participation in a genotypically heterogeneous pool.由于新生变异在基因型异质性群体中各自罕见,通常需要汇集大型病例队列才能在同一位点发现新生变异。
Genes identified using this strategy are more likely to overlap by chance as the number of family trios sequenced increases.随着测序的家系三人家组数量增加,使用该策略鉴定的基因更可能偶然重叠。
Thorough investigation of genotype and phenotype correlation is imperative to establish disease gene associations.彻底研究基因型与表型的相关性对于建立疾病基因关联至关重要。
Each of the strategies above can be hindered by genetic and phenotypic heterogeneity and the rarity of the disorders that remain unsolved.上述每种策略都可能受到遗传和表型异质性以及尚未解决疾病的罕见性的阻碍。
Indeed, much of the low-hanging fruit for rare disease associations has been discovered, and the majority of cases that undergo genome-wide sequencing are unsolved.事实上,对于罕见疾病关联的许多易于获取的成果已被发现,而大多数接受全基因组测序的病例仍未得到解决。
It is clear that for study of exceedingly rare disorders or new disease mechanisms, international collaboration is necessary, as no one cohort will likely have the numbers to make robust disease associations.显然,对于极罕见疾病或新疾病机制的研究,国际合作是必要的,因为单一队列不太可能拥有足够的数量来做出可靠的疾病关联。
One such effort is the Matchmaker Exchange (MME), which uses a federated network of data sharing to solve undiagnosed cases and facilitates publication of case series describing new disease genes .其中一项努力是Matchmaker Exchange(MME),它利用联邦式数据共享网络解决未确诊病例,并促进描述新疾病基因的病例系列发表。
The MME provides a genetic Gene 1 Gene 5 Gene 4 Gene 7 Gene 6 Gene 2 Gene 3 Gene 3 Gene 3 Matchmaker exchange Matchmaker Exchange Statistics (last updated February 2022) MME Node DECIPHER (UK) Gene Matcher (USA) IRUD (Japan) My Gene 2 (USA) Patient Matcher (Sweden) Phenome Central (Canada) RD-Connect GPAP (Europe) 80,000 42,534 3,578 2,521 8,945 12,118 12,114 7,929 40,544 64,852 62 1,599 14 8,904 4,901 1,174 8,967 13,436 55 1,302 25 3,014 821 1,224 seqr (USA) Patients/Cases Total Patients/Cases in MME Unique Genes For rare mendelian disorders, federated networks of data sharing can facilitate solutions for undiagnosed cases and establish genotype-phenotype correlations.MME提供了一个遗传的基因1基因5基因4基因7基因6基因2基因3基因3基因3 Matchmaker exchange Matchmaker Exchange统计数据(最后更新于2022年2月)MME节点 DECIPHER(英国) Gene Matcher(美国) IRUD(日本) My Gene 2(美国) Patient Matcher(瑞典) Phenome Central(加拿大) RD-Connect GPAP(欧洲) 80,000 42,534 3,578 2,521 8,945 12,118 12,114 7,929 40,544 64,852 62 1,599 14 8,904 4,901 1,174 8,967 13,436 55 1,302 25 3,014 821 1,224 seqr(美国) 患者/病例 MME中总患者/病例 独特基因 对于罕见孟德尔疾病,联邦式数据共享网络可以促进未确诊病例的解决并建立基因型-表型相关性。
Candidate genes without established disease associations, sequenced as either part of research or clinical service, can be submitted to MME via several projects.尚未建立疾病关联的候选基因,无论是作为研究还是临床服务的一部分进行测序,都可以通过多个项目提交至MME。
Matching genes can be linked between submitters for more detailed phenotype-genotype correlation.匹配的基因可以在提交者之间进行关联,以进行更详细的表型-基因型相关性分析。
In the example here, researchers and clinicians from three centers have submitted candidate genes (solid lines) for affected individuals to MME, with gene 3 in common.在此示例中,来自三个中心的研究人员和临床医生已将患病个体的候选基因(实线)提交至MME,其中基因3是共同的。
Direct connections between submitters (dotted bidirectional arrows) can then be established to further study potential causal relationships.然后可以在提交者之间建立直接联系(虚线双向箭头),以进一步研究潜在的因果关系。
The insert includes a snapshot of federated databases that link into MME with the number of cases and genes submitted.插图中包含链接到MME的联邦式数据库的快照,显示了提交的病例和基因数量。
24/46
matchmaking service that connects multiple researchers or clinicians who have patients with similar phenotypes and varia…
Ch11 — Segment 24
matchmaking service that connects multiple researchers or clinicians who have patients with similar phenotypes and variants in the same candidate gene.一种匹配服务,连接多位研究人员或临床医生,这些人员拥有表型相似且在同一候选基因中存在变异(variant)的患者。
Multiple nodes feed into MME, with a growing collective dataset of more than 150,000 cases submitted from 99 countries.多个节点向MME(匹配交换平台)输入数据,集体数据集不断增长,已包含来自99个国家提交的超过15万例病例。
As sequencing becomes less expensive, these datasets will continue to grow.随着测序成本降低,这些数据集将继续增长。
The increased application of genome-wide sequencing coupled with sharing of data has played a pivotal role in our understanding of gene-disease associations, leading to a rapid increase in gene discovery over the last 10+ years.全基因组测序应用的增加以及数据共享,在我们理解基因-疾病关联方面发挥了关键作用,在过去十多年中促使基因发现迅速增加。
Since the application of GS or ES to rare mendelian disorders was first described in 2009, many hundreds of such disorders have been studied, and the causative variants found among hundreds of previously unrecognized disease genes .自2009年首次描述将基因组测序(GS)或外显子组测序(ES)应用于罕见孟德尔疾病以来,已有数百种此类疾病得到研究,并在数百个此前未被识别的疾病基因中发现了致病变异。
These discoveries feed back into diagnostic testing where they not only provide information useful for genetic counseling in the families involved, but may inform clinical management and the potential development of effective treatments.这些发现反馈到诊断检测中,不仅为相关家庭的遗传咨询提供有用信息,还可能为临床管理和有效治疗的潜在开发提供依据。
The application of genome-wide sequencing has grown in diagnostic testing, notably in individuals with genetically heterogenous disorders, leading to a diagnostic yield of ~30%.全基因组测序在诊断检测中的应用有所增长,尤其在具有遗传异质性疾病的个体中,诊断率约为30%。
The success rate of this approach will only increase as the costs of sequencing continue to fall and with improved ability to interpret the likely functional consequences of sequence changes in the genome.随着测序成本持续下降以及解读基因组序列变化可能功能后果的能力提高,这种方法成功率只会增加。
Example: Identification of the Gene Causing Postaxial Acrofacial Dysostosis The genome-wide sequencing approach just outlined was used in the study of a family in which two siblings affected with a rare congenital malformation, known as postaxial acrofacial dysostosis (POAD), were born to two unaffected, unrelated parents.示例:鉴定导致轴后肢端面骨发育不良的基因。上述全基因组测序方法被用于一个家庭的研究,该家庭中两名同胞患有罕见的先天性畸形——轴后肢端面骨发育不良(POAD),其父母均未患病且无血缘关系。
Individuals with this disorder have small jaws, missing or poorly developed digits on the ulnar sides of their hands, underdevelopment of the ulna, cleft lip, and clefts (colobomas) of the eyelids.患有此病的个体表现为下颌小、手尺侧手指缺失或发育不良、尺骨发育不全、唇裂以及眼睑缺损(colobomas)。
The disorder was thought to be autosomal recessive, because some parents of an affected child are consanguineous, and because a few families are like the one here, with multiple affected siblings born to unaffected parents—both findings that are hallmarks of recessive inheritance (see Chapter 7).该疾病曾被认为是常染色体隐性遗传,因为一些患病儿童的父母为近亲结婚,且少数家庭与此处所述家庭相似,未患病父母生育多名患病同胞——这两项发现均为隐性遗传的特征(见第7章)。
This small family alone was clearly inadequate for linkage analysis.仅凭这个小型家庭显然不足以进行连锁分析。
Instead, all four members of the family had their entire genomes sequenced and analyzed.相反,该家庭所有四名成员的全基因组均进行了测序和分析。
From an initial list of more than 4 million variants and assuming autosomal recessive inheritance of the disorder in both affected children, a filtering scheme similar to that described earlier (see Figs. 11. 14 and 11. 15) yielded only four possible candidate genes.从最初超过400万个变异列表出发,假设两名患病儿童均患常染色体隐性遗传病,采用与前述(见图11.14和11.15)相似的过滤方案,仅得到四个可能的候选基因。
One of these, DHODH, had rare damaging variants in two other unrelated individuals with POAD, thereby identifying this gene as responsible for the disorder in these families.其中之一,DHODH,在另外两名无血缘关系的POAD患者中具有罕见的致病变异,从而确定该基因为这些家族中疾病的致病基因。
DHODH encodes dihydroorotate dehydrogenase, a mitochondrial enzyme involved in pyrimidine biosynthesis, and was not suspected on biologic grounds to be the gene responsible for this malformation syndrome.DHODH编码二氢乳清酸脱氢酶,一种参与嘧啶生物合成的线粒体酶,从生物学角度此前并未被怀疑为该畸形综合征的致病基因。
Limitations of Genome-Wide Sequencing and Future Outlook Although the genome-wide sequencing approach has proved powerful for both gene discovery and diagnosis of rare mendelian disease, it still has limitations.全基因组测序的局限性及未来展望。尽管全基因组测序方法已被证明对基因发现和罕见孟德尔病诊断均十分有效,但它仍有局限性。
Most groups Growth of Gene-phenotype Relationships 31 Dec 2021 7500 6750 6000 5250 4500 3750 Year Phenotypes with known molecular basis Genes with phenotype-causing variant 3000 2250 1500 750 0 1987 1990 1993 1996 1999 2002 2005 2008 2011 2014 2017 2020 Growth of gene association with disease phenotype has steadily increased due to the advent of genome-wide sequencing combined with data sharing.多数研究小组。基因-表型关系增长:由于全基因组测序的出现及数据共享,基因与疾病表型关联的增长稳步提升。
25/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 227 report diagnostic yields in the 20 to 40% range, depending on clinic…
Ch11 — Segment 25
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 227 report diagnostic yields in the 20 to 40% range, depending on clinical indication, which leaves the majority of cases without an identified causative variant and answer.识别人类疾病的遗传基础 227 报告诊断率在20%到40%之间,具体取决于临床指征,这导致大多数病例无法找到确定的致病变异和答案。
There are several reasons for this observation.造成这一观察结果有几个原因。
First, some disorders are intractable to standard genome-wide sequencing, including methylation disorders (e. g., Prader-Willi and Angelman syndromes) or those involving certain types of Uniparental disomy (UPD) if parents are not sequenced (i. e., heterodisomy).首先,某些疾病对标准全基因组测序难以处理,包括甲基化疾病(例如,普拉德-威利综合征和天使综合征)或涉及某些类型的单亲二倍体(UPD)的疾病(如果父母未进行测序,即异源二倍体)。
Second (and related), genome-wide sequencing may miss certain classes of variation that are difficult to detect by routine short-read sequencing alone (e. g., balanced changes, repeat expansions).其次(且与此相关),全基因组测序可能遗漏某些难以通过常规短读长测序单独检测的变异类别(例如,平衡性变异、重复扩增)。
As discussed earlier, whole genome sequencing has technological advantages over exome sequencing, with the ability to detect a broader range of variation; however, there are still limitations in accurately detecting or resolving complex regions of the genome.如前所述,全基因组测序在外显子组测序方面具有技术优势,能够检测更广泛的变异;然而,在准确检测或解析基因组复杂区域方面仍存在局限性。
Third, variation is detected that is difficult to interpret with our current understanding of the genome.第三,检测到的变异在当前对基因组的理解下难以解释。
This is particularly true of genome sequencing and the rare variation detected in noncoding and regulatory genomic regions.这对于基因组测序以及在非编码和调控基因组区域中检测到的罕见变异尤其如此。
Finally, at present, whole genome sequencing cannot actually sequence the entire genome in any one individual.最后,目前全基因组测序实际上无法对任何个体进行全基因组测序。
Cost limits most of the sequencing to short-reads and aligning back to a reference genome (termed resequencing) to detect variation.成本限制使得大部分测序采用短读长,并回帖到参考基因组(称为重测序)以检测变异。
Some sequence is too complex and/or highly homologous or repetitive to be technically sequenced or mapped back to the reference genome.有些序列过于复杂和/或高度同源或重复,技术上无法测序或回帖到参考基因组。
Additionally, sequence novel to one’s genome, i. e., not in the reference assembly, will be filtered out even though it may be pathogenic.此外,个体基因组中特有的序列(即不在参考组装中的序列)即使可能具有致病性,也会被过滤掉。
Such limitations are starting to be addressed in several ways.这些局限性正开始通过多种方式得到解决。
Improvements in both sequencing technology and informatics algorithms for variant detection allow a more accurate catalogue of genomic variation.测序技术和用于变异检测的信息学算法的改进,使得基因组变异目录更加准确。
The increasing number of genomes sequenced and the subsequent aggregation of data allow more accurate interpretation of variation.测序基因组数量的增加以及随后数据的聚合,使得对变异的解释更加准确。
In addition, the advancement of other -omic technologies, such as RNA sequencing or methylation experiments, are proving to be valuable tests ancillary to genome-wide sequencing for the interpretation of noncoding variation.此外,其他组学技术的进步,如RNA测序或甲基化实验,正被证明是辅助全基因组测序解释非编码变异的有价值检测方法。
The emergence of long-read sequencing technology (reads up to 10–100 kb in length) has provided insight into regions of the genome that have been intractable to short-read sequencing, including highly homologous or repetitive regions, with detection of variation previously unseen.长读长测序技术(读长达10–100 kb)的出现,为短读长测序难以处理的基因组区域(包括高度同源或重复区域)提供了见解,能够检测到以前未见的变异。
Long-read technology can also be used to advance de novo assembly of genomes, rather than reference-based assembly, to get a more complete picture of the genome.长读长技术还可用于推进基因组的从头组装,而非基于参考的组装,以获取更完整的基因组图像。
These advances will lead to a better understanding of our genome, its variation, and its relation to disease.这些进展将有助于更好地理解我们的基因组、其变异及其与疾病的关系。
ACKNOWLEDGMENT We wish to thank and Gregory Costain for contributing to this chapter.致谢 我们感谢Gregory Costain对本章的贡献。
GENERAL REFERENCES Altshuler D, Daly MJ, Lander ES: Genetic mapping in human disease, Science 322:881–888, 2008.一般参考文献 Altshuler D, Daly MJ, Lander ES:人类疾病的遗传定位,Science 322:881–888, 2008。
Boycott KM, Azzariti DR, Hamosh A, et al: Seven years since the launch of the Matchmaker Exchange: The evolution of genomic matchmaking, Hum Mutat, 2022.Boycott KM, Azzariti DR, Hamosh A, 等:Matchmaker Exchange启动七年:基因组匹配的演变,Hum Mutat, 2022。
Online ahead of print.在线提前发表。
PMID: 35537081.PMID: 35537081。
Boycott KM, Vanstone MR, Bulman DE, et al: Rare-disease genetics in the era of next-generation sequencing: Discovery to translation, Nat Rev Genet 14:681–691, 2013.Boycott KM, Vanstone MR, Bulman DE, 等:下一代测序时代的罕见病遗传学:从发现到转化,Nat Rev Genet 14:681–691, 2013。
Gilissen C, Hoischen A, Brunner HG, et al: Disease gene identification strategies for exome sequencing, Eur J Hum Genet 20:490–497, 2012.Gilissen C, Hoischen A, Brunner HG, 等:外显子组测序的疾病基因鉴定策略,Eur J Hum Genet 20:490–497, 2012。
Manolio TA: Genomewide association studies and assessment of the risk of disease, NEJM 363:166–176, 2010.Manolio TA:全基因组关联研究及疾病风险评估,NEJM 363:166–176, 2010。
Risch N, Merikangas K: The future of genetic studies of complex human diseases, Science 273:1516–1517, 1996.Risch N, Merikangas K:复杂人类疾病遗传学研究的未来,Science 273:1516–1517, 1996。
Sullivan PF, Daly MJ, O’Donovan M: Genetic architectures of psychiatric disorders: The emerging picture and its implications, Nat Rev Genet 13:537–551, 2012.Sullivan PF, Daly MJ, O'Donovan M:精神疾病的遗传结构:新兴图景及其意义,Nat Rev Genet 13:537–551, 2012。
Terwilliger JD, Ott J: Handbook of human genetic linkage, Baltimore, 1994, Johns Hopkins University Press.Terwilliger JD, Ott J:人类遗传连锁手册,巴尔的摩,1994, Johns Hopkins University Press。
REFERENCES FOR SPECIFIC TOPICS Abecasis GR, Auton A, Brooks LD, et al: An integrated map of genetic variation from 1,092 human genomes, Nature 491:56–65, 2012.特定主题参考文献 Abecasis GR, Auton A, Brooks LD, 等:来自1092个人类基因组的遗传变异整合图谱,Nature 491:56–65, 2012。
Bainbridge MN, Wiszniewski W, Murdock DR, et al: Whole-genome sequencing for optimized patient management, Science Transl Med 3:87re 3, 2011.Bainbridge MN, Wiszniewski W, Murdock DR, 等:全基因组测序以优化患者管理,Science Transl Med 3:87re3, 2011。
Bush WS, Moore JH: Genome-wide association studies, PLo S Computational Biol 8:e 1002822, 2012.Bush WS, Moore JH:全基因组关联研究,PLoS Computational Biol 8:e1002822, 2012。
Denny JC, Bastarache L, Ritchie MD, et al: Systematic comparison of phenome-wide association study of electronic medical record data and genome-wide association data, Nat Biotechnol 31:1102–1110, 2013.Denny JC, Bastarache L, Ritchie MD, 等:电子病历数据全表型组关联研究与全基因组关联研究的系统比较,Nat Biotechnol 31:1102–1110, 2013。
Fritsche LG, Chen W, Schu M, et al: Seven new loci associated with age-related macular degeneration, Nat Genet 17:1783–1786, 2013.Fritsche LG, Chen W, Schu M, 等:与年龄相关性黄斑变性相关的七个新位点,Nat Genet 17:1783–1786, 2013。
Gonzaga-Jauregui C, Lupski JR, Gibbs RA: Human genome sequencing in health and disease, Ann Rev Med 63:35–61, 2012.Gonzaga-Jauregui C, Lupski JR, Gibbs RA:健康与疾病中的人类基因组测序,Ann Rev Med 63:35–61, 2012。
Hindorff LA, Mac Arthur J, Morales J, et al: A catalog of published genome-wide association studies, 2015. www. genome. gov/ gwastudies International Hap Map Consortium: A second generation human haplotype map of over 3. 1 million SNPs, Nature 449:851–861, 2007.Hindorff LA, MacArthur J, Morales J, 等:已发表全基因组关联研究目录,2015. www.genome.gov/gwastudies International HapMap Consortium:超过310万个SNP的第二代人类单倍型图谱,Nature 449:851–861, 2007。
Kircher M, Witten DM, Jain P, et al: A general framework for estimating the relative pathogenicity of human genetic variants, Nat Genet 46:310–315, 2014.Kircher M, Witten DM, Jain P, 等:评估人类遗传变异相对致病性的通用框架,Nat Genet 46:310–315, 2014。
Koboldt DC, Steinberg KM, Larson DE, et al: The next-generation sequencing revolution and its impact on genomics, Cell 155:27–38, 2013.Koboldt DC, Steinberg KM, Larson DE, 等:下一代测序革命及其对基因组学的影响,Cell 155:27–38, 2013。
Lionel AC, Costain G, Monfared N, et al: Improved diagnostic yield compared with targeted gene sequencing panels suggests a role for whole-genome sequencing as a first-tier genetic test, Genet Med 20:435–443, 2018.Lionel AC, Costain G, Monfared N, 等:与靶向基因测序 panel 相比诊断率提高,提示全基因组测序可作为一线遗传检测,Genet Med 20:435–443, 2018。
Manolio TA: Bringing genome-wide association findings into clinical use, Nat Rev Genet 14:549–558, 2014.Manolio TA:将全基因组关联研究结果应用于临床,Nat Rev Genet 14:549–558, 2014。
Marshall CR, Howrigan DP, Merico D, et al: Contribution of copy number variants to schizophrenia from a genome-wide study of 41,321 subjects, Nat Genet 49:27–35, 2016.Marshall CR, Howrigan DP, Merico D, 等:来自41,321名受试者全基因组研究的拷贝数变异对精神分裂症的贡献,Nat Genet 49:27–35, 2016。
Matise TC, Chen F, Chen W, et al: A second-generation combined linkage-physical map of the human genome, Genome Res 17:1783– 1786, 2007.Matise TC, Chen F, Chen W, 等:第二代人类基因组联合连锁-物理图谱,Genome Res 17:1783–1786, 2007。
Roach JC, Glusman G, Smit AF, et al: Analysis of genetic inheritance in a family quartet by whole-genome sequencing, Science 328:636– 639, 2010.Roach JC, Glusman G, Smit AF, 等:通过全基因组测序分析一个四成员家族的遗传模式,Science 328:636–639, 2010。
Robinson PC, Brown MA: Genetics of ankylosing spondylitis, Mol Immunol 57:2–11, 2014.Robinson PC, Brown MA:强直性脊柱炎的遗传学,Mol Immunol 57:2–11, 2014。
SEARCH Collaborative Group: SLCO1B1 variants and statin-induced myopathy—A genomewide study, NEJM 359:789–799, 2008.SEARCH Collaborative Group:SLCO1B1变异与他汀类药物引起的肌病——一项全基因组研究,NEJM 359:789–799, 2008。
Stahl EA, Wegmann D, Trynka G, et al: Bayesian inference analyses of the polygenic architecture of rheumatoid arthritis, Nat Genet 44:4383–4391, 2012.Stahl EA, Wegmann D, Trynka G, 等:类风湿关节炎多基因架构的贝叶斯推断分析,Nat Genet 44:4383–4391, 2012。
Yang Y, Muzny DM, Reid JG, et al: Clinical whole-exome sequencing for the diagnosis of mendelian disorders, NEJM 369:1502–1511, 2013.Yang Y, Muzny DM, Reid JG, 等:临床全外显子组测序用于孟德尔疾病诊断,NEJM 369:1502–1511, 2013。
Yuen RK, Merico D, Bookman M, et al: Whole genome sequencing resource identifies 18 new candidate genes for autism spectrum disorder, Nat Neurosci 20:602–611, 2017.Yuen RK, Merico D, Bookman M, 等:全基因组测序资源识别出18个自闭症谱系障碍的新候选基因,Nat Neurosci 20:602–611, 2017。
26/46
PROBLEMS 1.
Ch11 — Segment 26
PROBLEMS 1.问题1。
In the early days of gene mapping the Huntington disease (HD) locus was found to be tightly linked to a DNA common variant on chromosome 4.在基因定位的早期,亨廷顿病(HD)位点被发现与4号染色体上的一个常见DNA变异紧密连锁。
In the same study, however, linkage was ruled out between HD and the locus for the polymorphic MNSs blood group, which also maps to chromosome 4.然而,在同一研究中,HD与同样定位于4号染色体的多态性MNSs血型位点之间的连锁被排除。
What is the explanation?如何解释?
LOD scores (Z) between a common variant in the α-globin locus on the short arm of chromosome 16 and an autosomal dominant disease (e. g., polycystic kidney disease) was analyzed in a series of British and Dutch families, with the following data: θ 0. 00 0. 01 0. 10 0. 20 0. 30 0. 40 Z −∞ 23. 4 24. 6 19. 5 12. 85 5. 5 Z 25. 85 at 0. 05 max max How would you interpret these data?在一系列英国家庭和荷兰家庭中,分析了16号染色体短臂上α-珠蛋白位点的一个常见变异与一种常染色体显性遗传病(例如多囊肾病)之间的LOD分数(Z),数据如下:θ 0.00 0.01 0.10 0.20 0.30 0.40;Z −∞ 23.4 24.6 19.5 12.85 5.5;Z最大值25.85在θ=0.05处。你如何解释这些数据?
In a subsequent study, a large family from Sicily with what looks like the same disease was also investigated for linkage to α-globin, with the following results: θ 0. 00 0. 10 0. 20 0. 30 0. 40 LOD scores (Z) −∞ −8. 34 −3. 34 −1. 05 −0. 02 How would you interpret the data in this second study?在随后的一项研究中,一个来自西西里岛的看起来患有相同疾病的大规模家庭也接受了与α-珠蛋白的连锁分析,结果如下:θ 0.00 0.10 0.20 0.30 0.40;LOD分数(Z)−∞ −8.34 −3.34 −1.05 −0.02。你如何解释第二项研究中的数据?
This pedigree was obtained in a study designed to determine whether a pathogenic variant in one of the genes coding for a γ-crystallin protein, CRYGD, may be responsible for an autosomal dominant form of cataract.该家系是在一项旨在确定编码γ-晶状体蛋白的基因之一CRYGD中的致病变异是否可能导致常染色体显性遗传性白内障的研究中获得的。
The filled-in symbols in the pedigree indicate family members with cataracts.家系中的实心符号表示患有白内障的家庭成员。
The letters indicate three alleles at the CRYGD locus on chromosome 2.字母表示2号染色体上CRYGD位点的三个等位基因。
If you examine each affected person who has passed on the cataract to his or her children, how many of these represent a meiosis that is informative for linkage between the cataract and CRYGD?如果你检查每个将白内障遗传给子女的患病个体,其中有多少代表对白内障与CRYGD之间的连锁具有信息性的减数分裂?
In which individuals is the phase known between the cataract pathogenic variant and the CRYGD alleles?在哪些个体中,白内障致病变异与CRYGD等位基因之间的相是已知的?
Are there any meioses in which a crossover must have occurred to explain the data?是否存在必须发生交换才能解释数据的减数分裂?
What would you conclude about linkage between the cataract and the CRYGD gene from this study?从这项研究中,你对白内障与CRYGD基因之间的连锁会得出什么结论?
What additional studies might be performed to confirm or reject the hypothesis? continued AC AC BC BD BD AC AB BC Pedigree for question 3 I II III IV V AB AB BC CC AB AB AB AB AB BC BC BB 4.可以执行哪些额外的研究来确认或否定该假设?
Genome-wide sequencing has become an effective strategy for the discovery of genes causing rare mendelian disorders and refers to either whole exome (ES) or whole genome sequencing (GS). a.全基因组测序已成为发现导致罕见孟德尔疾病的基因的有效策略,其指的是全外显子组测序(ES)或全基因组测序(GS)。
Define whole exome sequencing and describe conceptually how it is performed? b.定义全外显子组测序并从概念上描述其执行方式?
What are the advantages and disadvantages of GS compared to ES?与ES相比,GS有哪些优点和缺点?
Review the pedigree in and the normal allele (H) with respect to variant alleles M and m in the mother of the two affected boys? h M h M H/h M/m h M H m回顾家系中两位患病男孩的母亲关于变异等位基因M和m的正常等位基因(H)?h M h M H/h M/m h M H m
27/46
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 229 PROBLEMS—CONT’D Pedigree of X-linked hemophilia.
Ch11 — Segment 27
IDENTIFYING THE GENETIC BASIS FOR HUMAN DISEASE 229 PROBLEMS—CONT’D Pedigree of X-linked hemophilia.识别人类疾病的遗传基础 229 问题——续 X连锁血友病的家系图。
The affected grandfather in the first generation has the disease (allele h) and allele M at a polymorphic locus on the X chromosome.第一代中患病的祖父患有该疾病(等位基因h),并且在X染色体上的一个多态性位点处具有等位基因M。
Relative risk calculations are used for cohort studies and not case-control studies.相对风险计算用于队列研究,而非病例对照研究。
To demonstrate why, imagine a case-control study for the effect of a genetic variant on disease susceptibility.为说明原因,设想一项关于遗传变异对疾病易感性影响的病例对照研究。
The investigator has ascertained as many affected individuals (a + c) as possible and then arbitrarily chooses a set of (b + d) controls.研究者已尽可能多地确定了患病个体(a + c),然后任意选择一组(b + d)对照。
They are genotyped as to whether a variant is present: a/(a + c) of the affected have the variant, whereas b/(b + d) of the controls have the variant.对其进行基因分型以确定是否存在变异:患病者中有a/(a + c)携带变异,而对照者中有b/(b + d)携带变异。
Disease Present Disease Absent Variant present a b Variant absent c d a+c b+d Calculate the odds ratio (OR) and relative risk (RR) for the association between the variant being present and the disease being present.疾病存在 疾病不存在 变异存在 a b 变异不存在 c d a+c b+d 计算变异存在与疾病存在之间关联的比值比(OR)和相对风险(RR)。
Now, imagine the investigator arbitrarily decided to use three times as many unaffected individuals, 3 × (b + d), as controls.现在,设想研究者任意决定使用三倍数量的未患病个体,即3 × (b + d),作为对照。
The investigator has every right to do so because it is a case-control study and the numbers of affected and unaffected are not determined by the prevalence of the disease in the population being studied, as they would be in a cohort study.研究者完全有权这样做,因为这是一项病例对照研究,患病与未患病个体的数量并非由所研究人群中疾病的患病率决定,而在队列研究中则会如此。
Assume the distribution of the variant remains the same in this control group as with the smaller control group—that is, 3b/[3 × (b + d)] = b/(b + d) carrying the allele.假设该对照组中变异的分布与较小对照组相同——即,携带该等位基因的比例为3b/[3 × (b + d)] = b/(b + d)。
Disease Present Disease Absent Variant present a 3b Variant absent c 3d a+c 3 × (b + d) Recalculate the OR and RR with this new control group.疾病存在 疾病不存在 变异存在 a 3b 变异不存在 c 3d a+c 3 × (b + d) 使用这个新对照组重新计算OR和RR。
Do the same when an arbitrary control group is an n-tuple of the original control group—that is, the size of the control group is n × (b + d).当任意对照组为原始对照组的n倍时——即对照组大小为n × (b + d)——做同样的计算。

Genetic Counseling and Risk Assessment

28/46
The Molecular Basis of Genetic Disease Gregory Costain GENERAL PRINCIPLES AND LESSONS FROM THE HEMOGLOBINOPATHIES The te…
Ch12 — Segment 28
The Molecular Basis of Genetic Disease Gregory Costain GENERAL PRINCIPLES AND LESSONS FROM THE HEMOGLOBINOPATHIES The term molecular disease, introduced in 1949, refers to disorders in which the primary disease-­causing event is an alteration, either inherited or acquired, affecting a gene(s), its structure, and/­or its expression.遗传疾病的分子基础 Gregory Costain 从血红蛋白病中得出的总体原则与教训 分子疾病这一术语于1949年提出,指原发性致病事件为影响一个或多个基因、其结构和/或表达的遗传性或获得性改变的疾病。
In this chapter we first outline the basic DNA variant types and associated mechanisms underlying monogenic (single-­gene) disorders.在本章中,我们首先概述单基因(单个基因)疾病的基础DNA变异类型及其相关机制。
We then illustrate their molecular and clinical consequences using inherited diseases of hemoglobin—­the ­hemoglobinopathies—­as examples.然后我们以遗传性血红蛋白疾病——血红蛋白病——为例,说明其分子和临床后果。
This overview of mechanisms is expanded on in Chapter 13 to include other genetic diseases that illustrate key principles of genetics in medicine.本机制概述将在第13章中进一步扩展,纳入其他遗传疾病,以说明医学遗传学中的关键原则。
A genetic disease occurs when an alteration in the DNA of an essential gene changes the amount or function, or both, of the gene products—­typically messenger RNA (mRNA) and protein, but occasionally specific noncoding RNAs (ncRNAs) with structural or regulatory functions.当必需基因DNA的改变改变了基因产物(通常是信使RNA和蛋白质,偶尔也包括具有结构或调节功能的特定非编码RNA)的数量或功能或两者时,就会发生遗传疾病。
Although almost all known single-­gene disorders result from variants that affect the function of a protein, a few exceptions to this generalization are now known.尽管几乎所有已知的单基因疾病都是由影响蛋白质功能的变异引起的,但目前已知这一概括存在少数例外。
These exceptions are diseases due to variants in ncRNA genes, including micro RNA (miRNA) genes that regulate specific target genes, and mitochondrial genes that encode transfer RNAs (tRNAs; see Chapter 13).这些例外是由非编码RNA基因变异引起的疾病,包括调节特定靶基因的微小RNA基因以及编码转移RNA的线粒体基因(见第13章)。
In this chapter we restrict our attention to diseases caused by defects in protein-­coding genes.在本章中,我们将注意力限制在由蛋白质编码基因缺陷引起的疾病上。
It is essential to understand genetic disease at the molecular level because this knowledge is the foundation of rational therapy.在分子水平上理解遗传疾病至关重要,因为这一知识是合理治疗的基础。
By 2022, the online version of Mendelian Inheritance in Man listed over 7000 phenotypes for which the molecular basis is known.截至2022年,《人类孟德尔遗传》在线版列出了7000多种已知分子基础的表型。
Although it is impressive that the basic molecular defect has been found in so many disorders, it is sobering to realize that the pathophysiology is not entirely understood for any genetic disease.尽管在如此多的疾病中发现基本分子缺陷令人印象深刻,但认识到任何遗传疾病的病理生理学尚未完全被理解,则令人清醒。
Sickle cell disease (Case 42), discussed later in this chapter, was the first disease to be characterized at the molecular level and remains among the best characterized of all inherited disorders; even here, knowledge is incomplete.本章稍后讨论的镰状细胞病(病例42)是首个在分子水平上被表征的疾病,并且至今仍是所有遗传性疾病中特征最明确的疾病之一;即使如此,其知识仍不完整。
Genetically informed therapies for hemoglobinopathies are now emerging as a realistic prospect in the clinic, thanks in part to an increasingly sophisticated understanding of the genetic pathomechanisms.针对血红蛋白病的基因知情治疗如今正成为临床中现实可行的发展方向,部分归功于对遗传病理机制日益深入的理解。
EFFECT OF PATHOGENIC VARIANTS ON PROTEIN FUNCTION DNA variants within protein-­coding genes have been primarily found to cause disease through one of four different effects on protein function .致病性变异对蛋白质功能的影响 蛋白质编码基因内的DNA变异主要通过四种不同的蛋白质功能效应导致疾病。
The most common effect by far is a loss of function of the protein.迄今为止最常见的效应是蛋白质功能丧失。
Many important conditions arise, however, from other mechanisms: a gain of function, the acquisition of a novel property by the affected protein, or the expression of a gene at the wrong time (heterochronic expression) and/­or in the wrong place (ectopic expression).然而,许多重要疾病的产生源于其他机制:功能获得、受影响的蛋白质获得新特性,或基因在错误时间(异时表达)和/或错误位置(异位表达)的表达。
Loss-­of-­Function Variants The loss of function of a gene may result from alteration of its coding, regulatory, or other critical sequences due to nucleotide substitutions, deletions, insertions, or rearrangements.功能丧失性变异 基因的功能丧失可能源于其编码序列、调控序列或其他关键序列的改变,这些改变由核苷酸替换、缺失、插入或重排引起。
A loss of function due to deletion, leading to a reduction in gene dosage, is exemplified by the α-­thalassemias (Case 44), which are most commonly due to deletion of α-­globin genes (see later discussion); by chromosome-­loss diseases (Case 27), such as monosomies like Turner syndrome (see Chapter 6) (Case 47); and by acquired somatic variants that occur in tumor-­suppressor genes in many cancers, such as retinoblastoma (Case 39) (see Chapter 16).由缺失导致的功能丧失引起基因剂量减少,其典型例子包括α地中海贫血(病例44)(最常见原因是α珠蛋白基因缺失,见后文讨论)、染色体缺失疾病(病例27)如特纳综合征等单体(见第6章)(病例47),以及许多癌症中肿瘤抑制基因发生的获得性体细胞变异,例如视网膜母细胞瘤(病例39)(见第16章)。
Many other types of variants can also lead to a complete loss of function, and all are illustrated by the β-­thalassemias (see later discussion), a group of hemoglobinopathies that result from a reduction in the abundance of β-­globin, one of the major adult hemoglobin proteins in red blood cells.许多其他类型的变异也可导致完全功能丧失,所有这类情况均以β地中海贫血(见后文讨论)为例说明,该病是一组因红细胞中主要成人血红蛋白蛋白之一β珠蛋白丰度降低而引起的血红蛋白病。
The severity of a disease due to loss-­of-­function variants generally correlates with the amount of function lost.由功能丧失性变异引起的疾病严重程度通常与功能丧失的量相关。
In many instances, the retention of even a small percent of residual function by the abnormal protein greatly reduces the severity of the disease.在许多情况下,异常蛋白质即使保留少量残余功能,也能大大减轻疾病的严重程度。
Gain-­of-­Function Variants Variants may also enhance one or more of the normal functions of a protein; in a biologic system, however, more is not necessarily better, and disease may result.功能获得性变异 变异也可能增强蛋白质的一种或多种正常功能;然而,在生物系统中,更多不一定更好,并可能导致疾病。
It chapter 12它 第12章
29/46
is critical to recognize when a disease is due to a gain-­of-­ function variant because the treatment must necessarily d…
Ch12 — Segment 29
is critical to recognize when a disease is due to a gain-­of-­ function variant because the treatment must necessarily differ from disorders due to other mechanisms, such as loss-­of-­function variants.认识到疾病是由功能获得性变异引起至关重要,因为其治疗必须与由其他机制(如功能丧失性变异)引起的疾病有所不同。
Gain-­of-­function variants fall into two broad classes: Variants that increase the production of a normal protein.功能获得性变异可分为两大类:增加正常蛋白质生成的变异。
Some variants cause disease by increasing the synthesis of a normal protein in cells in which the protein is normally present.某些变异通过增加正常蛋白质在通常表达该蛋白质的细胞中的合成而引起疾病。
The most common variants of this type are due to increased gene dosage, which generally results from duplication of part or all of a chromosome.此类最常见的变异是由于基因剂量增加所致,通常由部分或整条染色体的重复引起。
As discussed in Chapter 6, the classic example is trisomy 21 (Down syndrome), which is due to the presence of three copies of chromosome 21.正如第6章所述,经典例子是21三体综合征(唐氏综合征),由21号染色体存在三个拷贝引起。
Other important diseases arise from the increased dosage of single genes, including one form of familial Alzheimer disease due to a duplication of the amyloid precursor protein (β APP) gene (see Chapter 13), and the peripheral nerve degeneration Charcot-­Marie-­Tooth disease type 1 A (Case 8), which generally results from duplication of the gene for peripheral myelin protein 22 (PMP22).其他重要疾病由单个基因的剂量增加引起,包括因淀粉样前体蛋白(β APP)基因重复所致的一种家族性阿尔茨海默病(见第13章),以及通常由外周髓鞘蛋白22(PMP22)基因重复引起的周围神经变性病——Charcot-Marie-Tooth病1A型(病例8)。
Variants that enhance one normal function of a protein.增强蛋白质某一正常功能的变异。
Rarely, a variant in the coding region may increase the ability of each protein molecule to perform one or more of its normal functions, even though this increase is detrimental to the overall physiologic role of the protein.少数情况下,编码区变异可能会增强每个蛋白质分子执行其一项或多项正常功能的能力,即使这种增强对蛋白质的整体生理作用有害。
For example, the missense variant that creates hemoglobin Kempsey locks hemoglobin into its high oxygen affinity state, thereby reducing oxygen delivery to tissues.例如,产生血红蛋白Kempsey的错义变异将血红蛋白锁定在高氧亲和力状态,从而减少向组织输送氧气。
Another example of this mechanism is the missense variation in the FGFR3 gene that causes achondroplasia (Case 2), the most common skeletal dysplasia.该机制的另一个例子是FGFR3基因中的错义变异导致软骨发育不全(病例2),这是最常见的骨骼发育不良。
Novel Property Variants In a few diseases, a change in the amino acid sequence confers a novel property on the protein, without necessarily altering its normal functions.新特性变异:在少数疾病中,氨基酸序列的改变赋予蛋白质一种新特性,而不一定改变其正常功能。
The classic example of this mechanism is sickle cell disease (Case 42), which, as we will see later in this chapter, is due to an amino acid substitution that has no effect on the ability of sickle hemoglobin to transport oxygen.该机制的经典例子是镰状细胞病(病例42),正如我们将在本章后面看到的,该病是由一种氨基酸替换引起的,这种替换对镰状血红蛋白运输氧气的能力没有影响。
Rather, unlike normal Variants disrupting RNA stability or RNA splicing Variants affecting gene regulation or dosage Variants in coding region Decreased amount CAUSE OF DISEASE Loss of protein function (the great majority) Gain of function Novel property (infrequent) Ectopic or heterochronic expression (uncommon, except in cancer) Inappropriate expression (wrong time, place) HPFH Many oncogenes Increased amount Trisomies Charcot-Marie-Tooth disease type 1A -Thalassemias Monosomies Tumor-suppressor variants Hb Hammersmith -Thalassemias Hb Kempsey Achondroplasia Hb S MUTATION Protein abnormal (if unstable decreased amount) Protein structure normal Variants in the coding region result in structurally abnormal proteins that have a loss or gain of function or a novel property that causes disease.相反,不同于正常,存在破坏RNA稳定性或RNA剪接的变异、影响基因调控或剂量的变异、编码区变异(减少量)——疾病原因包括:蛋白质功能丧失(绝大多数)、功能获得、新特性(罕见)、异位或异时表达(罕见,除癌症外)、不当表达(错误时间、位置)、HPFH、许多癌基因、增加量包括:三体、Charcot-Marie-Tooth病1A型、α-地中海贫血、单体、肿瘤抑制变异、Hb Hammersmith、β-地中海贫血、Hb Kempsey、软骨发育不全、Hb S——突变:蛋白质异常(如不稳定则减少)、蛋白质结构正常——编码区变异导致结构异常的蛋白质,这些蛋白质具有功能丧失、功能获得或新特性,从而引起疾病。
Variants in noncoding sequences are of two general types: those that alter the stability or splicing of the messenger RNA (mRNA) and those that disrupt regulatory elements or change gene dosage.非编码序列变异通常分为两类:改变信使RNA(mRNA)稳定性或剪接的变异,以及破坏调控元件或改变基因剂量的变异。
Variants in regulatory elements alter the abundance of the mRNA or the time or cell type in which the gene is expressed.调控元件变异改变mRNA的丰度,或改变基因表达的时间或细胞类型。
Variants in either the coding region or regulatory domains can decrease the amount of the protein produced.编码区或调控结构域的变异均可减少所生成蛋白质的量。
HPFH, Hereditary persistence of fetal hemoglobin.HPFH,即遗传性胎儿血红蛋白持续存在症。
30/46
The Molecular Basis of Genetic Disease 233 hemoglobin, sickle hemoglobin chains aggregate when they are deoxygenated and…
Ch12 — Segment 30
The Molecular Basis of Genetic Disease 233 hemoglobin, sickle hemoglobin chains aggregate when they are deoxygenated and form abnormal polymeric fibers that deform red blood cells.遗传疾病的分子基础 233 血红蛋白,镰状血红蛋白链在脱氧时聚集并形成异常多聚纤维,使红细胞变形。
That novel property variants are infrequent is not surprising because most amino acid substitutions are either neutral or detrimental to the function or stability of a protein that has been finely tuned by evolution.新特性变异罕见并不令人惊讶,因为大多数氨基酸替换要么是中性的,要么对经过进化精细调谐的蛋白质的功能或稳定性有害。
Variants Associated With Heterochronic or Ectopic Gene Expression An important class of variants includes those that lead to inappropriate expression of the gene at an abnormal time or place.与异时性或异位基因表达相关的变异 一类重要的变异包括那些导致基因在异常时间或位置不适当表达的变异。
These variants occur in the regulatory regions of the gene.这些变异发生在基因的调控区域。
Cancer can be driven by expression of a gene that normally promotes cell ­proliferation—­ a proto-oncogene—­in cells in which the gene is not normally expressed (see Chapter 16).癌症可由一种通常促进细胞增殖的基因(原癌基因)在其通常不表达的细胞中表达驱动(见第16章)。
Some variants in hemoglobin regulatory elements lead to the continued expression in adults of the γ-­globin gene, which is normally expressed at high levels only in fetal life.血红蛋白调控元件中的一些变异导致γ-珠蛋白基因在成人中持续表达,该基因通常仅在胎儿期高水平表达。
Such γ-­globin gene variants cause a benign phenotype called hereditary persistence of fetal hemoglobin (Hb F), as we explore later in this chapter.此类γ-珠蛋白基因变异引起一种称为遗传性胎儿血红蛋白(Hb F)持续存在的良性表型,我们将在本章后面探讨。
HOW VARIANTS DISRUPT THE FORMATION OF BIOLOGICALLY NORMAL PROTEINS Disruptions of the normal functions of a protein that result from the different types of variants outlined earlier can be well exemplified by the broad range of diseases due to variants in the globin genes, as we will discuss in the second part of this chapter.变异如何破坏生物学正常蛋白质的形成 由前面概述的不同类型变异导致的蛋白质正常功能破坏,可以通过珠蛋白基因变异引起的多种疾病得到很好的例证,我们将在本章第二部分讨论。
To form a biologically active protein (such as the hemoglobin molecule), information must be transcribed from the nucleotide sequence of the gene to the mRNA and then translated into the polypeptide, which then undergoes progressive stages of maturation (see Chapter 3).要形成具有生物学活性的蛋白质(如血红蛋白分子),信息必须从基因的核苷酸序列转录到mRNA,然后翻译成多肽,随后多肽经历逐步的成熟阶段(见第3章)。
Variants can disrupt any of these steps ( As we shall see next, abnormalities in five of these stages are illustrated by various hemoglobinopathies; the others are exemplified by diseases to be presented in Chapter 13.变异可以破坏这些步骤中的任何一步(接下来我们将看到,其中五个阶段的异常由各种血红蛋白病说明;其他阶段的异常由第13章将介绍的疾病例证)。
THE RELATIONSHIP BETWEEN GENOTYPE AND PHENOTYPE IN GENETIC DISEASE Key molecular concepts that can account for differences in the observed clinical phenotype associated with a genetic disease are: Allelic heterogeneity Locus heterogeneity Effect of modifier genes Each of these concepts is illustrated by variants in the α-­globin and/­or β-­globin genes ( Allelic Heterogeneity Genetic heterogeneity is most commonly due to the presence of multiple alleles at a single locus, a situation referred to as allelic heterogeneity (see Chapter 7 and In some instances there may be a clear genotype-­phenotype correlation between a specific allele and a specific phenotype.遗传疾病中基因型与表型之间的关系 能够解释与遗传疾病相关的观察到的临床表型差异的关键分子概念包括:等位基因异质性、位点异质性、修饰基因效应。这些概念均通过α-珠蛋白和/或β-珠蛋白基因的变异得以阐明(等位基因异质性 遗传异质性最常见的原因是单个位点上存在多个等位基因,这种情况称为等位基因异质性(见第7章),在某些情况下,特定等位基因与特定表型之间可能存在明确的基因型-表型相关性。
The most common explanation for the effect of allelic heterogeneity on the clinical g., Hb Hammersmith) Posttranslational modification I-­cell disease, a lysosomal storage disease that is due to a failure to add a phosphate group to mannose residues of lysosomal enzymes.等位基因异质性对临床表型影响的最常见解释(例如,Hb Hammersmith)是翻译后修饰;I-细胞病是一种溶酶体贮积症,是由于未能将磷酸基团添加到溶酶体酶的甘露糖残基上所致。
The mannose 6-­phosphate residues are required to target the enzymes to lysosomes (see Chapter 13) Assembly of monomers into a holomeric protein Types of osteogenesis imperfecta in which an amino acid substitution in a procollagen chain impairs the assembly of a normal collagen triple helix (see Chapter 13) Subcellular localization of the polypeptide or the holomer Familial hypercholesterolemia variants (class 4), in the carboxyl terminus of the LDL receptor, that impair the localization of the receptor to clathrin-­coated pits, preventing the internalization of the receptor and its subsequent recycling to the cell surface (see Chapter 13) Cofactor or prosthetic group binding to the polypeptide Types of homocystinuria due to poor or absent binding of the cofactor (pyridoxal phosphate) to the cystathionine synthase apoenzyme (see Chapter 13) Function of a correctly folded, assembled, and localized protein produced in normal amounts Diseases in which the altered protein is mostly normal but one of its critical biologic activities is altered by an amino acid substitution (e. g., in Hb Kempsey, impaired subunit interaction locks hemoglobin into its high oxygen affinity state) LDL, Low-­density lipoprotein; mRNA, messenger RNA.甘露糖-6-磷酸残基是将酶靶向溶酶体所必需的(见第13章);单体组装成全蛋白:某些成骨不全症类型中,前胶原链的氨基酸替换损害正常胶原三螺旋的组装(见第13章);多肽或全蛋白的亚细胞定位:家族性高胆固醇血症变异(第4类),位于LDL受体羧基端,损害受体向网格蛋白包被小窝的定位,阻止受体内化及其随后再循环至细胞表面(见第13章);辅因子或辅基与多肽的结合:某些同型胱氨酸尿症类型是由于辅因子(吡哆醛磷酸)与胱硫醚合酶脱辅基酶结合不良或缺失所致(见第13章);正确折叠、组装和定位且以正常量产生的蛋白质的功能:某些疾病中,改变的蛋白质基本正常,但其关键生物活性之一因氨基酸替换而改变(例如,在Hb Kempsey中,亚基相互作用受损使血红蛋白锁定在高氧亲和状态);LDL,低密度脂蛋白;mRNA,信使RNA。
31/46
phenotype is that alleles that confer more residual function on the altered protein are often associated with a milder f…
Ch12 — Segment 31
phenotype is that alleles that confer more residual function on the altered protein are often associated with a milder form of the principal phenotype associated with the disease.phenotype is that...
In some instances, however, alleles that confer some residual protein functions are associated with only one or a subset of the phenotypes seen with a missing or completely nonfunctional allele (frequently termed a null allele).然而,在某些情况下,赋予蛋白质一定残余功能的等位基因仅与缺失或完全无功能等位基因(常称为零效等位基因)所表现出的表型中的一种或一个子集相关。
As we will explore more fully in Chapter 13, this situation prevails with certain variants of the cystic fibrosis gene, CFTR, that lead to a phenotypically different condition—congenital absence of the vas deferens, but not to the other manifestations of cystic fibrosis.正如我们将在第13章中更详细探讨的那样,这种情况出现在囊性纤维化基因CFTR的某些变异体中,这些变异体导致一种表型不同的疾病——先天性输精管缺失,但不引起囊性纤维化的其他表现。
An important exception to this rule relates to variants that act in a dominant-­negative fashion, as exemplified by select missense variants in the genes encoding components of type I collagen that result in a more severe form of osteogenesis imperfecta than with a null allele (see Chapter 13).该规则的一个重要例外涉及以显性负效应方式起作用的变异体,例如编码I型胶原蛋白组分的基因中的特定错义变异体,它们导致的成骨不全症比零效等位基因更严重(见第13章)。
A second explanation for allele-­based differences in phenotype is that a specific property of the protein may be more perturbed by a particular variant.等位基因间表型差异的第二种解释是,蛋白质的某个特定性质可能更容易受到特定变异体的干扰。
This situation is well illustrated by Hb Kempsey, a β-­globin allele that maintains the hemoglobin in a high oxygen affinity structure.这种情况在Hb Kempsey中得到了很好的说明,这是一种β-珠蛋白等位基因,使血红蛋白保持在高氧亲和力结构状态。
This causes polycythemia because the reduced peripheral delivery of oxygen is misinterpreted by the hematopoietic system as being due to an inadequate production of red blood cells.这导致红细胞增多症,因为外周氧气输送减少被造血系统误判为红细胞生成不足。
The consequences of a specific variant on the function of a protein can be unpredictable.特定变异体对蛋白质功能的影响可能是不可预测的。
No one would have foreseen that the β-­globin allele associated with sickle cell disease would lead to the formation of globin polymers that deform erythrocytes to a sickle cell shape (see later in this chapter).没有人会预见到与镰状细胞病相关的β-珠蛋白等位基因会导致珠蛋白聚合物形成,使红细胞变形为镰状细胞形状(见本章后续内容)。
However, sickle cell disease is also unusual in that it results only from a single specific variant—­the p.然而,镰状细胞病也有其特殊性,因为它仅由单一的特定变异体引起——即p.
Glu 6Val substitution in the β-­globin chain—­whereas most genetic diseases can arise from any of a number of different DNA-­level variants in the corresponding gene.Glu6Val在β-珠蛋白链中的取代——而大多数遗传病可由相应基因中多种不同的DNA水平变异体中的任何一种引起。
Locus Heterogeneity Genetic heterogeneity also arises when variants at more than one locus can result in a specific clinical condition—a situation termed locus heterogeneity (see Chapter 7).位点异质性:当多个位点的变异体都能导致特定的临床状况时,也会产生遗传异质性——这种情况称为位点异质性(见第7章)。这种现象通过以下发现得到说明:地中海贫血可由α-珠蛋白或β-珠蛋白链基因的变异体引起(见一旦位点异质性被记录,对每个基因相关表型的仔细比较有时会揭示表型并非最初认为的那样均一)。
This phenomenon is illustrated by the finding that thalassemia can result from variants in either the α-­globin or β-­globin chain genes (see Once locus heterogeneity has been documented, careful comparison of the phenotype associated with each gene sometimes reveals that the phenotype is not as homogeneous as initially believed.修饰基因:有时,即使是最稳健的基因型-表型关联,也可能在特定个体中不成立。
Modifier Genes Sometimes even the most robust genotype-­phenotype relationships are found not to hold for a specific individual.这种表型变异原则上可归因于非遗传因素(例如环境、随机因素)或其他基因的作用,这些基因被称为修饰基因(见第9章)。
Such phenotypic variation can, in principle, be ascribed to nongenetic (e. g., environmental, stochastic) factors or to the action of other genes, termed modifier genes (see Chapter 9).针对特定人类单基因疾病已鉴定的修饰基因数量正在增加;然而,具有临床相关效应大小或治疗意义的例子仍然相对较少。
Identified modifier genes for specific human monogenic disorders are growing in number; however, there remain relatively few examples with clinically relevant effect sizes or therapeutic significance.如本章后面所述,伴有α-珠蛋白基因座缺失的β-地中海贫血患者可能表现出较轻的表型。
As described later in this chapter, individuals with β-­thalassemia who also have a deletion at the α-­globin locus can have a less severe phenotype.人类血红蛋白及相关疾病:为了更详细地说明本章第一部分介绍的概念,我们现在转向血红蛋白疾病。
HUMAN HEMOGLOBIN AND ASSOCIATED DISEASES To illustrate in greater detail the concepts introduced in the first section of this chapter, we now turn to disorders of hemoglobin.这些血红蛋白病总体上是人类最常见的单基因疾病,也是全球发病率的主要贡献者。
These hemoglobinopathies are collectively the most common monogenic diseases in humans, and major contributors to global morbidity.世界卫生组织估计,全球超过5%的人口是导致临床上重要血红蛋白疾病的遗传变异体的杂合携带者。
The World Health Organization estimates that more than 5% of the world’s population are heterozygous carriers of genetic variants associated with clinically important disorders of hemoglobin.血红蛋白病之所以重要,还因为其分子和生化病理学比其他任何一组遗传病都更易理解。
Hemoglobinopathies are also important because their molecular and biochemical pathology is better understood than perhaps that of any other group of genetic diseases.事实上,我们对基因基本结构(第3章)的理解在很大程度上源于对这些典型的单基因血红蛋白疾病的研究。
Indeed, our understanding of the basic anatomy of a gene (Chapter 3) arose in large part from studying these prototypical monogenic disorders of hemoglobin.在深入讨论血红蛋白病之前,有必要简要介绍珠蛋白基因和血红蛋白生物学的正常方面。
Before the hemoglobinopathies are discussed in depth, it is important to briefly introduce the normal aspects of the globin genes and hemoglobin biology.血红蛋白的结构与功能:血红蛋白是脊椎动物红细胞中的氧载体。
Structure and Function of Hemoglobin Hemoglobin is the oxygen carrier in vertebrate red blood cells.每个血红蛋白分子由四个亚基组成:两条α-(或α样)珠蛋白链和两条β-(或β样)珠蛋白链。
Each hemoglobin molecule consists of four subunits: two α-­ (or α-­like) globin chains and two β-­ (or β-­like) globin chains.每个亚基由一个多肽链(珠蛋白)和一个辅基(血红素)组成。
Each subunit is composed of a polypeptide chain, globin, and a prosthetic group, heme.后者是一种含铁色素,与遗传疾病相关的异质性类型:位点异质性定义:一个位点上存在多个等位基因,例如β-地中海贫血;位点异质性:多个位点与一个临床表型相关,例如地中海贫血可由α-珠蛋白或β-珠蛋白基因的变异体引起;临床或表型异质性:单个位点的变异体与多种表型相关,例如镰状细胞病和β-地中海贫血分别由不同的β-珠蛋白基因变异体引起。
The latter is an iron-­containing pigment that combines with 2 Types of Heterogeneity Associated With Genetic Disease Type of Heterogeneity Definition Example From the Hemoglobinopathies Genetic Allelic heterogeneity The occurrence of more than one allele at a locus β-­Thalassemia Locus heterogeneity The association of more than one locus with a clinical phenotype Thalassemia can result from variants in either the α-­globin or β-­globin genes Clinical or phenotypic The association of more than one phenotype with variants at a single locus Sickle cell disease and β-­thalassemia each result from distinct β-­globin gene variants后者是一种含铁色素,与遗传疾病相关的两种异质性类型:异质性类型、定义、来自血红蛋白病的例子:遗传等位基因异质性:一个位点上存在多个等位基因,例如β-地中海贫血;位点异质性:多个位点与一个临床表型相关,例如地中海贫血可由α-珠蛋白或β-珠蛋白基因的变异体引起;临床或表型异质性:单个位点的变异体与多种表型相关,例如镰状细胞病和β-地中海贫血分别由不同的β-珠蛋白基因变异体引起。
32/46
The Molecular Basis of Genetic Disease 235 oxygen to give the molecule its oxygen-­transporting ability .
Ch12 — Segment 32
The Molecular Basis of Genetic Disease 235 oxygen to give the molecule its oxygen-­transporting ability .遗传病的分子基础 235 氧气赋予分子其携氧能力。
The predominant adult human hemoglobin, Hb A, has an α2β2 structure in which the four chains are folded and fit together to form a globular tetramer.成年人的主要血红蛋白Hb A具有α2β2结构,其中四条链折叠并组装成球状四聚体。
As with all proteins that have been strongly conserved throughout evolution, the tertiary structure of globins is constant; virtually all globins have seven or eight helical regions (depending on the chain) .与所有在进化中高度保守的蛋白质一样,珠蛋白的三级结构是恒定的;几乎所有的珠蛋白都有七个或八个螺旋区域(取决于链)。
Variants that disrupt this tertiary structure invariably have pathologic consequences.破坏这种三级结构的变体必然导致病理后果。
In addition, variants that substitute a highly conserved amino acid or that replace one of the nonpolar residues—which form the hydrophobic shell that excludes water from the interior of the molecule, are likely to cause a hemoglobinopathy .此外,替代高度保守氨基酸或替换非极性残基(这些残基形成疏水壳层,阻止水分子进入分子内部)的变体很可能引起血红蛋白病。
Like all proteins, globin has sensitive areas, in which variants cannot occur without affecting function, and insensitive areas, in which variations are more freely tolerated.与其他蛋白质一样,珠蛋白既有敏感区域(在此区域的变体无法不影响功能),也有不敏感区域(在此区域的变体更容易被耐受)。
The Globin Genes In addition to Hb A, with its α2β2 structure, there are five other normal human hemoglobins, each of which has a tetrameric structure like that of Hb A, consisting of two α or α-­like chains and two non-­α chains .珠蛋白基因 除了具有α2β2结构的Hb A外,还有五种其他正常的人类血红蛋白,每种都具有类似Hb A的四聚体结构,由两条α链或类α链和两条非α链组成。
The genes for the α and α-­like chains are clustered in a tandem arrangement on chromosome 16.α链和类α链的基因在16号染色体上串联排列成簇。
Note that there are two identical α-­globin genes, designated α1 and α2, on each homologue.注意,每个同源染色体上有两个相同的α-珠蛋白基因,分别命名为α1和α2。
The β-­ and β-­like globin genes, located on chromosome 11, are close family members that, as described in Chapter 3, undoubtedly arose from a common ancestral gene .位于11号染色体上的β链和类β珠蛋白基因是密切相关的家族成员,如第3章所述,它们无疑源自一个共同的祖先基因。
Illustrating this close evolutionary relationship, the β-­ and δ-­globins differ in only 10 of their 146 amino acids.这种密切的进化关系体现在β-珠蛋白和δ-珠蛋白在它们的146个氨基酸中只有10个不同。
Developmental Expression of Globin Genes and Globin Switching The expression of the various globin genes changes during development, a process referred to as globin switching .珠蛋白基因的发育表达与珠蛋白转换 各种珠蛋白基因的表达在发育过程中发生变化,这个过程称为珠蛋白转换。
Note that the genes in the α-­ and β-­globin clusters are arranged in the same transcriptional orientation and, remarkably, the genes in each cluster are situated in the same order in which they are expressed during development.注意,α-珠蛋白簇和β-珠蛋白簇中的基因以相同的转录方向排列,而且值得注意的是,每个簇中的基因按照它们在发育中表达的顺序排列。
The temporal switches of globin synthesis are accompanied by changes in the principal site of erythropoiesis .珠蛋白合成的时序性转换伴随着红细胞生成主要部位的变化。
The three embryonic globins are made in the yolk sac from the third to eighth weeks of gestation, but at approximately the fifth week, hematopoiesis begins to move from the yolk sac to the fetal liver.三种胚胎珠蛋白在妊娠第3至第8周由卵黄囊生成,但大约在第5周,造血作用开始从卵黄囊转移到胎儿肝脏。
Hb F (α2γ2), the predominant hemoglobin throughout fetal life, constitutes ~70% of total hemoglobin at birth.Hb F(α2γ2)是胎儿期的主要血红蛋白,在出生时约占总血红蛋白的70%。
In adults, however, Hb F represents only a few percent of the total hemoglobin, although this can vary from less than 1% to ~5% in different individuals. β-­chain synthesis becomes significant near the time of birth, and by 3 months of age almost all hemoglobin is of the adult form: Hb A (α2β2) .然而,在成人中,Hb F仅占总血红蛋白的百分之几,尽管不同个体中这一比例可从低于1%变化至约5%。β链合成在出生前后变得显著,到3个月大时,几乎所有的血红蛋白都是成人型:Hb A(α2β2)。
In diseases due to variants that decrease the abundance of β-­globin, such as β-­thalassemia (see later section), strategies to increase the normally small amount of γ-­globin (and therefore of Hb F [α2γ2]) produced in adults are proving to be successful in ameliorating the disorder (see Chapter 14).在因β-珠蛋白减少的变体(如β-地中海贫血,见后续章节)引起的疾病中,通过增加成人中通常少量产生的γ-珠蛋白(从而增加Hb F [α2γ2])的策略已被证明能有效改善病情(见第14章)。
The Developmental Regulation of β-­Globin Gene Expression: The Locus Control Region Elucidation of the mechanisms that control expression of the globin genes has provided generalizable insights into both normal and pathologic biologic processes.β-珠蛋白基因表达的发育调控:位点控制区 对控制珠蛋白基因表达机制的阐明,为正常和病理生物学过程提供了普适性的见解。
The expression of the β-­globin gene is only partly controlled by the promoter and two enhancers in the immediate flanking DNA (see Chapter 3).β-珠蛋白基因的表达仅部分由启动子和两侧DNA中的两个增强子控制(见第3章)。
A requirement for additional regulatory elements was first suggested by the identification of individuals who had no gene expression from any of the genes in the β-­globin cluster, even though the genes themselves (including their individual regulatory elements) were intact.对额外调控元件需求的首次提示来自对个体的鉴定,这些个体尽管β-珠蛋白簇中的所有基因(包括它们各自的调控元件)完整,但没有任何基因表达。
These informative patients were found to have large deletions upstream of the β-­globin complex that removed an ~20 ­kb domain—now called the locus control region (LCR), located ~6 kb upstream of the ε-­globin gene .这些具有信息量的患者被发现携带β-珠蛋白复合体上游的大片段缺失,该缺失移除了一段约20 kb的区域——现在称为位点控制区(LCR),位于ε-珠蛋白基因上游约6 kb处。
The resulting disease, εγδβ-­ thalassemia, is described later in this chapter.由此导致的疾病——εγδβ-地中海贫血,将在本章后续部分描述。
These cases show us that the LCR is required for the expression of all genes in the β-­globin cluster.这些病例向我们表明,LCR是β-珠蛋白簇中所有基因表达所必需的。
The LCR is defined by five DNase I hypersensitive sites : genomic regions that are unusually open to certain proteins (including the enzyme DNase I) used experimentally to reveal potential regulatory sites.LCR由五个DNase I超敏位点定义:这些基因组区域对某些蛋白质(包括实验上用于揭示潜在调控位点的酶DNase I)异常开放。
Within the context of the epigenetic packaging of chromatin (see Chapter 3), these sites maintain an open chromatin Heme E B G b His 92 b Phe 42 H F Helix A D C Each subunit has eight helical regions, designated A to H.在染色质的表观遗传包装背景下(见第3章),这些位点维持开放的染色质:血红素 E B G b His 92 b Phe 42 H F 螺旋 A D C 每个亚基有八个螺旋区域,命名为A至H。
The two most conserved amino acids are shown: p.显示了两个最保守的氨基酸:p.
His 92, the histidine to which the iron of heme is covalently linked; and p.His 92,即与血红素铁共价连接的组氨酸;以及p.
Phe 42, the phenylalanine that wedges the porphyrin ring of heme into the heme “pocket” of the folded protein.Phe 42,即将血红素卟啉环楔入折叠蛋白质血红素“口袋”中的苯丙氨酸。
See discussion of Hb Hammersmith and Hb Hyde Park, which have substitutions for p.参见关于Hb Hammersmith和Hb Hyde Park的讨论,它们分别具有β-珠蛋白分子中p.
Phe 42 and p.Phe 42和p.
His 92, respectively, in the β-­globin molecule. .His 92的替代。
33/46
Developmental period Embryonic Fetal Adult Hemoglobins Hb Gower 2 α2ε2 Hb F α2γ2 β-like genes α-like genes Hb Gower 1 ζ2…
Ch12 — Segment 33
Developmental period Embryonic Fetal Adult Hemoglobins Hb Gower 2 α2ε2 Hb F α2γ2 β-like genes α-like genes Hb Gower 1 ζ2ε2 Hb A2 α2δ2 Hb A α2β2 Birth 5' 5' 3' 3' β Gγ α2 α1 ζ ε Aγ δ Hb Portland ζ2γ2 48 42 36 30 24 18 12 6 Postnatal age (weeks) Birth 36 30 24 18 12 6 Gestational age (weeks) 10 20 30 40 50 Site of erythropoiesis Percentage of total globin synthesis Liver Spleen Bone marrow α β α γ β ε ζ γ δ Yolk sac A B The α-­like genes are on chromosome 16, the β-­like genes on chromosome 11.发育阶段:胚胎期、胎儿期、成人期;血红蛋白:Hb Gower 2(α2ε2)、Hb F(α2γ2)、β样基因、α样基因、Hb Gower 1(ζ2ε2)、Hb A2(α2δ2)、Hb A(α2β2);出生;5' 5' 3' 3';β、Gγ、α2、α1、ζ、ε、Aγ、δ;Hb Portland(ζ2γ2);出生后周龄48、42、36、30、24、18、12、6周;胎龄周数36、30、24、18、12、6周;红细胞生成部位;总珠蛋白合成百分比;肝脏、脾脏、骨髓;α、β、α、γ、β、ε、ζ、γ、δ;卵黄囊;A和B;α样基因位于16号染色体,β样基因位于11号染色体。
The curved arrows refer to the switches in gene expression during development.曲线箭头表示发育过程中基因表达的转换。
(B) Development of erythropoiesis in the human fetus and infant.(B) 人类胎儿和婴儿的红细胞生成发育。
Types of cells responsible for hemoglobin synthesis, organs involved, and types of globin chain synthesized at successive stages are shown.展示了负责血红蛋白合成的细胞类型、涉及的器官以及连续阶段合成的珠蛋白链类型。
(A, Redrawn from Stamatoyannopoulos G, Nienhuis AW: Hemoglobin switching.(A,重绘自Stamatoyannopoulos G, Nienhuis AW: Hemoglobin switching. 见Stamatoyannopoulos G, Nienhuis AW, Leder P, 等编: The molecular basis of blood diseases, 费城, 1987, WB Saunders; B,重绘自Wood WG: Haemoglobin synthesis during fetal development, Br Med Bull 32:282–287, 1976.)
In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders; B, redrawn from Wood WG: Haemoglobin synthesis during fetal development, Br Med Bull 32:282–­287, 1976.) Normal 10 kb 4321 5 LCR Gγ Aγ ψβ δ β ε Gγ Aγ ψβ δ β ε Hispanic εγδβthalassemia Deletion 10 kb Each of the five regions of open chromatin (arrows) contains several consensus binding sites for both erythroid-­specific and ubiquitous transcription factors.正常10 kb 4321 5 LCR Gγ Aγ ψβ δ β ε Gγ Aγ ψβ δ β ε 西班牙裔 εγδβ地中海贫血 缺失10 kb 五个开放染色质区域(箭头)各自包含多个红系特异性和普遍性转录因子的共有结合位点。
The precise mechanism by which the LCR regulates gene expression is unknown.LCR调节基因表达的具体机制尚不清楚。
Also shown is a deletion of the LCR that has led to εγδβ-­thalassemia, which is discussed in the text.图中还显示了导致εγδβ-地中海贫血的LCR缺失,这在正文中讨论。
(Redrawn from Kazazian Jr HH, Antonarakis S: Molecular genetics of the globin genes.(重绘自Kazazian Jr HH, Antonarakis S: Molecular genetics of the globin genes.
In Singer M, Berg P, editors: Exploring genetic mechanisms, Sausalito, 1997, University Science Books.)见Singer M, Berg P 编: Exploring genetic mechanisms, Sausalito, 1997, University Science Books.)
34/46
The Molecular Basis of Genetic Disease 237 configuration that gives transcription factors access to the regulatory eleme…
Ch12 — Segment 34
The Molecular Basis of Genetic Disease 237 configuration that gives transcription factors access to the regulatory elements that mediate the expression of each of the β-­globin genes in erythroid cells (see Chapter 3).遗传病的分子基础237构型使转录因子能够接触调节元件,这些元件介导红细胞中每个β-珠蛋白基因的表达(见第3章)。
The LCR, along with its associated DNA-­binding proteins, interacts with the genes of the β-­globin locus to form a nuclear domain called the active chromatin hub, where β-­globin gene expression takes place.LCR与其相关的DNA结合蛋白相互作用,与β-珠蛋白基因座的基因形成称为活性染色质中心的核域,β-珠蛋白基因的表达在此发生。
The sequential switching of gene expression that occurs among the five members of the β-­globin gene complex during development results from the sequential association of the active chromatin hub with the different genes in the cluster, as the hub moves from the most proximal gene in the complex (the ε-­globin gene in embryos) to the most distal (the δ-­ and β-­globin genes in adults).发育过程中β-珠蛋白基因复合体五个成员间发生的基因表达顺序性切换,是由于活性染色质中心依次与簇内不同基因结合,该中心从复合体中最靠近的基因(胚胎中的ε-珠蛋白基因)移动到最远的基因(成人的δ-和β-珠蛋白基因)。
The clinical significance of the LCR could extend beyond those individuals with deletions of the LCR who fail to express the genes of the β-­globin cluster.LCR的临床意义可能超出那些LCR缺失且无法表达β-珠蛋白簇基因的个体。
Components of the LCR may prove relevant to gene therapy (see Chapter 14) for disorders of the β-­globin cluster, wherein a goal is for the therapeutic normal copy of the gene in question to be expressed at the correct time in life and in the appropriate tissue.LCR的组成部分可能被证明与β-珠蛋白簇疾病的基因治疗有关(见第14章),其中目标是在生命正确的时间和在适当的组织中表达相关基因的治疗性正常拷贝。
Knowledge of the molecular mechanisms that underlie globin switching may also make it feasible to up-­regulate the expression of the γ-­globin gene in those with β-­thalassemia (who have variants only in the β-­globin gene) because Hb F (α2γ2) is an effective oxygen carrier in adults who lack Hb A (α2β2) (see Chapter 14).对珠蛋白转换分子机制的了解也可能使得上调β-地中海贫血患者(仅在β-珠蛋白基因中有变异)中γ-珠蛋白基因的表达成为可能,因为Hb F(α2γ2)在缺乏Hb A(α2β2)的成人中是一种有效的氧载体(见第14章)。
Gene Dosage, Developmental Expression of the Globins, and Clinical Disease The differences, both in the gene dosage of the α-­ and β-­globins (four α-­globin and two β-­globin genes per diploid genome) and in their patterns of expression during development, are important to an understanding of the pathogenesis of many hemoglobinopathies.基因剂量、珠蛋白的发育表达与临床疾病:α-和β-珠蛋白的基因剂量(每个二倍体基因组有四个α-珠蛋白基因和两个β-珠蛋白基因)及其在发育过程中的表达模式的差异,对于理解许多血红蛋白病的发病机制很重要。
A variant in a β-­globin gene affects 50% of the β chains, whereas a single α-­chain variant affects only 25% of the α chains. β-­globin variants have no prenatal consequences because γ-­globin is the major β-­like globin before birth, with Hb F constituting 75% of the total hemoglobin at term .β-珠蛋白基因的变异影响50%的β链,而单个α链变异仅影响25%的α链。β-珠蛋白变异没有产前后果,因为γ-珠蛋白是出生前主要的β样珠蛋白,足月时Hb F占总血红蛋白的75%。
In contrast, because α chains are the only α-­like components of hemoglobin 6 weeks after conception, α-­globin variants cause severe disease in both fetal and postnatal life.相比之下,由于α链是受孕后6周血红蛋白中唯一的α样成分,α-珠蛋白变异在胎儿期和出生后都会导致严重疾病。
THE HEMOGLOBINOPATHIES Hereditary disorders of hemoglobin can be divided into the following three broad groups, which, in some rare instances, overlap: Structural alterations of the amino acid sequence of the globin polypeptide, altering properties such as its ability to transport oxygen, or reducing its stability.血红蛋白病:遗传性血红蛋白疾病可分为以下三大类,在少数罕见情况下会有重叠:珠蛋白多肽氨基酸序列的结构改变,改变其特性(如运输氧气的能力)或降低其稳定性。
An example is sickle cell disease (Case 42) due to a missense variant that makes deoxygenated β-­globin relatively insoluble, changing the shape of the red cell .一个例子是镰状细胞病(病例42),由错义变异引起,使脱氧β-珠蛋白相对不溶,改变红细胞的形状。
Thalassemias, which are diseases that result from the decreased abundance of one or more of the globin chains (Case 44).地中海贫血,是由于一种或多种珠蛋白链丰度降低导致的疾病(病例44)。
The decrease can result from reduced production of a globin chain or, less commonly, from a variant that destabilizes the chain.这种降低可能源于珠蛋白链生成减少,或较少见的源于使链不稳定的变异。
The resulting imbalance in the ratio of the α:β chains underlies the pathophysiology of these conditions.由此产生的α:β链比例失衡是这些疾病病理生理学的基础。
Examples include promoter variants that decrease expression of the β-­globin mRNA to cause β-­thalassemia.例子包括启动子变异降低β-珠蛋白mRNA的表达,导致β-地中海贫血。
Hereditary persistence of fetal hemoglobin, a group of clinically benign conditions that impair the perinatal switch from γ-­globin to β-­globin synthesis.遗传性胎儿血红蛋白持续存在,是一组临床良性疾病,损害围产期从γ-珠蛋白到β-珠蛋白合成的转换。
An example of a causal variant is a deletion that removes both the δ-­ and β-­globin genes but leads to continued postnatal expression of the γ-­globin genes, to produce Hb F, which is an effective oxygen transporter .一种致病变异的例子是缺失同时去除δ-和β-珠蛋白基因,但导致γ-珠蛋白基因在出生后持续表达,产生Hb F,这是一种有效的氧转运蛋白。
Hemoglobin Structural Alterations Most variant hemoglobins result from single nucleotide variants in one of the globin genes.血红蛋白结构改变:大多数变异血红蛋白源于一个珠蛋白基因中的单核苷酸变异。
More than 500 abnormal hemoglobins have been described, and approximately half of these are clinically significant.已描述超过500种异常血红蛋白,其中约一半具有临床意义。
The hemoglobin structural alterations can be separated B A Oxygenated cells are round and full.血红蛋白结构改变可分为B A氧合细胞呈圆形且饱满。
(B) The classic sickle cell shape is produced only when the cells are in the deoxygenated state.(B) 经典镰状细胞形状仅在细胞处于脱氧状态时产生。
(From Kaul DK, Fabry ME, Windisch P, et al: Erythrocytes in sickle cell anemia are heterogeneous in their rheological and hemodynamic characteristics, J Clin Invest 72:22, 1983.)(摘自Kaul DK, Fabry ME, Windisch P, 等:镰状细胞贫血中的红细胞在其流变学和血流动力学特征上具有异质性,J Clin Invest 72:22, 1983。)
35/46
into the following three classes, depending on the clinical phenotype ( Alterations with modified oxygen transport, due …
Ch12 — Segment 35
into the following three classes, depending on the clinical phenotype ( Alterations with modified oxygen transport, due to increased or decreased oxygen affinity or to the formation of methemoglobin—a form of globin incapable of reversible oxygenation.根据临床表型分为以下三类:因氧亲和力升高或降低或形成高铁血红蛋白(一种无法可逆结合氧的血红蛋白形式)而导致氧转运改变的变异。
Alterations due to variants in the coding region that cause thalassemia because they reduce the abundance of a globin polypeptide.由于编码区变异导致珠蛋白多肽丰度降低而引起地中海贫血的改变。
Most of these variants impair the rate of synthesis of the mRNA or otherwise affect the level of the encoded protein.大多数此类变异损害mRNA的合成速率或以其他方式影响编码蛋白的水平。
Hemolytic Anemias Hemoglobins With Novel Physical Properties: Sickle Cell Disease.溶血性贫血 具有新物理性质的血红蛋白:镰状细胞病。
Sickle cell hemoglobin is of great clinical importance in many parts of the world, affecting millions.镰状细胞血红蛋白在世界许多地区具有重要的临床意义,影响数百万人口。
The causal variant is a single nucleotide substitution that changes the codon of the sixth amino acid of β-­globin from glutamic acid to valine (GAG → GTG: p.致病性变异是一个单核苷酸替换,将β-珠蛋白第六位氨基酸的密码子从谷氨酸变为缬氨酸(GAG → GTG: p.
Glu 6Val) (see Homozygosity for this variant is the cause of sickle cell disease (Case 42).Glu6Val)(参见该变异的纯合性是镰状细胞病的原因(病例42)。
The disease has a characteristic geographic distribution: occurring most frequently in equatorial Africa and less commonly in the Mediterranean area, India, Spanish-­ speaking regions in the Western Hemisphere, or in countries to which people from these regions have migrated.该疾病具有特征性地理分布:最常见于赤道非洲,较少见于地中海地区、印度、西半球西班牙语地区,或来自这些地区的人群移民到的国家。
Approximately 1 in 400 Black persons in the United States is born with sickle cell disease.在美国,大约每400名黑人中就有1名出生时患有镰状细胞病。
Clinical Features.临床特征。
Sickle cell disease is a severe autosomal recessive hemolytic condition characterized by a tendency of the red blood cells to become grossly abnormal in shape (i. e., take on a sickle shape) under conditions of low oxygen tension .镰状细胞病是一种严重的常染色体隐性溶血性疾病,其特征是在低氧张力条件下红细胞形态趋于严重异常(即呈现镰刀形状)。
Heterozygotes—who are said to have sickle cell trait, are, generally, clinically unaffected, but their red cells can sickle when subjected to very low oxygen pressure.杂合子——即具有镰状细胞特征者,通常无临床症状,但其红细胞在极低氧压条件下可发生镰变。
Occasions when this occurs are uncommon, although heterozygotes appear to be at risk for splenic infarction, especially at high altitude (e. g., in airplanes with reduced cabin pressure) or when exerting themselves to extreme levels in athletic competition.这种情况发生的场合不常见,但杂合子似乎有脾梗死的风险,尤其是在高海拔地区(例如机舱压力降低的飞机)或在竞技运动中过度用力时。
The heterozygous state is present in ~8% of Black individuals in the United States, but in areas where the sickle cell allele (β S) frequency is high (e. g., West Central Africa), up to 25% of the newborn population is heterozygous for the allele.杂合状态在美国约8%的黑人个体中存在,但在镰状细胞等位基因(β S)频率高的地区(例如中西部非洲),高达25%的新生儿是该等位基因的杂合子。
The Molecular Pathology of Hb S.Hb S的分子病理学。
In the 1950s, Vernon Ingram discovered that the abnormality in sickle cell hemoglobin was a replacement of one of the 146 amino acids in the β chain of the hemoglobin molecule.在20世纪50年代,Vernon Ingram发现镰状细胞血红蛋白的异常是血红蛋白分子β链中146个氨基酸中的一个被替换。
All the clinical manifestations of sickle cell hemoglobin are consequences of this single change in the β-­globin gene.镰状细胞血红蛋白的所有临床表现都是β-珠蛋白基因中这一单一改变的结果。
Ingram’s discovery was the first demonstration in any organism that a variant in a structural gene could cause an amino acid substitution in the corresponding protein.Ingram的发现首次在任何生物体中证明了结构基因中的变异可导致相应蛋白质中的氨基酸替换。
Because the substitution is in the β-­globin chain, the formula for sickle cell hemoglobin is written as α2β2 S or, more precisely, α2Aβ2 S.由于替换发生在β-珠蛋白链,镰状细胞血红蛋白的分子式写为α2β2 S,或更精确地,α2Aβ2 S。
A heterozygote has a mixture of the two types of hemoglobin, A and S, summarized as α2Aβ2A/­α2Aβ2 S, as well as a hybrid hemoglobin tetramer, written as α2Aβ Aβ S.杂合子具有两种血红蛋白A和S的混合物,总结为α2Aβ2A/α2Aβ2 S,以及一个杂交血红蛋白四聚体,写为α2AβAβS。
Strong evidence indicates that the sickle cell variant arose in West Africa, but that it occurred independently elsewhere.有力证据表明,镰状细胞变异起源于西非,但在其他地区独立发生。
The β S allele has attained high frequency in malaria endemic areas of the world because it confers protection against malaria in heterozygotes (see Chapter 10).β S等位基因在世界疟疾流行地区达到高频率,因为它在杂合子中提供抗疟疾保护(见第10章)。
Sickling and Its Consequences.镰变及其后果。
The molecular and cellular pathology of sickle cell disease is summarized in , but in deoxygenated blood they are only one-­fifth as soluble as normal hemoglobin.镰状细胞病的分子和细胞病理学总结于,但在脱氧血液中,其溶解度仅为正常血红蛋白的五分之一。
Under conditions of low oxygen tension, this relative insolubility of deoxyhemoglobin S causes the sickle hemoglobin molecules to aggregate in the form of rod-­shaped polymers or fibers .在低氧张力条件下,脱氧血红蛋白S的这种相对不溶性导致镰状血红蛋白分子聚集形成棒状聚合物或纤维。
These molecular rods distort the α2β2 S erythrocytes to a sickle shape that prevents them from squeezing single file through capillaries—as do normal red cells, thereby blocking blood flow and causing local ischemia.这些分子棒将α2β2 S红细胞扭曲成镰刀形状,阻止它们像正常红细胞那样单列挤过毛细血管,从而阻塞血流并引起局部缺血。
They may also cause disruption of the red cell membrane Glu 6Val Deoxygenated Hb S polymerizes → sickle cells → vascular occlusion and hemolysis AR Hb Hammersmith β chain: p.它们还可能引起红细胞膜的破坏 Glu6Val 脱氧Hb S聚合 → 镰状细胞 → 血管闭塞和溶血 AR Hb Hammersmith β链: p.
Phe 42Ser An unstable Hb → Hb precipitation → hemolysis; also low oxygen affinity AD Hb M-­Hyde Park β chain: p.Phe42Ser 不稳定的Hb → Hb沉淀 → 溶血;同时低氧亲和力 AD Hb M-Hyde Park β链: p.
His 92Tyr The substitution makes oxidized heme iron resistant to methemoglobin reductase → Hb M, which cannot carry oxygen → cyanosis (asymptomatic) AD Hb Kempsey β chain: p.His92Tyr 该替换使氧化型血红素铁对高铁血红蛋白还原酶产生抗性 → Hb M无法携带氧 → 发绀(无症状) AD Hb Kempsey β链: p.
Asp 99Asn The substitution keeps the Hb in its high oxygen affinity structure → less oxygen to tissues → polycythemia AD Hb E β chain: p.Asp99Asn 该替换使Hb保持高氧亲和力结构 → 组织获氧减少 → 红细胞增多症 AD Hb E β链: p.
Glu 26Lys The variant → an abnormal Hb and decreased synthesis (abnormal RNA splicing) → mild thalassemiab AR a Hemoglobin variants are often named after a location related to the first described patient(s). b Additional β-­chain gene variants that cause β-­thalassemia are depicted in AD, Autosomal dominant; AR, autosomal recessive; Hb M, methemoglobin (See text.)Glu26Lys 该变异 → 异常Hb和合成减少(异常RNA剪接) → 轻度地中海贫血 AR a 血红蛋白变异常以与首次描述患者相关的地点命名。b 导致β-地中海贫血的其他β链基因变异描述于 AD,常染色体显性;AR,常染色体隐性;Hb M,高铁血红蛋白(见正文)。
36/46
The Molecular Basis of Genetic Disease 239 (hemolysis) and release of free hemoglobin, which can have deleterious effect…
Ch12 — Segment 36
The Molecular Basis of Genetic Disease 239 (hemolysis) and release of free hemoglobin, which can have deleterious effects on the availability of vasodilators, such as nitric oxide, thereby exacerbating the ischemia.遗传病的分子基础239(溶血)和游离血红蛋白的释放,这可能对血管扩张剂(如一氧化氮)的可用性产生有害影响,从而加剧缺血。
Modifier Genes Determine the Clinical Severity of Sickle Cell Disease.修饰基因决定镰状细胞病的临床严重程度。
It has long been known that a strong modifier of the clinical severity of sickle cell disease is the patient’s level of Hb F (α2γ2), higher levels being associated with less morbidity and lower mortality.长期以来已知,镰状细胞病临床严重程度的一个强修饰因子是患者的Hb F(α2γ2)水平,较高水平与较低的发病率和死亡率相关。
The physiologic basis of the ameliorating effect of Hb F is clear: Hb F is a perfectly adequate oxygen carrier in postnatal life and inhibits the polymerization of deoxyhemoglobin S.Hb F改善效应的生理基础是明确的:Hb F在出生后是一种完全充足的氧载体,并抑制脱氧血红蛋白S的聚合。
Until recently, however, it was not certain whether the variation in Hb F expression was heritable.然而,直到最近,尚不确定Hb F表达的变异是否具有遗传性。
Genome-­ wide association studies (GWAS) (see Chapter 11) have demonstrated that single nucleotide variants (SNPs) at three polymorphic loci (SNPs)—­the γ-­globin gene and two genes that encode transcription factors, BCL11A and MYB—­account for 40 to 50% of the variation in the levels of Hb F in individuals with sickle cell disease.全基因组关联研究(GWAS)(见第11章)已证明,三个多态位点(SNP)——γ-珠蛋白基因和两个编码转录因子的基因BCL11A和MYB——上的单核苷酸变异(SNP)占据了镰状细胞病患者Hb F水平变异的40%至50%。
Moreover, the Hb F–­associated SNPs are associated with the painful clinical episodes thought to be due to capillary occlusion caused by sickled red cells .此外,与Hb F相关的SNP与被认为由镰状红细胞引起的毛细血管闭塞所致的疼痛临床发作相关。
Individuals with heterozygous loss-offunction variants in BCL11A (gene) have a rare neurogenetic disorder but also hereditary persistence of fetal hemoglobin.携带BCL11A(基因)杂合功能丧失变异的个体患有一种罕见的神经遗传性疾病,但也伴有遗传性胎儿血红蛋白持续存在。
The genetically driven variations in the level of Hb F are also associated with variation in the clinical severity of β-­thalassemia (discussed later) because the reduced abundance of β-­globin (and thus of Hb A [α2β2]) in that disease is partly alleviated by higher levels of γ-­globin and, thus, of Hb F (α2γ2).基因驱动的Hb F水平变异也与β-地中海贫血(将在后面讨论)的临床严重程度变异相关,因为该疾病中β-珠蛋白(从而Hb A [α2β2])丰度降低在一定程度上被较高水平的γ-珠蛋白和因此的Hb F (α2γ2)所缓解。
The discovery of these genetic modifiers of Hb F abundance not only explains much of the variation in the clinical severity of sickle cell disease and β-­thalassemia, but highlights a general principle introduced in Chapter 9: modifier genes can play a major role in determining the clinical and physiologic severity of a single-­gene disorder.这些Hb F丰度遗传修饰因子的发现不仅解释了镰状细胞病和β-地中海贫血临床严重程度的大部分变异,而且强调了第9章介绍的一个一般原则:修饰基因在决定单基因疾病的临床和生理严重程度中可发挥主要作用。
BCL11A, a Silencer of γ-­Globin Gene Expression in Adult Erythroid Cells.BCL11A,成人红系细胞中γ-珠蛋白基因表达的沉默因子。
The identification of genetic modifiers of Hb F levels, particularly BCL11A, has opened great therapeutic potential.鉴定Hb F水平的遗传修饰因子,特别是BCL11A,开辟了巨大的治疗潜力。
The product of the BCL11A gene is a transcription factor that normally silences γ-­globin expression, thus shutting down Hb F production postnatally.BCL11A基因的产物是一种转录因子,通常沉默γ-珠蛋白表达,从而在出生后关闭Hb F的产生。
Accordingly, drugs that suppress BCL11A activity postnatally, thereby increasing the expression of Hb F, might be of great benefit to those with sickle cell disease and β-­thalassemia (see Chapter 14).因此,在出生后抑制BCL11A活性从而增加Hb F表达的药物,可能对镰状细胞病和β-地中海贫血患者大有裨益(见第14章)。
In addition, preliminary clinical trial data suggest that post-transcriptional genetic silencing of BCL11A may be an effective treatment for sickle cell disease.此外,初步临床试验数据表明,BCL11A的转录后基因沉默可能是镰状细胞病的有效治疗方法。
Trisomy 13, Micro RNAs, and MYB—Another Silencer of γ-­Globin Gene Expression.13三体、微小RNA和MYB——γ-珠蛋白基因表达的另一个沉默因子。
The indication from GWAS that MYB is an important regulator of γ-­globin expression has received further support from an unexpected direction: studies investigating the basis for the persistent increased postnatal expression of Hb F that is observed in individuals with trisomy 13 (see Chapter 6).GWAS提示MYB是γ-珠蛋白表达的重要调节因子,这一提示得到了来自一个意想不到方向的支持:研究调查了在13三体个体中观察到的出生后Hb F持续增加表达的基础(见第6章)。
Two miRNAs, mi R-­15a and mi R-­ 16-­1, directly target the 3′ untranslated region (UTR) of the MYB mRNA, thereby reducing MYB expression.两种miRNA,miR-15a和miR-16-1,直接靶向MYB mRNA的3'非翻译区(UTR),从而降低MYB表达。
The genes for these two miRNAs are located on chromosome 13; their extra dosage in trisomy 13 is predicted to reduce MYB expression to below normal levels, thereby partly relaxing the postnatal suppression of γ-­globin gene expression normally mediated by the MYB protein.这两种miRNA的基因位于13号染色体上;它们在13三体中的额外剂量预计会将MYB表达降低至正常水平以下,从而部分解除通常由MYB蛋白介导的出生后γ-珠蛋白基因表达的抑制。
This leads to increased expression of Hb F .这导致Hb F表达增加。
Unstable Hemoglobins.不稳定性血红蛋白。
The unstable hemoglobins are due largely to single nucleotide variants that cause denaturation of the hemoglobin tetramer in mature red blood cells.不稳定性血红蛋白主要是由导致成熟红细胞中血红蛋白四聚体变性的单核苷酸变异引起的。
The denatured globin tetramers are insoluble and precipitate to form inclusions (Heinz bodies) that damage the red cell membrane and cause hemolysis of mature red blood cells in the vascular tree .变性的珠蛋白四聚体不溶并沉淀形成包涵体(海因茨小体),损伤红细胞膜,导致血管树中成熟红细胞的溶血。
Normal codon Sickle cell codon GAG GTG Amino acid substitution β6 Glu Val Hb S Cell heterogeneity Oxy Deoxy Hb S solution Hb S fiber Vaso-occlusion正常密码子 镰状细胞密码子 GAG GTG 氨基酸置换 β6 Glu Val Hb S 细胞异质性 氧合 脱氧 Hb S溶液 Hb S纤维 血管闭塞
37/46
The amino acid substitution in the unstable hemoglobin, Hb Hammersmith (β-­chain p.
Ch12 — Segment 37
The amino acid substitution in the unstable hemoglobin, Hb Hammersmith (β-­chain p.不稳定血红蛋白Hb Hammersmith(β链p.)中的氨基酸替换。
Phe 42Ser; see This variant is notable because the substituted phenylalanine residue is one of the two amino acids that are conserved in all globins in nature .Phe42Ser;请参见此变异值得注意,因为被替换的苯丙氨酸残基是在自然界所有球蛋白中保守的两个氨基酸之一。
It is, therefore, not surprising that substitutions of this phenylalanine produce serious alterations in hemoglobin function.因此,该苯丙氨酸的替换导致血红蛋白功能严重改变并不令人惊讶。
In normal β-­globin, the bulky phenylalanine wedges the heme into a “pocket” in the folded β-­globin monomer.在正常的β-珠蛋白中,庞大的苯丙氨酸将血红素楔入折叠的β-珠蛋白单体中的一个“口袋”中。
Its replacement by serine, a smaller residue, creates a gap that allows the heme to slip out of its pocket.将其替换为较小的丝氨酸残基会产生一个间隙,使血红素能够从其口袋中滑出。
In addition to its instability, Hb Hammersmith has a low oxygen affinity, which can cause cyanosis in heterozygotes carriers.除其不稳定性外,Hb Hammersmith还具有低氧亲和力,可在杂合子携带者中引起发绀。
In contrast to variants that destabilize the tetramer, other variants destabilize the globin monomer and never form the tetramer, causing chain imbalance and thalassemia (see following section).与使四聚体不稳定的变异体相反,其他变异体使珠蛋白单体不稳定且从不形成四聚体,导致链失衡和地中海贫血(参见下一节)。
Variants With Altered Oxygen Transport Variants that alter the ability of hemoglobin to transport oxygen, although rare, are of general interest because they illustrate how a variant can impair one function of a protein (in this case, oxygen binding and release) and yet leave the other properties of the protein relatively intact.具有改变氧运输能力的变异体虽然罕见,但普遍受到关注,因为它们说明了一个变异体如何损害蛋白质的一种功能(在此为氧结合和释放)而同时使蛋白质的其他性质相对完整。
For example, the variants that affect oxygen transport generally have little or no effect on hemoglobin stability.例如,影响氧运输的变异体通常对血红蛋白稳定性几乎没有影响。
Methemoglobins.高铁血红蛋白。
Oxyhemoglobin is the form of hemoglobin that is capable of reversible oxygenation; its heme iron is in the reduced (or ferrous) state.氧合血红蛋白是能够进行可逆氧合的血红蛋白形式;其血红素铁处于还原(或亚铁)状态。
The heme iron tends to oxidize spontaneously to the ferric Euploid erythroid progenitor Micro RNAs 15a and 16-1 MYB MYB Fetal hemoglobin Midgestation Birth Trisomy 13 progenitor The basal level of these micro RNAs moderates expression of targets such as the MYB gene during erythropoiesis.血红素铁倾向于自发氧化为三价铁。整倍体红系祖细胞微小RNA 15a和16-1 MYB MYB 胎儿血红蛋白 孕中期 出生 13三体 祖细胞 这些微小RNA的基础水平在红细胞生成过程中调节MYB基因等靶标的表达。
In the case of trisomy 13, elevated levels of these micro RNAs result in additional down-­regulation of MYB expression, which in turn results in a delayed switch from fetal to adult hemoglobin and persistent expression of fetal hemoglobin.在13三体的情况下,这些微小RNA水平升高导致MYB表达进一步下调,进而导致胎儿血红蛋白向成人血红蛋白的转换延迟以及胎儿血红蛋白持续表达。
(Redrawn from Orkin SH: Disorders of hemoglobin synthesis: The thalassemias.(重新绘制自Orkin SH:血红蛋白合成障碍:地中海贫血。
In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders, pp. 106–­126.) A B C Peripheral blood smear and Heinz body preparation.见Stamatoyannopoulos G, Nienhuis AW, Leder P等编:血液疾病的分子基础,费城,1987年,WB Saunders出版社,第106-126页。)A B C 外周血涂片和亨氏小体制备。
The peripheral smear (A) shows “bite” cells with pitted-­out semicircular areas of the red blood cell membrane as a result of removal of Heinz bodies by macrophages in the spleen, causing premature destruction of the red cell.外周血涂片(A)显示“咬痕”细胞,红细胞膜上存在凹陷的半圆形区域,这是由于脾脏中巨噬细胞清除亨氏小体所致,导致红细胞过早破坏。
The Heinz body preparation (B) shows increased Heinz bodies in the same specimen when compared to a control (C).亨氏小体制备(B)显示与对照(C)相比,同一标本中亨氏小体增多。
(From Hoffman R, Furie B, Mc Glave P, et al: Hematology: Basic principles and practice, ed 5, 2008, Elsevier.)(摘自Hoffman R, Furie B, Mc Glave P等:血液学:基本原理与实践,第5版,2008年,Elsevier出版社。)
38/46
The Molecular Basis of Genetic Disease 241 form and the resulting molecule—referred to as methemoglobin, is incapable of…
Ch12 — Segment 38
The Molecular Basis of Genetic Disease 241 form and the resulting molecule—referred to as methemoglobin, is incapable of reversible oxygenation.遗传病的分子基础241形式,由此产生的分子——称为高铁血红蛋白,不能进行可逆的氧合。
If significant amounts of methemoglobin accumulate in the blood, cyanosis results.如果血液中积累大量高铁血红蛋白,就会导致发绀。
Maintenance of the heme iron in the reduced state is the role of the enzyme, methemoglobin reductase.维持血红素铁处于还原状态是酶——高铁血红蛋白还原酶的作用。
In several altered globins (either α or β), substitutions in the region of the heme pocket affect the heme-­globin bond in a way that makes the iron resistant to the reductase.在几种改变的珠蛋白(α或β)中,血红素口袋区域的替换影响血红素-珠蛋白键,使得铁对还原酶产生抵抗。
Although heterozygotes for these abnormal hemoglobins are cyanotic (a sign), they are asymptomatic.尽管这些异常血红蛋白的杂合子出现发绀(一种体征),但他们没有症状。
The homozygous state is presumably lethal.纯合状态可能是致命的。
One example of a β-­chain methemoglobin is Hb Hyde Park (see His 92 in to which heme is covalently bound has been replaced by tyrosine (p.β链高铁血红蛋白的一个例子是Hb Hyde Park(参见His 92,其与血红素共价结合,已被酪氨酸取代(p.
His 92Tyr).His 92Tyr)。
Hemoglobins With Altered Oxygen Affinity.氧亲和力改变的血红蛋白。
Variants that alter oxygen affinity demonstrate the importance of subunit interaction for the normal function of a multimeric protein such as hemoglobin.改变氧亲和力的变体证明了亚基相互作用对多聚蛋白(如血红蛋白)正常功能的重要性。
In the Hb A tetramer, the α:β interface has been highly conserved throughout evolution.在Hb A四聚体中,α:β界面在进化过程中高度保守。
It is subject to significant movement between the chains when the hemoglobin shifts from the oxygenated (relaxed) to the deoxygenated (tense) form of the molecule.当血红蛋白从氧合(松弛)形式转变为脱氧(紧张)形式时,该界面在链之间会发生显著移动。
Substitutions in residues at this interface, exemplified by the β-­globin mutant Hb Kempsey (see Thalassemia: An Imbalance of Globin-­Chain Synthesis The thalassemias are collectively the most common human single-­gene disorders in the world (Case 44).该界面残基的替换,例如β-珠蛋白突变体Hb Kempsey(参见“地中海贫血:珠蛋白链合成失衡”,地中海贫血共同是世界上最常见的人类单基因疾病(案例44))。
They are a heterogeneous group of diseases of hemoglobin synthesis in which variants reduce the synthesis or stability of either the α-­globin or β-­globin chain to cause α-­thalassemia or β-­thalassemia, respectively.它们是一组异质性的血红蛋白合成疾病,其中变体减少α-珠蛋白或β-珠蛋白链的合成或稳定性,分别导致α-地中海贫血或β-地中海贫血。
The resulting imbalance in the ratio of the α:β chains underlies the pathophysiology.由此产生的α:β链比例失衡是病理生理学的基础。
The chain that is produced at the normal rate is in relative excess; in the absence of a complementary chain with which to form a tetramer, the excess normal chains eventually precipitate in the cell, damaging the membrane and leading to premature red blood cell destruction.以正常速率生成的链相对过剩;在没有互补链形成四聚体的情况下,过量的正常链最终在细胞内沉淀,损伤细胞膜,导致红细胞过早破坏。
The excess β or β-­like chains are insoluble and precipitate in both red cell precursors (causing ineffective erythropoiesis) and in mature red cells (causing hemolysis) because they damage the cell membrane.过量的β或β样链不溶,并在红细胞前体(导致无效红细胞生成)和成熟红细胞(导致溶血)中沉淀,因为它们损伤细胞膜。
The result is anemia (lack of red blood cells) in which the red cells are both hypochromic (i. e., pale red cells) and microcytic (i. e., small red cells).结果是贫血(红细胞缺乏),其中红细胞既呈低色素性(即苍白红细胞)又呈小细胞性(即小红细胞)。
The name thalassemia (from the Greek thalassa, sea) was first used to signify that the disease was discovered in persons of Mediterranean origin.地中海贫血(来自希腊语thalassa,大海)这一名称最初用于表示该病在地中海血统的人群中发现。
Both α-­thalassemia and β-­thalassemia, however, have a high frequency in many populations; α-­thalassemia is more prevalent and more widely distributed.然而,α-地中海贫血和β-地中海贫血在许多人群中都有高频率;α-地中海贫血更常见且分布更广。
The high frequency of thalassemia is due to the protective advantage against malaria that it confers on carriers, analogous to the heterozygote advantage of sickle cell hemoglobin carriers (see Chapter 10).地中海贫血的高频率归因于它赋予携带者对疟疾的保护优势,类似于镰状细胞血红蛋白携带者的杂合子优势(见第10章)。
There is a characteristic distribution of the thalassemias in a band around the Old World—­in the Mediterranean, the Middle East, and parts of Africa, India, and Asia.地中海贫血在旧大陆周围呈带状分布——在地中海、中东以及非洲、印度和亚洲的部分地区具有特征性分布。
An important clinical consideration is that alleles for both types of thalassemia, as well as for structural alterations in hemoglobin, do often coexist in an individual.一个重要的临床考虑是,两种类型地中海贫血的等位基因以及血红蛋白结构改变的等位基因常常在同一个体中共存。
As a result, interactions may occur among different alleles of the same globin gene, or among variant alleles of different globin genes.因此,同一珠蛋白基因的不同等位基因之间,或不同珠蛋白基因的变异等位基因之间可能发生相互作用。
The α-­Thalassemias Genetic disorders of α-­globin production disrupt the formation of both fetal and adult hemoglobins , causing intrauterine as well as postnatal disease.α-地中海贫血 α-珠蛋白生成的遗传疾病破坏胎儿和成人血红蛋白的形成,导致子宫内及出生后疾病。
In the absence of α-­globin chains with which to associate, the chains from the β-­globin cluster are free to form a homotetrameric hemoglobin.在没有α-珠蛋白链与之结合的情况下,β-珠蛋白簇的链可自由形成同源四聚体血红蛋白。
Hemoglobin with a γ4 composition is known as Hb Barts, and the β4 tetramer is called Hb H.具有γ4组成的血红蛋白称为Hb Barts,β4四聚体称为Hb H。
Because neither of these hemoglobins is capable of releasing oxygen to tissues under normal conditions, they are completely ineffective oxygen carriers.由于这些血红蛋白在正常情况下都不能向组织释放氧气,因此它们是完全无效的氧载体。
Consequently, fetuses with severe α-­thalassemia and high levels of Hb Barts suffer severe intrauterine hypoxia and develop massive generalized fluid accumulation: a condition called hydrops fetalis.因此,患有严重α-地中海贫血且Hb Barts水平高的胎儿会遭受严重的宫内缺氧,并发展为全身性大量积液:一种称为胎儿水肿的病症。
In milder α-­thalassemias, an anemia develops because of the gradual precipitation of the Hb H in the erythrocyte.在较轻度的α-地中海贫血中,由于Hb H在红细胞中逐渐沉淀,导致贫血。
The formation of Hb H inclusions in mature red cells and the removal of these inclusions by the spleen damages the cells, leading to their premature destruction.成熟红细胞中Hb H包涵体的形成以及脾脏对这些包涵体的清除会损伤细胞,导致其过早破坏。
Deletions of the α-­Globin Genes.α-珠蛋白基因的缺失。
The most common molecular causes of α-­thalassemia are gene deletions.α-地中海贫血最常见的分子原因是基因缺失。
The high frequency of deletions in the α-­chain genes, compared with, for example, the β-­chain genes, is a consequence of having two identical α-­globin genes on each chromosome 16 .与β链基因相比,α链基因缺失的高频率是由于每条16号染色体上有两个相同的α-珠蛋白基因。
This arrangement of tandem homologous α-­globin genes facilitates misalignment, due to homologous pairing and subsequent recombination between the α1 gene domain on one chromosome and the corresponding α2 gene region on the other .这种串联同源α-珠蛋白基因的排列促进了错配,这是由于同源配对以及随后一条染色体上的α1基因区域与另一条染色体上相应的α2基因区域之间的重组。
Evidence supporting this pathogenic mechanism of nonallelic homologous recombination is provided by reports of individuals with a triplicated α-­globin gene complex.支持这种非等位基因同源重组致病机制的证据来自具有三重α-珠蛋白基因复合体的个体的报告。
Deletions or other alterations of one, two, three, or all four copies of the α-­globin gene cause a proportionately severe hematologic abnormality ( Individuals with two normal and two abnormal α-­globin genes are said to have α-­thalassemia trait. .α-珠蛋白基因的一个、两个、三个或全部四个拷贝的缺失或其他改变会导致相应程度的严重血液学异常(具有两个正常和两个异常α-珠蛋白基因的个体被称为具有α-地中海贫血特征)。
39/46
This can result from either of two genotypes (−−/­αα or −α/­−α), differing in whether the deletions are in cis or in tra…
Ch12 — Segment 39
This can result from either of two genotypes (−−/­αα or −α/­−α), differing in whether the deletions are in cis or in trans.这可由两种基因型(−−/αα 或 −α/−α)导致,区别在于缺失是顺式(cis)还是反式(trans)。
The α-­thalassemia trait is distributed throughout the world.α-地中海贫血特征分布于全球。
However, heterozygosity for deletion of both copies of the α-­globin gene in cis (−−/­αα genotype) is largely restricted to Southeast Asians.然而,顺式(cis)缺失两个 α-珠蛋白基因拷贝的杂合状态(−−/αα 基因型)主要局限于东南亚人群。
Offspring of two carriers of this deletion allele may receive two −−/­ −− chromosomes, leading to Hb Barts (γ4) and hydrops fetalis.两位携带此缺失等位基因的携带者所生子女可能获得两条 −−/−− 染色体,导致 Hb Barts(γ4)和胎儿水肿。
In other populations, however, α-­thalassemia trait is usually the result of the trans −α/­−α genotype, which cannot give rise to −−/­−− offspring.但在其他人群中,α-地中海贫血特征通常是反式(trans)−α/−α 基因型的结果,这种基因型无法产生 −−/−− 后代。
In addition to α-­thalassemia variants that result in deletion of the α-­globin genes, variants that delete only the LCR of the α-­globin complex can cause α-­thalassemia.除了导致 α-珠蛋白基因缺失的 α-地中海贫血变异型,仅删除 α-珠蛋白复合体 LCR 的变异也可引起 α-地中海贫血。
In fact, similar to the observations with respect to the β-­globin LCR, such deletions were critical for demonstrating the existence of this regulatory element at the α-­globin locus.事实上,与 β-珠蛋白 LCR 的观察结果类似,此类缺失对于证明 α-珠蛋白位点存在该调控元件至关重要。
Other Forms of α-­Thalassemia.α-地中海贫血的其他形式。
In all the classes of α-­thalassemia described earlier, deletions in the α-­globin genes or variants in their cis-­acting sequences account for the reduction of α-­globin synthesis.在上述所有类别的 α-地中海贫血中,α-珠蛋白基因的缺失或其顺式作用序列的变异可解释 α-珠蛋白合成的减少。
Other types of α-­thalassemia occur much less commonly.其他类型的 α-地中海贫血发生频率低得多。
One important rare form of α-­thalassemia is ATR-­X syndrome, which is associated with both α-­thalassemia and intellectual disability.α-地中海贫血的一种重要罕见形式是 ATR-X 综合征,该综合征同时伴有 α-地中海贫血和智力障碍。
It illustrates the importance of epigenetic packaging of the genome in the regulation of gene expression (see Chapters 3 and 8).这说明了基因组表观遗传包装在基因表达调控中的重要性(见第 3 章和第 8 章)。
The X chromosome ATRX gene encodes a chromatin remodeling protein that functions, in trans, to activate the expression of the α-­globin genes.X 染色体上的 ATRX 基因编码一种染色质重塑蛋白,该蛋白以反式方式发挥功能,激活 α-珠蛋白基因的表达。
The ATRX protein belongs to a family of proteins that function within large multiprotein complexes to change DNA topology.ATRX 蛋白属于一个在大型多蛋白复合体中发挥功能以改变 DNA 拓扑结构的蛋白质家族。
ATR-­X syndrome is one of several monogenic diseases that result from variants in chromatin remodeling proteins (chromatinopathies; see Chapter 8).ATR-X 综合征是由染色质重塑蛋白变异引起的几种单基因疾病之一(染色质病;见第 8 章)。
ATR-­X syndrome was initially recognized as unusual because the first families in which it was identified were northern Europeans, a population in which the deletion forms of α-­thalassemia are uncommon.ATR-X 综合征最初被认为不寻常,因为首次发现该病的家族是北欧人,而在该人群中 α-地中海贫血的缺失形式并不常见。
All affected individuals were males with severe intellectual disability, together with a wide range of other abnormalities, including characteristic facial features, skeletal defects, and urogenital malformations.所有受累个体均为男性,患有严重智力障碍,并伴有多种其他异常,包括特征性面容、骨骼缺陷和泌尿生殖畸形。
This diversity of phenotypes suggests that ATRX regulates the expression of numerous other genes besides the α-­globins.这种表型的多样性提示 ATRX 除 α-珠蛋白外还调控众多其他基因的表达。
In those with ATR-­X syndrome, the reduction in α-­globin synthesis is due to accumulation at the α-­globin gene cluster of a histone variant (see Chapter 3) called macro H2A.在患有 ATR-X 综合征的个体中,α-珠蛋白合成的减少是由于一种称为巨 H2A 的组蛋白变体(见第 3 章)在 α-珠蛋白基因簇处积聚所致。
This accumulation reduces α-­globin gene expression and causes α-­thalassemia.这种积聚降低了 α-珠蛋白基因表达并导致 α-地中海贫血。
To date, all variants in the ATRX gene associated with ATR-­X syndrome involve partial loss of function, leading to hematologic defects that are mild, compared with those seen in the classic forms of α-­thalassemia.迄今为止,所有与 ATR-X 综合征相关的 ATRX 基因变异均涉及部分功能缺失,导致的血液学缺陷相较于经典形式的 α-地中海贫血更为轻微。
Individuals with ATR-­X syndrome have abnormalities in DNA methylation patterns to indicate that the Single-gene complex Triple-gene complex Homologous pairing and unequal crossover ψα1 α2 α1 ψα1 α2 α1 ψα1 α2 α1 α ψα1 α Misalignment, homologous pairing, and recombination between the α1 gene on one chromosome and the α2 gene on the homologous chromosome result in the deletion of one α-­globin gene.患有 ATR-X 综合征的个体存在 DNA 甲基化模式的异常,这表明单基因复合体、三基因复合体、同源配对和不等交换 ψα1 α2 α1 ψα1 α2 α1 ψα1 α2 α1 α ψα1 α 的错排、同源配对以及一条染色体上的 α1 基因与同源染色体上的 α2 基因之间的重组导致一个 α-珠蛋白基因的缺失。
(Redrawn from Kazazian HH: The thalassemia syndromes: Molecular basis and prenatal diagnosis in 1990, Semin Hematol 27:209–­228, 1990.) 4 Clinical States Associated With α-­Thalassemia Genotypes Clinical Condition Number of Functional α Genes α-­Globin Gene Genotype α-­Chain Production “Normal” 4 αα/­αα 100% Silent carrier 3 αα/­α− 75% α-­Thalassemia trait (mild anemia, microcytosis) 2 α−/­α− or αα/­−− 50% Hb H (β4) disease (moderately severe hemolytic anemia) 1 α−/­−− 25% Hydrops fetalis or homozygous α-­thalassemia (Hb Barts: γ4) 0 −−/­−− 0%(重绘自 Kazazian HH: The thalassemia syndromes: Molecular basis and prenatal diagnosis in 1990, Semin Hematol 27:209–228, 1990。)与 α-地中海贫血基因型相关的四种临床状态:临床状况、功能性 α 基因数量、α-珠蛋白基因基因型、α 链产量;“正常”为 4 个 αα/αα、100%;静止型携带者为 3 个 αα/α−、75%;α-地中海贫血特征(轻度贫血、小细胞增多症)为 2 个 α−/α− 或 αα/−−、50%;Hb H(β4)病(中度至重度溶血性贫血)为 1 个 α−/−−、25%;胎儿水肿或纯合型 α-地中海贫血(Hb Barts: γ4)为 0 个 −−/−−、0%。
40/46
The Molecular Basis of Genetic Disease 243 ATRX protein is also required to establish or maintain the methylation patter…
Ch12 — Segment 40
The Molecular Basis of Genetic Disease 243 ATRX protein is also required to establish or maintain the methylation pattern in certain domains of the genome.遗传病的分子基础 243 ATRX蛋白也是建立或维持基因组特定区域甲基化模式所必需的。
This may be by modulating access of the DNA methyltransferase enzyme to its binding sites.这可能是通过调节DNA甲基转移酶与其结合位点的可及性来实现的。
This finding is noteworthy because variants in MECP2, which encodes a protein that binds to methylated DNA, cause Rett syndrome (Case 40) by disrupting the epigenetic regulation of genes in regions of methylated DNA, leading to neurodevelopmental regression.这一发现值得注意,因为编码一种结合甲基化DNA的蛋白质的MECP2基因变异,通过破坏甲基化DNA区域基因的表观遗传调控,导致Rett综合征(案例40),进而引起神经发育倒退。
Normally, ATRX and the Me CP2 protein interact; impairment of this interaction due to ATRX variants may contribute to the intellectual disability seen in ATR-­X syndrome.正常情况下,ATRX与MeCP2蛋白相互作用;由ATRX变异导致的这种相互作用受损,可能促成了ATR-X综合征中所见的智力障碍。
The β-­Thalassemias The β-­thalassemias share many features with α-­thalassemia.β-地中海贫血 β-地中海贫血与α-地中海贫血有许多共同特征。
In β-­thalassemia, the decrease in β-­globin production causes a hypochromic, microcytic anemia.在β-地中海贫血中,β-珠蛋白生成减少导致低色素性、小细胞性贫血。
An imbalance in globin synthesis is due to the excess of α chains.珠蛋白合成失衡是由于α链过多所致。
The latter are insoluble and precipitate in both red cell precursors (causing ineffective erythropoiesis) and mature red cells (causing hemolysis) because they damage the cell membrane.后者不溶,并在红细胞前体(导致无效红细胞生成)和成熟红细胞(导致溶血)中沉淀,因为它们损伤细胞膜。
In contrast to α-­globin, however, the β chain is important only in the postnatal period.然而,与α-珠蛋白相比,β链仅在出生后时期重要。
Consequently, the onset of β-­thalassemia is not apparent until a few months after birth, when β-­globin normally replaces γ-­globin as the major non–­α chain .因此,β-地中海贫血的起病在出生后几个月才显现,此时β-珠蛋白通常取代γ-珠蛋白成为主要的非α链。
Only the synthesis of the major adult hemoglobin, Hb A, is reduced.只有主要成人血红蛋白Hb A的合成减少。
The level of Hb F is increased in β-­thalassemia, not because of a reactivation of the γ-­globin gene expression that was switched off at birth, but because of selective survival and perhaps increased production of the minor population of adult red blood cells that contain Hb F.β-地中海贫血中Hb F水平升高,并非由于出生时被关闭的γ-珠蛋白基因表达重新激活,而是由于含有Hb F的少量成人红细胞的优先存活以及可能生成增加。
In contrast to α-­thalassemia, the β-­thalassemias are usually due to single nucleotide variants rather than to deletions ( In many regions of the world where β-­thalassemia is common, there are so many different β-­thalassemia variants that individuals with this condition are more likely to be compound heterozygotes (i. e., carrying two different β-­thalassemia alleles) than to be homozygotes for one allele.与α-地中海贫血不同,β-地中海贫血通常由单核苷酸变异而非缺失引起(在世界许多β-地中海贫血常见地区,存在大量不同的β-地中海贫血变异,以至于患有此病的个体更可能是复合杂合子(即携带两个不同的β-地中海贫血等位基因),而非单一等位基因的纯合子)。
Most individuals with two β-­thalassemia alleles have thalassemia major: a condition characterized by severe anemia and the need for lifelong medical management.大多数携带两个β-地中海贫血等位基因的个体患有重型地中海贫血:一种以严重贫血和需要终身医疗管理为特征的疾病。
When the β-­thalassemia alleles allow so little production of β-­globin that no Hb A is present, the condition is designated β0-­thalassemia.当β-地中海贫血等位基因导致β-珠蛋白生成极少以至于无Hb A存在时,该病被命名为β0-地中海贫血。
Clinically, affected individuals are dependent on red blood cell transfusions.临床上,受累个体依赖红细胞输注。
If some Hb A is detectable, the affected individual has β+-­thalassemia.如果可检测到少量Hb A,则受累个体患有β+-地中海贫血。
Although the severity of the clinical disease depends on the combined effect of the two alleles present, until recently, survival into adult life was unusual.尽管临床疾病的严重程度取决于所存在的两个等位基因的联合效应,但直到最近,存活至成年期仍不常见。
Infants with homozygous β-­thalassemia present with anemia once the postnatal production of Hb F decreases—generally before 2 years of age.纯合子β-地中海贫血婴儿在出生后Hb F生成减少时(通常在2岁前)出现贫血。
At present, in most countries, treatment of the thalassemias is based on correction of the anemia and the increased marrow expansion by blood transfusion; the consequent excess iron accumulation is controlled by administration of chelating agents.目前,在大多数国家,地中海贫血的治疗基于通过输血纠正贫血和骨髓过度增生;随之而来的铁过载通过给予螯合剂来控制。
Bone marrow transplantation is effective, but this is an option only if an HLA-­matched family member can be found.骨髓移植有效,但这仅适用于能找到HLA匹配的家庭成员的情况。
Gene therapy options are now emerging in clinical practice (see Chapter 14).基因治疗选择现已在临床实践中出现(见第14章)。
Carriers of one β-­thalassemia allele are clinically well and are said to have thalassemia minor.携带一个β-地中海贫血等位基因的个体临床状况良好,称为轻型地中海贫血。
Such individuals have hypochromic, microcytic red blood cells 12. 11C) Abnormal acceptor site of intron 1: AG → GG β0 Promoter variants Variant in the ATA box β+ −31 −30 −29 −28 −31 −30 −29 −28 A T A A → G T A A Abnormal RNA cap site A → C transversion at the mRNA cap site β+ Polyadenylation signal defects AATAAA → AACAAA β+ Nonsense variants Codon 39 gln → stop β0 CAG → UAG Codon 16 (1 bp deletion) Frameshift variants Normal trp gly lys val asn β0 15 16 17 18 19 UGG GGC AAG GUG AAC UGG GCA AGG UGA Variant trp ala arg stop Synonymous variants Codon 24 β+ gly → gly GGU → GGA Derived in part from Weatherall DJ, Clegg JB, Higgs DR, et al: The hemoglobinopathies.此类个体具有低色素性、小细胞性红细胞 12. 11C) 内含子1的异常受体位点:AG → GG β0 启动子变异 ATA盒变异 β+ −31 −30 −29 −28 −31 −30 −29 −28 A T A A → G T A A 异常RNA帽位点 mRNA帽位点A→C颠换 β+ 多聚腺苷酸化信号缺陷 AATAAA → AACAAA β+ 无义变异 第39号密码子谷氨酰胺→终止 β0 CAG → UAG 第16号密码子(1 bp缺失)移码变异 正常 trp gly lys val asn β0 15 16 17 18 19 UGG GGC AAG GUG AAC UGG GCA AGG UGA 变异 trp ala arg stop 同义变异 第24号密码子 β+ gly → gly GGU → GGA 部分引自Weatherall DJ, Clegg JB, Higgs DR, 等: 血红蛋白病。
In Scriver CR, Beaudet AL, Sly WS, et al, editors: The metabolic and molecular bases of inherited disease, ed 7, New York, 1995, Mc Graw-­Hill, pp. 3417–­3484; Orkin SH: Disorders of hemoglobin synthesis: the thalassemias.见Scriver CR, Beaudet AL, Sly WS, 等 编辑: 遗传病的代谢与分子基础, 第7版, 纽约, 1995, McGraw-­Hill, 第3417–3484页; Orkin SH: 血红蛋白合成障碍: 地中海贫血。
In Stamatoyannopoulos G, Nienhuis AW, Leder P, et al, editors: The molecular basis of blood diseases, Philadelphia, 1987, WB Saunders, pp. 106–­126. a One other hemoglobin structural variant that causes β-­thalassemia is shown in mRNA, Messenger RNA.见Stamatoyannopoulos G, Nienhuis AW, Leder P, 等 编辑: 血液疾病的分子基础, 费城, 1987, WB Saunders, 第106–126页。a 另一种导致β-地中海贫血的血红蛋白结构变异显示在mRNA中,信使RNA。
41/46
and may have a slight anemia that can be misdiagnosed initially as iron deficiency.
Ch12 — Segment 41
and may have a slight anemia that can be misdiagnosed initially as iron deficiency.并且可能患有轻度贫血,最初可能被误诊为缺铁性贫血。
The diagnosis of thalassemia minor can be supported by hemoglobin electrophoresis, which generally reveals an increase in the level of Hb A2 (α2δ2).轻型地中海贫血的诊断可得到血红蛋白电泳的支持,后者通常显示Hb A2(α2δ2)水平升高。
In many countries, thalassemia minor is sufficiently common to require diagnostic distinction from iron deficiency anemia and to be a frequent source of referral for prenatal diagnosis of affected homozygous fetuses (see Chapter 18). α-­Thalassemia Alleles as Modifier Genes of β-­Thalassemia.在许多国家,轻型地中海贫血十分常见,需与缺铁性贫血进行鉴别诊断,并且常需转诊进行受累纯合子胎儿的产前诊断(见第18章);α-地中海贫血等位基因作为β-地中海贫血的修饰基因。
In human genetics, one of the best examples of a modifier gene comes from co-existence of β-­thalassemia and α-­thalassemia alleles in a population.在人类遗传学中,修饰基因的最佳实例之一来自人群中β-地中海贫血与α-地中海贫血等位基因的共存。
In such populations, β-­thalassemia homozygotes may also inherit an α-­thalassemia allele.在此类人群中,β-地中海贫血纯合子也可能遗传一个α-地中海贫血等位基因。
The clinical severity of the β-­thalassemia is sometimes ameliorated by the presence of the α-­thalassemia allele, which acts as a modifier.β-地中海贫血的临床严重程度有时因α-地中海贫血等位基因(作为修饰基因)的存在而减轻。
The imbalance of globin chain synthesis that occurs in β-­thalassemia (due to the relative excess of α chains) is reduced by the decrease in α-­chain production that results from the α-­thalassemia gene deletion. β-­Thalassemia, Complex Thalassemias, and Hereditary Persistence of Fetal Hemoglobin.由于α-地中海贫血基因缺失导致α链生成减少,从而降低了β-地中海贫血中(因α链相对过剩而引起的)珠蛋白链合成失衡。
Almost every type of DNA variant known to reduce the synthesis of an mRNA or protein has been identified as a cause of β-­thalassemia.几乎所有已知能降低mRNA或蛋白质合成的DNA变异类型均被确定为β-地中海贫血的病因。
The following overview of these genetic defects is, therefore, instructive about variant mechanisms in general, by describing the molecular basis of one of the most common and severe genetic diseases in the world.因此,以下对这些遗传缺陷的概述通过描述世界上最常见且最严重的遗传病之一的分子基础,对一般的变异机制具有指导意义。
Variants of the β-­globin gene complex are separated into two broad groups with different clinical phenotypes.β-珠蛋白基因复合体的变异分为两大类,其临床表型不同。
One group, which accounts for the great majority of patients, impairs the production of β-­globin alone and causes simple β-­thalassemia.第一类占患者绝大多数,仅损害β-珠蛋白的生成,引起单纯型β-地中海贫血。
The second group consists of large deletions that cause the complex thalassemias, in which the β-­globin gene is removed, as well as one or more of the other genes—­or the LCR—­in the β-­globin cluster.第二类由大片段缺失组成,导致复杂型地中海贫血,其中β-珠蛋白基因以及β-珠蛋白簇中的一个或多个其他基因或LCR被缺失。
Finally, we are informed about the regulation of globin gene expression through some deletions within the β-­globin cluster that do not cause thalassemia, but rather, a benign phenotype termed the hereditary persistence of fetal hemoglobin (i. e., the persistence of γ-­globin gene expression throughout adult life).最后,通过β-珠蛋白簇内的一些缺失(这些缺失不引起地中海贫血,反而导致一种称为遗传性胎儿血红蛋白持续存在症的良性表型,即γ-珠蛋白基因表达持续至成年期),我们得以了解珠蛋白基因表达的调控。
Molecular Basis of Simple β-­Thalassemia.单纯型β-地中海贫血的分子基础。
Simple β-­thalassemia results from a remarkable diversity of molecular defects, predominantly single nucleotide variants, in the β-­globin gene , mRNA capping or tailing variants, and frameshift or nonsense variants that introduce premature termination codons within the coding region of the gene.单纯型β-地中海贫血由β-珠蛋白基因中多种多样的分子缺陷引起,主要为单核苷酸变异、mRNA加帽或加尾变异,以及在基因编码区内引入提前终止密码子的移码或无义变异。
A few hemoglobin structural alterations also impair processing of the β-­globin mRNA, as exemplified by Hb E (described later).少数血红蛋白结构改变也会损害β-珠蛋白mRNA的加工,如Hb E(后文描述)所示。
RNA Splicing Variants.RNA剪接变异。
Most β-­thalassemia cases with a decreased abundance of β-­globin mRNA have abnormalities in RNA splicing.大多数β-珠蛋白mRNA丰度降低的β-地中海贫血病例存在RNA剪接异常。
Dozens of defects of * * ** ** Transcription RNA splicing Cap site RNA cleavage Initiator codon Small deletion Unstable globin Nonsense codon Frameshift 100 bp 5' 3' 1 2 3 Note the distribution of variants throughout the gene and that the variants affect virtually every process required for the production of normal β-­globin.数十种缺陷涉及转录、RNA剪接、加帽位点、RNA切割、起始密码子、小缺失、不稳定珠蛋白、无义密码子、移码(100 bp 5' 3' 1 2 3)。注意变异在基因中的分布,以及这些变异几乎影响正常β-珠蛋白生成所需的每一个过程。
More than 100 different β-­globin point variants are associated with simple β-­thalassemia. .超过100种不同的β-珠蛋白点变异与单纯型β-地中海贫血相关。
42/46
The Molecular Basis of Genetic Disease 245 this type have been described, and their combined clinical burden is substant…
Ch12 — Segment 42
The Molecular Basis of Genetic Disease 245 this type have been described, and their combined clinical burden is substantial.遗传性疾病的分子基础 245 这类病变已有描述,其综合临床负担十分重大。
These variants have acquired high visibility because their effects on splicing are often unexpectedly complex, and analysis of the altered mRNAs has contributed extensively to knowledge of the sequences critical to normal RNA processing (introduced in Chapter 3).这些变异因其对剪接的影响往往异常复杂而备受关注,对 altered mRNA 的分析极大促进了对正常 RNA 加工关键序列的认识(参见第3章)。
The splice defects are separated into three groups depending on the region of the unprocessed RNA in which the variant is located.剪接缺陷根据未加工 RNA 中变异所在区域分为三组。
Splice junction variants include those at the canonical 5′ donor or 3′ acceptor splice junctions of the introns or in the consensus sequences surrounding the junctions.剪接接头变异包括位于内含子经典 5' 供体或 3' 受体剪接接头处及其周围共有序列中的变异。
The critical nature of the conserved GT dinucleotide at the 5′ intron donor site and of the AG at the 3′ intron acceptor site (see Chapter 3) is demonstrated by the complete loss of normal splicing that results from variants in these dinucleotides .内含子 5' 供体位点的保守 GT 二核苷酸和 3' 受体位点的 AG 二核苷酸的关键性(见第3章),通过这些二核苷酸变异导致正常剪接完全丧失而得到证实。
Inactivation of the normal acceptor site elicits the use of other acceptor-­like sequences elsewhere in the RNA precursor molecule.正常受体位点的失活会诱发 RNA 前体分子中其他类似受体的序列被启用。
These alternative sites are termed cryptic splice sites because they are not used by the splicing apparatus if the correct site is available.这些替代位点被称为隐性剪接位点,因为在正确位点可用时剪接装置不会使用它们。
Cryptic donor or acceptor splice sites can be found in either exons or introns.隐性供体或受体剪接位点可存在于外显子或内含子中。
Intronic variants enhance the use of a cryptic splice site by making it more similar or identical to the normal splice site.内含子变异通过使隐性剪接位点更接近或等同于正常剪接位点来增强其使用。
The activated cryptic site then competes with the normal site, with variable effectiveness.激活的隐性位点随后与正常位点竞争,效果不一。
This reduces the abundance of the normal mRNA by decreasing splicing from the correct site, which remains perfectly intact .这会减少从正确位点(该位点保持完整)的剪接,从而降低正常 mRNA 的丰度。
Cryptic splice site variants are often leaky, which means that some use of the normal site occurs, producing a β+-­ thalassemia phenotype.隐性剪接位点变异常呈漏出性,即仍有一定程度的正常位点被使用,产生 β+-地中海贫血表型。
Coding sequence changes that affect splicing result from variants in the open reading frame that activate a cryptic splice site in an exon, whether or not they also change the amino acid sequence .影响剪接的编码序列改变由开放阅读框中的变异引起,这些变异激活外显子中的隐性剪接位点,无论是否改变氨基酸序列。
For example, a mild form of β+-­thalassemia results from a variant in codon 24 (see 1]); this is an example of a synonymous variant that is not neutral in its effect.例如,一种轻型 β+-地中海贫血由第 24 位密码子的变异导致(见[1]);这是同义变异并非中性效应的一个例子。
Nonfunctional mRNAs.无功能的 mRNA。
Some mRNAs are nonfunctional and cannot direct the synthesis of a complete polypeptide, because the variant generates a premature stop codon, which prematurely terminates translation.某些 mRNA 无功能,无法指导完整多肽的合成,因为变异产生提前终止密码子,提前终止翻译。
Two β-­thalassemia variants near the amino terminus exemplify this effect (see In one (p.近氨基端的两个 β-地中海贫血变异说明了这一效应(见一种为 p.
Gln 39Ter), the failure in translation is due to a single nucleotide substitution that creates a nonsense variant.Gln39Ter),翻译失败是由于单个核苷酸替换产生无义变异。
In the other, a frameshift variant results from a single base pair deletion early in the open reading frame, removing the first nucleotide from codon 16, which normally encodes glycine.另一种是移码变异,由开放阅读框早期单个碱基对缺失引起,移除通常编码甘氨酸的第 16 位密码子的第一个核苷酸。
In the reading frame that results, a premature stop codon is quickly encountered downstream, well before the normal termination signal.在由此产生的阅读框中,下游很快遇到提前终止密码子,远早于正常终止信号。
Because no β-­globin is made from these alleles, both types of nonfunctional mRNA variants cause β0-­thalassemia in the homozygous state.由于这些等位基因不产生 β-珠蛋白,两种无功能 mRNA 变异在纯合状态下均导致 β0-地中海贫血。
In some instances, frameshifts near the carboxyl terminus of the protein allow most of the mRNA to be translated normally or to produce elongated globin chains, resulting in a variant hemoglobin rather than null alleles.某些情况下,接近蛋白质羧基末端的移码使大部分 mRNA 能正常翻译或产生延长的珠蛋白链,从而产生变异血红蛋白而非无效等位基因。
In addition to ablating the production of the β-­globin polypeptide, premature stop variants, including the two described earlier, often lead to reduced abundance of the abnormal mRNA; indeed, the mRNA may be undetectable.除消除 β-珠蛋白多肽的产生外,提前终止变异(包括前述两种)常导致异常 mRNA 丰度降低;事实上,mRNA 可能检测不到。
The mechanism underlying this phenomenon—called nonsense-­mediated mRNA decay, appears to be restricted to nonsense codons located more than 50 bp upstream of the final exon-­exon junction.这一现象背后的机制——称为无义介导的 mRNA 衰变——似乎仅限于位于最终外显子-外显子连接上游超过 50 bp 的无义密码子。
Defects in Capping and Tailing of β-­Globin mRNA.β-珠蛋白 mRNA 加帽与加尾缺陷。
Several β+-­thalassemia variants highlight the critical nature of post-transcriptional modifications of mRNAs.几种 β+-地中海贫血变异凸显了 mRNA 转录后修饰的关键性。
For example, the 3′ UTR of almost all mRNAs ends with a poly A sequence, and if this sequence is not added, the mRNA is unstable.例如,几乎所有 mRNA 的 3' UTR 末端均含 poly A 序列,若该序列未被添加,则 mRNA 不稳定。
As introduced in Chapter 3, polyadenylation of mRNA first requires enzymatic cleavage of the mRNA, which occurs in response to a signal for the cleavage site, AAUAAA, that is found near the 3′ end of most eukaryotic mRNAs.如第3章所述,mRNA 的多聚腺苷酸化首先需要 mRNA 的酶切切割,该切割响应于靠近大多数真核 mRNA 3' 末端的切割信号 AAUAAA 而发生。
Individuals with a substitution that changes the signal sequence to AACAAA produce only a minor fraction of correctly polyadenylated β-­globin mRNA.将信号序列变为 AACAAA 的替换携带者仅产生少量正确多聚腺苷酸化的 β-珠蛋白 mRNA。
Hemoglobin E: A Structurally Altered Hemoglobin With Thalassemia Phenotypes Hb E is probably the most common structurally abnormal hemoglobin in the world, occurring at high frequency in Southeast Asia, where there are at least 1 million homozygotes and 30 million heterozygotes.血红蛋白 E:一种结构异常且具有地中海贫血表型的血红蛋白 Hb E 可能是世界上最常见的结构异常血红蛋白,在东南亚高发,那里至少有 100 万纯合子和 3000 万杂合子。
Hb E is a β-­globin variant (p.Hb E 是一种 β-珠蛋白变异(p.
Glu 26Lys) that reduces the rate of synthesis of the abnormal β chain.Glu26Lys),该变异降低了异常 β 链的合成速率。
It is another example of a coding sequence variant that impairs normal splicing by activating a cryptic splice site .这是编码序列变异通过激活隐性剪接位点而损害正常剪接的又一例证。
Although Hb E homozygotes are asymptomatic and only mildly anemic, individuals who are genetic compounds of Hb E and another β-­thalassemia allele have clinically relevant phenotypes that are largely determined by the severity of the other allele.尽管 Hb E 纯合子无症状且仅轻度贫血,但 Hb E 与另一 β-地中海贫血等位基因形成遗传复合物的个体具有临床相关表型,其严重程度主要取决于另一等位基因。
Complex Thalassemias and the Hereditary Persistence of Fetal Hemoglobin As mentioned earlier, large deletions that cause the complex thalassemias remove the β-­globin gene plus one or more other genes—­or the LCR—­from the β-­globin cluster.复杂地中海贫血与胎儿血红蛋白持续遗传 如前所述,导致复杂地中海贫血的大片段缺失从 β-珠蛋白簇中移除 β-珠蛋白基因及一个或多个其他基因——或 LCR。
Thus, affected individuals have reduced expression因此,受累个体表达降低。
43/46
Exon 1 Intron 1 Exon 2 Exon 3 Intron 2 Intron 2 donor site: GT Intron 2 acceptor site: AG Normal splicing pattern Intron…
Ch12 — Segment 43
Exon 1 Intron 1 Exon 2 Exon 3 Intron 2 Intron 2 donor site: GT Intron 2 acceptor site: AG Normal splicing pattern Intron 2 Intron 1 bp 110 β+ mutation in a cryptic acceptor site reduced use of unaffected normal site preferred use of mutant site Exon 1 Exon 2 Exon 3 β+ Mutation Consensus acceptor site Normal sequence CCTATTAG T YYYYNYAG G CCTATTGG T 90% 10% Normal splice site unaffected New splice site in intron Mutation creating a new splice acceptor site in an intron Exon 1 Exon 2 Intron 2 Exon 3 40% Hb E: Exon 1 mutation in a cryptic donor site reduced use of normal site moderate use of cryptic site New splice site, in a codon 60% Codon β+ Mutation Donor consensus Normal exon 1 sequence24 25 26 27 Hb E codon 26 GAG-&gt;AAG glu-&gt;lys Mutation enhancing a cryptic splice donor site in an exon Intron 2 Exon 1 Exon 2 Exon 3 Intron 2 cryptic acceptor site Consensus acceptor site 3' part of intron 2 Intron 2 acceptor site β0 mutation no splicing from the mutant site use of an intron 2 cryptic site TTTCTTTCAG G YYYYYYNYAG G Intron 2 Exon 3 Intron 2 Exon 3 β0 Mutation.....外显子1 内含子1 外显子2 外显子3 内含子2 内含子2供体位点:GT 内含子2受体位点:AG 正常剪接模式 内含子2 内含子1 bp110 β+突变在隐蔽受体位点导致未受影响正常位点使用减少,突变位点优先使用 外显子1 外显子2 外显子3 β+突变 共有受体位点 正常序列CCTATTAG T YYYYNYAG G CCTATTGG T 90% 10% 正常剪接位点未受影响 内含子中新剪接位点 突变在内含子中创建新的剪接受体位点 外显子1 外显子2 内含子2 外显子3 40% Hb E:外显子1突变在隐蔽供体位点导致正常位点使用减少,隐蔽位点中度使用 新剪接位点,位于密码子 60% 密码子 β+突变 供体共有序列 正常外显子1序列24 25 26 27 Hb E密码子26 GAG→AAG glu→lys 突变增强外显子中的隐蔽剪接供体位点 内含子2 外显子1 外显子2 外显子3 内含子2隐蔽受体位点 共有受体位点 内含子2的3'部分 内含子2受体位点 β0突变 无来自突变位点的剪接 使用内含子2隐蔽位点 TTTCTTTCAG G YYYYYYNYAG G 内含子2 外显子3 内含子2 外显子3 β0突变……
CGG CTC.....CGG CTC……
Normal:.....正常:……
CAG CTC.....CAG CTC……
Mutation destroying a normal splice acceptor site and activating a cryptic site A B C D Normal splicing pattern.破坏正常剪接受体位点并激活隐蔽位点的突变 A B C D 正常剪接模式。
(B) An intron 2 variant (IVS2-­2A&gt;G) in the normal splice acceptor site aborts normal splicing.(B) 正常剪接受体位点中的内含子2变异(IVS2-2A>G)终止正常剪接。
This variant results in the use of a cryptic acceptor site in intron 2.此变异导致使用内含子2中的隐蔽受体位点。
The cryptic site conforms perfectly to the consensus acceptor splice sequence (where Y is either pyrimidine, T or C).该隐蔽位点完全符合共有受体剪接序列(其中Y为嘧啶,T或C)。
Because exon 3 has been enlarged at its 5′ end by inclusion of intron 2 sequences, the abnormal alternatively spliced messenger RNA (mRNA) made from this mutant gene has lost the correct open reading frame and cannot encode β-­globin.由于外显子3在其5'端因包含内含子2序列而增大,由此突变基因产生的异常选择性剪接信使RNA(mRNA)失去了正确的开放阅读框,无法编码β-珠蛋白。
(C) An intron 1 variant (G &gt; A in nucleotide 110 of intron 1) activates a cryptic acceptor site by creating an AG dinucleotide and increasing the resemblance of the site to the consensus acceptor sequence.(C) 内含子1中的变异(内含子1第110位核苷酸G>A)通过创建AG二核苷酸并增加该位点与共有受体序列的相似性,激活了一个隐蔽受体位点。
The globin mRNA thus formed is elongated (19 extra nucleotides) at the 5′ side of exon 2; a premature stop codon is introduced into the transcript.由此形成的珠蛋白mRNA在外显子2的5'侧延长(增加19个核苷酸);转录本中引入了提前终止密码子。
A β+ thalassemia phenotype results because the correct acceptor site is still used, although at only 10% of the wild-­type level.由于正确受体位点仍被使用,尽管仅处于野生型水平的10%,因此产生β+地中海贫血表型。
(D) In the Hb E defect, the missense variant (p.(D) 在血红蛋白E缺陷中,错义变异(p.
Glu 26Lys) in codon 26 in exon 1 activates a cryptic donor splice site in codon 25 that competes effectively with the normal donor site.外显子1中密码子26的Glu26Lys)激活了密码子25中的隐蔽供体剪接位点,该位点与正常供体位点有效竞争。
Moderate use is made of this alternative splicing pathway, but the majority of RNA is still processed from the correct site, and mild β+ thalassemia results.该替代剪接途径被中度使用,但大多数RNA仍从正确位点加工,导致轻度β+地中海贫血。
In Stamatoyannopoulos G, Majerus PW, Perlmutter RM, et al, editors: The molecular basis of blood diseases, ed 3, Philadelphia, 2001, WB Saunders.)引自Stamatoyannopoulos G, Majerus PW, Perlmutter RM等编:《血液疾病的分子基础》,第3版,费城,2001年,WB Saunders出版。)
44/46
The Molecular Basis of Genetic Disease 247 of β-­globin and one or more of the other β-­like chains.
Ch12 — Segment 44
The Molecular Basis of Genetic Disease 247 of β-­globin and one or more of the other β-­like chains.β-球蛋白和一个或多个其他β样链的遗传性疾病分子基础247。
These disorders are named according to the genes deleted (e. g., [δβ]0-­thalassemia or [Aγδβ]0-­thalassemia) .这些疾病根据缺失的基因命名(例如,[δβ]0-地中海贫血或[Aγδβ]0-地中海贫血)。
Deletions that remove the β-­globin LCR start ~50 to 100 kb upstream of the β-­globin gene cluster and extend 3′ to varying degrees.缺失β-球蛋白LCR的缺失从β-球蛋白基因簇上游约50至100 kb处开始,并向3′端延伸不同程度。
Although some of these deletions (such as the Hispanic deletion shown in leave all or some of the genes at the β-­globin locus completely intact, they ablate expression from the entire cluster to cause (εγδβ)0-­thalassemia.尽管其中一些缺失(例如所示的西班牙裔缺失)使β-球蛋白基因座上的全部或部分基因完全完整,但它们会消除整个簇的表达,导致(εγδβ)0-地中海贫血。
Such variants demonstrate the total dependence of gene expression from the β-­globin gene cluster on the integrity of the LCR .此类变异表明β-球蛋白基因簇的基因表达完全依赖于LCR的完整性。
A second group of large β-­globin gene cluster deletions of medical significance are those that leave at least one of the γ genes intact (such as the English deletion in .第二组具有医学意义的大片段β-球蛋白基因簇缺失是那些至少保留一个γ基因完整的缺失(例如所示的英语缺失)。
Individuals carrying such variants have one of two clinical manifestations, depending on the deletion: either δβ0-­thalassemia, or a benign condition called hereditary persistence of fetal hemoglobin (HPFH) that is due to disruption of the perinatal switch from γ-­globin to β-­globin synthesis.携带此类变异的个体根据缺失的不同有两种临床表现之一:要么是δβ0-地中海贫血,要么是一种称为遗传性胎儿血红蛋白持续存在症(HPFH)的良性状态,这是由于围产期从γ-球蛋白合成转换为β-球蛋白合成的过程被破坏所致。
Homozygotes with either of these conditions are viable because the remaining γ gene(s) are still active after birth, instead of switching off as would normally occur.患有这两种情况之一的纯合子均可存活,因为剩余的γ基因在出生后仍然活跃,而不是像正常情况下那样关闭。
As a result, Hb F (α2γ2) synthesis continues postnatally at a high level and compensates for the absence of Hb A.因此,Hb F(α2γ2)的合成在出生后持续高水平,并补偿了Hb A的缺失。
The clinically innocuous nature of HPFH that results from the substantial production of γ chains is due to a higher level of Hb F in heterozygotes (17–­35% Hb F) than is generally seen in δβ0-­thalassemia heterozygotes (5–­18% Hb F).HPFH(由于大量产生γ链)的临床无害性是因为杂合子中Hb F水平(17–35% Hb F)高于δβ0-地中海贫血杂合子中通常所见到的水平(5–18% Hb F)。
Because the deletions that cause δβ0-­ thalassemia overlap with those that cause HPFH , it is not clear why patients with HPFH have higher levels of γ gene expression.由于导致δβ0-地中海贫血的缺失与导致HPFH的缺失重叠,尚不清楚为何HPFH患者具有更高水平的γ基因表达。
One possibility is that some HPFH deletions bring enhancers closer to the γ-­globin genes.一种可能性是一些HPFH缺失将增强子更靠近γ-球蛋白基因。
Insight into the role of regulators of Hb F expression, such as BCL11A and MYB (see earlier discussion), has been partly derived from the study of individuals with complex deletions of the β-­globin gene cluster.对Hb F表达调节因子(如BCL11A和MYB,见前述讨论)作用的深入了解,部分源于对具有复杂β-球蛋白基因簇缺失的个体的研究。
For example, the study of several individuals with HPFH due to rare deletions of the β-­globin gene cluster identified a 3. 5 kb region, near the 5′ end of the δ-­globin gene, that contains binding sites for BCL11A, the critical silencer of Hb F expression in the adult.例如,对几名因β-球蛋白基因簇罕见缺失而患HPFH的个体的研究,在δ-球蛋白基因5′端附近识别出一个3.5 kb区域,该区域包含BCL11A的结合位点,BCL11A是成人Hb F表达的关键沉默因子。
Public Health Approaches to Preventing Thalassemia Large-­Scale Population Screening.预防地中海贫血的公共卫生方法——大规模人群筛查。
The clinical severity of many forms of thalassemia, combined with their high frequency, imposes a tremendous health burden on many societies.多种地中海贫血的临床严重性及其高发病率,给许多社会带来了巨大的健康负担。
To reduce the high incidence of the disease in some parts of the world, governments have introduced successful thalassemia control programs based on offering or requiring thalassemia carrier screening of individuals of childbearing age in the population (see 1).为降低该疾病在世界某些地区的高发病率,政府实施了成功的地中海贫血控制计划,这些计划基于向育龄人群提供或要求进行地中海贫血携带者筛查(见1)。
As a result of such programs, in many parts of the Mediterranean, the birth rate of affected newborns has been reduced by as much as 90%, through programs of education directed both to the general population and to health care providers.由于这些计划,在地中海许多地区,通过面向普通人群和医疗保健提供者的教育计划,受影响新生儿的出生率已降低了多达90%。
African American Indian Sicilian HPFH Turkish Thai (δβ)0 thalassemia German Italian (Αγδβ)0 thalassemia Hispanic English (εγδβ)0 thalassemia 5' HS ε δ β Aγ Gγ LCR –20 5' 3' –10 0 10 20 30 40 50 60 70 120 130 140 150 kb Chromosome 11p15 Note that deletions of the locus control region (LCR) abrogate the expression of all genes in the β-­globin cluster.非裔美国人、印第安人、西西里人、HPFH、土耳其人、泰国人、(δβ)0地中海贫血、德国人、意大利人、(Αγδβ)0地中海贫血、西班牙裔、英国人、(εγδβ)0地中海贫血、5' HS、ε、δ、β、Aγ、Gγ、LCR、–20、5'、3'、–10、0、10、20、30、40、50、60、70、120、130、140、150 kb、染色体11p15。注意:基因座控制区(LCR)的缺失会消除β-球蛋白簇中所有基因的表达。
The deletions responsible for δβ-­thalassemia, Aγδβ-­thalassemia, and HPFH overlap (see text).导致δβ-地中海贫血、Aγδβ-地中海贫血和HPFH的缺失重叠(见正文)。
HPFH, Hereditary persistence of fetal hemoglobin; HS, hypersensitive sites.HPFH,遗传性胎儿血红蛋白持续存在症;HS,超敏位点。
45/46
Screening Restricted to Extended Families.
Ch12 — Segment 45
Screening Restricted to Extended Families.[TL:failed]
The initiation of screening programs for thalassemia can be a major economic and logistical challenge.[TL:failed]
However, work in Pakistan and Saudi Arabia has demonstrated the effectiveness of a screening strategy that may be broadly applicable in countries where consanguineous marriages are common.[TL:failed]
In the Rawalpindi region of Pakistan, β-­thalassemia was found to be largely restricted to a specific group of families that came to attention because there was an identifiable index case (see Chapter 7).[TL:failed]
In 10 extended families with such an index case, testing of almost 600 persons established that ~8% of the married couples examined consisted of two carriers; outside of these 10 families, no couple at risk was identified among 350 randomly selected pregnant people and their partners.[TL:failed]
All carriers reported that the information provided was used to avoid further pregnancy if they already had two or more healthy children or, for couples with only one or no healthy children, for prenatal diagnosis.[TL:failed]
Although the long-­term impact of this program must be established, extended family screening of this type may contribute importantly to the control of recessive diseases in parts of the world where a cultural preference for consanguineous marriage is present.[TL:failed]
In other words, because of consanguinity, disease gene variants are trapped within extended families, so that an affected child indicates an extended family at high risk for the disease.[TL:failed]
The initiation of carrier testing and prenatal diagnosis programs for thalassemia requires not only the education of the public and of physicians but the establishment of skilled central laboratories and the consensus of the population to be screened (see Box).[TL:failed]
Whereas population-­wide programs to control thalassemia are inarguably less expensive than the cost of lifetime care for a large population of affected individuals, the temptation for governments or physicians to pressure individuals into accepting such programs must be avoided.[TL:failed]
The autonomy of the individual in reproductive decision making—a bedrock of modern bioethics, and the cultural and religious views of their communities must be respected.[TL:failed]
GENERAL REFERENCES Higgs DR, Engel JD, Stamatoyannopoulos G: Thalassaemia, Lancet 379:373–­383, 2012.[TL:failed]
Higgs DR, Gibbons RJ: The molecular basis of α-­thalassemia: a model for understanding human molecular genetics, Hematol Oncol Clin North Am 24:1033–­1054, 2010.[TL:failed]
Mc Cavit TL: Sickle cell disease, Pediatr Rev 33:195–­204, 2012.[TL:failed]
Roseff SD: Sickle cell disease: a review, Immunohematol 25:67–­74, 2009.[TL:failed]
Taher AT, Musallam KM, Cappellini MD: β-­thalassemias, N Engl J Med 384:727–­743, 2021.[TL:failed]
Weatherall DJ: The role of the inherited disorders of hemoglobin, the first “molecular diseases,” in the future of human genetics, Annu Rev Genomics Hum Genet 14:1–­24, 2013.[TL:failed]
REFERENCES FOR SPECIFIC TOPICS Bauer DE, Orkin SH: Update on fetal hemoglobin gene regulation in hemoglobinopathies, Curr Opin Pediatr 23:1–­8, 2011.[TL:failed]
Ingram VM: Gene mutations in human haemoglobin: the chemical difference between normal and sickle cell haemoglobin, Nature 180:326–­328, 1957..[TL:failed]
ETHICAL AND SOCIAL ISSUES RELATED TO POPULATION SCREENING FOR β-­THALASSEMIAa Worldwide, approximately 70,000 infants are born each year with β-­thalassemia, at high economic cost to health care systems and at great emotional cost to affected families.[TL:failed]
To identify individuals and families at increased risk for the disease, screening is done in many countries.[TL:failed]
National and international guidelines recommend that screening not be compulsory and that education and genetic counseling should inform decision making.[TL:failed]
Widely differing cultural, religious, economic, and social factors significantly influence the adherence to guidelines.[TL:failed]
For example: In Sardinia, a program initiated in 1975 involves voluntary screening, followed by testing of the extended family once a carrier is identified.[TL:failed]
In Greece, screening is voluntary, is available both premaritally and prenatally, requires informed consent, is widely advertised by the mass media and in military and school programs, and is accompanied by genetic counseling for carrier couples.[TL:failed]
In Iran and Turkey, these practices differ only in that screening is mandatory premaritally (but in all countries with mandatory screening, carrier couples have the right to marry if they wish).[TL:failed]
Major obstacles to more effective population screening for β-­thalassemia.[TL:failed]
The principal obstacles include the facts that pregnant individuals may feel overwhelmed by the array of tests offered to them, many health professionals have insufficient knowledge of genetic disorders, appropriate education and counseling are costly and time consuming, it is commonly misunderstood that informing an individual about a test is equivalent to obtaining consent, and the effectiveness of mass education varies greatly, depending on the community or country.[TL:failed]
The effectiveness of well-­executed β-­thalassemia screening programs.[TL:failed]
In populations where β-­thalassemia screening has been effectively implemented, the reduction in the incidence of the disease has been striking.[TL:failed]
For example, in Sardinia, screening between 1975 and 1995 reduced the incidence from 1 per 250 to 1 per 4000 individuals.[TL:failed]
Similarly, in Cyprus, the incidence of affected births fell from 51 in 1974 to none up to 2007. a[TL:failed]
46/46
The Molecular Basis of Genetic Disease 249 Ingram VM: Specific chemical difference between the globins of normal human a…
Ch12 — Segment 46
The Molecular Basis of Genetic Disease 249 Ingram VM: Specific chemical difference between the globins of normal human and sickle-­cell anaemia haemoglobin, Nature 178:792–­ 794, 1956.[TL:failed]
Kervestin S, Jacobson A: NMD, a multifaceted response to premature translational termination, Nat Rev Mol Cell Biol 13:700–­712, 2012.[TL:failed]
Pauling L, Itano HA, Singer SJ, et al: Sickle cell anemia, a molecular disease, Science 110:543–­548, 1949.[TL:failed]
Sankaran VG, Lettre G, Orkin SH, et al: Modifier genes in mendelian disorders: the example of hemoglobin disorders, Ann N Y Acad Sci 1214:47–­56, 2010.[TL:failed]
Steinberg MH, Sebastiani P: Genetic modifiers of sickle cell disease, Am J Hematol 87:795–­803, 2012.[TL:failed]
Weatherall DJ: The inherited diseases of hemoglobin are an emerging global health burden, Blood 115:4331–­4336, 2010.[TL:failed]
PROBLEMS 1.[TL:failed]
A newborn female dies of hydrops fetalis attributable to α-thalassemia.[TL:failed]
Draw a pedigree with genotypes illustrating to the biological parents the genetic basis of this disease.[TL:failed]
Explain why a Melanesian couple, whom they met in the hematology clinic and who both also have α-thalassemia trait, are unlikely to have a similarly affected child.[TL:failed]
Why are most individuals with β-thalassemia compound heterozygotes for causal DNA variants in the β-globin gene?[TL:failed]
In what situation(s) might you anticipate that an individual with β-thalassemia would likely have two identical β-globin alleles (i. e., to be homozygous for the causal DNA variant)?[TL:failed]
Tony, a young male of self-identified Italian ancestry, is found to havehas non-transfusion-dependent β-thalassemia, with a hemoglobin concentration of 7 g/d L (normal, amounts are 10 to 13 g/d L).[TL:failed]
When you perform a Northern blot of his reticulocyte RNA, you unexpectedly find three β-globin mRNA bands, one of normal size, one larger than normal, and one smaller than normal.[TL:failed]
What variant mechanism(s) could account for the presence of three bands like this observation in an individual with β-thalassemia?[TL:failed]
In this patient, the fact that the anemia is mild suggests that a significant fraction of normal β-globin mRNA is being made.[TL:failed]
What type(s) of variants would allow this to occur?[TL:failed]
A man is heterozygous for Hb M Saskatoon, a hemoglobin missense structural alteration in which the normal amino acid His is replaced by Tyr at position 63 of the β chain.[TL:failed]
His mate is heterozygous for Hb M Boston, in which His is replaced by Tyr at position 58 of the α chain.[TL:failed]
Heterozygosity for either of these mutant variant alleles produces methemoglobinemia.[TL:failed]
Outline the possible genotypes and phenotypes of their offspring.[TL:failed]
A child has a paternal uncle and a maternal aunt with sickle cell disease; both of her parents do not have sickle cell disease.[TL:failed]
What is the probability that the child has sickle cell disease?[TL:failed]
A woman has sickle cell trait, and her mate is heterozygous for Hb C.[TL:failed]
What is the probability that their child has no abnormal hemoglobin?[TL:failed]
Match the following: 8.[TL:failed]
Exome sequencing is organized for a child with unexplained intellectual disability, who also has non-transfusion-dependent β-thalassemia.[TL:failed]
Although this reveals a genetic cause for the intellectual disability, the molecular basis of the β-thalassemia phenotype is not elucidated, as only a single heterozygous pathogenic variant in the β-globin gene is identified.[TL:failed]
List possible explanations.[TL:failed]
What are some possible explanations for the fact that thalassemia control programs, such as the successful one in Sardinia, have not reduced the birth rate of newborns with severe thalassemia to zero?[TL:failed]
For example, in Sardinia from 1999 to 2002, approximately two to five such infants were born each year. ___________ complex β-thalassemia ___________ β+-thalassemia ___________ number of α-globin genes missing in Hb H disease ___________ two different variant alleles at a locu ___________ ATR-X syndrome ___________ insoluble β chains ___________ number of α-globin genes missing in hydrops fetalis with Hb Barts ___________ locus control region ___________ α−/α− genotype ___________ increased Hb A2 1. detectable Hb A 2. three 3. β-thalassemia 4. α-thalassemia 5. high-level β-chain expression 6. α-thalassemia trait 7. compound heterozygote 8. δβ genes deleted 9. four 10. intellectual disability[TL:failed]