Part 5

PART 5: Complex Inheritance and Population Genetics

← Back to Genetics Contents
Ch9 — Complex Inheritance of Common Multifactorial Disorders (28) Ch10 — Population Genetics for Genomic Medicine (25)

Population Genetics for Genomic Medicine

1/53
Population Genetics for Genomic Medicine INTRODUCTION TO POPULATION GENETICS FOR GENOMIC MEDICINE We have explored in pr…
Ch10 — Segment 1
Population Genetics for Genomic Medicine INTRODUCTION TO POPULATION GENETICS FOR GENOMIC MEDICINE We have explored in previous chapters the nature of genetic and genomic variation, mechanisms and types of mutation, and the inheritance of alleles (genetic variants) in families.基因组医学的群体遗传学 基因组医学群体遗传学导论 在前几章中,我们探讨了遗传和基因组变异的本质、突变的机制和类型,以及等位基因(遗传变异)在家族中的遗传。
Throughout, we have alluded to observed differences in allele frequencies across the globe, whether assessed by examining single nucleotide variants (SNVs), insertions and deletions (indels), or copy number variants (CNVs) in the genomes of many thousands of individuals.在整个过程中,我们提到了全球等位基因频率的观察差异,无论是通过检查数千个个体的基因组中的单核苷酸变异(SNV)、插入缺失(indel)还是拷贝数变异(CNV)来评估。
We have also discussed how the incidence and prevalence of some genetic disorders may differ among populations, and how discoveries have been made by selectively sampling individuals with specific phenotypes.我们还讨论了一些遗传疾病的发病率和患病率在不同人群中可能存在的差异,以及如何通过选择性抽样具有特定表型的个体来取得发现。
Here we take a deeper look into the assumptions and limitations of how populations are defined in genomics research and medicine, considering in greater detail the underlying forces that shift or maintain allele frequencies over time.在此,我们更深入地探讨基因组学研究和医学中如何定义人群的假设和局限性,更详细地考虑随时间推移改变或维持等位基因频率的潜在力量。
Identifying genetic susceptibilities to diseases is a key objective of medical genetics, with a critical role in clinical diagnosis and genetic counseling.识别疾病的遗传易感性是医学遗传学的一个关键目标,在临床诊断和遗传咨询中起着关键作用。
Concepts and observations from population genetics inform our understanding of the genetic architecture of health and disease by observing how variants may differ in frequency, effect size, and phenotypic expression in human populations.群体遗传学的概念和观察结果通过观察变异在人群中的频率、效应大小和表型表达如何不同,为我们理解健康和疾病的遗传结构提供了信息。
When we refer to a population frequency, we consider a hypothetical gene pool as a collection of all the alleles at a particular locus for the entire population.当我们提到群体频率时,我们将一个假想的基因库视为整个群体中特定基因座所有等位基因的集合。
The frequency of an allele is thus its proportion among all alleles at the same locus in a population.因此,等位基因的频率是其在群体中同一基因座所有等位基因中所占的比例。
In this chapter, we describe a central organizing concept of population genetics, Hardy-Weinberg equilibrium (HWE), and its utility for genomics and clinical genetics in understanding the relationship between allele and genotype frequencies.在本章中,我们描述了群体遗传学的一个核心组织概念——哈迪-温伯格平衡(HWE),以及它在理解等位基因频率和基因型频率之间关系方面对基因组学和临床遗传学的实用性。
Assumptions of the HardyWeinberg principle are considered in the context of factors that cause true or apparent deviation from equilibrium in real, as opposed to idealized, populations.在导致真实群体(而非理想化群体)出现真正或明显偏离平衡的因素背景下,考虑了哈迪-温伯格原理的假设。
Clinical genetics is primarily concerned with rare, often de novo mutations that cause genetic conditions and may severely impact function, with similar incidences across populations.临床遗传学主要关注罕见的、通常是新发突变,这些突变导致遗传状况并可能严重损害功能,且在不同人群中发病率相似。
Genetic variants that impact function and reproduction are rare in nearly all human populations because they are eliminated from the gene pool through natural selection.影响功能和繁殖的遗传变异在几乎所有人类群体中都很罕见,因为它们通过自然选择从基因库中被淘汰。
These variants are most often de novo – not inherited – so there is typically no link between genetic ancestral origins and the incidence of these pathogenic mutations.这些变异大多数是新发的(非遗传的),因此遗传祖先起源与这些致病突变的发生率之间通常没有关联。
Most human genomic variation is shared among all groups, so there are exceedingly rare cases of clinically relevant variants that are found exclusively in one category or group of patients.大多数人类基因组变异在所有群体中共享,因此,仅在某一类别或一组患者中发现的临床相关变异极其罕见。
The exception to this rule is when a defined ancestral group has experienced a bottleneck that reduces and then regrows the population with a subset of its original genetic variants.这条规则的例外情况是,当某个特定的祖先群体经历了瓶颈效应,导致群体规模缩减,然后利用其原始遗传变异的一个子集重新增长。
This increases the chance of maintaining a pathogenic variant in the population, as selective pressure is weaker than in a more genetically diverse population.这增加了致病变异在群体中持续存在的机会,因为选择压力比在遗传多样性更高的群体中更弱。
Population genetics is the quantitative study of the distribution of genetic variation in populations, including trends in the frequencies of genes and genotypes over time.群体遗传学是对群体中遗传变异分布的定量研究,包括基因和基因型频率随时间变化的趋势。
Differences in the frequencies of alleles that cause genetic disease are of particular interest to the medical geneticist and genetic counselor because they contribute to differences in disease risk among populations.引起遗传疾病的等位基因频率的差异对医学遗传学家和遗传咨询师尤其重要,因为它们导致了不同人群之间疾病风险的差异。
Nevertheless, it is important to consider what we do not know and examine our baseline assumptions about genetic etiologies of health and disease.尽管如此,考虑我们所不知道的内容并审视我们关于健康和疾病遗传病因的基本假设仍然很重要。
We present clinical examples to illustrate how the creation of genomic knowledge, as in all fields, depends on who is included in the underlying research, how their attributes are represented, and what categories are used to classify or stratify people into groups.我们通过临床实例来说明,正如在所有领域一样,基因组知识的产生取决于基础研究中包含了哪些人,他们的属性如何被代表,以及使用哪些类别将人们分类或分层为不同群体。
The distributions of genetic variants and attributes in families, communities, and geographic regions are driven by a combination of social, environmental, and biological factors.遗传变异和特征在家庭、社区和地理区域的分布是由社会、环境和生物因素共同驱动的。
Social scientists, anthropologists, and evolutionary biologists employ mathematical descriptions of shifting allele frequencies across geographic regions and time to reconstruct our evolutionary histories.社会科学家、人类学家和进化生物学家利用跨地理区域和时间的等位基因频率变化的数学描述来重建我们的进化史。
Knowing about differences in allele frequencies across populations may be important for physicians to predict an increased likelihood of certain conditions in patients with specific ancestral origins linked to a pathogenic了解不同人群之间等位基因频率的差异可能对医生预测具有特定祖先起源(与致病性相关)的患者患某些疾病的可能性增加很重要。
2/53
However, it is important to keep in mind that most genetic variation that is used to describe ancestral origins is selec…
Ch10 — Segment 2
However, it is important to keep in mind that most genetic variation that is used to describe ancestral origins is selectively neutral or has an unknown biological function.然而,需谨记,用于描述祖先起源的大多数遗传变异是选择中性的,或具有未知的生物学功能。
With today’s access to genome sequencing technology, considering all types of variation, we know that all humans share an average of 99% of our DNA sequences across the genome.借助当今的基因组测序技术,考虑到所有类型的变异,我们知道所有人类在整个基因组中平均共享99%的DNA序列。
The relatively small proportion of our genomes that do vary (polymorphism) do not differentiate us into discrete social categories; on the contrary, greater genomic variation has been observed within broad groupings (e. g., based on “race” or “ethnicity”) than between them.我们基因组中确实存在变异的较小部分(多态性)并未将我们区分为离散的社会类别;相反,在宽泛的分组(例如基于“种族”或“民族”)内部观察到的基因组变异大于组间差异。
Nevertheless, humans tend to make meaning out of anecdotal observations, attributing physical or health-related differences we observe to underlying attributes of broad semantic categories that reinforce social stereotypes.尽管如此,人类倾向于从轶事观察中赋予意义,将我们观察到的身体或健康相关差异归因于宽泛语义类别的基础属性,从而强化社会刻板印象。
This is one mechanism by which biological or scientific racism is enacted in genetics.这是生物学或科学种族主义在遗传学中得以实施的一种机制。
Data subjects are often stratified into groups to make statistical comparisons, which (with sufficient sample sizes) may yield average differences that offer a biased confirmation of natural classes.数据主体常被分层为组以进行统计比较,这(在样本量足够时)可能产生平均差异,从而对自然类别提供有偏见的确认。
Some have used these findings to argue for the inevitability of systemic inequities.一些人利用这些发现来论证系统性不平等的不可避免性。
This chapter touches on the history and impact of classifying humans into nominal groupings, making the case for genetics researchers and clinicians to understand the origins and assumptions underlying conceptual and functional models used to create and apply knowledge.本章涉及将人类分类为名义分组的歷史及影响,论证遗传学研究人员和临床医生有必要理解用于创造和应用知识的概念与功能模型的起源和假设。
Statistical methods have been used to identify variants that can differentiate many individuals into broad “continental ancestry” groupings based on allele frequencies observed across geographic regions.统计方法已被用于识别基于地理区域观察到的等位基因频率、能将许多个体区分为宽泛“大陆祖先”分组的变异。
Such ancestry informative markers (AIMs) often lack functional relevance, but they have been used as a proxy for genomic background, which may inflate estimated genetic distances between groups.此类祖先信息标记通常缺乏功能相关性,但已被用作基因组背景的代理,这可能夸大估计的组间遗传距离。
When such variants are associated with complex, multifactorial traits (with both genetic and environmental influences) such as obesity, diabetes, heart disease, and asthma, it is difficult to determine the true causal factors driving statistically significant associations, because social determinants of health influence the environment differently across sociocultural groups.当此类变异与复杂、多因素性状(受遗传和环境共同影响)如肥胖、糖尿病、心脏病和哮喘相关联时,难以确定驱动统计学显著关联的真正因果因素,因为健康的社会决定因素在不同社会文化群体中对环境影响不同。
A central challenge for human genomics research is thus teasing apart causal factors that drive phenotypic variation from confounders that are associated with patterns of both genotype and phenotype variation.因此,人类基因组研究的一个核心挑战是,从与基因型和表型变异模式相关的混杂因素中,分离出驱动表型变异的因果因素。
Stratification is not inherently problematic as a statistical strategy to reduce the complexity of genomic and environmental backgrounds, but there are deep scientific and ethical flaws in approaches that collapse these nuanced dimensions of variation into static population descriptors such as race, ethnicity, and ancestry.分层作为降低基因组和环境背景复杂性的统计策略本身并无固有问题,但若将变异的这些细微维度简化为种族、民族和祖先等静态人口描述符,则存在深刻的科学和伦理缺陷。
European colonization, slavery, and eugenics have institutionalized conceptual frameworks tied to categories of difference, preserving hierarchies of power.欧洲殖民、奴隶制和优生学已将与差异类别相关的概念框架制度化,维持了权力等级。
These frameworks have persisted in colonized societies and science, with implications for analytic approaches still used today in biomedical research.这些框架在殖民化社会和科学中持续存在,对至今仍用于生物医学研究的分析方法产生影响。
Racial categories are constructed using poorly defined criteria that subdivide humankind using physical appearance (i. e., skin color, hair texture, and facial features) combined with characteristic identities that have their origins in the geographical, historical, cultural, religious, and linguistic backgrounds of the communities in which an individual was born and raised.种族类别使用定义不清的标准构建,这些标准通过外貌(即肤色、发质和面部特征)结合具有地理、历史、文化、宗教和语言背景起源的特征身份来细分人类,这些背景源于个体出生和成长的社区。
Physical traits linked to racial and ethnic stereotypes may be influenced by genetics.与种族和民族刻板印象相关的身体特征可能受遗传影响。
However, genetic variation and diverse phenotypes exist across all populations, so attempts to equate continental origins with race are misguided.然而,所有人群中均存在遗传变异和多样化表型,因此试图将大陆起源等同于种族的做法是错误的。
DNA alone cannot be used to assign someone to a social identity group, and it is primarily nongenetic, social and systemic factors that create the conditions for observed group-level differences in health outcomes.仅凭DNA无法将个体归入社会身份群体,主要是非遗传的、社会和系统性因素造成了观察到的群体水平健康结果差异的条件。
Concept of race and racism are important in discussions of social and health policy, tracking disparities in outcomes and social determinants of health.种族和种族主义概念在社会和卫生政策讨论、追踪结果差异及健康社会决定因素中至关重要。
Longstanding systemic and structural factors create differences in health outcomes among racial groups, which can contribute to confounding and disparities in the quality of care, such as timely screening and test referrals.长期存在的系统性和结构性因素导致种族群体间健康结果差异,这可能造成混杂并加剧护理质量(如及时筛查和检测转诊)方面的差距。
Thus, race may be used as a proxy for the effects of racism, but not as a proxy for genetic background or to calculate genetic disease risk.因此,种族可作为种族主义影响的代理,但不可作为遗传背景的代理或用于计算遗传疾病风险。
Building on themes from previous chapters, we paint a detailed picture of global genomic diversity that is largely shared among all human populations, and the factors that influence shifts in relative frequencies of genetic variants over time.基于前几章的主题,我们描绘了全球基因组多样性的详细图景——该多样性在很大程度上为所有人类群体所共享——以及随时间推移影响遗传变异相对频率变化的因素。
We demonstrate the importance of African genomic diversity for all populations, highlighting the need for clinically relevant data on genetic heterogeneity, structural variation, and haplotype diversity across the globe.我们展示了非洲基因组多样性对所有群体的重要性,强调需要全球范围内关于遗传异质性、结构变异和单倍型多样性的临床相关数据。
These data will inform our understanding of genetic etiologies of health and disease in everyone and must be included in public databases.这些数据将有助于我们理解每个人健康和疾病的遗传病因,且必须纳入公共数据库。
Finally, we seek to clarify common misconceptions about human populations and diversity that could exacerbate health disparities in genomics and medicine.最后,我们力求澄清关于人类群体和多样性的常见误解,这些误解可能加剧基因组学和医学中的健康差距。
HUMAN ORIGINS AND FOUNDER EFFECTS OF SERIAL MIGRATION Our species (Homo sapiens) is comprised of more than 7 billion members, all of whom share a common ancestral lineage on the African continent dating back ~200,000 years ago.人类起源与连续迁徙的奠基者效应 我们智人物种由超过70亿成员组成,所有人共享可追溯至约20万年前非洲大陆的共同祖先谱系。
3/53
Population Genetics for Genomic Medicine 177 of variants, or alleles, across a barrier.
Ch10 — Segment 3
Population Genetics for Genomic Medicine 177 of variants, or alleles, across a barrier.《基因组医学中的群体遗传学》:177个变异或等位基因跨越屏障。
Gene flow usually involves a large population and a gradual change in allele frequencies.基因流通常涉及大群体和等位基因频率的逐渐变化。
The allele frequencies of migrant populations gradually merge into the gene pool of the population into which they have migrated, a process referred to as genetic admixture.迁移群体的等位基因频率逐渐融入其迁入群体的基因库,这一过程称为遗传混合。
The term migration is used here in the broad sense of crossing a reproductive barrier, which may be social or cultural (not necessarily geographic) and does not require physical movement from one region to another.此处“迁移”一词是广义上的穿越生殖屏障,该屏障可能是社会或文化性的(不一定是地理性的),且无需从一个区域到另一个区域的物理移动。
During serial migration events a subset of the genomic variation from the original population travels with the migratory group into the new population.在系列迁移事件中,来自原始群体的基因组变异子集随迁移群体进入新群体。
These events create founder effects, which are marked by a reduction in genomic diversity from the original, ancestral population to the new subpopulation.这些事件产生奠基者效应,其标志是从原始祖先群体到新亚群体的基因组多样性减少。
Founder events such as bottlenecks, the reduction in size and subsequent regrowth of a population whereby it loses a large portion of its original diversity, create so-called founder populations.奠基者事件,例如瓶颈效应(群体规模缩减后重新增长,从而丧失大部分原始多样性),产生了所谓的奠基者群体。
In populations that have experienced a bottleneck (e. g., European ancestry groups), we observe greater genetic homogeneity (lower heterozygosity) than would be expected, based on the overall estimated population size.在经历过瓶颈效应的群体(例如欧洲血统群体)中,我们观察到比基于总体估计群体规模所预期的更高的遗传同质性(更低的杂合度)。
Founder populations thus have lower effective population sizes (Ne) relative to their census populations, and various estimators for genetic diversity use this concept to model human population history.因此,奠基者群体的有效群体规模(Ne)相对于其普查群体规模更低,各种遗传多样性估计量利用这一概念来模拟人类群体历史。
Founder populations with small Ne are relevant for clinical genetics, because individuals with these ancestries are generally at a greater risk of accumulating genetic diseases that are rare in other populations.具有小Ne的奠基者群体与临床遗传学相关,因为这些血统的个体通常面临更高的累积其他群体中罕见的遗传疾病的风险。
Dominant deleterious mutations can rise to high frequency more easily when there are fewer potential reproductive partners, as this reduces competition in the form of alternative alleles.当潜在生殖伴侣较少时,显性有害突变更容易升至高频率,因为这减少了以替代等位基因形式存在的竞争。
For example, the high incidence of Huntington disease (Case 24) among inhabitants around Lake Maracaibo, Venezuela, resulted from genetic isolation following the introduction of a (likely European) genetic variant that causes Huntington disease.例如,委内瑞拉马拉开波湖周围居民中亨廷顿病(案例24)的高发病率,是由于引入了一个(可能来自欧洲的)导致亨廷顿病的遗传变异后发生遗传隔离所致。
Similarly, French-Canadians in the Saguenay-LacSaint-Jean region of Quebec are at greater risk for the autosomal recessive condition hereditary type I tyrosinemia.类似地,魁北克省萨格奈-圣让湖地区的法裔加拿大人患常染色体隐性遗传病——遗传性I型酪氨酸血症的风险更高。
If untreated, this condition causes hepatic failure and renal tubular dysfunction due to deficiency of fumarylacetoacetase, an enzyme in the degradative pathway of tyrosine.若不治疗,该病因延胡索酰乙酰乙酸酶(酪氨酸降解途径中的一种酶)缺乏而导致肝衰竭和肾小管功能障碍。
The disease incidence was estimated (in 1994) at 1 in 1846, with the pathogenic allele carrier frequency estimated at 1 in 22.该疾病发病率估计(1994年)为1/1846,致病等位基因携带者频率估计为1/22。
Nearly all pathogenic alleles observed in Saguenay-Lac-Saint-Jean patients are due to the same variant inherited through a shared lineage of recent common ancestor(s).在萨格奈-圣让湖患者中观察到的几乎所有致病等位基因均源于通过近期共同祖先的共享谱系遗传的同一变异。
This observation is consistent with the history of this population, which is a known founder population or genetic isolate, meaning there was reduced genetic variation among individuals who originally founded the population, followed by population growth and a paucity of new alleles flowing into it.这一观察结果与该群体的历史一致,该群体是一个已知的奠基者群体或遗传隔离群,意味着最初奠基该群体的个体间遗传变异减少,随后群体增长且新等位基因流入匮乏。
As a result of serial founder events throughout our history, nearly all human genomic variation in modern populations is only a subset of the variation that exists in Africa, the ancestral homeland of our species, and the current home of modern human populations with 50-60Kya 15Kya 60-100Kya 45Kya 45Kya Founder effect Source of founder effect Migration path 35-40Kya This map highlights demic events that began with a source population in southern Africa 60 to 100 kya and concluded with the settlement of South America approximately 12 to 14 kya.由于人类历史上的系列奠基者事件,现代群体中几乎所有人类基因组变异都只是存在于非洲(我们物种的祖先家园)的变异的一个子集,而非洲也是现代人类群体的当前家园,伴随有50-60Kya、15Kya、60-100Kya、45Kya、45Kya、奠基者效应、奠基者效应来源、迁移路径35-40Kya,这张图强调了始于60至100 kya的南部非洲源群体、结束于大约12至14 kya的南美洲定居的群体扩张事件。
Wide arrows indicate major founder events during the demographic expansion into different continental regions.宽箭头表示在人口扩张至不同大陆区域过程中的主要奠基者事件。
Colored arcs indicate the putative source for each of these founder events.彩色弧线表示每个奠基者事件的推定来源。
Thin arrows indicate potential migration paths.细箭头表示潜在的迁移路径。
Many additional migrations occurred during the Holocene.在全新世期间发生了许多额外的迁移。
From Henn BM, Cavalli-Sforza LL, Feldman MW.引自Henn BM, Cavalli-Sforza LL, Feldman MW。
The great human expansion.《伟大的人类扩张》。
Proceedings of the National Academy of Sciences of the United States of America. 109:17758–64.《美国国家科学院院刊》109:17758–64。
PMID 23077256.PMID 23077256。
4/53
more genetic diversity compared to anywhere else in the world.
Ch10 — Segment 4
more genetic diversity compared to anywhere else in the world.比世界上任何其他地方都更多的遗传多样性。
African populations and individuals with recent African ancestries have the largest effective population sizes, with genomic evidence pointing to Indigenous South Africans having the most diverse and anciently diverged genomes.非洲人群以及近期有非洲血统的个体拥有最大的有效群体规模,基因组证据表明南非原住民拥有最多样化且古老分化的基因组。
It is a common error to think of these ancestral genomes as an artifact of the past; they have in fact continued to accumulate variation and to experience effects of natural selection over time.认为这些祖先基因组是过去的产物是一个常见错误;实际上,它们随着时间的推移不断积累变异并经历自然选择的影响。
In addition to standing ancestral variation, new mutations arise – albeit slowly – and other sources of variation have been introduced in different ancestral lineages.除了已有的祖先变异外,新的突变也会出现——尽管缓慢——并且其他变异来源也被引入了不同的祖先谱系。
Evidence (mostly from studies of European ancestry populations) suggests Homo sapiens encountered other early hominid species such as Neanderthals and Denisovans – whose genetic fragments were preserved in cold, dry climates.证据(主要来自欧洲血统人群的研究)表明,智人曾遇到其他早期人科物种,如尼安德特人和丹尼索瓦人——其遗传片段在寒冷干燥的气候中得以保存。
Segments of these archaic hominids’ chromosomes appear to have flowed into human genomes through introgression approximately ~50 to 90k years ago.这些古人类的染色体片段似乎通过基因渗入在大约5万至9万年前流入人类基因组。
By comparing human genomes to DNA extracted from excavated bone fragments, ancient DNA researchers estimate that 1% to 4% of the DNA of modern humans may be derived from these species.通过将人类基因组与从出土骨碎片中提取的DNA进行比较,古DNA研究人员估计,现代人类1%至4%的DNA可能源自这些物种。
While there is little evidence of archaic hominid introgression into human populations in regions of the world marked by hot and humid climates, (which are not conducive to the preservation of DNA in ancient remains), it is likely that such introgression events occurred wherever other hominid species met and mingled with ancient Homo sapiens.尽管在炎热潮湿气候(不利于古代遗骸中DNA保存)的世界地区,几乎没有古人类基因渗入人类群体的证据,但很可能在其他人科物种与古代智人相遇并混杂的任何地方都发生过此类渗入事件。
DEFINING HUMAN POPULATIONS TO CHARACTERIZE DIVERSITY There are many ways to define a population – whether criteria are based on geography, societal and cultural factors, geopolitical or national borders, or any other features that may be considered characteristic of individuals within a group.定义人类群体以描述多样性 定义群体有多种方式——标准可基于地理、社会文化因素、地缘政治或国家边界,或任何其他可能被视为群体内个体特征的特征。
Whether an individual or group is designated as belonging to a “population” depends on the classification framework, and how it is implemented in practice.一个个体或群体是否被指定属于某个"群体"取决于分类框架及其在实际中的应用方式。
Many designated “populations” have nothing to do with genetics – such as the census population of a large, urban metropolis.许多被指定的"群体"与遗传学毫无关系——例如一个大都市的人口普查群体。
The specific criteria that are used to define a population often flow from certain research or clinical objectives and may also be influenced by social, cultural, and political factors that privilege certain approaches.用于定义群体的具体标准通常源于特定的研究或临床目标,也可能受到偏重某些方法的社会、文化和政治因素的影响。
Who has the authority to determine such classification criteria also plays an important role in defining human populations.谁有权确定此类分类标准也在定义人类群体中扮演重要角色。
Decisions made by researchers and clinicians about how to classify participants or patients into groups strongly influence population-level estimates such as allele frequencies and disease prevalence.研究人员和临床医生关于如何将受试者或患者分组到不同群体的决策强烈影响群体水平估计,如等位基因频率和疾病患病率。
Methods used to detect and interpret combinations and mixtures of these alleles, and to investigate how they impact our health, play an important role in what is generally accepted about average group-level differences.用于检测和解释这些等位基因的组合与混合,以及研究它们如何影响健康的方法,在关于平均群体水平差异的普遍认知中起着重要作用。
For example, population-level statistics such as point estimates that serve to represent averages across an entire group (e. g., allele frequencies) are influenced by who is included and excluded from the group, and their level of genetic relatedness.例如,用于代表整个群体平均值的点估计(如等位基因频率)等群体水平统计量,受群体中纳入和排除的个体及其遗传相关程度的影响。
The magnitude of genetic similarities and differences among individuals defined by some category or population descriptor thus impacts the accuracy of any given point estimate for that category.因此,由某种类别或群体描述符定义的个体间遗传相似性和差异的程度,影响该类别任何给定点估计的准确性。
It also influences the precision with which such a point estimate can be used to make predictions about individuals in the group.它还影响此类点估计用于对群体中个体进行预测的精确度。
For example, population categories such as “Asian” and “African” are so broad that they describe more than half of the world population and land mass.例如,"亚洲人"和"非洲人"等群体类别非常宽泛,涵盖了世界一半以上的人口和陆地面积。
The vast genetic and environmental heterogeneity within such groups call into question the reliability of point estimates based on such groupings.这些群体内部巨大的遗传和环境异质性使基于此类分组的点估计的可靠性受到质疑。
In contrast, populations with less total variability are more likely to yield group-level estimates that reflect a greater number of individuals in the category.相比之下,总变异性较小的群体更可能产生反映该类别中更多个体的群体水平估计。
For example, populations that have experienced bottlenecks have less overall diversity in the underlying gene pool and population mean and median values may be more informative.例如,经历过瓶颈效应的群体在潜在基因库中整体多样性较低,群体平均值和中位数值可能更具信息量。
This is important because it means that research methods or approaches to estimating disease risk based on point estimates have differential validity and utility across population groups, benefiting those with less variability (i. e., European ancestries).这一点很重要,因为这意味着基于点估计的研究方法或疾病风险评估方法在不同群体间具有不同的有效性和实用性,有利于变异性较小的群体(即欧洲血统)。
Many clinical case studies present population allele frequencies as point estimates corresponding to different race, ethnicity, or ancestry categories.许多临床病例研究将群体等位基因频率作为对应于不同种族、民族或血统类别的点估计呈现。
Unfortunately, these terms are conceptualized and used in vastly different ways among clinical genetics professionals and researchers, so summary statistics that are based on such poorly or broadly defined groupings may be unreliable and inconsistent, depending on the underlying characteristics of the groups.不幸的是,这些术语在临床遗传学专业人员和研究人员中的概念化和使用方式大相径庭,因此基于这些定义模糊或宽泛分组的汇总统计可能不可靠且不一致,具体取决于各群体的潜在特征。
Genetic ancestry and ethnicity represent different kinds of information, so direct comparisons cannot be made between them (despite misguided attempts to ‘valide’ self-identified race or ethnicity using estimated proportions of DNA shared with ‘known’ population reference datasets).遗传血统和民族代表不同类型的信息,因此它们之间不能直接比较(尽管有人误导性地试图利用与"已知"群体参考数据集共享DNA的估计比例来"验证"自我认定的种族或民族)。
Most people are not limited to one ancestry or ethnic identity, and these identity concepts can change over time.大多数人不限于单一血统或民族身份,这些身份概念可能随时间变化。
As such, ontological frameworks and methodologies in population genetics that force a single-group assignment are conceptually inconsistent, insufficient for characterizing the spectrum of human cultural and genomic diversity, and are becoming increasingly obsolete.因此,群体遗传学中强制单一群体归属的本体框架和方法在概念上不一致,不足以描述人类文化和基因组多样性的范围,并且正变得越来越过时。
For example, while “Ashkenazi Jewish” represents a cultural or ethnic designation without specific geographic constraints, its related definition of having “ancestry from Africa, the Middle East, or the Mediterranean” is geographically diverse and includes例如,虽然"阿什肯纳兹犹太人"代表一种没有特定地理限制的文化或民族称谓,但其相关定义"血统来自非洲、中东或地中海"在地理上是多样化的,并且包括
5/53
Population Genetics for Genomic Medicine 179 a vast array of social and cultural identities.
Ch10 — Segment 5
Population Genetics for Genomic Medicine 179 a vast array of social and cultural identities.基因组医学的群体遗传学179:广泛的社会和文化身份。
For many years, carrier screening for autosomal recessive diseases such as for Tay-Sachs disease used a risk-based strategy that relied on self-described ethnicity; this approach is now known to introduce inaccuracy into the screening process.多年来,针对常染色体隐性遗传病(如泰-萨克斯病)的携带者筛查采用基于风险的策略,该策略依赖于自我描述的种族;现在已知这种方法会在筛查过程中引入不准确性。
Thus, recent recommendations from the American College of Medical Genetics and Genomics (ACMG) note that carrier screening paradigms should be agnostic to race, ethnicity, and ancestry.因此,美国医学遗传学与基因组学学会(ACMG)近期的建议指出,携带者筛查范式应对种族、民族和血统保持不可知论态度。
Carrier screening, referrals for genetic testing, and reporting risk predictions to patients or other care providers have relied on patient self-reported information about race, ethnicity, or ancestry.携带者筛查、基因检测转诊以及向患者或其他医疗服务提供者报告风险预测,一直依赖于患者自我报告的种族、民族或血统信息。
Certain populations may be at higher risk for conditions that are not routinely screened for, but only individuals who have self-identified with (or for whom a provider has determined their membership in) these populations would have access to those screenings.某些人群可能对未常规筛查的疾病具有更高的风险,但只有那些自我认同(或由医疗服务提供者确定其属于)这些人群的个体才能获得这些筛查。
In some circumstances, physicians may not think to refer patients for genetic testing unless they fit the stereotype of the population known to be at high risk, and insurance companies may not cover genetic testing for certain conditions unless a patient has disclosed information about their background that is consistent with population-specific elevated risk of disease.在某些情况下,除非患者符合已知高风险人群的刻板印象,否则医生可能不会考虑转诊患者进行基因检测;而保险公司可能不会为某些疾病的基因检测提供报销,除非患者披露的背景信息与特定人群的疾病风险升高一致。
Thus, whether and how patients are classified into population groups or categories is of great importance here.因此,患者是否以及如何被归入人群群体或类别在此至关重要。
Expanded carrier screening has received much attention and has the potential to alleviate some of the pitfalls of relying on proxy measures and poorly defined population groups to guide clinical decision-making.扩展性携带者筛查已受到广泛关注,并有可能减轻依赖代理指标和定义不清的人群群体来指导临床决策的一些弊端。
However, once the results of genetic tests are received, there are additional uses for population-level information about patients that inform the interpretation and curation of findings.然而,一旦收到基因检测结果,关于患者的人群水平信息还有额外用途,可用于指导结果的解读和整理。
A survey of clinical genetics professionals and researchers, conducted mainly in the United States, revealed that patient self-reported race and ethnicity information may be used to make decisions about ordering tests, interpreting results, and reporting findings to patients.一项主要在美国进行的针对临床遗传学专业人员和研究人员的调查显示,患者自我报告的种族和民族信息可能被用于决定检测项目的选择、结果的解读以及向患者报告发现。
At least 18% of respondents reported that this information may have been entered into a patient’s medical record without asking them directly.至少18%的受访者报告称,这些信息可能是在未直接询问患者的情况下被录入患者病历的。
Fewer than 5% of clinical genetics professionals who responded to the survey reported that they conduct ancestry analyses or have access to ancestry estimates based on patients’ DNA for the purpose of clinical variant interpretation and gene curation.在参与调查的临床遗传学专业人员中,不到5%的人报告称他们进行血统分析或能够获取基于患者DNA的血统估计,以用于临床变异解读和基因整理。
This means that social categories are used as a proxy for genomic background in clinical genetics.这意味着在临床遗传学中,社会类别被用作基因组背景的代理指标。
Cultural and political contexts shape the way we collect and use information about patients’ social/cultural identities and ancestral backgrounds – and asking for a person’s race or ethnicity is illegal in some parts of the world.文化和政治背景决定了我们收集和使用患者社会/文化身份及祖先背景信息的方式——而在世界某些地区,询问一个人的种族或民族是非法的。
It is important for clinical genetics professionals to know about the risks of relying too heavily on this information to make predictions or decisions in the clinical setting. 1 offers descriptions of “race,” “ethnicity,” and “ancestry” that have been useful for genomics researchers..临床遗传学专业人员了解在临床环境中过度依赖这些信息进行预测或决策的风险至关重要。1提供了对“种族”、“民族”和“血统”的描述,这些描述对基因组学研究人员很有用。
DESCRIPTIONS OF RACE, ETHNICITY, AND ANCESTRY “Race,” “ethnicity,” and “ancestry” are often used interchangeably, yet they have no universal definitions.种族、民族和血统的描述 “种族”、“民族”和“血统”这三个词经常互换使用,但它们没有通用的定义。
We provide brief descriptions of our usage below.我们在下面提供对用法的简要描述。
For extensive discussion in context of genomics, including recommendations from professional organizations, see Banda et al.关于基因组学背景下的广泛讨论,包括专业组织的建议,参见Banda等人。
(2015); Mersha and Abebe (2015); Race, Ethnicity, and Genetics Working Group (2005).(2015); Mersha and Abebe (2015); 种族、民族与遗传学工作组 (2005)。
Race: A culturally and politically charge term, for which definitions and meaning are context-specific.种族:一个具有文化和政治敏感性的术语,其定义和含义因情境而异。
Race is related to individual and/or group identity and is often linked to stereotypes of visible physical attributes such as skin and hair pigmentation.[TL:missing]
The concept of race is tightly linked to social power dynamics and has historically been used to justify hierarchies of power, discrimination, and oppression in an unequal society.[TL:missing]
Social and cultural conditions may differ among racial groups, on average, and these differences may lead to environmental effects such as chronic stress and unequal access to goods and services, including healthcare and nutrition.[TL:missing]
These inequities can affect environmental risk for complex diseases and/or potentially interact with genetics to affect risk.[TL:missing]
Ethnicity: Describes people as belonging to cultural groups, usually on the basis of shared language, traditions, foods, etc.[TL:missing]
Ethnicity has often been used interchangeably with race and is similarly ambiguous.[TL:missing]
To the extent that traits are affected by social and environmental differences, ethnicity has previously served as a proxy for health and disease risk at the population level as a result of social, cultural and community effects described above.[TL:missing]
There is no universal agreement on a system of “ethnic” groupings worldwide.[TL:missing]
Some ethnic groups may share genetic factors due to similar ancestral origins; other groups may be more social and cultural in nature.[TL:missing]
Ancestry: Meaning varies by context.[TL:missing]
Here, we use the term to denote genetic ancestry, a description of the population(s) from which an individual’s recent biological ancestors originated, as reflected in the DNA inherited from those ancestors.[TL:missing]
Genetic ancestry can be estimated via comparison of participant’s genotypes to global reference populations, so incomplete availability of these references can create biased estimates.[TL:missing]
We note that different methods of calculating genetic ancestry can yield different results.[TL:missing]
Thus, discreet labeling of ancestral populations oversimplifies the complexity of human genetic variation and demography.[TL:missing]
Nevertheless, accounting for systemic differences in allele frequencies and linkage disequilibrium is necessary for genetic analyses.[TL:missing]
In this paper, diversity in genomics is described primarily in terms of ancestry.[TL:missing]
From Peterson RE, Kuchenbaecker K, Walters RK, et al: Genome-wide association studies in ancestrally diverse populations: Opportunities, methods, pitfalls, and recommendations, Cell 179(3):589-603, 2019.[TL:missing]
6/53
History and Influence of Eugenics on Population Genetics Countless atrocities have been committed in the name of perceiv…
Ch10 — Segment 6
History and Influence of Eugenics on Population Genetics Countless atrocities have been committed in the name of perceived differences among human beings, from oppression, discrimination, and displacement to slavery and genocide.优生学的历史及其对群体遗传学的影响——无数暴行以人类之间所谓差异的名义被犯下,从压迫、歧视、流离失所到奴隶制和种族灭绝。
In the United States, Jim Crow laws were enacted upon the abolition of slavery and persisted overtly through the 1960s, which forced segregation of “Black” and “white” people to preserve exploitative power dynamics and justify economic and social injustice.在美国,废除奴隶制后颁布了吉姆·克劳法,并公开持续到20世纪60年代,这些法律强制实行“黑人”与“白人”的隔离,以维持剥削性权力动态,并为经济和社会不公提供正当理由。
The ideological underpinning of segregating hospitals and clinical care was that white and Black people needed different medical care due to perceived differences in biology.隔离医院和临床护理的意识形态基础在于,认为白人和黑人因生物学上的所谓差异需要不同的医疗护理。
However, there are no unique, fundamental genetic differences between groups of people who identify as “white” or “Black,” and the persistence of this racial binary in genomics research has contributed to misconceptions and harm.然而,自认为“白人”或“黑人”的人群之间并不存在独特的、根本性的遗传差异,这种种族二元论在基因组学研究中的持续存在导致了误解和伤害。
Many in human genomics may not realize how the field’s history is rooted in the American Eugenics Movement, which provided a false sense of scientific legitimacy to Nazi propaganda during WWII.人类基因组学领域的许多人可能没有意识到,该领域的历史根植于美国优生运动,而这场运动在二战期间为纳粹宣传提供了虚假的科学合法性。
Indeed, the Annals of Human Genetics was originally called the Annals of Eugenics, founded in 1925 by an English eugenics thought leader, Francis Galton.事实上,《人类遗传学年鉴》最初名为《优生学年鉴》,由英国优生学思想领袖弗朗西斯·高尔顿于1925年创立。
Mentees of Galton’s became the journal’s editors, including Karl Pearson and R.高尔顿的学生们成为该期刊的编辑,包括卡尔·皮尔逊和R.
Fisher, who developed statistical methods we still use today.费舍尔,他们开发了至今仍在使用的统计方法。
They include the chi-square test, p-values, analysis of variance (ANOVA), and principal component analysis (PCA).这些方法包括卡方检验、p值、方差分析(ANOVA)和主成分分析(PCA)。
Another of Galton’s mentees, Charles Davenport, used taxonomic frameworks classify humans, the scientific endeavor of biological racism.高尔顿的另一位学生查尔斯·达文波特利用分类学框架对人类进行分类,这是生物种族主义的科学尝试。
The idea that racial groups were genetically distinct predates eugenics, back to early European slave traders who devized categories that dehumanized people with darker skin to justify their enslavement and exploitation.种族群体在遗传上截然不同的观念早于优生学,可追溯到早期的欧洲奴隶贩子,他们编造了使深色皮肤人群非人化的分类,以便为奴役和剥削他们提供正当理由。
While it may feel uncomfortable to engage the history of these ideologies researchers and physicians must understand how they persist in our scientific and clinical methodologies so that we can heal the wounds of the past, moving forward with greater care and precision.虽然探讨这些意识形态的历史可能令人不适,但研究人员和医生必须理解它们如何持续存在于我们的科学和临床方法中,从而能够治愈过去的创伤,以更谨慎和精确的方式向前迈进。
Aristotle’s Scala Naturae laid the foundation for Carl Linnaeus’s 1737 Systema Naturae, a taxonomy of humans subdivided into four continental groups based on skin color: “whitish European,” “reddish American,” “tawny Asian,” and “blackish African”; a fifth category described “wild and monstrous humans, unknown groups, and more or less abnormal people.” Such taxonomic classes strike a shuddering resemblance to continent- or race-based categories that persist in human genomics.亚里士多德的《自然阶梯》为卡尔·林奈1737年的《自然系统》奠定了基础,后者是一种基于肤色将人类分为四个大陆群体的分类法:“白肤色欧洲人”、“红肤色美洲人”、“黄肤色亚洲人”和“黑肤色非洲人”;第五类描述了“野生和怪异的人类、未知群体以及或多或少异常的人”。这种分类与人类基因组学中持续存在的基于大陆或种族的分类惊人地相似。
It is everyone’s responsibility to reflect on this to ensure our studies are robust to the conceptual and mathematical influence of constructed social hierarchies rooted in biological racism.每个人都有责任反思这一点,以确保我们的研究能够抵御植根于生物种族主义中构建的社会等级所带来的概念和数学影响。
It is a powerful and pervasive myth that genetics are primarily responsible for differences we observe across geographic regions, cultural contexts, and social or political identities.一个强大而普遍的迷思是,遗传因素主要导致了我们在不同地理区域、文化背景以及社会或政治身份之间观察到的差异。
Eugenic scientists invoked natural class theory, but also relied on genetic essentialism–the notion that genetics are necessary and sufficient for all phenotypes.优生科学家援引自然等级理论,但也依赖遗传本质主义——即遗传因素对所有表型都是必要且充分的观念。
Early in the field’s history, statistical geneticist R.在该领域早期,统计遗传学家R.
Fisher was commissioned by Leonard Darwin to develop mathematical models that could provide a biological explanation for trait differences between groups of individuals.费舍尔受伦纳德·达尔文委托,开发能够为个体群体间性状差异提供生物学解释的数学模型。
It is widely accepted today that the social and culture-bound traits that eugenicists focused on (e. g., imbecility and promiscuity) are not caused by genetics.如今广泛认为,优生学家关注的社会和文化束缚性性状(如低能与滥交)并非由遗传因素引起。
However, concepts and statistical methods that Fisher developed to explain how such traits occurred more frequently within families are still used today.然而,费舍尔为解释此类性状在家族中更频繁出现而开发的概念和统计方法至今仍在使用。
Twin studies of inherited contributions to disease – using Fisher’s Heritability – were conducted on children imprisoned in concentration camps during WWII.利用费舍尔的遗传度概念,对二战期间被囚禁在集中营的儿童进行了关于疾病遗传贡献的双生子研究。
Heritability estimates from twin studies have motivated GWAS of social phenomena (e. g., educational attainment) despite their problematic assumptions and historical abuses.来自双生子研究的遗传度估计值推动了社会现象(如教育成就)的全基因组关联研究(GWAS),尽管这些估计值存在有问题的假设和历史滥用。
Heritability estimates have curiously remained central to certain branches of human genetics and social science.奇怪的是,遗传度估计值在人类遗传学和某些社会科学分支中仍保持核心地位。
Causal relationships between exposures and outcomes can rarely be proven, because there is no way to disentangle the impact of shared family or cultural environment (twins reared apart are often raised by relatives or other community members) and the impact of shared genetic variation within those extended family communities.暴露与结局之间的因果关系很少能被证明,因为无法区分共享家庭或文化环境的影响(分开抚养的双胞胎通常由亲戚或其他社区成员抚养)与这些大家庭社区内共享遗传变异的影响。
In particular, for complex, or multifactorial, traits (in which the environment plays a role), it is important to acknowledge the limitations of this heuristic.特别是对于复杂或多因素性状(其中环境起作用),承认这种启发式方法的局限性很重要。
Heritability estimates cannot truly distinguish between genetic and nongenetic effects, and there must be a proposed mechanism justify claims that genetics could influence purely social outcomes.遗传度估计值无法真正区分遗传效应和非遗传效应,并且必须提出一种机制来证明遗传因素可能影响纯社会结局的说法。
There is no scientific basis for using broad social categories or sweeping statistical assumptions to infer anyone’s precise genomic variation, so we need to develop better methods and conceptual frameworks to bolster our interpretation of genomics.没有科学依据利用宽泛的社会类别或笼统的统计假设来推断任何人的精确基因组变异,因此我们需要开发更好的方法和概念框架来加强我们对基因组学的解读。
Researchers, clinicians, professional societies, and funding agencies are dealing with problems introduced by the legacy of using race, ethnicity, and ancestry in biomedical research and medicine.研究人员、临床医生、专业学会和资助机构正在应对因在生物医学研究和医学中使用种族、民族和血统而遗留的问题。
Eliminating race-based clinical algorithms may reduce health disparities, but in the absence of robust data from diverse populations, the standard of care is biased toward white people of European descent.消除基于种族的临床算法可能减少健康差异,但在缺乏来自多样性人群的可靠数据的情况下,护理标准偏向于欧洲血统的白人。
Collecting this information is illegal in some countries; thus rendering invisible the influence of social stigma and discrimination on health outcomes.在某些国家收集此类信息是非法的,从而使社会污名和歧视对健康结局的影响变得不可见。
We must be able to track health and healthcare disparities, while avoiding unscientific and unethical applications of broad social categories.我们必须能够追踪健康和医疗保健差异,同时避免对大范围社会类别进行不科学和不道德的应用。
Complex Disease and Confounding It is important to recognize the role of confounding in the causal pathway from exposures to health outcomes.复杂疾病与混杂——认识混杂因素在从暴露到健康结局的因果路径中的作用至关重要。
7/53
Population Genetics for Genomic Medicine 181 Confounding can create spurious associations (e. g., between continental an…
Ch10 — Segment 7
Population Genetics for Genomic Medicine 181 Confounding can create spurious associations (e. g., between continental ancestry and disease risk), even if there is no genetic basis for the condition.群体遗传学与基因组医学 181 混杂因素可以产生虚假关联(例如,大陆祖先与疾病风险之间的关联),即使该疾病没有遗传基础。
Statistical associations observed between outcomes and exposures are confounded when there is another, unmeasured or hidden factor that is both causally linked to the outcome of interest and associated with the exposure.当存在另一个未测量或隐藏的因素,该因素既与感兴趣的结局存在因果关系,又与暴露因素相关时,观察到的结局与暴露之间的统计关联会受到混杂。
This generates statistical correlations between measured exposure(s) and outcome(s) despite the absence of a causal relationship. illustrates confounding by social determinants of health (SDOH) and group categories (race, ethnicity, and ancestry) in a causal pathway for complex traits.这导致测量到的暴露与结局之间产生统计相关性,尽管不存在因果关系,并说明了健康社会决定因素(SDOH)和群体分类(种族、民族和祖先)在复杂性状因果路径中的混杂作用。
The true cause(s) of these traits could be nongenetic, environmental factors at a higher prevalence in the population, and/or genetic factors that are difficult to identify due to small effect sizes across multiple loci.这些性状的真正原因可能是非遗传的、在人群中患病率较高的环境因素,和/或由于多个位点效应较小而难以识别的遗传因素。
Linkage disequilibrium (LD), or the non-independence of allele frequencies along segments of chromosomes that are inherited together, creates additional statistical challenges.连锁不平衡(LD),即染色体片段上等位基因频率的非独立性(这些片段共同遗传),带来了额外的统计挑战。
In an association study, the true causal variant cannot be distinguished from other variants that are in LD with it.在关联研究中,真正的因果变异无法与与其存在连锁不平衡的其他变异区分开来。
If there is confounding by SDOH associated with genetic ancestry and an outcome of interest, ancestral patterns of LD may create group-specific genetic signals that are not causal.如果存在与遗传祖先和感兴趣的结局相关的健康社会决定因素混杂,祖先的连锁不平衡模式可能会产生非因果性的群体特异性遗传信号。
One example of uncontrolled confounding leading to harmful and false conceptions is the belief that genetics primarily drives racial or ethnic differences in health outcomes, or other complex traits with environmental components.未控制的混杂导致有害且错误观念的一个例子是,认为遗传因素主要驱动健康结局中的种族或民族差异,或驱动具有环境成分的其他复杂性状。
Because some racial and ethnic groupings are associated with genetic ancestry at the global, or continental, level (i. e., African, Asian, European, Latin American, etc.), genetic variation that happens to be more prevalent in some groups than others (by chance) may be mistaken for causal genetic factors contributing to intergroup differences in traits.由于某些种族和民族分类在全球或大陆层面(如非洲、亚洲、欧洲、拉丁美洲等)与遗传祖先相关,那些恰好在某些群体中比在其他群体中更常见(偶然)的遗传变异可能被误认为是导致群体间性状差异的因果遗传因素。
Instead, these are caused by environmental factors that also differ among groups.相反,这些是由群体间也存在差异的环境因素引起的。
In these cases, nongenetic factors confound the relationship between ancestry-associated traits or outcomes and genetic variation shared in ancestry-associated cultural groups.在这些情况下,非遗传因素混杂了祖先相关性状或结局与祖先相关文化群体中共有的遗传变异之间的关系。
Identity by Descent Versus Allele Sharing in Unrelated Individuals Throughout recorded history and still today, humans have migrated around the world and exchanged DNA with one another such that our ancestral roots live in our genomes as mixtures of chromosomal segments (haplotypes) that have been inherited through our maternal and paternal lineages.不相关个体中的血缘同一性与等位基因共享 在有记录的历史中直至今日,人类在世界各地迁徙并相互交换DNA,因此我们的祖先根源以染色体片段(单倍型)的混合形式存在于我们的基因组中,这些片段通过母系和父系遗传。
When these lineages are related, segments of maternal and paternal chromosomes in an individual will be shared in the manner of being identical by descent (IBD), meaning they descend from a recent common ancestor.当这些谱系相关时,个体中的母系和父系染色体片段将以血缘同一性(IBD)的方式共享,意味着它们来自一个最近的共同祖先。
In contrast, genomic loci that have shared variation in a large population with multiple ancestries, or between ancestral groups, are considered identical by state (IBS) and do not necessarily share a recent common ancestral origin.相比之下,在具有多种祖先的大群体中或祖先群体之间共享变异的基因组位点被认为是状态同一性(IBS),不一定共享最近的共同祖先起源。
IBD sharing between maternal and paternal chromosomes is a result of consanguinity, or reproduction between genetically related individuals.母系和父系染色体之间的血缘同一性共享是近亲结婚(即遗传相关个体之间的繁殖)的结果。
This, like founder effects, leads to an increase in the frequency of autosomal recessive traits because carriers of the same recessive allele are more likely to meet and reproduce.这与奠基者效应类似,导致常染色体隐性性状的频率增加,因为相同隐性等位基因的携带者更可能相遇并繁殖。
The kinds of recessive disorders seen in the offspring of related parents may be very rare and unusual in the general population.在近亲父母的后代中出现的隐性疾病的种类在一般人群中可能非常罕见且不常见。
This is because consanguineous mating allows an uncommon allele inherited from a heterozygous common ancestor to become homozygous.这是因为近亲交配使得从杂合共同祖先遗传的罕见等位基因能够变为纯合。
A B Mendelian trait: Autosomal dominant, monogenic, rare Pathogenic variant: Rare, de novo, highly penetrant Social determinant(s) of health (SDOH) Complex trait: Multifactorial, polygenic, highly prevalent Candidate variant(s): Common, inherited, variable penetrance Correlation observed between ancestry and complex trait variation as a result of confounding by SDOH and social categories linked to ancestry.A B 孟德尔性状:常染色体显性、单基因、罕见 致病变异:罕见、新生、高外显率 健康社会决定因素(SDOH) 复杂性状:多因素、多基因、高度流行 候选变异:常见、遗传、可变外显率 观察到的祖先与复杂性状变异之间的相关性是由于健康社会决定因素和与祖先相关的社会分类的混杂所致。
No correlation between ancestry and mendelian trait variation because causal variant(s) are very rare, and often de novo.祖先与孟德尔性状变异之间无相关性,因为因果变异非常罕见,且通常是新生的。
PATHWAY n A = 10,000 n B = 10,000 n A = 100 n B = 100 Ancestry A Ancestry A Ancestry B Ancestry B CONFOUNDING Complex trait variation Social category with ancestries A and B (racial or ethnic group) CAUSAL Complex traits are common, caused by combined genetic, social, and other environmental factors.通路 n A=10,000 n B=10,000 n A=100 n B=100 祖先A 祖先A 祖先B 祖先B 混杂 复杂性状变异 具有祖先A和B的社会分类(种族或民族群体) 因果 复杂性状常见,由遗传、社会和其他环境因素共同引起。
Ancestry may appear to influence disease risk if it is associated with the trait of interest and social determinants of health.如果祖先与感兴趣的性状和健康社会决定因素相关,那么祖先可能显得影响疾病风险。
8/53
IBD sharing is more prevalent in genetic isolates, small populations derived from a limited number of common ancestors w…
Ch10 — Segment 8
IBD sharing is more prevalent in genetic isolates, small populations derived from a limited number of common ancestors who tended to mate only among themselves.IBD共享在遗传隔离群中更为普遍,这些群体是由少数共同祖先衍生而来的小种群,其成员倾向于仅在内部交配。
Reproduction between two apparently “unrelated” individuals in a genetic isolate may have the same risk for certain recessive conditions as that observed in consanguineous reproduction because the individuals are both carriers via inheritance from common ancestors within the isolate. ; in contrast, to a most recent common ancestor (MRCA).在遗传隔离群中,两个表面上“无关”的个体之间的繁殖可能对某些隐性疾病的患病风险与近亲繁殖相同,因为这些个体均通过遗传从隔离群内的共同祖先成为携带者;这与最近共同祖先(MRCA)形成对比。
Coalescent theory is used to estimate t for two copies of an allele or haplotype (collection of linked alleles) that are shared IBD .溯祖理论用于估计两个共享IBD的等位基因或单倍型(连锁等位基因的集合)拷贝的t值。
In contrast, identical alleles are IBS when they arise through independent processes such as mutation or inheritance through different ancestral lineages.相反,当相同等位基因通过独立过程(如突变或通过不同祖先谱系的遗传)产生时,它们属于IBS。
The relative number of alleles (or average lengths of stretches of chromosomes) shared IBD among individuals in a population is proportional to the effective population size.群体中个体间共享IBD的等位基因相对数量(或染色体片段的平均长度)与有效群体大小成正比。
Smaller Ne leads to greater IBD sharing, on average, because there are fewer haplotypes circulating in the population.较小的有效群体大小(Ne)平均会导致更大的IBD共享,因为群体中循环的单倍型较少。
This leads to a greater prior probability that two recessive deleterious alleles will be inherited by chance through closely or distantly related ancestral lineages.这导致两个隐性有害等位基因通过近亲或远亲祖先谱系偶然遗传的先验概率更高。
Assortative mating describes a phenomenon in which certain groups of individuals have – for a variety of historical, cultural, or religious reasons – remained relatively genetically separate during modern times.选型交配描述了一种现象,即某些个体群体由于各种历史、文化或宗教原因,在现代时期仍保持相对的遗传隔离。
When mate selection in a population is restricted for any reason to members of a particular group, and that group happens to have a variant with a higher frequency than the population as a whole, the result will be an apparent excess of homozygotes in the overall population beyond what one would predict.当群体中的择偶因任何原因仅限于某一特定群体的成员,且该群体恰好拥有一个频率高于整体群体的变异时,结果将导致整体群体中纯合子显著多于预期。
A clinically important aspect of assortative mating is the tendency to choose partners with similar traits, such as congenital deafness or blindness.选型交配在临床上重要的一个方面是倾向于选择具有相似性状的伴侣,例如先天性耳聋或失明。
In such cases, the genotypes of two reproductive partners at loci influencing the trait are not predicted from population allele frequencies.在这种情况下,影响性状的位点上两个生殖伴侣的基因型无法从群体等位基因频率预测。
For example, consider achondroplasia (Case 2), an autosomal dominant form of skeletal dysplasia with a population incidence of 1 per 15,000 to 1 per 40,000 live births.例如,考虑软骨发育不全(案例2),这是一种常染色体显性遗传的骨骼发育不良,群体发病率为每15,000至40
Offspring homozygous for the achondroplasia variant have a severe, lethal form of skeletal dysplasia that is almost never seen unless both parents have achondroplasia and are thus heterozygous for the variant.[TL:missing]
This would be highly unlikely to occur by chance, except for assortative mating among those with achondroplasia.[TL:missing]
When reproductive partners have autosomal recessive disorders caused by the same pathogenic variant or by allelic variants in the same gene, all their offspring will also have the disease.[TL:missing]
Even if there is locus heterogeneity with assortative mating, the chance that two individuals carry pathogenic variants in the same locus is increased over what it would be under true random mating, and therefore likelihood of the trait in their offspring is also increased.[TL:missing]
Genetic Ancestry and Population Structure The concept of genetic ancestry may seem more scientifically valid and concrete than self-reported measures like race and ethnicity because it is based on genetic information; however, it also lacks a firm “ground truth.” Genetic ancestry is a dynamic and relative measure; estimates can change over time depending on the reference data and methods used, all of which have inherent assumptions and limitations.[TL:missing]
Often reported as percentages of the genome that can be traced back to ancestral populations, ancestry estimates are usually based on the MRCA A B t generations cryptic relatedness unrelated lineages mutation inheritance mutation IBD IBS - recent ancestry - consanguinity Blue diamonds are individuals and the edges connecting them indicate their relations in a pedigree; orange circles and arrows represent pathogenic variants, their origins traced forward being inheritance and backward being coalescence in lineages.[TL:missing]
9/53
Population Genetics for Genomic Medicine 183 average proportion of an individual’s DNA that most closely matches a given…
Ch10 — Segment 9
Population Genetics for Genomic Medicine 183 average proportion of an individual’s DNA that most closely matches a given reference dataset with assigned population labels (compared to all other reference data available for the analysis).群体遗传学在基因组医学中的应用 183 个体DNA中与给定参考数据集(带有指定群体标签)最匹配的平均比例(与可用于分析的所有其他参考数据相比)。
This means that the more robust and geographically specific reference datasets are, the higher the resolution achieved for reported proportions of population-specific ancestry.这意味着参考数据集越稳健且地理特异性越强,所报告的群体特异性祖源比例的分辨率就越高。
For example, most direct-to-consumer (DTC) genetic ancestry companies can predict specific geographic origins for individuals’ European ancestry components (e. g., a small village in Northern Ireland) but often report ancestry components at the continental level for regions that are underrepresented among reference datasets, such as “Sub-Saharan Africa.” Representation in genomic reference datasets disproportionately excludes the most genetically diverse ancestries while including mostly Europeans.例如,大多数直接面向消费者(DTC)的遗传祖源公司可以预测个体欧洲祖源成分的具体地理起源(如北爱尔兰的一个小村庄),但对于参考数据集中代表性不足的地区(如“撒哈拉以南非洲”),通常仅报告大陆级别的祖源成分。基因组参考数据集中的代表性不成比例地排除了遗传多样性最高的祖源,而主要包含欧洲人。
Since genetic ancestry is estimated using reference data, individuals with ancestries from regions of the world that have yet to be broadly included in genomics research may not receive accurate, detailed, or informative ancestry proportion results.由于遗传祖源是使用参考数据估算的,来自尚未广泛纳入基因组学研究的世界地区的祖源个体,可能无法获得准确、详细或有信息量的祖源比例结果。
Similarly, imputation (the process of filling in of alleles that are missing from genotype data using some reference dataset), is less accurate in groups with greater genetic diversity that is missing from reference resources.类似地,基因型填补(利用参考数据集填充基因型数据中缺失等位基因的过程)在参考资源中缺失的遗传多样性较高的群体中准确性较低。
Alleles that differ in their frequencies among preconstructed ancestry groupings are referred to as ancestry informative markers (AIMs).在预先构建的祖源分组中频率存在差异的等位基因被称为祖源信息标记(AIMs)。
AIMs have been identified to differentiate among broad geographical groupings (e. g., African, East Asian, South Asian, European, Middle Eastern, Native American, and Pacific Islanders).已鉴定出可区分广泛地理分组(如非洲、东亚、南亚、欧洲、中东、美洲原住民和太平洋岛民)的AIMs。
Such markers have been used for charting human migration patterns, for documenting historical admixture between or among populations, and for determining the degree of genetic diversity among ancestral population groups.此类标记已被用于绘制人类迁徙模式、记录群体间或群体内的历史混合,以及确定祖先群体间的遗传多样性程度。
Studies of hundreds of thousands of AIMs from across the genome have been used to distinguish and determine the genome-wide relationships among many different populations.对来自全基因组数十万个AIMs的研究已被用于区分和确定许多不同群体之间的全基因组关系。
Since AIMs are selected to maximize differences between predefined groups, they should not be considered representative of genome-wide variation among those groups.由于AIMs的选择旨在最大化预设组间的差异,因此不应将其视为这些组间全基因组变异的代表性标记。
Polymorphisms that happen to occur at higher frequencies in certain groups of individuals may be identified as AIMs by chance, even if the grouping scheme is not otherwise biologically meaningful.在某些个体群体中偶然出现较高频率的多态性可能被偶然鉴定为AIMs,即使分组方案在生物学上并无意义。
Predefined groups are often based on sociocultural categories or their respective broad, continental groupings; they influence the loci that are selected, then those AIMS may be treated as a proxy for genomic background.预设分组通常基于社会文化类别或其相应的大陆级广泛分组;它们影响所选位点,随后这些AIMs可能被用作基因组背景的代理。
This has downstream implications for research and quality control measures that rely on AIMs (e. g., to identify population-specific reference panels, impute missing genomic information, or test the accuracy of various analytic methods).这对依赖AIMs的研究和质量控制措施(例如,识别群体特异性参考面板、填补缺失的基因组信息或测试各种分析方法的准确性)具有下游影响。
In 2008, population geneticist John Novembre and colleagues famously published results of a principal component analysis (PCA) that appear to reconstruct the geography of Europe using genotype data sampled from countries across the continent .2008年,群体遗传学家John Novembre及其同事发表了主成分分析(PCA)的著名结果,该分析似乎利用从欧洲各国采样的基因型数据重构了欧洲的地理分布。
Subsequently, many genomics researchers have used PCA and other statistical clustering methods like admixture analysis that reduce the complexity of data to visualize population structure.随后,许多基因组研究人员使用PCA和其他统计聚类方法(如混合分析)来降低数据复杂性,以可视化群体结构。
While these approaches may provide some insight, they also have limitations and may be misleading.尽管这些方法可能提供一些见解,但它们也存在局限性,可能具有误导性。
Novembre and colleagues have published concerns about the potential for confusion about global genetic population structure because of these methods.Novembre及其同事已发表对因这些方法可能导致的全球遗传群体结构混淆的担忧。
For example, geographic clusters that appear distinct in the PCA plot of Europe are genetically very similar, but the figure may mislead people into thinking they have large overall genetic differences.例如,在欧洲PCA图中看似不同的地理聚类在遗传上非常相似,但该图可能误导人们认为它们具有巨大的整体遗传差异。
In the case of Europe, clines or gradients in allele frequencies are roughly aligned with latitude and longitude, as humans migrated EastWest and South-North across the continent.就欧洲而言,等位基因频率的渐变群或梯度大致与纬度和经度对齐,因为人类在该大陆上进行了东西向和南北向的迁徙。
This pattern of allele frequencies mirroring geography is unique to Europe, and the structure shown only emerges after >100k loci are included in the analysis because allele frequency differences are so small.这种等位基因频率反映地理分布的模式是欧洲独有的,并且由于等位基因频率差异非常小,这种结构仅在分析中包含超过10万个位点后才显现。
Genetic variation among populations may be falsely perceived as divided among regional groups, though few alleles are restricted to just one region of the world.群体间的遗传变异可能被错误地视为按区域群体划分,尽管很少有等位基因局限于世界上的某一个区域。
Data from the US Population Architecture using Genomics and Epidemiology (PAGE) study were used to visualize population structure with PCA .来自美国人群基因组学与流行病学建筑(PAGE)研究的数据被用于通过PCA可视化群体结构。
Each colored dot in the plot of principal components (PCs) 1 and 2 represents an individual research participant, and its color corresponds to the self-reported race or ethnicity category selected by the participant.在主成分(PC)1和2的图中,每个彩色点代表一名研究参与者,其颜色对应于参与者自报的种族或民族类别。
The dots are positioned relative to one another according to similarities and differences in genotypes across the genome.这些点根据全基因组基因型的相似性和差异相互定位。
We can see from the spread of these individuals across PCs that there is a complex spectrum of shared variation that cannot be adequately represented by categorical data structures.从这些个体在PC上的分布可以看出,存在一个复杂的共享变异谱,无法通过分类数据结构充分表示。
This illustrates why sociocultural categories used in the US Census cannot be considered genetically differentiated; there is a continuous distribution of genomic variation within and among groups such that most variants are shared and their frequency distributions overlap.这说明了为什么美国人口普查中使用的社会文化类别不能被视为遗传分化;群体内和群体间的基因组变异呈连续分布,使得大多数变异是共享的,且其频率分布相互重叠。
Misclassification results from individuals being assigned to the wrong analytic group in a study (e. g., case being classified incorrectly as a control, or vice versa).错误分类源于个体在研究中被分配到错误的分析组(例如,病例被错误分类为对照,反之亦然)。
In clinical algorithms that rely on racial and ethnic classification, misclassification could be considered the incorrect attribution of a category-based mean value to an individual patient.在依赖种族和民族分类的临床算法中,错误分类可被视为将基于类别的平均值错误地归因于个体患者。
Furthermore, classification error is almost certain due to genetic heterogeneity within racial and ethnic groups that violate baseline assumptions and motivations for applying population- or group-level adjustments.此外,由于种族和民族群体内部的遗传异质性违反了应用群体或群体水平调整的基线假设和动机,分类错误几乎不可避免。
If genetic ancestry were used instead of self-reported measures of racial or ethnic identity, the preselected nature of如果使用遗传祖源而非自我报告的种族或民族身份测量,那么预先选择的性质
10/53
AIMs and then using those to classify individuals into discrete groupings may create the exact same problems as race and…
Ch10 — Segment 10
AIMs and then using those to classify individuals into discrete groupings may create the exact same problems as race and ethnicity categories.AIMs,然后利用它们将个体分类为离散分组,可能会产生与种族和民族类别完全相同的问题。
That is, ancestry categories that are semantic variations of racial and ethnic groupings do no better at describing genomic variation because human genetic diversity is a narrow spectrum of [mostly] shared alleles at gradually changing, relative frequencies.也就是说,作为种族和民族群体语义变体的祖先类别,在描述基因组变异方面并不更好,因为人类遗传多样性是一系列狭窄的、大部分共享的等位基因,其相对频率逐渐变化。
Humans from different regions of the world have genetic similarities and differences that may have nothing to do with geography.来自世界不同地区的人类具有可能与地理无关的遗传相似性和差异。
Whereas there may be geographic trends in aggregate, an allele frequency cannot be used to determine an individual’s genotype.尽管总体上可能存在地理趋势,但等位基因频率不能用于确定个体的基因型。
Even Duffy blood group alleles, which are classically known for differences in frequencies by geography, can be seen on every continent, with clines or gradients showing clear variation within those regions.甚至经典上已知频率因地理而异的Duffy血型等位基因,在每个大陆上都能观察到,并呈现出在这些区域内具有明显变异的渐变群或梯度。
Earlier versions of this very textbook have described blood group variation according to geographic, racial, and ethnic categories (which change over time and across cultural contexts) but today, we have enough genetic data to establish that neither is genomic variation restricted to such categories nor is it easily characterized by any categorical framework.这本教科书的早期版本根据地理、种族和民族类别(这些类别随时间及文化背景变化)描述了血型变异,但如今,我们有足够的遗传数据可以确定:基因组变异既不受限于这些类别,也不容易被任何分类框架简单描述。
To illustrate this point, , which have historically been used as the canonical example of geographic differentiation.为了说明这一点,,这些区域历史上曾被用作地理分化的经典例子。
We highlight here that not all parts of Africa have a high frequency of the FY*BES allele that confers protection against malarial infection via the Plasmodium parasites, and some other parts of the world (not in Africa) have elevated frequencies as well.我们在此强调,并非非洲所有地区都具有高频率的FY*BES等位基因(该等位基因通过疟原虫提供抗感染保护),而世界其他一些地区(非非洲)也具有升高的频率。
In geographic locations where Plasmodium species causing malaria in humans are endemic, the protective allele is highly prevalent.在导致人类疟疾的疟原虫物种流行的地理位置,保护性等位基因高度普遍。
All three common Duffy alleles are present at a range of frequencies across the Americas, and none are restricted to a single continent or large geographic region that corresponds to continental ancestry groupings.所有三种常见的Duffy等位基因在美洲各地以一系列频率出现,且没有一种局限于单一大陆或对应大陆祖先分组的大型地理区域。
There are many different factors that can explain how differences in disease incidence, prevalence, and allele frequencies arise among biogeographic populations.有许多不同因素可以解释疾病发病率、患病率和等位基因频率在生物地理人群间的差异如何产生。
In cases where genetic variants are contributing Small colored labels represent individuals and large colored points represent median PC1 and PC2 values for each country.在遗传变异有贡献的情况下,彩色小标签代表个体,彩色大点代表每个国家的主成分1和主成分2的中位数。
The inset map provides a key to the labels.插图地图提供了标签的图例。
The PC axes are rotated to emphasize the similarity to the geographic map of Europe.主成分轴经过旋转,以强调与欧洲地理地图的相似性。
AL, Albania; AT, Austria; BA, Bosnia-Herzegovina; BE, Belgium; BG, Bulgaria; CH, Switzerland; CY, Cyprus; CZ, Czech Republic; DE, Germany; DK, Denmark; ES, Spain; FI, Finland; FR, France; GB, Great Britain; GR, Greece; HR, Croatia; HU, Hungary; IE, Ireland; IT, Italy; KS, Kosovo; LV, Latvia; MK, Macedonia; NO, Norway; NL, Netherlands; PL, Poland; PT, Portugal; RO, Romania; RS, Serbia and Montenegro; RU, Russia; Sct, Scotland; SE, Sweden; SI, Slovenia; SK, Slovakia; TR, Turkey; UA, Ukraine; YG, Yugoslavia.AL,阿尔巴尼亚;AT,奥地利;BA,波斯尼亚和黑塞哥维那;BE,比利时;BG,保加利亚;CH,瑞士;CY,塞浦路斯;CZ,捷克共和国;DE,德国;DK,丹麦;ES,西班牙;FI,芬兰;FR,法国;GB,英国;GR,希腊;HR,克罗地亚;HU,匈牙利;IE,爱尔兰;IT,意大利;KS,科索沃;LV,拉脱维亚;MK,马其顿;NO,挪威;NL,荷兰;PL,波兰;PT,葡萄牙;RO,罗马尼亚;RS,塞尔维亚和黑山;RU,俄罗斯;Sct,苏格兰;SE,瑞典;SI,斯洛文尼亚;SK,斯洛伐克;TR,土耳其;UA,乌克兰;YG,南斯拉夫。
From Novembre J, Johnson T, Bryc K. et al.来自Novembre J, Johnson T, Bryc K.等人。
Genes mirror geography within Europe.基因映射欧洲内部的地理分布。
Nature 456:98–101, 2008. 07331《自然》456:98–101, 2008. 07331
11/53
Population Genetics for Genomic Medicine 185 to disease etiology, it is possible that inheritance of an ancestral haplot…
Ch10 — Segment 11
Population Genetics for Genomic Medicine 185 to disease etiology, it is possible that inheritance of an ancestral haplotype containing a pathogenic variant is more likely, given reported or estimated ancestry of an individual patient.群体遗传学在基因组医学中的应用 185:对于疾病病因学而言,考虑到个体患者的报告或推定祖先,携带致病性变异祖先单倍型的遗传可能性更大。
However, disease-causing alleles that reduce the fitness of an individual tend to be rare in populations that are sufficiently large (such that other haplotypes are frequent enough to outcompete one that is pathogenic).然而,降低个体适应度的致病等位基因在足够大的群体中往往较为罕见(因为其他单倍型频率足以竞争过致病单倍型)。
In smaller populations that have been genetically isolated from others, or in situations where environmental conditions enhance the fitness of carriers of pathogenic variants and create a heterozygote advantage, it is possible to see disease-causing alleles at frequencies higher than would be expected given disease prevalence.在与其他群体遗传隔离的较小群体中,或当环境条件增强致病变异携带者的适应度并产生杂合子优势时,致病等位基因的频率可能高于根据疾病患病率预期的水平。
Other factors include genetic drift, which applies to benign variants that rise to high frequency due to physical proximity to fitness-enhancing alleles, without having direct impact on the phenotype.其他因素包括遗传漂变,这适用于由于物理邻近适应度增强等位基因而升至高频、但对表型无直接影响的良性变异。
If the overall population is sufficiently large, the frequency of an allele in a small subset of the population will not change the total population allele frequency.若整个群体足够大,则群体小子集中某等位基因的频率不会改变总群体的等位基因频率。
However, the larger the subset of carriers relative to the total population, the greater the chance that it will alter the population allele frequency.然而,相对于总群体,携带者子集越大,其改变群体等位基因频率的可能性就越大。
Therefore, classification of individuals into population categories important because the relationship between the numerator (number of observations) and the denominator (total population under investigation) determines the allele frequency.因此,将个体划分为群体类别至关重要,因为分子(观察数)与分母(调查总群体)的关系决定了等位基因频率。
In practice, estimates of allele frequencies are used in combination with disease prevalence and incidence to determine genotype frequencies, given their modes of inheritance.实践中,等位基因频率估计值与疾病患病率和发病率结合,用于根据遗传方式确定基因型频率。
The population substructure that is present in the multiethnic sample of PAGE (n = 49,839) reveals complex patterns preventing meaningful stratification.PAGE多民族样本(n = 49,839)中存在的群体亚结构揭示了阻碍有意义分层的复杂模式。
PC1 and PC2 show major patterns of variation, stratified by self-identified race/ethnicity.PC1和PC2显示主要变异模式,按自我认同的种族/民族分层。
Individuals denoted by orange self-identified as “Other.” From Wojcik GL, Graff M, Nishimura KK, et al.橙色标记的个体自我认同为“其他”。来源:Wojcik GL, Graff M, Nishimura KK, 等。
Genetic analyses of diverse populations improves discovery for complex traits, Nature 570:514–518, 2019. 41586-019-1310-4多样本群体的遗传分析促进复杂性状的发现,Nature 570:514–518, 2019。41586-019-1310-4。
12/53
Most sections in this chapter have thus far dealt with the complexity of defining human populations, characterizing ance…
Ch10 — Segment 12
Most sections in this chapter have thus far dealt with the complexity of defining human populations, characterizing ancestry and genetic population structure, and accounting for nongenetic factors associated with systemic inequities that may create erroneous correlations among group-level phenotypic differences, ancestry estimates, and sociocultural categories.本章中的大部分章节迄今为止讨论了定义人类群体、描述祖源和遗传群体结构以及解释与系统性不平等相关的非遗传因素的复杂性,这些因素可能在群体水平的表型差异、祖源估计和社会文化类别之间产生错误的相关性。
While it is important to recognize limitations of point estimates and other discrete measures to characterize continuous variables (e. g., genomic variant distributions in human populations), there is utility in calculating allele and genotype frequencies to inform genomic research and medicine.尽管认识到点估计和其他离散测量在表征连续变量(例如,人类群体中的基因组变异分布)方面的局限性很重要,但计算等位基因和基因型频率为基因组研究和医学提供信息仍然具有实用性。
In the next section, we describe how to calculate these frequencies and offer clinically relevant examples.在下一节中,我们将描述如何计算这些频率并提供临床相关示例。
ALLELE AND GENOTYPE FREQUENCIES For autosomal loci, the size of the gene pool at one locus is twice the number of individuals (2 N) in the population because each autosomal genotype consists of two alleles.等位基因和基因型频率 对于常染色体位点,一个位点的基因库大小是群体中个体数(2 N)的两倍,因为每个常染色体基因型由两个等位基因组成。
Consider a population that is structured by recent ancestry or migration from another population, such that 10% of the total population contains a group in which the minor allele frequency (MAF) for a biallelic, autosomal recessive disease is 5% (q = 0. 05).考虑一个由近期祖源或从另一群体迁移而形成的群体,其中总人口的10%包含一个亚群,在该亚群中,一个双等位基因、常染色体隐性遗传病的次要等位基因频率(MAF)为5%(q = 0.05)。
Since the trait has two alleles whose frequencies sum to 1 in the population (p + q = 1), we can infer the other allele’s frequency to be p = 0. 95 in this group by subtracting the frequency of q from 1.由于该性状有两个等位基因,其频率在群体中之和为1(p + q = 1),我们可以通过从1中减去q的频率推断出该亚群中另一个等位基因的频率为p = 0.95。
In the remaining 90% of the total population, the frequency of this pathogenic variant is so small that it is not observed and presumed to be absent such that q≈0andp≈1.在总人口的其余90%中,该致病性变异频率极低,未被观察到并假定其不存在,因此q≈0且p≈1。
Let us continue with the Duffy blood group example to illustrate the relationship between allele and genotype frequencies in populations.让我们继续以达菲血型为例来说明群体中等位基因和基因型频率之间的关系。
Consider the gene ACKR1 (atypical chemokine receptor 1) on chromosome 1 (1q23. 2), which encodes the major subunit of the Duffy blood group system and serves as an entry point for Plasmodium vivax, an important parasite causing malaria.考虑位于1号染色体(1q23.2)上的ACKR1基因(非典型趋化因子受体1),该基因编码达菲血型系统的主要亚基,并作为间日疟原虫(一种引起疟疾的重要寄生虫)的入侵点。
Hundreds of variants have been observed in this gene with varying functional consequences.在该基因中已观察到数百个具有不同功能影响的变异。
We will focus on a single nucleotide variant (SNV) in the promoter region of ACKR1 (rs 2814778 in Ensembl), which confers protection against malarial infection and exhibits allele frequency differences across global populations.我们将关注ACKR1启动子区域中的一个单核苷酸变异(SNV)(Ensembl数据库中的rs2814778),该变异赋予对疟疾感染的保护作用,并在全球人群中表现出等位基因频率差异。
We draw on data from the 1000 Genomes (1KG) Project to calculate allele frequencies from observed genotype FY*A frequency FY*B frequency FY*BES frequency FY*A IQR FY*B IQR 0–5% FY*BES IQR 50–70% 0–5% 5–10% 10–20% 20–30% 50–60% 0–5% 5–10% 10–20% 20–30% 30–50% 50–70% 30–50% 70–80% 80–90% 90–95% 95–100% 70–85% 50–70% 30–50% 30–50% 20–30% 10–20% 5–10% 0–5% 50–70% 70–80% 80–90% 90–95% 95–100% 0–5% 5–10% 10–20% 20–30% 30–50% 50–70% 70–90% 20–30% 10–20% 5–10% 0–5% 5–10% 10–20% 20–30% 30–50% B C A D E F A, B, and C correspond to FY*A, FY*B, and FY*BES allele frequency maps, respectively (median values of the prediction posterior distributions); D–F show the respective interquartile ranges (IQR) of each allele frequency map (25–75% interval).我们利用千人基因组(1KG)计划的数据,根据观察到的基因型计算等位基因频率:FY*A频率、FY*B频率、FY*BES频率、FY*A的IQR、FY*B的IQR(0–5%)、FY*BES的IQR(50–70%)、0–5%、5–10%、10–20%、20–30%、50–60%、0–5%、5–10%、10–20%、20–30%、30–50%、50–70%、30–50%、70–80%、80–90%、90–95%、95–100%、70–85%、50–70%、30–50%、30–50%、20–30%、10–20%、5–10%、0–5%、50–70%、70–80%、80–90%、90–95%、95–100%、0–5%、5–10%、10–20%、20–30%、30–50%、50–70%、70–90%、20–30%、10–20%、5–10%、0–5%、5–10%、10–20%、20–30%、30–50%;B、C、A、D、E、F:A、B和C分别对应FY*A、FY*B和FY*BES等位基因频率图(预测后验分布的中位数);D–F显示每个等位基因频率图相应的四分位距(IQR)(25–75%区间)。
Predictions are made on a 5 × 5 km grid in Africa and a 10 × 10 km grid elsewhere.预测在非洲采用5×5公里网格,在其他地区采用10×10公里网格进行。
From Howes R, Patil A, Piel F. et al.源自Howes R, Patil A, Piel F. 等。
The global distribution of the Duffy blood group.达菲血型的全球分布。
Nat Commun 2:266, 2011. ncomms 1265Nat Commun 2:266, 2011。ncomms1265
13/53
Population Genetics for Genomic Medicine 187 frequencies.
Ch10 — Segment 13
Population Genetics for Genomic Medicine 187 frequencies.群体遗传学与基因组医学:187个频率
Because each homozygous individual has two copies of the same allele, and heterozygous individuals have one copy of each allele, the frequency of each allele is twice the number of individuals in the population homozygous for the allele, plus the number of heterozygotes, divided by the total number of alleles in the population or 2N: T: 2 1793 1 88 2504 2 0. 734 × () + × () × = C: 2 623 1 88 2504 2 0. 266 × () + × () × = Rather than calculating the frequency of each allele independently, the calculated frequency of one allele can simply be subtracted from one (e. g., 1 – 0. 734 = 0. 266) because the frequencies of the two alleles must add up to 1 (recall p + q = 1).由于每个纯合个体拥有同一等位基因的两个拷贝,而杂合个体拥有每个等位基因的一个拷贝,因此每个等位基因的频率等于该等位基因纯合个体数目的两倍加上杂合个体数目,除以群体中等位基因总数(即2N):T: 2×1793+1×88 / 2×2504 = 0.734;C: 2×623+1×88 / 2×2504 = 0.266。无需独立计算每个等位基因的频率,只需用1减去已计算出的一个等位基因的频率即可(例如1-0.734=0.266),因为两个等位基因的频率之和必为1(回顾p+q=1)。
Now consider that genotype and allele frequencies differ across geographic regions due to the protective nature of this allele in regions with malarial parasites.现在考虑,由于该等位基因在疟原虫流行地区具有保护性质,不同地理区域的基因型和等位基因频率存在差异。
Data from the 1KG Project can be used to assess global allele and genotype frequencies for discrete geographic regions that have been sampled for large-scale genome sequencing, but it is important to keep in mind that this sampling scheme does not account for all the global genomic variation that exists (most of which has been unsampled to date); it simply tells us about the subset of variation among project participants. are shared across continents.来自千人基因组项目的数据可用于评估已采样进行大规模基因组测序的离散地理区域的全球等位基因和基因型频率,但需注意该采样方案并未涵盖所有存在的全球基因组变异(其中大部分至今尚未采样);它仅反映了项目参与者之间的部分变异情况,且这些变异在各大洲间共享。
Hardy-Weinberg Equilibrium As we have shown, a sample of individuals with known genotypes in a population can be used to derive estimates of allele frequencies by simply counting the alleles in individuals with each genotype.哈迪-温伯格平衡 如我们所示,通过简单计数每个基因型个体中的等位基因,可以利用群体中已知基因型的个体样本推导出等位基因频率的估计值。
How about the converse?反过来呢?
Can we calculate the proportion of the population with various genotypes once we know the allele frequencies?一旦我们知道等位基因频率,能否计算出具有各种基因型的群体比例?
Deriving genotype frequencies from allele frequencies is not as straightforward as counting because we do not know in advance how the alleles are distributed among homozygotes and heterozygotes.从等位基因频率推导基因型频率并不像计数那样直接,因为事先不知道等位基因是如何在纯合子和杂合子之间分配的。
If a population meets certain assumptions, however, there is a simple mathematical equation for calculating genotype frequencies from allele frequencies.然而,如果某个群体满足特定假设,则存在一个简单的数学方程,可以从等位基因频率计算出基因型频率。
This equation is known as Hardy-Weinberg equilibrium (HWE), named after Godfrey Hardy, an English mathematician, and Wilhelm Weinberg, a German physician, who formulated it in 1908.该方程称为哈迪-温伯格平衡(HWE),以英国数学家Godfrey Hardy和德国医生Wilhelm Weinberg的名字命名,他们于1908年提出了该方程。
The Hardy-Weinberg principle has two critical components.哈迪-温伯格原理包含两个关键组成部分。
The first is that under certain ideal conditions (see 2), a simple relationship exists between allele frequencies and genotype frequencies in a population.首先,在特定理想条件下(见第2条),群体中等位基因频率与基因型频率之间存在简单关系。
To illustrate the utility of Hardy-Weinberg for understanding the relationship between allele and genotype frequencies, we offer here a step-by-step visual mathematical proof that under the right conditions, allele and genotype frequencies remain constant in each generation for a population in HWE.为说明哈迪-温伯格平衡在理解等位基因与基因型频率关系中的用途,我们在此提供一个逐步的可视化数学证明:在适当条件下,处于HWE的群体中,等位基因和基因型频率在每一代保持不变。
First, we start with genotypes that are observed in a population and simulate all possible reproductive pairings between individuals of each possible genotype.首先,我们从群体中观察到的基因型出发,模拟每种可能基因型的个体之间所有可能的生殖配对。
In this example, we consider a biallelic locus with alleles A and a, such that there are three possible genotypes: two are homozygous (AA or aa) and one is heterozygous (Aa)..在本例中,我们考虑一个具有等位基因A和a的双等位基因位点,因此存在三种可能基因型:两种纯合子(AA或aa)和一种杂合子(Aa)。
C|T 88 0. 035 T 0. 734 C|C 623 0. 249 C 0. 266 Total 2504 1. 000 From The 1000 Genomes Project Consortium: A global reference for human genetic variation, Nature 526:68–74, 2015. doi:10. 1038/nature 15393.C|T 88 0.035 T 0.734 C|C 623 0.249 C 0.266 总计 2504 1.000 数据来源:千人基因组项目联合体:人类遗传变异的全球参考,Nature 526:68–74, 2015。doi:10.1038/nature15393。
Genotype frequencies were obtained online from Ensembl ( Ensembl. org) in October 2022..基因型频率于2022年10月从Ensembl(Ensembl.org)在线获取。
C 170 0. 885 African Ancestry in Southwest US (ASW) T C 25 97 0. 205 0. 795 Esan in Nigeria (ESN) T C 0 198 0. 0 1. 0 Total (2N = 512) T C 47 465 0. 092 0. 908 From The 1000 Genomes Project Consortium: A global reference for human genetic variation, Nature 526:68–74, 2015. doi:10. 1038/nature 15393.C 170 0.885 美国西南部非裔(ASW) T C 25 97 0.205 0.795 尼日利亚埃桑人(ESN) T C 0 198 0.0 1.0 总计(2N=512) T C 47 465 0.092 0.908 数据来源:千人基因组项目联合体:人类遗传变异的全球参考,Nature 526:68–74, 2015。doi:10.1038/nature15393。
Allele counts and frequencies were obtained online from Ensembl (https:// www.等位基因计数和频率于2022年10月从Ensembl(https:// www.
Ensembl. org) in October 2022.Ensembl.org)在线获取。
14/53
Parent Generation (P1,2) Possible Mating Pairs P1 Genotype: AA P2 Genotype: AA A AA AA AA Aa Aa Aa AA Aa Aa Aa AA AA AA …
Ch10 — Segment 14
Parent Generation (P1,2) Possible Mating Pairs P1 Genotype: AA P2 Genotype: AA A AA AA AA Aa Aa Aa AA Aa Aa Aa AA AA AA AA AA Aa Aa Aa Aa aa aa aa Aa Aa Aa Aa Aa aa aa aa Aa aa aa aa Aa Aa A A A a a a A A a a a P2 Genotype: Aa P2 Genotype: aa P1 Genotype: Aa P1 Genotype: aa F1 F1 F1 Predicted offspring (F1) genotypes, given genotypes of possible parent mating pairs (P1,2) in a population F1 F1 F1 F1 F1 F1 The table above uses Punnett squares to illustrate all possible genotypes present in a generation of offspring (F1) resulting from all possible combinations of genotypes in the parent generation (P), given two ACB ASW BEB CDX CEU CHB CHS CLM ESN FIN GBR GIH GWD IBS ITU JPT KHV LWK MSL MXL PEL PJL PUR STU TSI YRI Private to population Private to continent Shared across continents Shared across all continents 18 million 12 million 24 million Individual Variant sites per genome (million) 3. 8 4 4. 2 4. 4 4. 6 4. 8 5 MSL ESN LWK YRI GWD ACB ASW PUR CLM MXL PEL BEB ITU PJL STU GIH KHV JPT CHB CDX CHS TSI IBS CEU GBR FIN Singletons per genome (×1,000) 0 2 4 6 8 10 12 14 16 18 20 LWK GWD MSL ACB ASW YRI ESN BEB STU ITU PJL GIH CHB KHV CHS JPT CDX TSI CEU IBS GBR FIN PEL MXL CLM PUR A B c Polymorphic variants within sampled populations.亲代(P₁,₂)可能的交配对 P₁ 基因型:AA P₂ 基因型:AA A AA AA AA Aa Aa Aa AA Aa Aa Aa AA AA AA AA AA Aa Aa Aa Aa aa aa aa Aa Aa Aa Aa Aa aa aa aa Aa aa aa aa Aa Aa A A A a a a A A a a a P₂ 基因型:Aa P₂ 基因型:aa P₁ 基因型:Aa P₁ 基因型:aa F₁ F₁ F₁ 预测后代(F₁)基因型,给定种群中可能的亲代交配对(P₁,₂)的基因型 F₁ F₁ F₁ F₁ F₁ F₁ 上表使用庞纳特方格展示了在给定等位基因 A 和 a(AA、Aa 和 aa)的情况下,由亲代(P)所有可能的基因型组合产生的子代(F₁)中所有可能的基因型,同时标注了 ACB、ASW、BEB、CDX、CEU、CHB、CHS、CLM、ESN、FIN、GBR、GIH、GWD、IBS、ITU、JPT、KHV、LWK、MSL、MXL、PEL、PJL、PUR、STU、TSI、YRI 等群体中私有(深色,群体特有)、洲内共有(浅色,洲内群体共有)、洲际共有(浅灰色)和全球共有(深灰色)的变异位点数量(单位:百万),以及每个基因组的个体变异位点数量(图 B)和每个基因组的单例变异位点数量(图 C,单位:千),虚线表示在其祖先洲外采样的群体。多态性变异位于采样群体内。
The area of each pie is proportional to the number of polymorphisms within a population.每个饼图的面积与群体内多态性的数目成正比。
Pies are divided into four slices, representing variants private to a population (darker color unique to population), private to a continental area (lighter color shared across continental group), shared across continental areas (light gray), and shared across all continents (dark gray).饼图被分为四个扇区,分别代表群体私有变异(深色,群体特有)、洲内私有变异(浅色,洲内群体共有)、洲际共有变异(浅灰色)和全球共有变异(深灰色)。
Dashed lines indicate populations sampled outside of their ancestral continental region.虚线表示在其祖先洲外采样的群体。
(B) The number of variant sites per genome.(B) 每个基因组的变异位点数量。
(C) The average number of singletons per genome.(C) 每个基因组的平均单例变异位点数量。
From The 1000 Genomes Project Consortium.数据来自千人基因组计划联盟。
A global reference for human genetic variation.人类遗传变异的全球参考。
Nature 526:68–74, 2015. https:// doi. org/10. 1038/nature 15393 alleles A and a: AA, Aa, and aa.《自然》526:68–74, 2015. https://doi.org/10.1038/nature15393 等位基因 A 和 a:AA、Aa 和 aa。
Assuming that there is no mutation or other violations of HWE conditions, the proportions of genotypes in F1 can be calculated by plugging in the P1 genotype pairings as such:假设没有突变或其他违反哈代-温伯格平衡条件的情况,F₁ 中基因型的比例可通过代入 P₁ 基因型配对计算如下:
15/53
Population Genetics for Genomic Medicine 189 Suppose p is the frequency of allele A, and q is the frequency of allele a …
Ch10 — Segment 15
Population Genetics for Genomic Medicine 189 Suppose p is the frequency of allele A, and q is the frequency of allele a in the gene pool, such that p + q = 1 for a hypothetical biallelic trait.基因组医学中的群体遗传学 189 假设p是等位基因A的频率,q是基因库中等位基因a的频率,对于一个假设的双等位基因性状,满足p + q = 1。
Substituting in the frequency variables p and q (below) for their Aa F1 (AA x aa)p(1. 0) Aa F1 (AA x Aa)p(0. 5) AAF1 (AA x Aa)p(0. 5) Aa F1 (Aa x aa)p(0. 5) Aa F1 (Aa x Aa)p(0. 25) AAF1 (Aa x Aa)p(0. 25) AAF1 (AA x AA)p(1. 0) AAF1 (Aa x AA)p(0. 5) aa F1 (Aa x aa)p(0. 5) aa F1 (Aa x Aa)p(0. 25) Aa F1 (Aa x Aa)p(0. 25) Aa F1 (Aa x AA)p(0. 5) aa F1 (aa x aa)p(1. 0) aa F1 (aa x Aa)p(0. 5) Aa F1 (aa x Aa)p(0. 5) Aa F1 (aa x AA)p(1. 0) AA Aa aa AA Parent (p) Genotypes Offspring genotype frequencies given parental genotypes, assuming Hardy-Weinberg conditions (no mutation) Aa aa Now let us assume alleles combine into genotypes randomly; that is, mating in the population is completely random with respect to the genotypes at this locus.将频率变量p和q(如下)代入其Aa F1(AA × aa)p(1.0)、Aa F1(AA × Aa)p(0.5)、AA F1(AA × Aa)p(0.5)、Aa F1(Aa × aa)p(0.5)、Aa F1(Aa × Aa)p(0.25)、AA F1(Aa × Aa)p(0.25)、AA F1(AA × AA)p(1.0)、AA F1(Aa × AA)p(0.5)、aa F1(Aa × aa)p(0.5)、aa F1(Aa × Aa)p(0.25)、Aa F1(Aa × Aa)p(0.25)、Aa F1(Aa × AA)p(0.5)、aa F1(aa × aa)p(1.0)、aa F1(aa × Aa)p(0.5)、Aa F1(aa × Aa)p(0.5)、Aa F1(aa × AA)p(1.0) 等项,假设哈迪-温伯格条件(无突变)下,根据亲本基因型的后代基因型频率表(AA、Aa、aa亲本基因型)如下所示。现在让我们假设等位基因随机组合成基因型;即,群体中在该位点的基因型上的交配是完全随机的。
The chance that two A alleles will pair up to give the AA genotype is then p 2; the chance that two a alleles will come together to give the aa genotype is q 2; and the chance of having one A and one a pair, resulting in the Aa genotype, is 2pq (the factor 2 comes from the fact that the A allele could be inherited from one parent and the a allele from the other, or vice versa).两个A等位基因配对形成AA基因型的概率为p²;两个a等位基因结合形成aa基因型的概率为q²;而一个A和一个a配对产生Aa基因型的概率为2pq(系数2源于A等位基因可来自一个亲本而a等位基因来自另一个亲本,反之亦然)。
This applies to all autosomal loci and to the X chromosome in females, but not to X-linked loci in males who have just one X chromosome. respective alleles in F1 genotype proportion equations (above), the table below expresses offspring genotype proportions for all possible mating pairs in P in terms of p and q.这适用于所有常染色体位点和女性的X染色体,但不适用于只有一条X染色体的男性的X连锁位点;而上述F1基因型比例方程中的相应等位基因,下表以p和q表示亲本代P中所有可能配对的后代基因型比例。
Genotype frequencies in offspring generation (F1), given parental genotype frequencies; assuming two alleles (A and a) are in HWE with population allele frequencies Freq[A] = p and Freq[a] = q; such that p + q = 1.在给定亲本基因型频率的条件下,子代(F1)的基因型频率;假设两个等位基因(A和a)处于哈迪-温伯格平衡,群体等位基因频率为Freq[A] = p和Freq[a] = q,且满足p + q = 1。
Parent (P) Genotype Frequencies AA = p 2 AAF1 (p 2 x p 2)(1. 0) = p 4 AAF1 (2pq x p 2)(0. 5) = p 3q Aa F1 (2pq x p 2)(0. 5) = p 3q Aa F1 (q 2 x p 2)(1. 0) = p 2q2 AAF1 (p 2 x 2pq)(0. 5) = p 3q AAF1 (2pq x 2pq)(0. 25) = p 2q2 Aa F1 (2pq x 2pq)(0. 25) = p 2q2 Aa F1 (q 2 x 2pq)(0. 5) = pq 3 Aa F1 (p 2 x 2pq)(0. 5) = p 3q Aa F1 (2pq x 2pq)(0. 25) = p 2q2 aa F1 (2pq x 2pq)(0. 25) = p 2q2 aa F1 (q 2 x 2pq)(0. 5) = pq 3 Aa F1 (p 2 x q 2)(1. 0) = p 2q2 Aa F1 (2pq x q 2)(0. 5) = pq 3 aa F1 (2pq x q 2)(0. 5) = pq 3 aa F1 (q 2 x q 2)(1. 0) = q 4 AA = p 2 Aa = 2pq Aa = 2pq aa = q 2 aa = q 2 The Hardy-Weinberg principle states that the frequency of the three genotypes AA, Aa, and aa is given by the terms of the binomial expansion of (p + q)2 = p 2 + 2pq + q 2 = 1.亲本代(P)基因型频率 AA = p²;AA F1 (p² × p²)(1.0) = p⁴;AA F1 (2pq × p²)(0.5) = p³q;Aa F1 (2pq × p²)(0.5) = p³q;Aa F1 (q² × p²)(1.0) = p²q²;AA F1 (p² × 2pq)(0.5) = p³q;AA F1 (2pq × 2pq)(0.25) = p²q²;Aa F1 (2pq × 2pq)(0.25) = p²q²;Aa F1 (q² × 2pq)(0.5) = pq³;Aa F1 (p² × 2pq)(0.5) = p³q;Aa F1 (2pq × 2pq)(0.25) = p²q²;aa F1 (2pq × 2pq)(0.25) = p²q²;aa F1 (q² × 2pq)(0.5) = pq³;Aa F1 (p² × q²)(1.0) = p²q²;Aa F1 (2pq × q²)(0.5) = pq³;aa F1 (2pq × q²)(0.5) = pq³;aa F1 (q² × q²)(1.0) = q⁴;AA = p²;Aa = 2pq;Aa = 2pq;aa = q²;aa = q²。哈迪-温伯格原理指出,三种基因型AA、Aa和aa的频率由(p + q)² = p² + 2pq + q² = 1的二项展开式各项给出。
If allele frequencies do not change from generation to generation, the proportion of genotypes will not change either; that is, the genotype frequencies from generation to generation will remain constant (at equilibrium) in the population if the allele frequencies (p and q) remain constant.如果等位基因频率逐代保持不变,则基因型的比例也不会改变;也就是说,如果等位基因频率(p和q)保持不变,群体中逐代的基因型频率将保持恒定(处于平衡状态)。
When there is random mating in a population at HWE equilibrium, and genotypes AA, Aa, and aa are present in the proportions p 2:2pq:q 2, then genotype frequencies in the next generation will remain in the same relative proportions, p 2:2pq:q 2.当群体处于哈迪-温伯格平衡并进行随机交配,且基因型AA、Aa和aa以p²:2pq:q²的比例存在时,下一代基因型频率将保持相同的相对比例p²:2pq:q²。
This is proven as follows:证明如下:
16/53
Genotype frequencies in the parent generation (P) can be used to predict genotype frequencies in the first generation of…
Ch10 — Segment 16
Genotype frequencies in the parent generation (P) can be used to predict genotype frequencies in the first generation of offspring (F1) by adding up the resulting genotype frequencies from all possible mating pairs (P1,2).[TL:failed]
Constant Genotype frequency Proportions for a Population in Hardy-Weinberg Equilibrium (HWE) Genotype frequency Proportions remain constant in a population under Hardy-Weinberg conditions.[TL:failed]
AA x AA: AA x Aa: AA x aa: Aa x Aa: Aa x aa: = p 2(p 2 + 2pq + q 2) = p 2(p + q)2 = p 2(1)2 aa x aa: F1 AA p 2p2 = p 4 p 2q + p 3q = 2p3q 0 p 2q2 p 4 + 2p3p + p 2p2 = 2pq(p 2 + 2pq + q 2) = 2pq(p + q)2 = 2pq(1)2 2p3q + 2p2q2 + 2p2q2 + 2pq 3 = q 2(p 2 + 2pq + q 2) = q 2(p + q)2 = q 2(1)2 p 2p2 + 2pq 3 + q 4 0 0 0 p 3q + p 3q = 2p3q p 2q2 + p 2q2 = 2p2q2 p 2q2 + p 2q2 = 2p2q2 pq 3 + pq 3 = 2pq 3 0 0 0 0 p 2q2 pq 3 + pq 3 = 2pq 3 q 4:::::::::::::::: F2 Aa aa AA P Aa aa p 2 AA = p 2 Aa = 2pq aa = q 2 q 2 2pq AA p 2 Aa + + 2pq aa = 1 q 2 2 HARDY-WEINBERG EQUILIBRIUM ASSUMPTIONS The principle of Hardy-Weinberg equilibrium rests on the assumption that genotype frequency proportions remain constant over time because of the following: Random mating.[TL:failed]
Reproductive pairings are random with respect to the locus in question.[TL:failed]
No genetic drift.[TL:failed]
The population under study is sufficiently large such that alleles are not likely to dramatically rise or drop in frequency by random chance.[TL:failed]
No mutation.[TL:failed]
Rate of mutation is low such that allele frequencies are not impacted.[TL:failed]
No selection.[TL:failed]
Individuals are equally capable of passing on their genes, regardless of genotype, preserving equality between each allele frequency and its chance of being inherited.[TL:failed]
No gene flow.[TL:failed]
There has been no significant migration of individuals between populations with significantly different allele frequencies.[TL:failed]
A population that appears to meet these assumptions is in Hardy-Weinberg equilibrium.[TL:failed]
It is important to note that HWE does not require any particular values for p and q; whatever allele frequencies happen to be present in the population will result in genotype frequencies of p 2:2pq:q 2, and these relative genotype frequencies will remain constant from generation to generation as long as the allele frequencies remain constant and the other conditions introduced in 2 are met.[TL:failed]
This principle can be adapted for genes with more than two alleles.[TL:failed]
For example, if a locus has three alleles, with frequencies p, q, and r, the genotypic distribution can be determined from (p + q + r)2 = 1.[TL:failed]
In general terms, the genotype frequencies for any known number of alleles an with allele frequencies p 1, p 2, … pn can be derived from the terms of the expansion of (p 1 + p 2 + … pn)2.[TL:failed]
Applying HWE to the ACKR1 (Duffy blood group) example given earlier, with relative frequencies of the two alleles in the 1KG Project dataset 0. 734 (for the T allele) and 0. 266 (for the C allele), the relative proportions of the three combinations of alleles (genotypes) are Pr[T|T] = p 2 = 0. 734 × 0. 734 = 0. 539 (for an individual having two T alleles), Pr[C|C] = q 2 = 0. 266 × 0. 266 = 0. 071 (for two C alleles), and Pr[C|T] = 2pq = (0. 734 × 0. 266) + (0. 734 × 0. 266) = 0. 39 (for individuals with one T and one C allele).[TL:failed]
These genotype frequencies were calculated assuming HWE; so, when applied to a population of 2504 individuals, the derived numbers of people with the three different genotypes (TT:CT:CC) should be equivalent to the observed genotype frequencies from the 1KG dataset.[TL:failed]
However, when we do this calculation (total population size × genotype frequency), the proportions of individuals with each genotype are 1350:977:177.[TL:failed]
This is very different from the actual observed proportions in This[TL:failed]
17/53
Population Genetics for Genomic Medicine 191 makes sense, because the sampling scheme of the project was meant to identi…
Ch10 — Segment 17
Population Genetics for Genomic Medicine 191 makes sense, because the sampling scheme of the project was meant to identify individuals from different parts of the world to enable comparisons of average frequencies across the globe and cannot be considered a single population in HWE.[TL:failed]
Two key HWE assumptions that are violated in this scenario are: (1) Random mating, because we would not expect individuals to meet and reproduce with people from across the globe with equal chance compared to those close by; and (2) selection, since we know this allele is under positive selective pressure in locations that have a high incidence of malaria.[TL:failed]
Now, let us try this exercise again with a different variant in the same gene (e. g., ACKR1 rs 36007769; Ensembl) that is synonymous and therefore not predicted to change the protein, such that its frequency is not under the influence of natural selection (unless it is in strong LD with a functionally relevant variant that is under selection and carries it along).[TL:failed]
We can again use HWE to calculate predicted genotype frequencies from observed allele frequencies in 1KG and compare those results to the observed genotype frequencies reported for this dataset in Ensembl.[TL:failed]
Recall that the HWE equation p 2 + 2pq + q 2 = 1 can be used to calculate expected genotype frequencies, given observed allele frequencies such that the first and third terms (p 2 and q 2) are the expected frequencies of homozygous genotypes for alleles p and q, respectively, and the second term (2pq) is the expected frequency of heterozygotes.[TL:failed]
Comparing these derived frequencies (0. 992:0. 008: 0. 0) to observed genotype frequencies in the 1KG dataset (0. 993:0. 007:0. 0), we can see that they are roughly equivalent.[TL:failed]
Given that the assumptions of HWE hold, we would expect these genotype frequencies to remain constant generation after generation.[TL:failed]
Although the 1KG dataset is not a genetically homogeneous population and it includes samples from many different parts of the world, the absence of selection acting on this variant and the fact that it is rare in a relatively large population (defined as the dataset) mean that genotype frequency calculations based on HWE are still useful.[TL:failed]
Upon further inspection of the incidence of this variant, 5 occurrences are in Central and South American ancestry (AMR) populations and 13 occurrences are in European ancestry (EUR) populations.[TL:failed]
Using the reported allele frequencies in those populations from 1KG, AMR (G: 0. 993, A: 0. 007), and EUR (G: 0. 987, A: 0. 013) to calculate expected genotype frequencies with HWE, they are equivalent to reported frequency proportions in AMR (0. 986:0. 014:0) and EUR (0. 974:0. 026:0).[TL:failed]
This suggests that HWE holds for each of these populations for this specific variant, and thus their frequency proportions should persist in subsequent generations.[TL:failed]
Applying Hardy-Weinberg Equilibrium to Autosomal Recessive Traits The major practical application of the Hardy-Weinberg principle in medical genetics is in genetic counseling for autosomal recessive conditions.[TL:failed]
For a disease such as phenylketonuria (PKU), there are hundreds of different pathogenic alleles with frequencies that vary among different population groups defined by geography and/ or ethnicity (see Chapter 13).[TL:failed]
Affected individuals can be homozygoues for the same pathogenic allele, but they are often compound heterozygotes for different pathogenic variants (see Chapter 7).[TL:failed]
For many conditions, it is convenient to consider all disease-causing alleles together and treat them as a single pathogenic allele, with frequency q, even when there is significant allelic heterogeneity among pathogenic alleles.[TL:failed]
Similarly, the combined frequency of all benign or nonpathogenic alleles, p, is given by 1 − q.[TL:failed]
Suppose we would like to know the frequency of all disease-causing PKU alleles in a population for use in genetic counseling, for example, to inform couples of their risk for having a child with PKU.[TL:failed]
If we were to attempt to determine the frequency of disease-causing PKU alleles directly from genotype frequencies, we would need to know the frequency of heterozygotes in the population, a frequency that cannot be measured directly because of the recessive nature of PKU.[TL:failed]
This is because heterozygotes are asymptomatic silent carriers (see Chapter 7), and their frequency in the population (i. e., 2pq) cannot be reliably determined directly from phenotype observation..[TL:failed]
A 18 0. 004 Total 5008 1. 0 From The 1000 Genomes Project Consortium: A global reference for human genetic variation, Nature 526:68–74, 2015. doi:10. 1038/nature 15393.[TL:failed]
Accessed online via Ensembl. 993 p 2 = (0. 996)2 = 0. 992 A|G 18 0. 007 2pq = 2(0. 996) (0. 004) = 0. 008 A|A 0 0 q 2=(0)2=0[TL:failed]
18/53
However, the frequency of affected homozygotes/ compound heterozygotes for disease-causing alleles in the population (i.…
Ch10 — Segment 18
However, the frequency of affected homozygotes/ compound heterozygotes for disease-causing alleles in the population (i. e., q 2) could be determined directly, by counting the number of babies with PKU born over a given time period and identified through newborn screening (see Chapter 19), divided by the total number of babies screened during that same time period.然而,人群中致病等位基因的患病纯合子/复合杂合子的频率(即 q²)可直接通过计算特定时间段内出生并经新生儿筛查确认的苯丙酮尿症婴儿数量,除以同一时间段内接受筛查的婴儿总数来确定(见第19章)。
Now, using HWE, we can calculate the pathogenic allele frequency (q) from the observed frequency of homozygotes/ compound heterozygotes alone (q 2), thereby providing an estimate (2pq) of the frequency of heterozygotes for use in genetic counseling.现在,利用哈迪-温伯格平衡,我们可以仅从观察到的纯合子/复合杂合子频率(q²)计算出致病等位基因频率(q),从而提供杂合子频率的估计值(2pq),用于遗传咨询。
To illustrate this example further, consider a population in which the frequency of PKU is approximately 1 per 4500.为进一步说明此例,考虑一个苯丙酮尿症频率约为每4500人中1例的人群。
If we group all disease-causing alleles together and treat them as a single allele with frequency q, then the frequency of affected individuals q 2 = 1/4500.如果我们将所有致病等位基因归为一组,并将其视为频率为 q 的单个等位基因,则患病个体频率 q² = 1/4500。
From this, we calculate q = 0. 015, and thus 2pq = 0. 029.由此,我们计算出 q = 0.015,因此 2pq = 0.029。
The carrier frequency for all disease-causing alleles lumped together in this population is therefore approximately 3%.因此,该人群中所有致病等位基因合并后的携带者频率约为3%。
For an individual known to be a carrier of PKU through the birth of an affected child in the family, there would then be an approximately 3% chance that he or she would find a new mate from the same population who would also be a carrier, and this estimate could be used to provide genetic counseling.对于因家族中生育患病儿童而已知为苯丙酮尿症携带者的个体,其从同一人群中选择的新配偶同样为携带者的概率约为3%,此估计值可用于提供遗传咨询。
Note, however, that this estimate applies only to the population in question; if the new mate was from a different genetic ancestral population where the frequency of PKU is much lower (e. g., 1 per 200,000), their chance of being a carrier would be only 0. 6%.但需注意,此估计仅适用于所讨论的人群;若新配偶来自苯丙酮尿症频率低得多(例如每20万人中1例)的不同遗传祖先人群,则其为携带者的概率仅为0.6%。
In this example, all PKU-causing alleles are collapsed for the purpose of estimating q.在本例中,为估计 q 值,将所有导致苯丙酮尿症的等位基因合并处理。
For other conditions, however, such as hemoglobin disorders that we will consider in Chapter 12, different pathogenic alleles can lead to very different conditions, and therefore it would make no sense to group all pathogenic alleles together, even when the same locus is involved.然而,对于其他疾病,如我们将在
Instead, the frequencies of alleles leading to different phenotypes (e. g., sickle cell disease and β-thalassemia in the case of different pathogenic alleles at the β-globin locus) are calculated separately.[TL:missing]
Allele and Genotype Frequencies in X-Linked Conditions Recall from Chapter 7 that, for X-linked genes, there are three female genotypes but only two possible male genotypes.[TL:missing]
To illustrate the relationship between allele and genotype frequencies when a gene of interest is X linked, we use the trait known as X-linked red-green color blindness, which is caused by structural variants in genes encoding cell receptors that respond to photons of light at wavelengths we perceive as red and green (OPN1LW and OPNMW, respectively) that are adjacent to one another on the X chromosome.[TL:missing]
We use color blindness as an example because, as far as we know, it is not a deleterious trait (except for possible difficulties with traffic lights), and persons with color blindness are not subject to selection.[TL:missing]
In this example, we use the symbol cb to represent variants conferring some variation of color blindness and the symbol + for variants without color blindness, with frequencies q and p, respectively ( Because females have two X chromosomes, their genotypes are distributed like autosomal genotypes, but because color blindness variants are recessive, their homozygous and heterozygous genotypes are typically not distinguishable.[TL:missing]
In contrast, males with only one X chromosome will exhibit the trait with a single copy of a cb variant.[TL:missing]
As such, the frequency of color blindness in females is much lower than that in males (<1%).[TL:missing]
Frequencies of the two types of variants (cb and +) can be determined directly from the prevalence of the corresponding phenotypes in males.[TL:missing]
So, if the prevalence of colorblindness in a population of biologically male individuals (with an X and a Y chromosome) is roughly 8%, the frequency of cb variants in the population is likewise 0. 08.[TL:missing]
Genotypes of unaffected females (homozygous or heterozygous) cannot be determined by looking at phenotypes; frequencies of variants for the trait that were ascertained by looking at phenotype frequencies in male individuals can be used to determine approximate variant frequencies for females using HWE.[TL:missing]
As shown in Among these heterozygous unaffected individuals, those who are pregnant with a male fetus have a 50% chance of giving birth to a male child with colorblindness (because they have a 50% chance of passing on the Xcb variant and this is the only X chromosome a male will receive from either parent).[TL:missing]
In contrast, a female fetus receives two copies of the X chromosome, so the 50% probability of an unaffected carrier transmitting the cb variant is instead the chance that a female fetus will also be a silent carrier. 08 0. 92 q=0. 08 p=0. 92 Female (X/X) Colorblindness Unaffected cb <1% Xcb Xcb Xcb X+ X+X+ q 2 = (0. 08)2 = 0. 0064 2pq = 2(0. 08)(0. 92) = 0. 1472 p 2 = (0. 92)2 = 0. 8464[TL:missing]
19/53
Population Genetics for Genomic Medicine 193 Violating Assumptions of Hardy-Weinberg Equilibrium Underlying the principl…
Ch10 — Segment 19
Population Genetics for Genomic Medicine 193 Violating Assumptions of Hardy-Weinberg Equilibrium Underlying the principle of Hardy-Weinberg equilibrium and its use are several assumptions (see 2), not all of which can be met (or reasonably inferred to be met) by all populations.基因组医学中的群体遗传学 193 违背哈迪-温伯格平衡的假设 哈迪-温伯格平衡及其应用的基础是若干假设(见第2节),并非所有群体都能满足(或可合理推断满足)这些假设。
In this section, we provide a highlevel overview of the conditions and factors that contribute to violations of HWE assumptions: (1) nonrandom mating, (2) genetic drift, (3) mutation, (4) selection, and (5) gene flow.在本节中,我们概述导致违背HWE假设的条件和因素:(1) 非随机交配,(2) 遗传漂变,(3) 突变,(4) 选择,以及 (5) 基因流。
We have seen examples of these throughout the chapter, as population genetics is concerned with modeling and measuring shifts in allele frequencies.我们在本章中已见过这些因素的实例,因为群体遗传学关注的是等位基因频率变化的建模与测量。
First, we have seen nonrandom mating at work in small genetic isolates, or founder populations, which (by definition) are isolated from reproduction events outside the group.首先,我们观察到非随机交配在小型遗传隔离群或奠基者群体中发挥作用,这些群体(按定义)与群体外的繁殖事件相隔离。
In human populations, assortative mating, underlying population structure, or cryptic relatedness due to shared recent ancestry, and consanguinity can all lead to nonrandom mate choices.在人类群体中,选型交配、潜在群体结构、因近期共同祖先导致的隐匿亲缘关系以及近亲结婚均可导致非随机配偶选择。
Assortative mating is a type of nonrandom mating in which individuals in a population engage in preferential reproductive choices.选型交配是一种非随机交配,其中群体中的个体进行优先性的繁殖选择。
This may increase the frequency of variants contributing to traits that influence individuals in these reproductive choices and other variants in LD with them.这可能增加影响这些繁殖选择性状的变异以及与之连锁不平衡的其他变异的频率。
Genetic drift refers to changes in allele frequencies over time due to random chance, which occur more quickly in smaller populations.遗传漂变指因随机偶然事件导致等位基因频率随时间变化,在小群体中变化更为迅速。
When a new mutation occurs in a small population, its frequency is represented by only one copy among all the copies of that gene in the population.当小群体中出现新突变时,其频率仅由该基因在群体中所有拷贝中的一个拷贝代表。
Random effects of the environment or other chance occurrences that are independent of the genotype (i. e., events that occur for reasons unrelated to whether an individual is carrying a pathogenic variant) can produce significant changes in the frequency of the disease allele when the population is small.环境随机效应或独立于基因型的其他偶然事件(即与个体是否携带致病性变异无关的事件)可在群体较小时导致疾病等位基因频率发生显著变化。
Such chance occurrences disrupt Hardy-Weinberg equilibrium and cause the allele frequency to change from one generation to the next.此类偶然事件破坏哈迪-温伯格平衡,导致等位基因频率逐代改变。
During the next few generations, although the population size of the new group remains small, there may be considerable fluctuation until allele frequencies come to a new equilibrium as the population increases in size.在随后几代中,尽管新群体规模仍小,可能存在显著波动,直至群体规模增大后等位基因频率达到新平衡。
HWE assumes no genetic drift – which requires an absence of migration in and out of the population by groups whose allele frequencies at loci of interest differ drastically from those of the population under HWE assumptions.HWE假设无遗传漂变——这要求不存在群体迁入或迁出,且迁入或迁出群体在关注位点的等位基因频率与HWE假设下的群体显著不同。
This is a phenomenon called gene flow, whereby alleles are exchanged into and out of populations via migration and reproduction among individuals from different populations.此现象称为基因流,即通过不同群体个体间的迁移与繁殖,等位基因进出群体进行交换。
Gene flow disrupts HWE because allele frequencies can change when new alleles are introduced into a population.基因流破坏HWE,因为新等位基因引入群体时等位基因频率可能改变。
Similarly, mutation (see Chapter 4) introduces new allelic variants into a population at random, which can influence stability of allele frequencies.类似地,突变(见第4章)随机向群体引入新的等位基因变异,可影响等位基因频率的稳定性。
When these newly introduced alleles (either through gene flow or mutation) confer some evolutionary advantage over existing alleles in the population, they will naturally rise in frequency due to pressures of natural selection, thereby disrupting HWE.当这些新引入的等位基因(通过基因流或突变)比群体中现有等位基因具有某种进化优势时,由于自然选择压力,其频率自然上升,从而破坏HWE。
Changes in allele frequencies due to selection or mutation usually occur slowly, in small increments, and cause much less deviation from HWE, at least for recessive diseases.由选择或突变导致的等位基因频率变化通常缓慢、增量微小,且对HWE的偏离程度较小,至少针对隐性遗传病而言。
This is because rates of new mutations are generally well below the frequency of heterozygotes for autosomal recessive diseases.这是因为新突变率通常远低于常染色体隐性遗传病杂合子的频率。
The addition of new pathogenic alleles to the gene pool thus has little effect (in the short term) on allele frequencies for such diseases.因此,向基因库添加新的致病等位基因(短期内)对此类疾病的等位基因频率影响甚微。
In addition, most deleterious recessive alleles are hidden in asymptomatic heterozygotes and thus are not subject to selection.此外,大多数有害隐性等位基因隐藏于无症状杂合子中,不受选择作用。
Consequently, selection is not likely to have major short-term effects on allele frequencies of these recessive alleles.因此,选择不太可能对这些隐性等位基因的频率产生显著短期影响。
Therefore, to a first approximation, HWE may apply even for alleles that cause severe autosomal recessive disease.故作为一级近似,HWE甚至可能适用于导致严重常染色体隐性遗传病的等位基因。
Importantly, however, for dominant or X-linked conditions, mutation and selection do perturb allele frequencies from what would be expected under HWE, by substantially reducing or increasing certain genotypes in just a few generations.但重要的是,对于显性或X连锁疾病,突变和选择确实会扰动等位基因频率,使其偏离HWE预期值,可在短短几代内显著减少或增加某些基因型。
In practice, some violations of HWE we have discussed are more disruptive than others when applying the principle to human populations.在实践中,我们讨论的某些HWE违背情况在将该原则应用于人类群体时比其他情况更具破坏性。
For example, violating the assumption of random mating can cause large deviations from the expected frequency of individuals homozygous for an autosomal recessive condition.例如,违背随机交配假设可导致常染色体隐性遗传病纯合子个体的预期频率出现大幅偏差。
In contrast, changes in allele frequency due to mutation, selection, or migration usually cause more minor and gradual deviations from HWE.相反,由突变、选择或迁移引起的等位基因频率变化通常导致对HWE的偏离更轻微且更渐进。
When HWE assumptions do not hold for a particular disease allele at a particular locus, it may be instructive to investigate why the allele and its associated genotypes are not in equilibrium as this may provide clues about the pathogenesis of the condition or point to historical events that have affected the frequency of alleles in different population groups over time.当针对特定位点的特定疾病等位基因不满足HWE假设时,探究该等位基因及其相关基因型为何未达平衡可能具有启发意义,因为这可能为疾病的发病机制提供线索,或指示随时间推移影响不同群体等位基因频率的历史事件。
Mutation and Selection Balance in Traits With Different Modes of Inheritance In this section, we examine the concepts of mutation and selection through the lens of fitness, a heuristic device that indicates the likelihood of a mutation at a particular locus being eliminated, becoming stable, or becoming (over time) the predominant allele (or fixed) in a population.不同遗传方式性状中的突变与选择平衡 在本节中,我们通过适合度这一启发式工具来审视突变与选择的概念,适合度表示特定位点突变被消除、趋于稳定或(随时间)成为群体中主导等位基因(或固定)的可能性。
The frequency of an allele in a population at any given time represents a balance between the rate at which new alleles appear through mutation and the influence of selection on these alleles.某一等位基因在群体中任何时刻的频率,代表了通过突变出现的新等位基因速率与选择对这些等位基因影响之间的平衡。
If the mutation rate or the effectiveness of selection is altered, the allele frequency is expected to change.若突变率或选择的有效性发生改变,则等位基因频率预期也会变化。
More formally, whether an allele is transmitted to the succeeding generation depends on its fitness, ω, which is a quantitative measure for the expected (average) number of offspring of affected persons who survive to更正式地说,等位基因是否传递给下一代取决于其适合度ω,这是对受累者存活至生育年龄的后代预期(平均)数量的定量度量。
20/53
reproductive age.
Ch10 — Segment 20
reproductive age.生殖年龄。
This measure is called relative fitness when compared with that of an appropriate control group.这个指标在与适当对照组比较时称为相对适应度。
If a pathogenic allele is just as likely to be represented in the next generation compared to functionally neutral alleles, ω = 1.如果致病等位基因与功能中性等位基因相比,在下一代中出现的可能性相同,则ω = 1。
If an allele causes death or sterility, purifying, or negative, selection acts against it completely, and ω = 0.如果某个等位基因导致死亡或不育,纯化选择或负选择会完全对其起作用,且ω = 0。
Values between 0 and 1 indicate transmission of the variant, and values of ω > 1 indicate positive selection increasing the variant’s frequency in a population.介于0和1之间的值表示变异体的传递,而ω > 1的值表示正选择增加了变异体在人群中的频率。
A related parameter is the coefficient of selection, s, which is a measure of the loss of fitness and is defined as 1− ω, that is, the proportion of pathogenic alleles that are not passed on and are therefore lost due to negative selection.一个相关参数是选择系数s,它是适应度损失的度量,定义为1−ω,即未传递给后代因而因负选择而丢失的致病等位基因的比例。
When a genetic condition limits reproduction such that ω = 0 and s = 1, the variant conferring this trait is referred to as a genetic lethal.当某种遗传状况限制生殖,使得ω = 0且s = 1时,赋予该性状的变异体被称为遗传致死。
In the genetic sense, a variant that prevents reproduction by an adult is just as “lethal” as one that causes a very early miscarriage of an embryo, because in neither case is the variant transmitted to the next generation.在遗传学意义上,阻止成年个体生殖的变异体与导致胚胎极早期流产的变异体同样“致死”,因为这两种情况下变异体都不会传递给下一代。
Fitness is thus the outcome of the joint effects of survival and fertility.因此,适应度是生存和生育力共同作用的结果。
In the biological sense, relative fitness has no connotation of superior endowment but is simply a measure of comparative ability to contribute alleles to the next generation, on average.在生物学意义上,相对适应度并无优越禀赋的含义,而仅仅是衡量平均而言向下一代贡献等位基因的比较能力。
The frequency of pathogenic alleles in a population represents a balance between loss of pathogenic alleles through the effects of selection and gain of pathogenic alleles through recurrent mutation.人群中致病等位基因的频率代表了通过选择效应丢失致病等位基因与通过反复突变获得致病等位基因之间的平衡。
A stable allele frequency will then be reached at whatever level balances the two opposing forces: one (selection) that removes pathogenic alleles from the gene pool and one (de novo mutation) that adds new ones back.然后会在平衡两种对立力量的任何水平上达到稳定的等位基因频率:一种力量(选择)从基因库中移除致病等位基因,另一种力量(新生突变)重新添加新的致病等位基因。
The mutation rate per generation, µ, at a locus with pathogenic variants must be sufficient to account for the fraction of all pathogenic alleles that are lost by selection from each generation.具有致病变异体的位点每代突变率µ必须足以解释每一代通过选择丢失的所有致病等位基因的比例。
That is, µ=sq where µ is the mutation rate, s is the coefficient of selection, and q is the allele frequency.即µ=sq,其中µ是突变率,s是选择系数,q是等位基因频率。
If a pathogenic allele for a condition has a dominant mode of inheritance and is deleterious but not lethal, affected persons may reproduce but will nevertheless contribute fewer than the average number of offspring to the next generation; that is, 0 &lt; ω &lt; 1.如果某种疾病的致病等位基因呈显性遗传模式且有害但不致死,受累个体可能能够生殖,但向下一代贡献的后代数量仍将低于平均水平;即0 < ω < 1。
Such a variant will be lost through selection at a rate proportional to the reduced fitness of heterozygotes.此类变异体将以与杂合子适应度降低成比例的速度通过选择丢失。
For example, consider a phenotypic trait that reduces the fitness of affected persons to an average of one-fifth the number of children compared to unaffected persons in a population.例如,考虑一种表型性状,该性状将受累个体的适应度降低至人群中未受累个体平均子女数量的五分之一。
The relative fitness is thus ω = 0. 20, and the coefficient of selection, s = 1 − ω = 0. 80.因此相对适应度ω = 0.20,选择系数s = 1 − ω = 0.80。
This means that in the next generation of offspring, only 20% of pathogenic alleles are passed on.这意味着在下一代后代中,只有20%的致病等位基因被传递。
When the frequency of such conditions appears stable from one generation to the next, new mutations are most likely responsible for replacing 80% of the pathogenic alleles predicted to be lost through selection.当此类状况的频率在代际间看似稳定时,新突变最有可能负责替代预计通过选择丢失的80%的致病等位基因。
If the fitness of affected persons suddenly improves (e. g., because of medical advances or removal of other barriers to reproduction for individuals with the condition), the observed incidence in the population is predicted to increase and reach a new equilibrium.如果受累个体的适应度突然改善(例如由于医学进步或消除了该状况个体生殖的其他障碍),预计人群中观察到的发病率将增加并达到新的平衡。
Retinoblastoma (Case 39) and other dominant embryonic tumors with childhood onset are examples of conditions that now have a greatly improved prognosis compared to when they were first described, which may have increased the frequency of such conditions in the population.视网膜母细胞瘤(病例39)和其他儿童期发病的显性胚胎性肿瘤是预后较初次描述时显著改善的疾病示例,这可能增加了此类疾病在人群中的频率。
Selection against pathogenic alleles with a recessive mode of inheritance has less of an effect on population frequencies than selection against dominant variants, because only a small proportion of these alleles cooccur in homozygotes, with phenotypic presentation that exposes them selective pressure.针对隐性遗传模式的致病等位基因的选择对人群频率的影响小于针对显性变异体的选择,因为只有一小部分此类等位基因以纯合子形式共同出现,其表型呈现使其暴露于选择压力。
Even if there were complete selection against homozygotes (ω = 0), as in many lethal autosomal recessive conditions, it would take many generations to reduce the allele frequency appreciably because most pathogenic alleles are carried by unaffected carriers (heterozygotes).即使对纯合子进行完全选择(ω = 0),如同许多致死性常染色体隐性遗传病一样,也需要许多世代才能显著降低等位基因频率,因为大多数致病等位基因由未受累携带者(杂合子)携带。
For example, the frequency of pathogenic alleles causing Tay-Sachs disease, q, can be as high as 1. 5% in Ashkenazi Jewish populations.例如,导致泰-萨克斯病的致病等位基因频率q在德系犹太人群中可高达1.5%。
Given this value q=0. 015;p=1−0. 015=0. 95;thustheproportion of heterozygous individuals is expected to be 2pq = 2(0. 015)(0. 95) = 0. 029.给定该值q=0.015;p=1−0.015=0.95;因此杂合个体比例预期为2pq = 2(0.015)(0.95) = 0.029。
This means approximately 3% of individuals in such a population are expected to carry one copy of the pathogenic variant.这意味着该人群中约3%的个体预期携带一份致病变异体拷贝。
In contrast, only 1 individual per 4500 (q 2 = 0. 015 β 0. 015 = 0. 0002) is homozygous for the pathogenic allele, such that selection can act on the recessive phenotype.相比之下,每4500人中仅有1人(q² = 0.015 × 0.015 = 0.0002)为致病等位基因纯合子,因此选择可以对隐性表型起作用。
The proportion of all pathogenic variants among homozygotes in such a population is thus: 2 0 0002 2 0 0002 1 0 03 0 0132 × × () + × () ≈.... such that &lt;2% of all pathogenic variants in the population are in affected individuals, whose condition would be exposed to negative selection in the absence of effective treatment.因此在该人群中,纯合子中所有致病变异体的比例为:2 × 0.0002 / (2 × 0.0002 + 1 × 0.03) ≈ 0.0132,使得人群中不到2%的致病变异体存在于受累个体中,若无有效治疗,其状况将暴露于负选择。
We hope this example offers mathematical intuition as to why recessive variants are slow to influence allele frequencies in a population, due to the relatively small number of pathogenic variants that are subjected to selection at any given time.我们希望这个例子能提供数学上的直观理解,说明为何隐性变异体对人群中等位基因频率的影响缓慢,因为在任何给定时间受到选择的致病变异体数量相对较少。
Reduction or removal of selection against an autosomal recessive disorder by successful treatment (e. g., as in the case of PKU) would have just as slow an effect on increasing the allele frequency over many generations.通过成功治疗(例如苯丙酮尿症病例)减少或消除对常染色体隐性遗传病的选择,对增加等位基因频率的影响同样缓慢,需经过多个世代。
Thus, if mating is random, genotypes in autosomal recessive diseases are considered to be in Hardy-Weinberg equilibrium, despite selection against因此,如果交配是随机的,尽管存在选择,常染色体隐性遗传病中的基因型仍被视为处于哈代-温伯格平衡。
21/53
Population Genetics for Genomic Medicine 195 homozygotes for the recessive allele.
Ch10 — Segment 21
Population Genetics for Genomic Medicine 195 homozygotes for the recessive allele.基因组医学的群体遗传学 195 个隐性等位基因纯合子。
It follows that the mathematical relationship between genotype and allele frequencies described by HWE holds for most practical purposes in the case of an autosomal recessive disease.因此,由哈迪-温伯格平衡描述的基因型频率与等位基因频率之间的数学关系,就常染色体隐性遗传病而言,在大多数实际情况下成立。
In contrast to recessive pathogenic alleles, dominant pathogenic alleles are exposed directly to selection.与隐性致病等位基因不同,显性致病等位基因直接暴露于选择压力之下。
Consequently, the effects of selection and mutation are more obvious and can be more readily measured for dominant traits.因此,选择和突变对显性性状的影响更为明显,且更容易被测量。
A genetic lethal dominant allele, if fully penetrant, will be exposed to selection in heterozygotes, thus removing all alleles responsible for the disorder in a single generation.一个遗传致死性显性等位基因,若完全外显,将在杂合子中暴露于选择压力,从而在单代中清除所有导致该疾病的等位基因。
Several human diseases are thought to be autosomal dominant traits with zero or near-zero fitness and thus always result from de novo, rather than inherited, autosomal dominant variants.几种人类疾病被认为是适应性为零或接近零的常染色体显性性状,因此总是由新发而非遗传的常染色体显性变异引起。
This is a point of great significance for genetic counseling, and examples of these conditions are listed in In some of these conditions, the specific pathogenic alleles are known, and family studies have revealed de novo mutations in affected individuals that were not inherited from the parents.这对遗传咨询具有重要意义,这些疾病的例子列于……在某些此类疾病中,已知特定的致病等位基因,家系研究已揭示受累个体中存在并非遗传自父母的新发突变。
In other conditions, the responsible genes are not known, but paternal age effects (Chapter 4) have been observed, which suggests a possible mechanism of de novo mutations in the paternal germline.在其他疾病中,致病基因尚不明确,但已观察到父亲年龄效应(第4章),这提示父系生殖系中新发突变的一种可能机制。
The implication for genetic counseling is that parents of a child with an autosomal dominant (but genetically lethal) condition will typically have a very low risk of recurrence in subsequent pregnancies because the condition would require another independent de novo mutation.对遗传咨询的启示是,患常染色体显性(但遗传致死性)疾病儿童的父母,在后续妊娠中复发风险通常非常低,因为该疾病需要另一次独立的新发突变。
A caveat to keep in mind is the possibility of germline mosaicism (See and the possibility of abundant de novo mutations in a germline heavily exposed to mutagens.需牢记的一点是生殖系嵌合体的可能性(参见……)以及大量暴露于诱变剂的生殖系中可能出现大量新发突变。
In clinically relevant conditions that have an X-linked recessive mode of inheritance, selection acts on hemizygous males but not in heterozygous females, except for the small proportion of females who are manifesting heterozygotes with reduced fitness (see Chapter 7).在具有X连锁隐性遗传模式的临床相关疾病中,选择作用于半合子男性,而不作用于杂合子女性,除了一小部分表现为杂合子且适应性降低的女性(见第7章)。
In this brief discussion, we assume that heterozygous females do not have reduced fitness.在此简要讨论中,我们假设杂合子女性不具有降低的适应性。
Because males have one X chromosome and females have two, the pool of X-linked alleles in the entire population’s gene pool is partitioned, such that one-third of pathogenic alleles are in males and two-thirds are in females.由于男性有一条X染色体而女性有两条,整个群体基因库中的X连锁等位基因池被分割,使得三分之一的致病等位基因存在于男性,三分之二存在于女性。
As we saw in the case of autosomal dominant variants, pathogenic alleles lost through selection must be replaced by recurrent new mutations to maintain the observed disease incidence.正如我们在常染色体显性变异中所见,通过选择丢失的致病等位基因必须由反复发生的新突变来补充,以维持观察到的疾病发病率。
If the incidence of an X-linked condition is not changing, and selection is operating (only) against hemizygous males, the mutation rate, µ, must equal the coefficient of selection, s (i. e., the proportion of pathogenic alleles that are not passed on), times q (the pathogenic allele frequency), adjusted by a factor of 3, since selection is operating only on the third of pathogenic alleles in the population that are present in males.如果X连锁疾病的发病率保持不变,且选择(仅)作用于半合子男性,则突变率µ必须等于选择系数s(即未传递的致病等位基因比例)乘以q(致病等位基因频率),再乘以因子3进行调整,因为选择仅作用于群体中存在于男性的那三分之一的致病等位基因。
The equation is thus: µ=sq/3 For an X-linked genetic lethal condition, s = 1, and one-third of all copies of the pathogenic allele are lost from each generation; so, at equilibrium, these must be replaced by de novo mutations.因此方程为:µ = sq/3。对于X连锁遗传致死性疾病,s = 1,每代有三分之一的所有致病等位基因拷贝丢失;因此,在平衡状态下,这些丢失的拷贝必须由新发突变补充。
Thus, roughly one-third of all persons who have X-linked lethal disorders are predicted to carry a de novo mutation, and their unaffected mothers have a low risk of future pregnancies harboring the same disorder (in the absence of germline mosaicism).因此,预计大约三分之一的X连锁致死性疾病患者携带新发突变,而其未患病的母亲在后续妊娠中怀有相同疾病胎儿的风险较低(在无生殖系嵌合体的情况下)。
The remaining two-thirds of mothers of individuals with an X-linked lethal disorder are predicted to be carriers, each with a 50% future risk of conceiving an affected child, given that it is male.预计其余三分之二的X连锁致死性疾病患者的母亲为携带者,每人有50%的未来风险怀上患病孩子(前提是胎儿为男性)。
However, the prediction that two-thirds of mothers of individuals with an X-linked lethal disorder are carriers of a disease-causing variant assumes that mutation rates in males and in females are equal.然而,该预测——即三分之二的X连锁致死性疾病患者的母亲是致病性变异携带者——假设男性和女性的突变率相等。
Given that the germline mutation rate is higher in males with advanced paternal age than in females, the chance of a de novo mutation occurring in the egg is very low.鉴于高龄父亲生殖系突变率高于女性,卵子中发生新发突变的概率非常低。
(Note: the impact on genetic counseling related to these sex-dependent considerations of mutation rates will be discussed in Chapter 17).(注:这些与性别相关的突变率考虑因素对遗传咨询的影响将在第17章讨论。)
Most mothers of affected children are carriers, having most likely inherited novel variants from unaffected fathers, which they then have a 50% chance of passing on to their children.大多数患病儿童的母亲是携带者,她们很可能从未患病的父亲那里遗传了新发变异,然后有50%的概率将其传递给子女。
Taking advantage of the statistical property that the probability of two independent events occurring is equal to the product of probabilities of each separate event, Pr[A and B] Pr[A] Pr[B] = × where A and B are independent events, we can therefore calculate the total risk of an unaffected carrier of an X-linked condition having a child with the condition as: Pr[male child] Pr[passing on the pathogenic allele] (0. 5)( × = 0. 5) 0. 25 = 7. 6 C)利用两个独立事件同时发生的概率等于各自事件概率之积这一统计性质,Pr[A and B] = Pr[A] × Pr[B],其中A和B是独立事件,因此我们可以计算X连锁疾病未患病携带者生育患病孩子的总风险为:Pr[男性孩子] × Pr[传递致病等位基因] = 0.5 × 0.5 = 0.25。7.6 C)
22/53
such that the probability that an unaffected carrier will give birth to a child with an X-linked genetic lethal conditio…
Ch10 — Segment 22
such that the probability that an unaffected carrier will give birth to a child with an X-linked genetic lethal condition is 25% or one-fourth, given that the other parent is not affected.使得在另一方未受影响的情况下,未受影响的携带者生出患有X连锁遗传致死性疾病孩子的概率为25%或四分之一。
In less severe disorders, such as hemophilia A (Case 21), the proportion of affected individuals representing new mutations is less than one-third (~15%).在较轻的疾病中,如血友病A(病例21),由新突变所致的患病个体比例不足三分之一(约15%)。
Because the treatment of hemophilia has improved significantly, the total frequency of pathogenic alleles can be expected to rise rapidly and to reach a new equilibrium.由于血友病的治疗显著改善,致病等位基因的总频率预计将迅速上升并达到新的平衡。
Assuming that the mutation rate at this locus stays the same over time, the proportion of those with hemophilia whose pathogenic variant arises de novo will decrease, but the overall incidence of the disease will increase.假设该位点的突变率随时间保持不变,血友病患者中致病变异为新发突变的比例将下降,但该病的总体发病率将上升。
Such a change would have significant implications for genetic counseling for this condition (see Chapter 17).这种变化将对本病的遗传咨询产生重要影响(见第17章)。
Although certain pathogenic alleles may be deleterious in homozygotes, there may be environmental conditions in which heterozygotes for some conditions have increased fitness relative to homozygotes for both the pathogenic allele and the reference allele.尽管某些致病等位基因在纯合子中可能有害,但在某些环境条件下,某些疾病的杂合子相对于致病等位基因和参考等位基因的纯合子可能具有更高的适应度。
This is called heterozygote advantage because even a slightly greater relative fitness of heterozygotes can lead to an increase in frequency of an allele that is severely detrimental in homozygotes.这称为杂合子优势,因为即使杂合子的相对适应度略高,也可能导致在纯合子中严重有害的等位基因频率增加。
This is because heterozygotes greatly outnumber homozygotes in the population.这是因为在群体中杂合子数量远多于纯合子。
A situation in which selective forces operate to both maintain a deleterious allele and remove it from the gene pool is often referred to as balancing selection.选择力同时维持有害等位基因并将其从基因库中移除的情况通常称为平衡选择。
A well-known example of heterozygote advantage is resistance to malaria in individuals who are heterozygous for the pathogenic allele that causes sickle cell disease (Case 42).杂合子优势的一个著名例子是,对导致镰状细胞病(病例42)的致病等位基因杂合的个体具有疟疾抗性。
This pathogenic variant in the β-globin gene HBB has reached its highest frequency in certain regions of Africa and Southeast Asia, where malaria is endemic and heterozygotes have greater relative fitness than either type of homozygote, due to their resistance to malarial infection.β-珠蛋白基因HBB中的这一致病变异已在非洲和东南亚某些地区达到最高频率,这些地区疟疾流行,且杂合子因对疟疾感染具有抗性,其相对适应度高于任何一种纯合子。
In the presence of (mosquito) vectors that carry malaria-inducing parasites, homozygotes without the trait allele are highly susceptible; they may become infected and are severely (even fatally) affected.在携带致疟原虫的(蚊子)媒介存在下,不携带该性状等位基因的纯合子高度易感;他们可能被感染并受到严重(甚至致命)影响。
Homozygotes for the pathogenic allele are even more disadvantaged, with a relative fitness that approaches zero due to severely debilitating hematological disease (see Chapter 12).致病等位基因的纯合子更为不利,其相对适应度因严重致残的血液疾病(见第12章)而趋近于零。
Heterozygotes, on the other hand, have red blood cells that are inhospitable to the malarial parasite but do not typically undergo the characteristic sickling that leads to pain crises in active sickle cell disease.另一方面,杂合子的红细胞对疟原虫不适宜,但通常不会发生导致活动性镰状细胞病疼痛危象的特征性镰变。
As such, these heterozygotes have a much greater relative fitness than homozygotes for the typical β-globin allele.因此,这些杂合子的相对适应度远高于典型β-珠蛋白等位基因的纯合子。
Over time, the pathogenic allele for sickle cell disease has reached a frequency as high as 0. 15 in some areas of the world that are endemic for malaria, far higher than could be accounted for by recurrent mutation alone.随着时间的推移,镰状细胞病的致病等位基因在世界某些疟疾流行地区的频率已高达0.15,远高于仅靠反复突变所能解释的水平。
Heterozygote advantage in sickle cell disease offers a clear example of how assumptions of HWE are violated when the mathematical relationship between allele and genotype frequencies diverges from expected values, (in this case) due to the effects of balancing selection.镰状细胞病中的杂合子优势提供了一个清晰的例子,说明当等位基因频率与基因型频率之间的数学关系偏离期望值时,HWE的假设是如何被违反的,(在此情况下)是由于平衡选择的影响。
To firmly ground this example, let us consider the sickle cell allele in the β-globin gene HBB, rs 334 (c. 20 A&gt;T [p.为了牢固地确立这个例子,让我们考虑β-珠蛋白基因HBB中的镰状细胞等位基因rs334(c.20 A>T [p.
Glu 7Val]), in which the pathogenic allele β S is under balancing selection due to heterozygote advantage.Glu7Val]),其中致病等位基因βS因杂合子优势而处于平衡选择之下。
Let us define the benign (nonpathogenic) allele β+ such that the two alleles β S and β+ give rise to three genotypes: β+|β+ (unaffected; homozygous), β+|β S (unaffected carriers; heterozygous), and β S|β S (affected).让我们定义良性(非致病)等位基因β+,使得两个等位基因βS和β+产生三种基因型:β+|β+(未受影响;纯合子)、β+|βS(未受影响的携带者;杂合子)和βS|βS(受影响)。
In a study of whole-genome sequence data from 2932 individuals aggregated across the 1000 Genomes Project, the African Genome Variation Project, and Qatar, balancing selection on the sickle cell variant β S is estimated to have conferred strong heterozygote advantage, having reached its equilibrium frequency of 12% after 87 generations (while the initial mutation is dated back 259 generations, or ~7300 years ago).在一项对来自1000基因组计划、非洲基因组变异项目和卡塔尔的2932个体全基因组序列数据的汇总研究中,镰状细胞变异βS上的平衡选择估计赋予了强大的杂合子优势,在87代后达到了12%的平衡频率(而初始突变可追溯至259代前,约7300年前)。
Using the β S allele’s reported equilibrium frequency of 0. 12 (q), we can calculate the expected ratio of genotypes under HWE (p 2:2pq:q 2) and compare this to observed genotype frequencies in the 1KG dataset for populations that have the β S allele at or near its equilibrium frequency.利用βS等位基因报告平衡频率0.12(q),我们可以计算HWE下预期的基因型比例(p²:2pq:q²),并将其与1KG数据集中βS等位基因处于或接近平衡频率的群体的观察基因型频率进行比较。
Given that q = 0. 12 and p = 1 – q, the frequency of p = 1 – 0. 12 = 0. 88.已知q=0.12且p=1-q,则p的频率为1-0.12=0.88。
From this, we can use HWE to calculate expected genotype frequencies as follows: Pr[affected homozygotes (β S|β S)] = q 2 = (0. 12) * (0. 12) = 0. 014; Pr[unaffected homozygotes (β+|β+)] = p 2 = (0. 88) * (0. 88) = 0. 774; and Pr[heterozygotes (β S|β+)] = 2pq = 2*(0. 88) * (0. 12) = 0. 211.由此,我们可以利用HWE计算预期基因型频率如下:患病纯合子(βS|βS)的概率 = q² = (0.12)*(0.12)=0.014;未患病纯合子(β+|β+)的概率 = p² = (0.88)*(0.88)=0.774;杂合子(βS|β+)的概率 = 2pq = 2*(0.88)*(0.12)=0.211。
Thus the expected genotype proportions of p 2:2pq:q 2 are 0. 774:0. 211:0. 014.因此,p²:2pq:q²的预期基因型比例为0.774:0.211:0.014。
Investigating rs 334 allele frequencies in the 1KG dataset through Ensembl, two populations sampled from Africa appear to have the β S allele at or near its equilibrium frequency (0. 12): the Esan in Nigeria (ESN) with β S at 12. 1% and the Mende in Sierra Leone (MSL) with β S at 12. 4% in the population.通过Ensembl调查1KG数据集中rs334等位基因频率,来自非洲的两个样本群体似乎拥有处于或接近平衡频率(0.12)的βS等位基因:尼日利亚的埃桑人(ESN)的βS频率为12.1%,塞拉利昂的门德人(MSL)的βS频率为12.4%。
Observed genotype frequencies for these populations are as follows: 0. 758:0. 242:0 (for ESN) and 0. 753:0. 247:0 (for MSL).这些群体的观察基因型频率如下:ESN为0.758:0.242:0,MSL为0.753:0.247:0。
In both these populations, the observed proportions of heterozygous (β S|β+) individuals exceed what was predicted assuming HWE, whereas the observed number of unaffected homozygotes (β+|β+) and affected homozygotes (β S|β S) are below what was predicted.在这两个群体中,观察到的杂合子(βS|β+)个体的比例超过了假定HWE下的预测值,而观察到的未患病纯合子(β+|β+)和患病纯合子(βS|βS)的数量则低于预测值。
This trend reflects balancing selection at this locus, illustrating how forces of selection, operating both negatively on the relatively rare β S|β S genotype and positively on the more common β S|β+genotype, cause deviation from HWE.这一趋势反映了该位点的平衡选择,说明了选择力量如何通过对相对稀少的βS|βS基因型发挥负向作用,以及对更常见的βS|β+基因型发挥正向作用,导致偏离HWE。
The effects of balancing selection on malaria resistance are also apparent in other infectious diseases.平衡选择对疟疾抗性的影响在其他感染性疾病中也显而易见。
For example, many people with the severe renal disease known as focal segmental glomerulosclerosis are homozygotes for certain alleles in the coding region of the APOL1 gene that encodes the apolipoprotein L1.例如,许多患有称为局灶节段性肾小球硬化的严重肾脏疾病的人,是编码载脂蛋白L1的APOL1基因编码区中某些等位基因的纯合子。
Apolipoprotein L1 is a serum factor that kills the trypanosome parasite载脂蛋白L1是一种杀死锥虫寄生虫的血清因子。
23/53
Population Genetics for Genomic Medicine 197 Trypanosoma brucei, which causes trypanosomiasis (sleeping sickness).
Ch10 — Segment 23
Population Genetics for Genomic Medicine 197 Trypanosoma brucei, which causes trypanosomiasis (sleeping sickness).基因组医学的群体遗传学197 布氏锥虫,它引起锥虫病(昏睡病)。
The same variants that increase one’s risk of severe kidney disease in homozygotes tenfold over the rest of the population protect heterozygotes carrying these variants against trypanosomes (e. g., T. brucei rhodesiense) that have developed resistance to wild-type apolipoprotein L1.那些使纯合子患严重肾脏疾病的风险比其余人群高十倍的相同变异,保护携带这些变异的杂合子免受对野生型载脂蛋白L1产生抗性的锥虫(例如,罗得西亚布氏锥虫)的侵害。
As a result, the frequency of heterozygous carriers for these alleles can be as high as ~45% in parts of the world in which the rhodesiense trypanosomiasis is endemic.因此,在罗得西亚锥虫病流行的世界部分地区,这些等位基因的杂合携带者频率可高达约45%。
GENOMIC VARIATION AND BIASES IN POPULATION DATASETS The type and amount of information available to researchers and clinicians for our work is the foundation for discoveries, diagnostics, treatment regimens, and approaches to measuring outcomes.群体数据集中的基因组变异与偏倚 研究人员和临床医生工作中可用的信息类型和数量是发现、诊断、治疗方案和结果测量方法的基础。
Missing data is inevitable – but uncertainty arises in the interpretation and portability of findings when the amount and type of data available is missing or of differential quality in a nonrandom ways, that is, if certain groups or patient populations are better represented in databases than others, for example.数据缺失是不可避免的——但当可用数据的数量和类型缺失或以非随机方式存在质量差异时,即在某些群体或患者人群在数据库中比其他群体有更好代表性等情况下,就会在结果的解释和可移植性中产生不确定性。
Ascertainment bias refers to systemic biases in observations that steer our research or interpretations in a direction based on what is observed – without having any knowledge of that which remains unobserved (and could potentially change the result or interpretation if revealed).确定偏倚是指观察中的系统性偏倚,这些偏倚基于所观察到的事物将我们的研究或解释导向某个方向——而对那些未被观察到的事物(如果揭示出来可能改变结果或解释)一无所知。
For example, &gt;80% of genomics research to date has been conducted on people of mostly European ancestries, so population reference data and genomic databases are heavily biased toward a European genomic background. .例如,迄今为止超过80%的基因组研究是在主要具有欧洲血统的人群中进行的,因此人群参考数据和基因组数据库严重偏向欧洲基因组背景。
Genetic or locus heterogeneity in a trait means that variation in different regions of the genome can be pathogenic for the same trait.一个性状的遗传异质性或位点异质性意味着基因组不同区域的变异可能对同一性状具有致病性。
The ascertainment bias toward prevalence in European ancestries contributes to higher rates of variants of unknown or uncertain significance (VUS) in non-Europeans, as well as higher rates of false negative diagnoses, due to missing genetic heterogeneity.由于缺少遗传异质性,欧洲血统中流行的确定偏倚导致了非欧洲人中未知或不确定意义变异(VUS)的较高发生率,以及较高的假阴性诊断率。
Higher false positive rates have also been documented, as in the case of a variant associated with hypertrophic cardiomyopathy (and curated as pathogenic based on information from European ancestry individuals), which was later shown to be benign in African American patients – after many had already undergone an invasive prophylactic intervention.较高的假阳性率也有记录,例如一个与肥厚型心肌病相关的变异(基于欧洲血统个体的信息被归类为致病性),后来在非裔美国患者中被证明是良性的——而许多患者已经接受了侵入性预防性干预。
Exclusion of individuals with recent African ancestries is a mistake for any study of the genetic underpinnings of health and disease, due to the wealth of genetic variation that exists in African populations. but also how biased representations of global genomic diversity can be when restricted to continental ancestry groupings that seek to “balance” representation across the canonical discrete categories, as in排除近期有非洲血统的个体对于任何关于健康和疾病遗传基础的研究都是一个错误,因为非洲人群中存在丰富的遗传变异,但也表明了当局限于寻求在经典离散类别间“平衡”代表性的大陆血统分组时,全球基因组多样性的代表性会有多么偏倚,如。
24/53
Population names are sample providers’ collected the samples. self-identifications, or the descriptions used by those wh…
Ch10 — Segment 24
Population names are sample providers’ collected the samples. self-identifications, or the descriptions used by those who Six geographical labels correspond to the recent origins of the populations that fall mainly into the seven ancestry clusters produced by an algorithm.† A Individual genomes Each horizontal bar corresponds to the genome of a single person, in this case from a population identified or self-identified as African American.群体名称来源于样本提供者的自我认同,或样本收集者所使用的描述。六个地理标签对应于主要落入算法生成的七个祖先聚类中的群体的近期起源。个体基因组:每条水平条代表一个人的基因组,此例中来自被识别或自我认同为非裔美国人的群体。
This person has 42% Eurasian ancestry, 5% African Great Lakes ancestry and 53% West African ancestry.该个体具有42%的欧亚祖先血统、5%的非洲大湖区祖先血统和53%的西非祖先血统。
Individuals Populations from Africa sampled: 85% (unpublished) Populations from Africa Adjusting the sampling and using an African-centric data set creates a more representative view. sampled: 13. 5%.来自非洲的样本群体:85%(未发表)。调整采样并使用非洲中心数据集可获得更具代表性的视图。来自非洲的样本群体:13.5%。
East and South Asia, Americas (Indigenous) North Africa, Europe, Americas South Asia**, Horn of Africa African Great Lakes Africa Western Nama Mende Zulu Africa Middle East Europe Central and South Asia East Asia America Oceania Esan Yoruba African Mandinka African American Caribbean Banyarwanda Bakiga Barundi Banyankole Baganda Rwandese Gumuz Somali Luhya Oromo Wolayta Amhara Egyptian Gujarati Mexican British Italian Han Chinese Southern Africa ‡ B A C B An analysis from 2008* suggested significant genetic differences between seven continental populations (B) But only 13. 5% of the populations represented were from Africa.东亚与南亚、美洲(土著)、北非、欧洲、美洲、南亚**、非洲之角、非洲大湖区、西非、纳马族、门德族、祖鲁族、非洲、中东、欧洲、中亚与南亚、东亚、美洲、大洋洲、埃桑族、约鲁巴族、非洲曼丁哥族、非裔美国人、加勒比地区、班亚旺达族、巴基加族、巴伦迪族、巴尼安科莱族、巴甘达族、卢旺达族、古穆兹族、索马里族、卢希亚族、奥罗莫族、沃莱塔族、阿姆哈拉族、埃及人、古吉拉特族、墨西哥人、英国人、意大利人、汉族、南部非洲‡ B A C B 2008年的一项分析*表明七个大陆群体之间存在显著遗传差异(B),但其中仅有13.5%的群体来自非洲。
Boosting representation to 85% and sampling more broadly across the continent (C) underlines that the level of genetic variation within Africa is equivalent to that seen between continents.将代表性提升至85%并在非洲大陆更广泛地采样(C)强调,非洲内部的遗传变异水平相当于洲际间的遗传变异水平。
Figure adapted from Carlson J, Henn BM, Al-Hindi DR, Ramachandran S.图改编自Carlson J, Henn BM, Al-Hindi DR, Ramachandran S。
Counter: the weaponization of genetics research by extremists, Nature 610: 2022. **South Asia appears twice because Gujarati people in India have intermediate allele frequencies.评论:极端主义者对遗传学研究的武器化,《自然》610:2022。**南亚出现两次,因为印度的古吉拉特人具有中等等位基因频率。
25/53
Population Genetics for Genomic Medicine 199 the population level due to structural racism that disproportionately impac…
Ch10 — Segment 25
Population Genetics for Genomic Medicine 199 the population level due to structural racism that disproportionately impacts Black patients of all ancestries (e. g., the well-documented undertreatment of pain in African Americans).群体遗传学在基因组医学中的应用 199 由于结构性种族主义对各个祖先的非裔美国人患者(例如,非裔美国人疼痛治疗不足的充分记录)产生了不成比例的影响,因此在群体层面上存在这一问题。
We will ground this discussion in the practical context of clinical variant interpretation.我们将以临床变异解读的实际背景为基础来讨论这一问题。
Guidelines for clinical variant interpretation protocols published by Association for Molecular Pathology (AMP) and the American College of Medical Genetics and Genomics (ACMG) include population-level information (e. g., allele frequency in population databases), among other data (e. g., functional evidence) to guide decisions about whether to designate a variant as benign, likely benign, likely pathogenic, pathogenic, or uncertain significance.美国分子病理学会和美国医学遗传学与基因组学学院发布的临床变异解读方案指南中包含了群体水平信息(例如,群体数据库中的等位基因频率),以及其他数据(例如,功能证据),以指导关于将变异指定为良性、可能良性、可能致病、致病或意义未明的决策。
For variants detected in European populations, the genomic knowledgebase is robust and reliable enough such that the absence of a variant from population databases can be considered evidence for pathogenicity.对于在欧洲人群中检测到的变异,基因组知识库足够稳健和可靠,以至于变异在群体数据库中的缺失可被视为致病性的证据。
This is because these populations are well represented among reference datasets, so the absence of a variant can be reasonably interpreted as evidence that it is not tolerated in the population.这是因为这些人群在参考数据集中有充分的代表性,因此变异的缺失可以合理地解释为它在群体中不被耐受的证据。
In contrast, most global populations have not been included in genomic datasets at a comparable rate or magnitude, so this interpretation may be subject to limited certainty.相比之下,大多数全球人群并未以可比的比例或规模被纳入基因组数据集,因此这种解读可能具有有限的可信度。
Instead, it may be that absence or very low frequency of a variant in population databases reflects insufficient representation, and more rigorous sampling would reveal higher population-allele frequencies than would be consistent with predicted pathogenicity of the variant.相反,变异在群体数据库中缺失或频率极低可能反映了代表性不足,更严格的采样可能会揭示出比与预测的变异致病性一致的种群等位基因频率更高的频率。
Several efforts are underway to improve diversity and inclusion in genomic databases; 3 details two of these examples.目前有几项努力正在改善基因组数据库的多样性和包容性;第3项详细介绍了其中两个例子。
As biomedical researchers and clinicians, we must acknowledge that the ways in which we collect, 3 “A TALE OF TWO INITIATIVES” (BY LAURA ARBOUR) The Genome Aggregation Database (gnom AD) is an international collaborative effort of researchers utilizing available data sources of genomic variation compiled through the Broad Institute ( It is used frequently as a clinical tool in the diagnosis of rare, severe, genetic disease.作为生物医学研究人员和临床医生,我们必须承认我们收集数据的方式,3“两个倡议的故事”(作者:劳拉·阿布尔)基因组聚集数据库(gnomAD)是一项国际合作的努力,研究人员利用通过博德研究所汇编的可用基因组变异数据源(),它经常被用作诊断罕见严重遗传疾病的临床工具。
The frequency of variants present in the dataset and reported geographical ancestry of origin are openly available to clinicians and researchers, which aids in the first steps of consideration for pathogenicity of variants (common variants are unlikely to cause severe, early onset disease).数据集中存在的变异频率以及报告的地理祖先来源对临床医生和研究人员公开可用,这有助于在考虑变异致病性的第一步中提供帮助(常见变异不太可能引起严重的早发性疾病)。
The combination of whole genomes and exomes of more than 140,000 unrelated individuals contributes to the genomic reference database.超过14万名无关个体的全基因组和外显子组组合构成了基因组参考数据库。
Although there are ongoing efforts to increase diversity in public genomic databases including gnom AD, the problem is that not all populations are represented in available datasets for a multitude of reasons, therefore not all children or families with rare genetic conditions will have the same opportunity for a precise diagnosis in a timely manner.尽管包括gnomAD在内的公共基因组数据库在持续增加多样性,但问题在于,由于多种原因,并非所有人群都在现有数据集中得到代表,因此并非所有患有罕见遗传病的儿童或家庭都能同样有机会及时获得精准诊断。
The lack of genomic reference data increases the “genomic divide,” where those with the greatest health disparities benefit least from genomic advances.基因组参考数据的缺乏加剧了“基因组鸿沟”,即那些健康差距最大的人群从基因组进展中获益最少。
The impact of lack of Indigenous genomic data is staggering when it is considered that there are more than 370 million Indigenous people spanning 90 countries worldwide ( documents/5session_factsheet 1. pdf.).当考虑到全球90个国家有超过3.7亿土著人时,缺乏土著基因组数据的影响是惊人的( documents/5session_factsheet 1. pdf.)。
Two initiatives aim to address this issue for Indigenous patients with genetic conditions.两项举措旨在解决患有遗传病的土著患者的这一问题。
In parallel, the Silent Genomes Project (Canada) and the Aotearoa Variome (New Zealand) are developing genomic data reference databases that are prioritized for genomic health care ( articles/10. 3389/fpubh. 2020. 00111/full) and may also be used for health research.与此同时,沉默基因组项目(加拿大)和奥特亚罗瓦变异组(新西兰)正在开发优先用于基因组医疗保健的基因组数据参考数据库( articles/10. 3389/fpubh. 2020. 00111/full),也可用于健康研究。
These initiates are led or coled by Indigenous scholars, and variant use and release mechanisms are being developed and are informed by long-standing ethical frameworks from within their respective countries (“DNA on Loan” and the Te Mata Ira guidelines for medical genomics with Māori) and are consistent with the recent International Indigenous Data Sovereignty Interest Group “CARE” principles, CARE being the acronym for “Collective Benefit, Authority to Control, Responsibility and Ethics” ( codata. org/articles/10. 5334/dsj-2020-043/).这些举措由土著学者领导或共同领导,变异的使用和发布机制正在制定中,并受到各自国家长期存在的伦理框架(“DNA借用”和针对毛利人的医学基因组学Te Mata Ira指南)的指导,并与最近的国际土著数据主权兴趣小组“CARE”原则一致,CARE是“集体利益、控制权、责任和伦理”的缩写( codata. org/articles/10. 5334/dsj-2020-043/)。
The CARE principles are Indigenous focused but are meant to complement the “FAIR principles (Findable, Accessible, Interoperable, Reusable) which are Guiding Principles for scientific data management and stewardship” (https:// www. nature. com/articles/sdata 201618).CARE原则以土著为重点,但旨在补充“FAIR原则(可发现、可访问、可互操作、可重用),这些是科学数据管理和管理的指导原则”(https:// www. nature. com/articles/sdata 201618)。
The CARE principles support the notion of benefit for, and selfdetermination of, Indigenous people, consistent with the United Nations Declaration of the Rights of Indigenous Peoples (adopted by the UN General Assembly in 2007 and endorsed by law in Canada [Bill C-15] in 2021).CARE原则支持土著人民的福祉和自决理念,与《联合国土著人民权利宣言》(2007年联合国大会通过,2021年加拿大通过法律[Bill C-15]认可)相一致。
Both the Silent Genomes Project and Aotearoa Variome will start with sequencing the samples of consented individuals within their countries.沉默基因组项目和奥特亚罗瓦变异组都将从对其国家内同意者的样本进行测序开始。
Storage, use, and release of variants for clinical and possibly research purposes is being informed by local Indigenous perspectives and governance mechanisms.用于临床以及可能的研究目的的变异的存储、使用和发布正在受到当地土著观点和治理机制的指导。
The Silent Genomes Project will also assess the efficacy of the Indigenous Background Variant Library (IBVL) in a cohort of Indigenous children who have gone through diagnosis without the IBVL.沉默基因组项目还将评估土著背景变异库(IBVL)在一组未使用IBVL进行诊断的土著儿童中的有效性。
The primary goal of both initiatives is to reduce, and not increase, health disparities with genomic advances.这两项举措的主要目标是通过基因组进展减少而非增加健康差距。
References (also integrated above): documents/5session_factsheet 1. pdf aotearoa-nz-genomic-variome Indigenous genomic databases: Pragmatic considerations and cultural contexts NR Caron, M Chongo, M Hudson, L Arbour, WW Wasserman, S Robertson, S Correard, P Wilcox: Front Public Health 8:111, 2020. doi:10. 3389/fpubh. 2020. 00111; 201618 dsj-2020-043/参考文献(也整合于上文): documents/5session_factsheet 1. pdf aotearoa-nz-genomic-variome Indigenous genomic databases: Pragmatic considerations and cultural contexts NR Caron, M Chongo, M Hudson, L Arbour, WW Wasserman, S Robertson, S Correard, P Wilcox: Front Public Health 8:111, 2020. doi:10. 3389/fpubh. 2020. 00111; 201618 dsj-2020-043/
26/53
analyze, visualize, and report or publish our data on human population genetics are crucially important.
Ch10 — Segment 26
analyze, visualize, and report or publish our data on human population genetics are crucially important.分析、可视化并报告或发布我们关于人类群体遗传学数据至关重要。
Our analytic approaches are heavily influenced by the historical, cultural, social, and political contexts in which we conduct our research and clinical practices.我们的分析方法深受我们开展研究和临床实践的历史、文化、社会及政治背景影响。
How we think about and represent categories of difference and similarity within and among populations matters – both for science and medicine, but also for society.我们如何思考并表征群体内部及群体间的差异与相似性类别,对科学和医学以及社会都至关重要。
Members of the public with harmful political agendas have weaponized figures illustrating admixture mapping that are published in peer-reviewed journals (e. g., , claiming such clustering methods support their ideologies that are steeped in biological racism.持有有害政治议程的公众人士已将同行评审期刊中发表的描绘混血图谱的图表武器化(例如,声称此类聚类方法支持他们植根于生物种族主义的意识形态)。
We must counter these efforts and work to prevent further misconceptions from spreading, through responsible and trustworthy research and reporting.我们必须通过负责任且可信的研究与报告,抵制这些行为并努力防止更多误解的传播。
Until we have a more robust and complete picture of global genomic variation and how it contributes to complex disease etiology, it is imperative that we exercise caution when implementing existing (and new) methodologies to analyze genome-wide data.在我们对全球基因组变异及其对复杂疾病病因学的贡献有更全面、更完整的认识之前,在采用现有(及新)方法分析全基因组数据时必须谨慎行事。
For example, polygenic risk scores, or polygenic scores (PRS), have the potential to exacerbate both conceptual and practical issues related to equity in human population genetics and medicine.例如,多基因风险评分或多基因评分(PRS)有可能加剧与人类群体遗传学和医学公平性相关的概念及实践问题。
Figs. 10. 10 and 10. 11 provide a highlevel overview of how PRS are constructed, and the LD LD LD tag SNP2 causal SNP2 tag SNP3 causal SNP3 predicted phenotype Total number of SNPs m Y j=1 gj j = ∑ tag SNP genotype SNP weight (adjusted effect size) Model Adjustments tag SNP1 causal SNP1 effect sizes '3 GWAS SIGNAL IN DISCOVERY POPULATION SNPs and effect sizes identified by GWAS in discovery population to predict phenotype or trait of interest in a target population.图10.10和图10.11提供了PRS构建方式的高级概述,以及LD LD LD标签SNP2因果SNP2标签SNP3因果SNP3预测表型 总SNP数量 m Y j=1 gj j = ∑ 标签SNP基因型 SNP权重(调整后的效应量)模型调整 标签SNP1因果SNP1效应量 '3 发现群体的GWAS信号 通过GWAS在发现群体中鉴定的SNP和效应量,用于预测目标群体的表型或感兴趣性状。
PHENOTYPE PREDICTION IN TARGET POPULATION β β '2 β '1 β Variants identified by GWAS in the discovery population are not necessarily causing or contributing to variation of the trait in the discovery population, but these “tag SNPs” are linked to causal SNP(s) such that they signal a candidate region to be further investigated.目标群体的表型预测 β β '2 β '1 β 通过GWAS在发现群体中鉴定的变异未必是导致或促成该群体表型变异的原因,但这些“标签SNP”与因果SNP相关联,从而指示一个需要进一步研究的候选区域。
ALLELE FREQUENCIES IN DISCOVERY POPULATION ALLELE FREQUENCIES IN TARGET POPULATION Tag SNPs from discovery GWAS signal candidate loci associated with a trait; causal SNPs unknown LD structure, gene x environment, and gene x gene interactions may limit prediction accuracy LD LD LD LD Gx E Gx G LD LD tag SNP2 causal SNP2 tag SNP2 causal SNP2 tag SNP3 tag SNP3 causal SNP3 causal SNP3 causal SNP4 tag SNP1 tag SNP1 = 0. 002 causal SNP1 tag SNP1 causal SNP1 No LD causal SNP1 = 0. 002 tag SNP2 = 0. 0003 causal SNP2 = 0. 0003 tag SNP1 = 0. 01 causal SNP1 = 0. 01 tag SNP1 = 0. 005 causal SNP1 = 0. 002 tag SNP2 = 0. 0008 causal SNP2 = 0. 0008 tag SNP3 = 0. 01 causal SNP3 = 0. 01 When LD structure differs between discovery and target populations, associations between tag SNPs and causal SNPs in the discovery GWAS may not be replicated in the target population, limiting the predictive power of this model.发现群体的等位基因频率 目标群体的等位基因频率 来自发现GWAS的标签SNP指示与性状相关的候选位点;因果SNP未知 LD结构、基因×环境以及基因×基因相互作用可能限制预测精度 LD LD LD LD GxE GxG LD LD 标签SNP2因果SNP2标签SNP2因果SNP2标签SNP3标签SNP3因果SNP3因果SNP3因果SNP4标签SNP1标签SNP1 = 0.002 因果SNP1标签SNP1因果SNP1无LD 因果SNP1 = 0.002 标签SNP2 = 0.0003 因果SNP2 = 0.0003 标签SNP1 = 0.01 因果SNP1 = 0.01 标签SNP1 = 0.005 因果SNP1 = 0.002 标签SNP2 = 0.0008 因果SNP2 = 0.0008 标签SNP3 = 0.01 因果SNP3 = 0.01 当发现群体与目标群体的LD结构不同时,发现GWAS中标签SNP与因果SNP之间的关联可能无法在目标群体中重现,从而限制该模型的预测能力。
Differences in environmental factors and gene-by-environment interactions (Gx E) between discovery and target populations may impact the accuracy of prediction in the target population, particularly for multifactorial traits.发现群体与目标群体之间环境因素及基因-环境相互作用(GxE)的差异可能影响目标群体预测的准确性,尤其对于多因素性状。
The presence of other variants in the causal pathway with gene-by-gene interactions (Gx G, or epistatic effects) in one population, but not the other, may also limit the portability of PRS between populations.在一个群体中存在而另一个群体中不存在的因果通路中包含基因-基因相互作用(GxG,或上位效应)的其他变异,也可能限制PRS在不同群体间的可移植性。
27/53
Population Genetics for Genomic Medicine 201 conceptual as well as technical pitfalls of trying to predict phenotypic va…
Ch10 — Segment 27
Population Genetics for Genomic Medicine 201 conceptual as well as technical pitfalls of trying to predict phenotypic variation in one (target) population using GWAS results from another (discovery) population.基因组医学的人群遗传学 201 试图利用来自另一个(发现)群体的GWAS结果预测某个(目标)群体的表型变异时存在的概念性及技术性陷阱。
In this simplified model, PRS construction involves several assumptions about homogeneity in genetic contributions to disease, allele frequencies, LD structure, Gx G, and Gx E between a discovery population and the target (prediction) population.在此简化模型中,PRS构建涉及关于发现群体与目标(预测)群体之间疾病遗传贡献、等位基因频率、LD结构、GxG和GxE同质性的若干假设。
Allele frequencies and LD structure may differ substantively between populations (depending on how they are defined and their underlying characteristics).等位基因频率和LD结构在不同群体之间可能存在显著差异(取决于群体如何定义及其潜在特征)。
Similarly, differential Gx E and Gx G between populations can dramatically affect the results of genomic investigations and obscure the role of genetic variants in disease etiology.同样,群体间差异性的GxE和GxG可显著影响基因组研究结果,并掩盖遗传变异在疾病病因学中的作用。
Clinical professionals must be aware that PRS still have a long way to go before one could argue that they offer enhanced utility beyond the current standards of care and could make matters worse in the meantime.临床专业人员必须认识到,在PRS能被论证其效用超越当前标准护理并可能在此期间使情况恶化之前,PRS仍有很长的路要走。
What is true of PRS is the same for all approaches we use in genomic research and medicine.PRS的情况适用于我们在基因组研究与医学中使用的所有方法。
We must critically examine our underlying assumptions about what factors contribute most to health and disease, consider the impact of those assumptions, investigate the history and biases of approaches we seek to use, and question the foundations of what we think we know – to make room for more curiosity and innovation that will lead to novel discoveries and more precise genomic medicine.我们必须批判性地审视我们关于哪些因素最有助于健康与疾病的潜在假设,考虑这些假设的影响,调查我们试图使用的方法的历史与偏见,并质疑我们认为已知的基础——从而为更多好奇与创新腾出空间,这将带来新的发现和更精准的基因组医学。
ACKNOWLEDGMENT Some language and sections included in this chapter were inherited from previous (published) versions of the textbook.致谢 本章中包含的部分语言和章节源自本教科书的先前(已出版)版本。
Laura Arbour contributed “A Tale of Two Initiatives”.Laura Arbour 撰写了“两个倡议的故事”一节。
Conversations with clinical genetics professionals and other interdisciplinary collaborations through the NIH-funded Clinical Genome Resource (Clin Gen) Ancestry &amp; Diversity Working Group helped motivate the development of new content for this revision.通过美国国立卫生研究院资助的临床基因组资源(Clin Gen)祖先与多样性工作组的临床遗传学专业人员对话及其他跨学科合作,推动了本修订版新内容的开发。
We thank Sonja Rasmussen for contributing to this chapter.我们感谢Sonja Rasmussen对本章的贡献。
GENERAL REFERENCES Li CC: First course in population genetics, Pacific Grove, 1975, Boxwood Press.一般参考文献 Li CC: 群体遗传学入门, 太平洋格罗夫, 1975, Boxwood出版社.
Nielsen R, Slatkin M: An introduction to population genetics, Sunderland, 2013, Sinauer Associates, Inc.Nielsen R, Slatkin M: 群体遗传学导论, 桑德兰, 2013, Sinauer联合公司.
Dorothy R: Fatal invention: How Science, Politics, and Big Business Re-Create Race in the Twenty-First Century, New York, 2011, New Press.Dorothy R: 致命的发明:科学、政治与商业如何在二十一世纪重塑种族, 纽约, 2011, 新出版社.
Popejoy AB, Crooks KR, Fullerton SM, et al: Clinical Genome Resource (Clin Gen) Ancestry and Diversity Working Group: Clinical genetics lacks standard definitions and protocols for the collection and use of diversity measures, Am J Hum Genet 107(1):72–82, 2020.Popejoy AB, Crooks KR, Fullerton SM, 等: 临床基因组资源(Clin Gen)祖先与多样性工作组:临床遗传学缺乏多样性指标收集与使用的标准定义与方案, Am J Hum Genet 107(1):72–82, 2020.
Royal CD, Novembre J, Fullerton SM, et al: Inferring genetic ancestry: opportunities, challenges and implications, Am J Hum Genet 86:661–673, 2010.Royal CD, Novembre J, Fullerton SM, 等: 推断遗传祖先:机遇、挑战与意义, Am J Hum Genet 86:661–673, 2010.
REFERENCES FOR SPECIFIC TOPICS American Society of Human Genetics: ASHG denounces attempts to link genetics and racial supremacy, Am J Hum Genet 103:636, 2018.特定主题参考文献 美国人类遗传学学会:ASHG谴责将遗传学与种族优越论联系起来的企图, Am J Hum Genet 103:636, 2018.
Behar DM, Yunusbayev B, Metspalu M, et al: The genome-wide structure of the Jewish people, Nature 466:238–242, 2010.Behar DM, Yunusbayev B, Metspalu M, 等: 犹太人群的基因组结构, Nature 466:238–242, 2010.
Borrell LN, Elhawary JR, Fuentes-Afflick E, et al: Race and genetic ancestry in medicine – A time for reckoning with racism, N Engl J Med 384:474–480, 2021. 2029562 Corona E, Chen R, Sikora M, et al: Analysis of the genetic basis of disease in the context of worldwide human relationships and migration, PLo S Genet 9:e 1003447, 2013.Borrell LN, Elhawary JR, Fuentes-Afflick E, 等: 医学中的种族与遗传祖先——正视种族主义的时刻, N Engl J Med 384:474–480, 2021. 2029562 Corona E, Chen R, Sikora M, 等: 在全球人类关系与迁徙背景下分析疾病的遗传基础, PLoS Genet 9:e1003447, 2013.
Gregg AR, Aarabi M, Klugman S, et al: ACMG Professional Practice and Guidelines Committee: Screening for autosomal recessive and X-linked conditions during pregnancy and preconception: a practice resource of the American College of Medical Genetics and Genomics (ACMG, Genet Med 23(10):1793–1806, 2021. https:// doi. org/10. 1038/s 41436-021-01203-z Henn BM, Cavalli-Sforza LL, Feldman MW: The great human expansion.Gregg AR, Aarabi M, Klugman S, 等: ACMG专业实践与指南委员会:孕期及孕前常染色体隐性和X连锁疾病筛查:美国医学遗传学与基因组学学会实践资源, Genet Med 23(10):1793–1806, 2021. https://doi.org/10.1038/s41436-021-01203-z Henn BM, Cavalli-Sforza LL, Feldman MW: 伟大的人类扩张.
Proceedings of the National Academy of Sciences of the United States of America 109:17758–64.美国国家科学院院刊 109:17758–64.
PMID 23077256. org/10. 1007/S12045-019-0830-4 Howes R, Patil A, Piel F, et al: The global distribution of the Duffy blood group.PMID 23077256. doi:10.1007/S12045-019-0830-4 Howes R, Patil A, Piel F, 等: Duffy血型的全球分布.
Nat Commun 2:266, 2011. ncomms 1265 Kaseniit KE, Haque IS, Goldberg JD, Shulman LP, Muzzey D: Genetic ancestry analysis on &gt;93,000 individuals undergoing expanded carrier screening reveals limitations of ethnicity-based medical guidelines, Genet Med 22(10):1694–1702, 2020. s 41436-020-0869-3 Kumar R, Seibold MA, Aldrich MC, et al: Genetic ancestry in lungfunction predictions, N Engl J Med 363:321–330, 2010.Nat Commun 2:266, 2011. ncomms1265 Kaseniit KE, Haque IS, Goldberg JD, Shulman LP, Muzzey D: 对超过93,000名接受扩展携带者筛查个体的遗传祖先分析揭示了基于种族的医学指南的局限性, Genet Med 22(10):1694–1702, 2020. s41436-020-0869-3 Kumar R, Seibold MA, Aldrich MC, 等: 肺功能预测中的遗传祖先, N Engl J Med 363:321–330, 2010.
Lewontin RC: The apportionment of human diversity.Lewontin RC: 人类多样性的分配.
In: Dobzhansky T, Hecht MK, Steere WC, editors: Evolutionary biology: volume, 6, New York, 1972, Springer.载: Dobzhansky T, Hecht MK, Steere WC, 编: 进化生物学: 第6卷, 纽约, 1972, Springer.
Martin A, Kanai M, Kamatani Y, Okada Y, Neale BM, Daly MJ: Clinical use of current polygenic risk scores may exacerbate health disparities, Nat Genet 51:584–591, 2019. s 41588-019-0379-x Novembre J, Johnson T, Bryc K, et al: Genes mirror geography within Europe.Martin A, Kanai M, Kamatani Y, Okada Y, Neale BM, Daly MJ: 当前多基因风险评分的临床应用可能加剧健康不平等, Nat Genet 51:584–591, 2019. s41588-019-0379-x Novembre J, Johnson T, Bryc K, 等: 基因反映欧洲内部地理.
Nature 456:98–101, 2008. nature 07331 Novembre J, Peter BM: Recent advances in the study of fine-scale population structure in humans.Nature 456:98–101, 2008. nature07331 Novembre J, Peter BM: 人类精细尺度群体结构研究的最新进展.
Curr Opin Genet Dev 41:98–105, 2016.Curr Opin Genet Dev 41:98–105, 2016.
Peterson RE, Kuchenbaecker K, Walters RK, et al: Genome-wide association studies in ancestrally diverse populations: Opportunities, methods, pitfalls, and recommendations, Cell 179(3):589–603, 2019.Peterson RE, Kuchenbaecker K, Walters RK, 等: 祖先多样性人群的全基因组关联研究:机遇、方法、陷阱与建议, Cell 179(3):589–603, 2019.
Richards S, Aziz N, Bale S, et al: Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology, Genet Med 17:405–423, 2015.Richards S, Aziz N, Bale S, 等: 序列变异解读的标准与指南:美国医学遗传学与基因组学学会与分子病理学协会联合共识推荐, Genet Med 17:405–423, 2015.
Sankararaman S, Mallick S, Dannemann M, et al: The genomic landscape of Neanderthal ancestry in present-day humans, Nature 507:354–357, 2014.Sankararaman S, Mallick S, Dannemann M, 等: 现代人类中尼安德特人祖先的基因组景观, Nature 507:354–357, 2014.
Shriner D, Rotimi CN: Whole-genome-sequence-based haplotypes reveal single origin of the sickle allele during the Holocene Wet Phase, Am J Hum Genet 102(4):547–556, 2018. ajhg. 2018. 02. 003.Shriner D, Rotimi CN: 基于全基因组序列的单倍型揭示镰状等位基因在全新世湿润期的单一起源, Am J Hum Genet 102(4):547–556, 2018. ajhg.2018.02.003.
Epub 2018 Mar 8.电子版2018年3月8日.
PMID: 29526279; PMCID: PMC5985360.PMID: 29526279; PMCID: PMC5985360.
Wastnedge E, Waters D, Patel S, et al: The global burden of sickle cell disease in children under five years of age: a systematic review and meta-analysis, J Global Health 8(2):021103, 2018. ncbi. nlm. nih. gov/pmc/articles/PMC6286674/ Wexler (Need REF for Venezuela finding) Wojcik GL, Graff M, Nishimura KK, et al: Genetic analyses of diverse populations improves discovery for complex traits, Nature 570:514– 518, 2019. 41586-019-1310-4Wastnedge E, Waters D, Patel S, 等: 五岁以下儿童镰状细胞病的全球负担:系统综述与荟萃分析, J Global Health 8(2):021103, 2018. ncbi.nlm.nih.gov/pmc/articles/PMC6286674/ Wexler(需要委内瑞拉发现参考文献)Wojcik GL, Graff M, Nishimura KK, 等: 多样性人群的遗传分析改善复杂性状的发现, Nature 570:514–518, 2019. 41586-019-1310-4
28/53
PROBLEMS 1.
Ch10 — Segment 28
PROBLEMS 1.问题1。
A short tandem repeat (STR) variant consists of 5 different alleles, each with a frequency of 0. 20 in a population. a.一个短串联重复序列(STR)变异由5个不同的等位基因组成,每个等位基因在人群中的频率为0.20。a.
What proportion of individuals in this population would you expect to be homozygous at this locus? b.在这个人群中,你预期有多少比例的个体在该位点是纯合子?b.
What proportion of the population is likely to be heterozygous at this locus? c.在这个人群中,有多少比例可能是杂合子?c.
What proportion of individuals would be homozygous, and what proportion would be heterozygous if the 5 alleles had different frequencies of 0. 40, 0. 30, 0. 15, 0. 10, and 0. 05?如果5个等位基因的频率不同,分别为0.40、0.30、0.15、0.10和0.05,那么纯合子的比例是多少?杂合子的比例是多少?
In a population with allele frequencies in Hardy-Weinberg equilibrium, three genotypes are present in the following proportions: A/A, 0. 81; A/a, 0. 18; a/a, 0. 01. a.在一个等位基因频率符合哈迪-温伯格平衡的人群中,三种基因型的比例分别为:A/A 0.81;A/a 0.18;a/a 0.01。a.
What are the allele frequencies of A and a? b.A和a的等位基因频率是多少?b.
What will their frequencies be in the next generation, assuming the conditions of Hardy-Weinberg equilibrium hold?假设哈迪-温伯格平衡条件成立,下一代中它们的频率是多少?
In a screening program designed to detect carriers of an autosomal recessive condition, the carrier frequency in a specific founder population was approximately 4%. a.在一个旨在检测常染色体隐性遗传病携带者的筛查项目中,某个特定奠基者人群的携带者频率约为4%。a.
Calculate the frequency of the pathogenic allele in this population (assuming only one). b.计算该人群中致病等位基因的频率(假设只有一个)。b.
If the genetic fitness of affected individuals is zero, what proportion of possible reproductive pairings in this population could produce an affected child? c.如果患病个体的遗传适合度为零,该人群中可能产生患病后代的生殖配对比例是多少?c.
What is the prevalence of unaffected carriers among offspring of couples in which both partners are heterozygous for the trait? d.在双方均为该性状杂合子的夫妇的后代中,未受影响的携带者的患病率是多少?d.
If the pathogenic allele for this condition had the same frequency but were instead inherited through an autosomal dominant mode of inheritance (fully penetrant), what would be the frequency of unaffected adult carriers?如果该疾病的致病等位基因频率相同,但改为通过常染色体显性遗传方式(完全外显)遗传,未受影响的成年携带者的频率是多少?
What if it were X-linked dominant?如果是X连锁显性遗传呢?
X-linked recessive?X连锁隐性遗传呢?
Which of the following populations is in Hardy-Weinberg equilibrium, based on the information provided? a.根据提供的信息,下列哪个人群符合哈迪-温伯格平衡?a.
A/A, 0. 70; A/a, 0. 21; a/a, 0. 09. b.A/A 0.70;A/a 0.21;a/a 0.09。b.
A/A, 0. 32; A/a, 0. 64; a/a, 0. 04. c.A/A 0.32;A/a 0.64;a/a 0.04。c.
A/A, 0. 64; A/a, 0. 32; a/a, 0. 04.A/A 0.64;A/a 0.32;a/a 0.04。
What explanations could you offer to explain the frequencies in those populations that are not in equilibrium?你能给出什么解释来说明那些不符合平衡的人群中的频率?
You are consulted by a couple, Meera and Arjun, who tell you that Meera’s sister has Hurler syndrome (a mucopolysaccharidosis) and that they are concerned that they might have a child with the same syndrome.一对夫妇Meera和Arjun前来咨询,他们告诉你Meera的姐姐患有Hurler综合征(一种黏多糖贮积症),他们担心自己可能生下有相同综合征的孩子。
Hurler syndrome is inherited as an autosomal recessive trait with an estimated prevalence of 1 in 90,000 individuals in a large population. a.Hurler综合征是一种常染色体隐性遗传病,在大型人群中估计患病率为1/90,000。a.
What is the chance that Arjun is heterozygous for the pathogenic allele? b.Arjun携带致病等位基因杂合子的概率是多少?b.
What is the chance that Meera is a carrier of the pathogenic allele for Hurler syndrome? c.Meera是Hurler综合征致病等位基因携带者的概率是多少?c.
If Meera and Arjun are not genetically related through their parents (no consanguinity), what is the risk that Meera and Arjun’s first child will have Hurler syndrome? d.如果Meera和Arjun没有通过父母遗传相关(无血缘关系),他们的第一个孩子患Hurler综合征的风险是多少?d.
If Meera and Arjun share the same population ancestry, or received similar continental-level ancestry results from a direct-to-consumer genetic testing company, what is the risk that their first child will have the syndrome?如果Meera和Arjun有相同的人群祖先,或从直接面向消费者的基因检测公司收到了类似的大陆级祖先结果,他们的第一个孩子患该综合征的风险是多少?
In a certain population, each of 3 serious neuromuscular conditions—autosomal dominant facioscapulohumeral muscular dystrophy, autosomal recessive Friedreich ataxia, and X-linked recessive Duchenne muscular dystrophy—has an incidence of approximately 1 in 25,000 individuals. a.在某个特定人群中,三种严重的神经肌肉疾病——常染色体显性面肩肱型肌营养不良症、常染色体隐性遗传弗里德赖希共济失调和X连锁隐性遗传杜氏肌营养不良症——每种疾病的发病率约为1/25,000。a.
What are the frequencies of pathogenic alleles for each of these conditions? b.每种疾病的致病等位基因频率是多少?b.
Suppose that each condition could be treated such that affected individuals could have children.假设每种疾病都可以治疗,使得患病个体能够生育孩子。
What would be the resulting effect on the incidence of each condition?这将对每种疾病的发病率产生什么影响?
As discussed in this chapter, the autosomal recessive condition tyrosinemia type I has an incidence of 1 in 685 individuals in one population in the province of Quebec, but approximately 1 in 100,000 elsewhere.正如本章所讨论的,常染色体隐性遗传病Ⅰ型酪氨酸血症在魁北克省某个人群中的发病率为1/685,而在其他地区约为1/100,000。
What is the frequency of the variant associated with tyrosinemia in these two groups?在这两组人群中,与酪氨酸血症相关的变异频率是多少?
Suggest possible explanations for the difference in allele frequencies between the population in Quebec and populations elsewhere.请提出可能的解释,说明魁北克人群与其他地区人群之间等位基因频率的差异。

Complex Inheritance of Common Multifactorial Disorders

29/53
Complex Inheritance of Common Multifactorial Disorders Cristen J.
Ch9 — Segment 29
Complex Inheritance of Common Multifactorial Disorders Cristen J.常见多因子疾病的复杂遗传 Cristen J.
Willer Gonçalo R.Willer Gonçalo R.
Abecasis Common diseases such as heart disease, cancer, diabetes, neuropsychiatric disease, and asthma cause morbidity and premature mortality in nearly two of every three individuals ( Many of these diseases “run in families” – cases seem to cluster among the relatives of affected individuals more frequently than in similarly situated unrelated individuals.Abecasis 常见疾病如心脏病、癌症、糖尿病、神经精神疾病和哮喘,在近三分之二的人群中导致发病和过早死亡。这些疾病许多“具有家族聚集性”——患者亲属中的病例出现频率高于类似情况下的无关个体。
The inheritance of these diseases generally does not follow one of the mendelian patterns seen in the single-gene disorders described in Chapter 7.这些疾病的遗传通常不遵循第七章所述单基因疾病中的孟德尔遗传模式之一。
This is because these common diseases rarely result simply from inheriting a specific genetic defect at a single gene (or locus), as is the case for classical dominant and recessive mendelian disorders.这是因为这些常见疾病很少仅仅源于单个基因(或位点)特定遗传缺陷的遗传,而经典显性和隐性孟德尔疾病则如此。
Instead, they result from the combined effects of multiple genetic variants and environmental risk factors.相反,它们是由多个遗传变异和环境风险因素的共同作用所致。
Together, these genetic and environmental risk factors alter susceptibility to disease.这些遗传和环境风险因素共同改变疾病的易感性。
For this reason, these disorders are considered to be multifactorial in origin, and the familial clustering generates a pattern of inheritance that is referred to as complex.因此,这些疾病被认为具有多因子起源,而家族聚集性产生了一种称为复杂的遗传模式。
The familial clustering seen with multifactorial disorders can be explained by recognizing that family members share a greater proportion of their genetic information and environmental exposures than individuals chosen at random in the population.多因子疾病中观察到的家族聚集性可以通过认识到家族成员共享更大比例的遗传信息和环境暴露来解释,相较于人群中随机选择的个体。
Thus the relatives of an affected individual are more likely to experience the same genes, as well as gene-gene and gene-environment interactions that led to disease susceptibility than are individuals who are unrelated to the proband.因此,与先证者无关的个体相比,患病个体的亲属更有可能经历导致疾病易感性的相同基因以及基因-基因和基因-环境相互作用。
The pattern is complex because the contribution of chance events and because the combined effects of multiple environmental and genetic risk factors are not easy to predict and can result in a wide variety of patterns of disease segregation across families.该模式之所以复杂,是因为偶然事件的作用,以及多种环境和遗传风险因素的共同影响难以预测,并可能导致家族间疾病分离模式的多样性。
In this chapter we first address the question of how we infer that genetic variants predispose to common diseases.在本章中,我们首先探讨如何推断遗传变异易感于常见疾病。
We then describe how studies of familial aggregation and twin studies are used by geneticists to quantify the relative contributions of genetic variation and environment and show how these tools have been applied to multifactorial diseases.然后,我们描述遗传学家如何使用家族聚集性研究和双生子研究来量化遗传变异和环境的相对贡献,并展示这些工具如何应用于多因子疾病。
We describe a few examples of complex disorders where information has emerged about the specific nature of the genetic and environmental contributions to disease.我们列举一些复杂疾病的例子,其中已出现关于遗传和环境因素对疾病贡献的具体性质的信息。
Finally, we discuss how understanding of the genetic basis of common diseases enables polygenic risk scores (PRS) and predictions of individual disease risk that may soon impact clinical care in prevention, diagnosis, and individualized therapeutics.最后,我们讨论理解常见疾病的遗传基础如何使多基因风险评分(PRS)和个体疾病风险预测成为可能,这些可能很快影响预防、诊断和个体化治疗的临床护理。
For most common diseases a number of genetic variants contributing to disease susceptibility have now been identified, mainly through large-scale genetic association studies ( Still, the specific genes, variants, and environmental factors that contribute to disease risk have not yet been fully identified for the vast majority of common multifactorial diseases.对于大多数常见疾病,现已通过大规模遗传关联研究识别出多种与疾病易感性相关的遗传变异。然而,对于绝大多数常见多因子疾病,导致疾病风险的具体基因、变异和环境因素尚未完全鉴定。
Completing the task and identifying all genetic and environmental factors that contribute to each disease is challenging work to which you, dear reader, might contribute – it is our hope that this chapter will give you a helpful introduction to the process.完成这项任务并识别导致每种疾病的所有遗传和环境因素是一项具有挑战性的工作,亲爱的读者,您可能为此做出贡献——我们希望本章能为您提供有益的过程介绍。
A more detailed understanding of the approaches that geneticists use to identify the genetic factors underlying complex disease first requires a full appreciation of the distribution of genetic variation in different populations.要更详细地理解遗传学家用于识别复杂疾病遗传因素的方法,首先需要充分认识不同人群中遗传变异的分布。
This topic is presented in Chapter 10, after which we will turn, in Chapter 11, to discussion of the specific population-based epidemiological approaches that geneticists are using to identify the particular genes and variants contributing to conditions with complex inheritance.这一主题将在第10章中介绍,之后我们将在第11章中讨论遗传学家用于识别具有复杂遗传条件的具体基因和变异的人群流行病学方法。
Ultimately, finding the genes and variants that interact with the environment to contribute to disease susceptibility will give us a better understanding of the underlying processes leading to common multifactorial diseases and, perhaps, better tools for prevention or treatment.最终,发现与环境相互作用导致疾病易感性的基因和变异,将使我们更好地理解导致常见多因子疾病的基本过程,或许还能提供更好的预防或治疗工具。
QUALITATIVE AND QUANTITATIVE TRAITS The first step in a genetic analysis of disease is often to summarize the disease state for each individual.定性和定量性状 疾病遗传分析的第一步通常是总结每个个体的疾病状态。
For convenience, disease states for multifactorial diseases with complex inheritance are most often summarized as discrete or binary traits (classifying individuals as为方便起见,具有复杂遗传的多因子疾病的疾病状态通常被总结为离散性或二分类性状(将个体分类为
30/53
affected cases or unaffected controls) but can also be summarized using continuous quantitative traits or even discrete …
Ch9 — Segment 30
affected cases or unaffected controls) but can also be summarized using continuous quantitative traits or even discrete ordinal scales.受影响病例或未受影响对照组)但也可通过连续定量性状甚至离散有序量表进行总结。
At first glance, a binary classification strategy is the simpler approach; a disease, such as asthma, obesity, or hearing loss, is classified as present or absent in each individual being studied.乍一看,二元分类策略是更简单的方法;对于所研究的每个个体,疾病(如哮喘、肥胖或听力损失)被分类为存在或不存在。
Distinguishing between individuals who have a disease and those who do not may not always be straightforward and may require detailed examination, specialized testing, or even arbitrary distinctions and cutoff points.区分患病个体与未患病个体并非总是简单明了,可能需要详细检查、专业检测,甚至主观划分和截断点。
As an alternative to detailed examination of each study participant, many contemporary studies use automated algorithms for assigning disease states to individuals based on electronically encoded information in their medical records.作为对每位研究参与者进行详细检查的替代方案,许多当代研究使用自动化算法,根据医疗记录中的电子编码信息为个体分配疾病状态。
A popular set of strategies in this class is the use of Phe Codes, a disease state definition based on presence or absence of one or more billing, diagnostic, or procedure codes in the medical record.这类策略中常用的一种是使用PheCodes,这是一种基于医疗记录中是否存在一个或多个计费、诊断或操作代码的疾病状态定义。
It’s worth noting that while these strategies can be extremely practical and have enabled many successful gene-mapping experiments, they can also result in arbitrary or inconsistent classification of individual individuals (e. g., depending on the choices their health care provider might have made in diagnosis, testing, billing, or treatment).值得注意的是,尽管这些策略非常实用,并促成了许多成功的基因定位实验,但它们也可能导致对个体(例如,取决于其医疗保健提供者在诊断、检测、计费或治疗中所做的选择)进行随意或不一致的分类。
In some cases, multiple instances of a code in the electronic health record are required for a more confident diagnosis, and individuals with unclear phenotypes are excluded from both the case and control groups.在某些情况下,需要电子健康记录中多次出现某个代码才能更确信地诊断,而表型不明确的个体则被排除在病例组和对照组之外。
Quantitative traits can often avoid arbitrary boundaries between cases and controls.定量性状通常可以避免病例与对照组之间的随意边界。
Instead of classifying individuals as asthmatic cases and nonasthmatic controls, we might measure their lung capacity or the number of emergency room visits per year.与其将个体分类为哮喘病例和非哮喘对照,我们可以测量他们的肺活量或每年急诊就诊次数。
Instead of classifying individuals as hard-of-hearing cases or normal-hearing controls, we might use a quantitative measure to quantify the loudness or pitch of sounds each individual can hear.与其将个体分类为听力受损病例或听力正常对照,我们可以使用定量测量来量化每个个体能听到的声音的响度或音调。
Instead of classifying individuals as obese or nonobese, we might measure their weight or body mass index.与其将个体分类为肥胖或非肥胖,我们可以测量他们的体重或身体质量指数。
Quantitative traits are also commonly used to summarize disease-related measurable physiologic or biochemical quantities such as blood pressure, serum cholesterol concentration, or activity levels that vary among individuals within a population.定量性状也常用于总结疾病相关的可测量生理或生化指标,如血压、血清胆固醇浓度或活动水平,这些指标在群体中个体间存在差异。
The Normal Distribution As is often the case with physiologic quantities, such as systolic blood pressure, a graph of the number (or the fraction) of individuals in the population (y-axis) having a particular quantitative value (x-axis) approximates the familiar, bell-shaped curve known as the normal (or gaussian) distribution .正态分布 与生理指标(如收缩压)常见的情况一样,群体中具有特定定量值(x轴)的个体数量(或比例)的图形(y轴)近似于熟悉的钟形曲线,即正态(或高斯)分布。
The position of the peak and the width of the curve of the normal distribution are governed by two quantities, the mean (µ) and the variance (σ2), respectively.正态分布曲线的峰值位置和宽度分别由两个量决定:均值(μ)和方差(σ²)。
The mean is the arithmetic average of the values, and because – for many traits – more people have values for the trait near the average, the curve ordinarily has its peak at the mean value.均值是数值的算术平均值,并且由于对于许多性状而言,更多人的性状值接近平均值,因此曲线通常在其平均值处达到峰值。
The variance (or its square root, σ, the standard deviation [SD]) is a measure of how much spread there is in the values to either side of the mean and therefore determines the breadth of the curve.方差(或其平方根σ,即标准差[SD])是衡量均值两侧数值离散程度的指标,因此决定了曲线的宽度。
Any physiologic quantity that can be measured in a sample of a population is a quantitative phenotype, and the mean and variance for that sample can be calculated and used to approximate the underlying mean and variance of the population from which the sample was drawn.在群体样本中可测量的任何生理指标都是定量表型,可以计算该样本的均值和方差,并用于近似估计该样本所来源的群体的真实均值和方差。
For example, the systolic blood pressure of thousands of men in two different age groups is shown in , with more individuals with systolic blood pressures above the mean than below.例如,两个不同年龄组数千名男性的收缩压如图所示,收缩压高于平均值的个体多于低于平均值的个体。
The normal distribution provides guidelines for setting the limits of the normal range.正态分布为设定正常范围的界限提供了指导。
A normal range is often defined as the values of a quantitative trait that are seen in ~95% of the population.正常范围通常定义为约95%群体中出现的定量性状值。
Statistical theory states that when the values of a quantitative trait in a population follow the bell-shaped normal curve (i. e., are normally distributed), ~5% of the population will have measurements more than 2 SD above or below the population mean.统计理论指出,当群体中定量性状的值遵循钟形正态曲线(即正态分布)时,约5%的群体测量值将高于或低于群体均值超过2个标准差。
It is important to note that an individual may be perfectly healthy (i. e., “normal”) despite having a trait value outside the normal range.重要的是要注意,尽管个体性状值超出正常范围,但仍可能完全健康(即“正常”)。
Furthermore, since 5% of the population will, by definition, be outside the normal range, when many traits are measured and thus classified, each individual will typically be extreme in several traits.此外,由于根据定义,5%的群体将处于正常范围之外,当测量并分类多个性状时,每个个体通常会在若干性状上表现极端。
It’s important to note that this concept of normal range should not be confused with health and disease.重要的是要注意,正常范围的概念不应与健康和疾病相混淆。
For example, because body mass is typically high in industrialized societies, many individuals within the normal range (e. g., within 2 SD of the mean) might be considered clinically obese.例如,由于工业化社会中体重通常偏高,许多处于正常范围内(例如均值±2个标准差内)的个体可能被视为临床肥胖。
As another example, because the human body can physiologically tolerate variation in platelet levels, the extreme platelet levels used to diagnose clinical conditions like thrombocytosis 8 3. 8 Disorders due to single-gene variants 10 3. 6 20 Disorders with multifactorial inheritance ≈50 ≈50 ≈600 Data from Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin's principles and practice of medical genetics, ed 3, Edinburgh, 1997, Churchill Livingstone.另一个例子是,由于人体在生理上能耐受血小板水平的变异,用于诊断血小板增多症等临床状况的极端血小板水平8 3. 8 单基因变异导致的疾病10 3. 6 20 多因子遗传疾病≈50 ≈50 ≈600 数据来源:Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin's principles and practice of medical genetics, 第3版, 爱丁堡, 1997, Churchill Livingstone。
31/53
Complex Inheritance of Common Multifactorial Disorders 151 or thrombocytopenia are far outside the normal range.
Ch9 — Segment 31
Complex Inheritance of Common Multifactorial Disorders 151 or thrombocytopenia are far outside the normal range.常见多因素疾病的复杂遗传 151 或血小板减少症远超出正常范围。
In general the normal distribution and the normal range are a statistical convenience but cannot be used directly to diagnose health and disease.通常,正态分布和正常范围是一种统计便利,但不能直接用于诊断健康和疾病。
For convenience, many genetic analyses assume that traits follow this bell-shaped normal distribution in the population.为方便起见,许多遗传分析假设性状在人群中遵循这种钟形正态分布。
Although this is approximately true for many traits (such as height, weight, total cholesterol, and blood pressure in younger individuals), it is clearly not the case in other cases (such as triglyceride levels, number of moles in skin, or number of children in a family).尽管这对许多性状(如身高、体重、总胆固醇以及年轻个体的血压)大致成立,但在其他情况(如甘油三酯水平、皮肤痣的数量或家庭中子女的数量)下显然并非如此。
For convenience, many genetic studies map the original measurements, which may not be normally distributed, to a new measurement scale that is normally distributed.为方便起见,许多遗传学研究将可能不服从正态分布的原始测量值映射到服从正态分布的新测量尺度上。
This can provide much more flexibility in the choice of analysis strategy.这可以在分析策略的选择上提供更大的灵活性。
The process typically works by rank-ordering the quantitative measurements (first highest among 100, second highest among 100, etc.) and then mapping these ranks to corresponding values in a simulated normal distribution (2. 58 SD, 2. 17 SD, etc.).该过程通常通过对定量测量值进行排序(例如,在100个中排名最高、第二高等),然后将这些排名映射到模拟正态分布中的对应值(例如2.58个标准差、2.17个标准差)。
The inverse normal distribution function is used to tabulate the expected values in the simulated normal distribution.逆正态分布函数用于列出模拟正态分布中的期望值。
This popular strategy handles outlier values or nonnormality in the original measured values and provides genetic results in SD units allowing for easier comparison between different studies.这一常用策略处理原始测量值中的异常值或非正态性,并以标准差单位提供遗传结果,便于不同研究之间的比较。
FAMILIAL AGGREGATION AND CORRELATION Allele Sharing Among Relatives The more closely related two individuals are, the more alleles they share in common, on average (see Chapter 7).家族聚集与相关性 亲属间的等位基因共享:两个个体亲缘关系越近,平均而言他们共享的等位基因越多(见第7章)。
The most extreme example of allele sharing is identical (monozygotic [MZ]) twins (see later in this chapter), who have the same alleles at every locus, with possibly a few small exceptions arising from somatic variants.等位基因共享的最极端例子是同卵(单卵[MZ])双胞胎(见本章后文),他们在每个位点拥有相同的等位基因,可能有少量由体细胞变异引起的例外。
The next most closely related individuals are typically first-degree relatives, such as a parent offspring or sibling pairs, including fraternal (dizygotic [DZ]) twins.次近亲缘关系的个体通常是一级亲属,例如父母与子女或兄弟姐妹对,包括异卵(双卵[DZ])双胞胎。
In a parent-child pair, the child shares at least one allele out of two (50% of alleles) in common with the parent at every genetic location.在亲子对中,子女在每个遗传位点上与父母共享至少两个等位基因中的一个(50%的等位基因)。
This shared allele is on the chromosome the child inherited from that parent; sharing on the chromosome inherited from the other parent can result from consanguinity, very distant relatedness, or chance.这个共享的等位基因位于子女从该父母遗传的染色体上;而在从另一方父母遗传的染色体上的共享可能源于近亲结婚、非常远的亲缘关系或偶然。
Siblings (including DZ twins) also share 50% or more of their alleles on average, but this can vary along the genome.兄弟姐妹(包括异卵双胞胎)平均也共享50%或更多的等位基因,但这种比例在基因组中可能存在变化。
This is because, at a given locus, a pair of siblings inherits the same two chromosomes from their parents ¼ of the time, inherits one chromosome in common ½ the time, and inherits no chromosomes in common the remaining ¼ of the time .这是因为在同一特定基因座,一对兄弟姐妹有1/4的概率从父母遗传到相同的两条染色体,1/2的概率遗传到一条共同染色体,其余1/4的概率遗传到没有共同染色体。
At any one locus, in the absence of consanguinity, the average number of chromosomes that a sibling pair is expected to share identical by descent (i. e., alleles that are identical because they are copies of the same ancestral chromosome) is: 14 14 12 () () () 2alleles 0allele 1allele 0. 5 0. 5 0 1allele + + = + + = The more distantly related two members of a family are, the fewer alleles they are expected to inherit from a common ancestor.在任何一位点,没有近亲结婚的情况下,一对兄弟姐妹预期通过血缘同源(即因同一祖先染色体复制而相同的等位基因)共享的染色体平均数量为:1/4×2 + 1/2×1 + 1/4×0 = 0.5 + 0.5 + 0 = 1个等位基因。家庭中两个成员的亲缘关系越远,他们预期从共同祖先遗传的等位基因就越少。
Familial Aggregation in Binary Traits If genetic variants modify disease risk, relatives of an affected individual will have a greater-than-expected rate of the disease than unrelated individuals with similar nongenetic risk profiles (familial aggregation of Percent of population Quantity being measured Systolic blood pressure 30 20 10 0 –2 SD –1 SD MEAN +1 SD +2 SD 100 120 140 160 180 200 134±40 127±34 A B For many traits, the “normal” range is considered the mean ±2 SD, as indicated by the shaded region.二元性状的家族聚集 如果遗传变异改变疾病风险,患病个体的亲属与具有相似非遗传风险特征的无关个体相比,疾病发生率会高于预期(人口百分比 测量量 收缩压 30 20 10 0 –2 标准差 –1 标准差 均值 +1 标准差 +2 标准差 100 120 140 160 180 200 134±40 127±34 A B 对于许多性状,“正常”范围被视为均值±2个标准差,如阴影区域所示)。
(B) Distribution of systolic blood pressure in ~3300 men aged 40–45 (solid line) and ~2200 men aged 50–55 (dotted line).(B)约3300名40–45岁男性(实线)和约2200名50–55岁男性(虚线)的收缩压分布。
The mean and ±2 SD are shown above double-headed arrows.均值和±2个标准差显示在双头箭头上方。
(B, Data from Sive PH, Medalie JH, Kahn HA, et al: Distribution and multiple regression analysis of blood pressure in 10,000 Israeli men, Am J Epidemiol 93:317–327, 1971.)(B,数据来自Sive PH, Medalie JH, Kahn HA, 等:10000名以色列男性血压的分布与多元回归分析,Am J Epidemiol 93:317–327, 1971。)
32/53
This is because closely related family members are expected to share, on average, some disease predisposing alleles.
Ch9 — Segment 32
This is because closely related family members are expected to share, on average, some disease predisposing alleles.这是因为近亲家属平均而言会共享一些疾病易感等位基因。
Next, we will discuss two approaches to measuring familial aggregation: relative risk ratios and family history case-control studies.接下来,我们将讨论测量家族聚集性的两种方法:相对风险比和家族史病例对照研究。
Relative Risk Ratio One way to measure familial aggregation of a disease is by comparing the frequency of the disease in the relatives of an affected proband with its disease frequency (prevalence) in the general population.相对风险比 衡量疾病家族聚集性的一种方法是将患病先证者亲属中的疾病频率与其在一般人群中的疾病频率(患病率)进行比较。
The relative risk ratio λr (where the subscript r refers to relatives) is defined as: λr Prevalence of the disease in the relatives of an affecte = d person Prevalence of the disease in the general population The value of λr as a measure of familial aggregation depends both on how frequently a disease occurs in relatives of an affected individual (the numerator) and on the population prevalence (the denominator); the larger λr is, the greater the familial aggregation.相对风险比λr(下标r指亲属)定义为:λr = 患者亲属中的疾病患病率 / 一般人群中的疾病患病率;λr作为家族聚集性衡量指标的值取决于疾病在患者亲属中发生的频率(分子)和人群患病率(分母);λr越大,家族聚集性越强。
This estimate is typically performed within a specific type of close relative (e. g., siblings, offspring, identical twins).该估计通常在特定类型的近亲中进行(例如,兄弟姐妹、子女、同卵双胞胎)。
The population prevalence enters into the calculation because the more common a disease is, the greater the likelihood that apparent aggregation may be a coincidence.引入人群患病率进行计算,是因为疾病越常见,表面上的聚集性就越可能是巧合。
A value of λr = 1 indicates that a relative is no more likely to develop the disease than is any individual in the population, whereas a value greater than 1 indicates that a relative is more likely to develop the disease.λr = 1表示亲属患该疾病的可能性不高于人群中的任何个体,而大于1的值则表示亲属更有可能患病。
Examples of relative risk ratios determined for various diseases in samples of siblings (λs where s is for siblings) are shown in Since many diseases have sexand/or age-specific prevalences, it is important to ensure relatives and reference populations are appropriately matched with respect to these factors.兄弟姐妹样本中各种疾病的相对风险比(λs,其中s代表兄弟姐妹)的实例已展示;由于许多疾病具有性别和/或年龄特异性患病率,因此确保亲属与参考人群在这些因素方面适当匹配至关重要。
Family History Case-Control Studies Another approach to estimating familial aggregation is the case-control study, in which individuals with a disease (the cases) are compared with suitably chosen individuals without the disease (the controls), with respect to family history of disease (as well as other factors, such as environmental exposures, occupation, geographic location, parity, and previous illnesses).家族史病例对照研究 评估家族聚集性的另一种方法是病例对照研究,其中将患有疾病个体(病例组)与适当选择的未患病个体(对照组)在疾病家族史(以及其他因素,如环境暴露、职业、地理位置、胎次和既往疾病)方面进行比较。
To assess a possible genetic contribution to familial aggregation of a disease, the frequency with which the disease is found in the extended families of the cases (positive family history) is compared with the frequency of positive family history among suitable controls, matched for age and ancestry.为了评估遗传因素对疾病家族聚集性的可能贡献,将病例的扩展家族中发现疾病(阳性家族史)的频率与经年龄和祖先匹配的适当对照组中阳性家族史的频率进行比较。
Spouses can be used as controls in this situation because they usually match the cases in age and ancestry and share the same household environment (provided the disease is of similar prevalence in males and females).在这种情况下,配偶可用作对照,因为他们通常在年龄和祖先方面与病例匹配,并共享相同的家庭环境(前提是该疾病在男性和女性中的患病率相似)。
Other frequently used controls are individuals with unrelated diseases matched for age, sex, and ancestry.其他常用的对照是患有无关疾病的个体,经年龄、性别和祖先匹配。
Thus, for example, in a study of multiple sclerosis (MS), ~3. 5% of first-degree relatives of patients with MS also had MS, a prevalence that was much higher than among first-degree relatives of matched controls without A1A3 A1A2 A3A4 A1A3 A1A4 A1A4 A2A3 A2A3 A2A4 A2A4 2 1 1 0 1 2 0 1 1 0 2 1 0 1 1 2 Sib #1 Sib #2 Genotype of sib #1 Genotype of sib #2 The parents’ genotypes are shown as A1A2 for the father and A3A4 for the mother.例如,在一项多发性硬化症(MS)研究中,约3.5%的MS患者一级亲属也患有MS,这一患病率远高于匹配对照的一级亲属,以下为基因型数据:A1A3 A1A2 A3A4 A1A3 A1A4 A1A4 A2A3 A2A3 A2A4 A2A4 2 1 1 0 1 2 0 1 1 0 2 1 0 1 1 2 同胞#1 同胞#2 同胞#1基因型 同胞#2基因型,父母的基因型显示为父亲A1A2,母亲A3A4。
All four possible genotypes for sib #1 are given across the top of the table, and all four possible genotypes for sib #2 are given along the left side of the table.同胞#1的所有四种可能基因型列于表格顶部,而同胞#2的所有四种可能基因型列于表格左侧。
The numbers inside the boxes represent the number of alleles both sibs have in common for all 16 different combinations of genotypes for both sibs.方框内的数字表示两个同胞在所有16种不同基因型组合中共有的等位基因数量。
For example, the upper left-hand corner has the number 2 because sib #1 and sib #2 both have the genotype A1A3 and so have both A1 and A3 alleles in common.例如,左上角数字为2,因为同胞#1和同胞#2均为A1A3基因型,因此共有的等位基因为A1和A3。
The bottom left-hand corner contains the number 0 because sib #1 has genotype A1A3, whereas sib #2 has genotype A2A4, so there are no alleles in common.左下角数字为0,因为同胞#1的基因型为A1A3,而同胞#2的基因型为A2A4,因此没有共有的等位基因。
33/53
Complex Inheritance of Common Multifactorial Disorders 153 MS (0. 2%).
Ch9 — Segment 33
Complex Inheritance of Common Multifactorial Disorders 153 MS (0. 2%).常见多因素疾病的复杂遗传 153 多发性硬化症(0.2%)。
That is, the odds of having a first-degree relative with MS were 18 times higher among, people with MS than among controls.也就是说,多发性硬化症患者拥有一级亲属患多发性硬化症的几率是对照组的18倍。
(In Chapter 11, we will discuss how one calculates odds ratios in case-control studies.) One can conclude therefore that substantial familial aggregation is occurring in MS, thereby providing evidence of a genetic predisposition to this disease.(在第11章中,我们将讨论如何在病例对照研究中计算比值比)因此可以得出结论,多发性硬化症存在显著的家族聚集性,从而为这种疾病的遗传易感性提供了证据。
These types of studies are somewhat vulnerable to recall bias, since diseased individuals are more likely to be aware of the disease status of similarly affected relatives.这类研究在一定程度上容易受到回忆偏倚的影响,因为患病个体更可能了解同样受影响的亲属的疾病状况。
Measuring the Genetic Contribution to Quantitative Traits Sharing of alleles that govern a particular quantitative trait affects the distribution of values of that trait in family members.测量对数量性状的遗传贡献:控制特定数量性状的等位基因共享会影响该性状在家族成员中的值分布。
The more sharing of alleles that govern a quantitative trait there is among relatives, the more similar values of the trait are expected to be.亲属之间共享控制数量性状的等位基因越多,该性状的值预计就越相似。
The effect of genetic variation on quantitative traits is often measured and reported in two related ways: correlation between relatives and heritability.遗传变异对数量性状的影响通常通过两种相关方式进行测量和报告:亲属间的相关性和遗传力。
Familial Correlation Just like the relative risk ratios are used to summarize aggregation of disease within families, there are analogous strategies to summarize whether quantitative traits aggregate within families.家族相关性:正如使用相对危险度来总结疾病在家族内的聚集性一样,也有类似的策略来总结数量性状是否在家族内聚集。
Checking for these patterns of familial aggregation provides important clues about the role of genetic variation in each trait.检查这些家族聚集模式为遗传变异在每个性状中的作用提供了重要线索。
Prior to studying the genetic factors underlying a trait, it is important to first establish that the trait has a genetic component.在研究一个性状的遗传因素之前,首先确定该性状具有遗传成分是重要的。
Geneticists answer this question in a few different ways.遗传学家通过几种不同的方式回答这个问题。
The tendency for the values of a physiologic measurement to be more similar among relatives is summarized through the correlation of these physiologic quantities among relatives.亲属之间生理测量值趋于更相似的趋势通过亲属间这些生理量的相关性来总结。
The coefficient of correlation (symbolized by the letter r) is a statistical measure of correlation applied to a pair of measurements, such as a child’s serum cholesterol level and that of a parent.相关系数(用字母r表示)是一种应用于一对测量值(如儿童的血浆胆固醇水平与其父母的血浆胆固醇水平)的相关性统计度量。
A positive correlation between the cholesterol measurements in two groups of relatives exists if it is found that a higher level in the first individual (e. g., the child) predicts a proportionately higher level in the relative (e. g., the parent).如果在两个亲属组中发现第一个个体(例如孩子)的较高水平预示着其亲属(例如父母)相应较高的水平,则存在胆固醇测量值的正相关性。
When a correlation exists, a graph of values in the proband and his or her relatives, in which each point represents a proband-relative pair of values, will tend to cluster around a straight line.当存在相关性时,先证者及其亲属的值图(其中每个点代表一个先证者-亲属值对)将倾向于聚集在一条直线附近。
The value of r can range from 0 when there is no correlation to +1 for perfect positive correlation.r的值范围从无相关时的0到完全正相关时的+1。
In the example of serum cholesterol, between serum cholesterol level of mothers 350 250 150 50 Cholesterol of son, mg% Cholesterol of mother, mg% 50 150 250 350 450 Each dot represents a mother-son pair of measurements.在血清胆固醇的例子中,母亲的血浆胆固醇水平(350、250、150、50)与儿子的胆固醇水平(mg%)及母亲的胆固醇水平(mg%)(50、150、250、350、450)之间,每个点代表一对母子的测量值。
The straight line is a “best fit” through the data points.这条直线是通过数据点的“最佳拟合”线。
(Data from Johnson BC, Epstein FH, Kjelsberg MO: Distributions and familial studies of blood pressure and serum cholesterol levels in a total community – Tecumseh, Michigan.(数据来自Johnson BC, Epstein FH, Kjelsberg MO: 分布及家族研究:一个完整社区中的血压和血浆胆固醇水平——密歇根州特库姆塞,
J Chronic Dis 18:147–160, 1965.)慢性病杂志 18:147–160, 1965.)
34/53
aged 30 to 39 and those of their male children aged 4 to 9.
Ch9 — Segment 34
aged 30 to 39 and those of their male children aged 4 to 9.年龄在30至39岁之间的人群及其4至9岁的男性子女。
A negative correlation exists when an increase in the first individual’s measurement predicts a lower measurement in relatives.当第一个个体的测量值增加预示着其亲属的测量值降低时,存在负相关。
The measurements are still correlated but in the opposite direction.这些测量值仍然相关,但方向相反。
In such a case, the value of r can range from 0 to −1 for a perfect negative correlation.在此情况下,r值可以从0到-1,表示完全负相关。
This is relatively rare in genetic studies of relatives, but it can occur in other settings where correlations are used to summarize relationships between measurements (e. g., an individual’s activity levels might be negatively correlated with body mass).这在亲属的遗传学研究中相对罕见,但在其他利用相关性总结测量间关系的情境中可能出现(例如,个体的活动水平可能与体重呈负相关)。
Heritability The concept of heritability of a quantitative trait (symbolized as H2) was developed to describe how much the genetic differences between individuals in a population contribute to variability of that trait in the population.遗传力 数量性状遗传力的概念(符号化为H²)被用来描述群体中个体间的遗传差异对该群体中该性状变异性的贡献程度。
H2 is defined as the fraction of the variation in a quantitative trait that is due to genetic variation in the broadest sense, regardless of the mechanism by which the various alleles affect the phenotype.H²定义为数量性状的变异中由最广义遗传变异(无论不同等位基因通过何种机制影响表型)所贡献的部分所占的比例。
The higher the heritability, the greater the contribution of genetic differences among people to the variability of the trait in the population.遗传力越高,人群中个体间的遗传差异对该性状变异性的贡献就越大。
The value of H2 varies from 0, if genotype contributes nothing to the total phenotypic variance in a population, to 1, if genotype is totally responsible for the phenotypic variance in that population.H²的取值范围为:若基因型对群体总表型方差无贡献,则H²=0;若基因型完全决定群体表型方差,则H²=1。
Heritability of a human trait is a theoretical quantity that is usually estimated from the correlation between measurements of that trait among relatives of known degrees of relatedness, such as parents and children, siblings, or, as we shall see later in this chapter, twins.人类性状的遗传力是一个理论量,通常通过该性状在已知亲缘关系程度的亲属(如父母与子女、兄弟姐妹,或本章后续将讨论的双胞胎)之间的测量相关性来估计。
DETERMINING THE RELATIVE CONTRIBUTIONS OF GENES AND ENVIRONMENT TO COMPLEX DISEASE Distinguishing Between Genetic and Environmental Influences Using Family Studies For both qualitative and quantitative traits, similarities among family members are most likely the result of shared genetics and shared environment, such as socioeconomic status, local environment, dietary habits, or cultural behaviors, all of which are frequently shared among family members but are generally considered to be of nongenetic origin.确定基因和环境对复杂疾病的相对贡献 利用家庭研究区分遗传和环境影响 对于定性和定量性状,家庭成员间的相似性很可能是共享遗传和共享环境(如社会经济地位、当地环境、饮食习惯或文化行为)的结果,这些因素通常在家庭成员间共享,但通常被认为源于非遗传因素。
Given evidence of familial aggregation of a disease or correlation of a quantitative trait, geneticists attempt to separate the relative contributions of genotype and environment to the phenotype using a variety of approaches.在获得疾病家族聚集性或数量性状相关性的证据后,遗传学家尝试使用多种方法分离基因型和环境对表型的相对贡献。
Historically, when sample sizes were smaller and genetic data were less available, researchers would collect pedigree information from cases (probands) and their family members.历史上,当样本量较小且遗传数据获取有限时,研究者会收集先证者及其家系成员的家系信息。
This would enable estimates of disease risk in different relatives, grouped by their shared genetics (e. g., 50% genetic sharing in parent-offspring pairs or full siblings, 25% genetic sharing in grandparent-grandchild pairs, avuncular or half-sibling pairs) and an evaluation of whether the degree of similarity between individuals attenuates in proportion to their genetic relatedness.这可以按共享遗传比例分组(例如,亲代-子代对或全同胞对共享50%遗传物质,祖父母-孙代对、叔侄对或半同胞对共享25%遗传物质),评估不同亲属的疾病风险,并检验个体间相似程度是否随遗传相关性的降低而衰减。
Another common attempt to control for shared environment is to compare the concordance of disease status in MZ or identical twins with that in same-sex DZ or fraternal twins.另一种控制共享环境的常用方法是比较同卵双生子与同性别异卵双生子的疾病状态一致性。
Since twins share much of their environment in utero and early childhood whether identical or fraternal, it is convenient to hypothesize that any difference in concordance is due to the difference in genetic sharing (100% sharing for identical twins and ~50% sharing for fraternal twins).由于同卵和异卵双生子在子宫内及儿童早期共享大部分环境,因此可以方便地假设一致性上的任何差异源于遗传共享的差异(同卵双生子共享100%,异卵双生子共享约50%)。
More recently, with the development of biobanks (i. e., very large-scale collections of research participants, coupled with genetic data and large-scale electronic health records), heritability of many phenotypes and diseases can be efficiently examined using very distant relative pairs identified by comparing their genetic data directly.近来,随着生物样本库(即大规模收集研究参与者,并结合遗传数据和大规模电子健康记录)的发展,可以通过直接比较遗传数据识别的远亲对,高效地检验许多表型和疾病的遗传力。
For example, we might hypothesize that for a genetic trait, concordance of disease status or correlation between quantitative measurements will be higher between individuals who share 4% of their genetic material than between individuals who share 2% of their genetic material.例如,我们可以假设:对于遗传性状,共享4%遗传物质的个体之间在疾病状态一致性或数量测量相关性上,会高于共享2%遗传物质的个体。
We next discuss some possibilities in more detail.接下来我们将更详细地讨论一些可能性。
Attenuation of Risk for Progressively More Distant Relatives One approach is to compare λr measurements or quantitative trait correlations between relatives of varying degrees of relatedness to the proband.风险随亲属关系渐远而衰减 一种方法是比较与先证者亲缘关系程度不同的亲属之间的λr测量值或数量性状相关性。
For example, if genes predispose to a disease, one would expect λr to be greatest for MZ twins, to be somewhat smaller for firstdegree relatives such as sibs or parent-child pairs, and to continue to decrease as allele-sharing decreases among the more distant relatives in a family .例如,若基因易感某种疾病,则预期λr在同卵双生子中最大,在一级亲属(如兄弟姐妹或父母-子女对)中稍小,并随家族中更远亲属间等位基因共享的减少而持续降低。
To illustrate this approach, consider cleft lip with or without cleft palate, or CL(P), one of the most common congenital malformations, affecting 1. 4 per 1000 newborns worldwide.为说明这种方法,考虑伴或不伴腭裂的唇裂(CL(P)),这是最常见的先天性畸形之一,全球每1000名新生儿中约1.4例受累。
CL(P) originates as a failure of fusion of embryonic tissues that will go to make up the upper lip and the hard palate at approximately day 35 of gestation.CL(P)起源于妊娠约第35天时,构成上唇和硬腭的胚胎组织未能融合。
It is a multifactorial disorder with complex inheritance; for reasons that are not well understood, ~60% to 80% of those affected with CL(P) are males.它是一种具有复杂遗传方式的多因素疾病;原因尚不明确,约60%至80%的CL(P)患者为男性。
Despite the similarity in names, CL(P) is usually etiologically distinct from isolated cleft palate (i. e., without cleft lip).尽管名称相似,CL(P)通常与孤立性腭裂(即无唇裂)在病因学上截然不同。
CL(P) is heterogeneous and includes forms in which the clefting is only one feature of a syndrome that includes other features, known as syndromic CL(P), as well as forms that are not associated with other birth defects, which are known as nonsyndromic CL(P).CL(P)具有异质性,包括仅以唇腭裂为特征之一的综合征型(称为综合征型CL(P)),以及与其它出生缺陷无关的非综合征型CL(P)。
Syndromic CL(P) can be inherited as a mendelian single-gene disorder or can be caused by chromosome disorders (especially trisomy 13 and 4p− deletion syndrome) (see Chapter 6) or teratogenic exposure (rubella embryopathy, thalidomide, or anticonvulsants) (see综合征型CL(P)可作为孟德尔单基因疾病遗传,或由染色体疾病(尤其是13三体和4p−缺失综合征)(见第6章)或致畸物暴露(风疹胚胎病、沙利度胺或抗惊厥药)引起(见
35/53
Complex Inheritance of Common Multifactorial Disorders 155 Chapter 15).
Ch9 — Segment 35
Complex Inheritance of Common Multifactorial Disorders 155 Chapter 15).常见多因素疾病的复杂遗传(第15章)。
Nonsyndromic CL(P) can also be inherited as a single-gene disorder but more commonly is a sporadic occurrence and demonstrates some degree of familial aggregation without an obvious mendelian inheritance pattern.非综合征性唇裂伴或不伴腭裂也可作为单基因疾病遗传,但更常见的是散发性发生,并表现出一定程度的家族聚集性而无明显的孟德尔遗传模式。
The risk for CL(P) in a child increases as a function of the number of relatives the child has who are affected with CL(P) and the more closely related they are to the child ( The simplest explanation for this is that the more closely related one is to the proband and, the more probands there are in the family, the more likely one is to inherit disease-susceptibility alleles, thus increasing risk for the disorder.儿童患唇裂伴或不伴腭裂的风险随其受影响亲属数量及与这些亲属的亲缘关系密切程度而增加(最简单的解释是,与先证者亲缘关系越近,且家庭中先证者数量越多,个体越可能遗传疾病易感等位基因,从而增加患病风险)。
Another approach is to compare the disease relative risk ratio in biological relatives of the proband with that in biologically unrelated family members (e. g., adoptees or spouses), all living in the same household environment.另一种方法是将先证者生物学亲属的疾病相对风险比与生活在同一家庭环境中的非生物学亲属(如领养子女或配偶)进行比较。
Returning to MS, for example, λr is 190 for MZ twins and 20 to 40 for first-degree biological relatives (parents, children, and sibs).以多发性硬化为例,同卵双生子的λr为190,而一级生物学亲属(父母、子女和兄弟姐妹)的λr为20至40。
In contrast, λr is 1 for the adopted siblings of an affected individual, suggesting that most of the familial aggregation in MS is genetic rather than the result of a shared environment.相比之下,患病个体的领养兄弟姐妹的λr为1,这表明多发性硬化的大部分家族聚集性源于遗传,而非共享环境的结果。
A similar analysis can be carried out for quantitative traits such as blood pressure: No correlation exists between a child’s blood pressure and that of the child’s adopted siblings, in contrast to the positive correlation for blood pressure of biological siblings, all living in the same household.对于血压等数量性状可进行类似分析:儿童血压与其领养兄弟姐妹的血压之间不存在相关性,而生活在同一家庭中的生物学兄弟姐妹的血压则呈正相关。
Distinguishing Between Genetic and Environmental Influences Using Twin Studies Of all methods used to separate genetic and environmental influences, geneticists have historically relied most heavily on twin studies.利用双生子研究区分遗传与环境影响 在用于区分遗传与环境影响的所有方法中,遗传学家历来最依赖双生子研究。
Twinning MZ and DZ twins are “experiments of nature” that provide an excellent opportunity to separate environmental and genetic influences on phenotypes in humans.同卵双生子和异卵双生子是“自然实验”,为区分环境与遗传对人类表型的影响提供了绝佳机会。
MZ twins arise from the cleavage of a single fertilized zygote into two separate zygotes early in embryogenesis (see Chapter 14).同卵双生子源于单个受精卵在胚胎发育早期分裂成两个独立的受精卵(见第14章)。
They occur in ~0. 3% of all births, without large differences among different populations.其在所有分娩中发生率约为0.3%,不同人群间无显著差异。
At the time the zygote cleaves in two, MZ twins start out with identical genotypes at every locus and are therefore often thought of as having identical genomes.当受精卵一分为二时,同卵双生子在每个基因座上起始具有相同的基因型,因此通常被认为拥有相同的基因组。
In contrast, DZ twins arise from the simultaneous fertilization of two eggs by two sperm; genetically, DZ twins are siblings who share a womb and, like all siblings, share, on average, 50% of the alleles at all loci.相比之下,异卵双生子源于两个精子同时使两个卵子受精;从遗传学上看,异卵双生子是共享子宫的兄弟姐妹,与所有兄弟姐妹一样,平均共享所有基因座上50%的等位基因。
DZ twins are of the same sex half the time and of opposite sex the other half.异卵双生子半数情况下性别相同,另一半情况下性别不同。
In contrast to MZ twins, DZ twins occur with a frequency that varies as much as fivefold in individuals with different ancestries - e. g., 0. 2% among individuals with Asian ancestry and 1% of births in parts of Africa.与同卵双生子不同,异卵双生子的发生率在不同祖先个体中差异可达五倍——例如,亚洲血统个体中为0.2%,而非洲部分地区出生中为1%。
Disease Concordance in MZ and DZ Twins When twins have the same disease, they are said to be concordant for that disorder.同卵与异卵双生子的疾病一致性 当双生子患有相同疾病时,称该疾病具有一致性。
Conversely, when only one member of the pair of twins is affected and the other is not, the relatives are discordant for the disease.相反,当双生子中仅一人患病而另一人未患病时,则称该疾病具有不一致性。
An examination of how frequently MZ twins are concordant for a disease is a powerful method for determining whether genotype alone is sufficient to produce a particular disease.检测同卵双生子对某一疾病的一致率是判断基因型本身是否足以导致特定疾病的强有力方法。
The differences between a disease that is mendelian from one that shows complex inheritance are immediately evident.孟德尔疾病与复杂遗传疾病之间的差异一目了然。
Using sickle cell disease (Case 42) as an example of a mendelian disorder, if one MZ twin has sickle cell disease, the other twin will always have the disease as well.以镰状细胞病(病例42)作为孟德尔疾病的例子,若一例同卵双生子患有镰状细胞病,另一例双生子也必然患病。
In contrast, as an example of a multifactorial disorder, when one MZ twin has type 1 diabetes mellitus (previously known as insulin-dependent or juvenile diabetes) the other twin will also have type 1 diabetes only ~40% of such twin pairs.相反,作为多因素疾病的例子,当一例同卵双生子患有1型糖尿病(先前称为胰岛素依赖型或青少年糖尿病)时,另一例双生子仅约40%的双生子对也会患有1型糖尿病。
Disease concordance of less than 100% in MZ twins is strong evidence that nongenetic factors play a role in the disease.同卵双生子疾病一致率低于100%是强烈证据,表明非遗传因素在该疾病中起作用。
Such factors could include environmental influences, such as exposure to infection or diet, as well as other effects, such as somatic variation, effects of aging, or epigenetic changes in gene expression in one twin compared with the other.这些因素可能包括环境影响(如感染暴露或饮食)以及其他效应,例如一例双生子相对于另一例的体细胞变异、衰老效应或基因表达的表观遗传改变。
MZ and same-sex DZ twins share a common intrauterine environment and biological sex and are usually reared together in the same household by the same parents.同卵双生子和同性异卵双生子共享相同的子宫内环境和生物性别,通常由同一父母在同一家庭中共同抚养。
Thus a comparison of concordance for a disease between MZ and same-sex DZ twins shows how frequently disease occurs when relatives who experience the same prenatal and often the same postnatal environment have the same alleles at every locus (MZ twins) and how often it occurs when they share 50% of alleles in common (DZ twins).因此,比较同卵双生子和同性异卵双生子之间的疾病一致率,可以揭示当经历相同产前并常为相同产后环境的亲属在每个基因座上有相同等位基因(同卵双生子)时疾病发生频率,以及当共享50%共同等位基因(异卵双生子)时疾病发生频率。
Greater concordance in MZ versus DZ twins is strong evidence of a genetic component to the disease, as shown in of Affected Parents 0 1 2 None 0. 1 3 34 One sibling 3 11 40 Two siblings 8 19 45 One sibling and one second-degree relative 6 16 43 One sibling and one third-degree relative 4 14 44 CL(P), Cleft lip with or without cleft palate.同卵双生子相较异卵双生子的一致率更高,是疾病具有遗传成分的强烈证据,如受影响父母数量为0、1、2时,无兄弟姐妹分别为0.1、3、34;一个兄弟姐妹分别为3、11、40;两个兄弟姐妹分别为8、19、45;一个兄弟姐妹加一位二级亲属分别为6、16、43;一个兄弟姐妹加一位三级亲属分别为4、14、44所示。CL(P)即唇裂伴或不伴腭裂。
36/53
Estimating Heritability from Twin Studies Just as twins are be used to separate the roles of genes and environment in qu…
Ch9 — Segment 36
Estimating Heritability from Twin Studies Just as twins are be used to separate the roles of genes and environment in qualitative disease traits, twins are also used to estimate the heritability of a quantitative trait using the correlation in the values of a physiologic measurement in MZ and DZ twins.从双生子研究中估计遗传度:如同双生子被用于在定性疾病的性状中分离基因和环境的作用一样,双生子也通过利用同卵双生(MZ)和异卵双生(DZ)双生子中某项生理测量值的相关性来估计数量性状的遗传度。
If one assumes that the alleles affecting the trait exert their effect additively (which is certainly overly simplistic and probably incorrect in many cases), MZ twins, who share 100% of their alleles, have twice the amount of allele sharing compared to DZ twins, who share 50% of their alleles on average.如果假设影响该性状的等位基因以相加方式发挥效应(这当然过于简单化,且在许多情况下可能不正确),那么共享100%等位基因的MZ双生子,其等位基因共享程度是平均共享50%等位基因的DZ双生子的两倍。
H2, introduced earlier in this chapter, can therefore be approximated by taking twice the difference in the correlation coefficient r for a quantitative trait between MZ twins (r MZ) and r between same-sex DZ twins (r DZ) (as given by Falconer’s formula): H r r MZ DZ 2 2 () If the variability of the trait is determined chiefly by environment, the correlation within pairs of DZ twins will be similar to that seen between pairs of MZ twins; there will be little difference in the value of r for MZ and DZ twins.因此,本章前面引入的H²可以通过取MZ双生子间数量性状相关系数r(r MZ)与同性别的DZ双生子间r(r DZ)之差的两倍来近似(如Falconer公式所示):H² = 2(r MZ - r DZ)。如果该性状的变异主要由环境决定,则DZ双生子对内的相关性将接近于MZ双生子对间的相关性;MZ与DZ双生子的r值差异很小。
Thus, r MZ − r DZ = ≈0, and H2 will approach 0.因此,r MZ − r DZ ≈ 0,H²将趋近于0。
At the other extreme, however, if the variability is determined exclusively by genetic makeup, the correlation coefficient r between MZ pairs will approach 1, whereas r between DZ twins will be half of that.然而,在另一个极端,如果变异完全由遗传构成决定,则MZ双生子对间的相关系数r将趋近于1,而DZ双生子之间的r则为该值的一半。
Now, r MZ − r DZ = ≈12, and therefore H2 will be approximately 2 × (12) = 1.此时,r MZ − r DZ ≈ 1/2,因此H²将约为2 × (1/2) = 1。
Twins Reared Apart Although a rare occurrence, twins are sometimes separated at birth for social reasons and placed in different homes, thus providing an opportunity to observe individuals of identical or half-identical genotypes reared in different environments.分开抚养的双生子:虽然罕见,但双生子有时因社会原因在出生时被分离并安置于不同家庭,从而提供了观察相同或半相同基因型个体在不同环境中成长的机会。
Such studies have been used primarily in research on psychiatric disorders, substance use, and eating disorders, in which strong environmental influences within the family are believed to play a role in the development of disease.此类研究主要用于精神疾病、物质使用和进食障碍的研究中,这些疾病被认为家庭内部的强烈环境影响在疾病发生中起作用。
For example, in one study of obesity, the body mass index (BMI; weight/height 2, expressed in kg/m 2) was measured in MZ and DZ twins reared in the same household versus those reared apart ( Although the average BMI among MZ or DZ twins was similar, regardless of whether they were reared together or apart, the pairwise correlation for BMI between a pair of twins was much higher for MZ than for DZ twins.例如,在一项关于肥胖的研究中,测量了在同一家庭抚养与分开抚养的MZ和DZ双生子的身体质量指数(BMI;体重/身高²,以kg/m²表示)。尽管MZ或DZ双生子的平均BMI相似,不论他们是共同抚养还是分开抚养,但MZ双生子对之间的BMI配对相关性远高于DZ双生子。
Also interesting is that the higher correlation between MZ and DZ twins was independent of whether the twins were reared together or apart, which suggests that genetics has a highly significant impact on adult weight and consequently on the risk for obesity and its complications.同样有趣的是,MZ与DZ双生子之间的较高相关性独立于双生子是否共同或分开抚养,这表明遗传学对成年人体重具有高度显著的影响,进而影响肥胖及其并发症的风险。
Limitations of Familial Aggregation and Heritability Estimates From Family and Twin Studies Potential Sources of Bias There are a number of difficulties in measuring and interpreting λs.来自家系和双生子研究的家族聚集性与遗传度估计的局限性——潜在偏倚来源:测量和解释λs存在诸多困难。
One is that studies of familial aggregation of disease are subject to various forms of bias.其一,疾病家族聚集性研究易受各种形式的偏倚影响。
There is ascertainment bias, which arises when families with more than one affected sibling are more likely to come to a researcher’s attention, thereby inflating the sibling recurrence risk λs.存在确定偏倚,即当家庭中有不止一个患病同胞时更易引起研究者的注意,从而高估了同胞复发风险λs。
Ascertainment bias is DZ, Dizygotic; MZ, monozygotic.确定偏倚是:DZ,异卵双生;MZ,同卵双生。
Data from Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin’s principles and practice of medical genetics, ed 3, Edinburgh, 1997, Churchill Livingstone; King RA, Rotter JI, Motulsky AG: The genetic basis of common diseases, Oxford, England, 1992, Oxford University Press; Tsuang MT: Recent advances in genetic research on schizophrenia, J Biomed Sci 5:28–30, 1998. of Pairs BMI* Pairwise Correlation No. of Pairs BMI* Pairwise Correlation Monozygotic Apart 49 24. 8 ± 2. 4 0. 70 44 24. 2 ± 3. 4 0. 66 Together 66 24. 2 ± 2. 9 0. 74 88 23. 7 ± 3. 5 0. 66 Dizygotic Apart 75 25. 1 ± 3. 0 0. 15 143 24. 9 ± 4. 1 0. 25 Together 89 24. 6 ± 2. 7 0. 33 119 23. 9 ± 3. 5 0. 27 *Mean ± 1 SD.数据来源:Rimoin DL, Connor JM, Pyeritz RE: Emery and Rimoin’s principles and practice of medical genetics, ed 3, Edinburgh, 1997, Churchill Livingstone; King RA, Rotter JI, Motulsky AG: The genetic basis of common diseases, Oxford, England, 1992, Oxford University Press; Tsuang MT: Recent advances in genetic research on schizophrenia, J Biomed Sci 5:28–30, 1998。 配对BMI* 配对相关性 配对数量 BMI* 配对相关性 同卵双生 分开 49 24.8 ± 2.4 0.70 44 24.2 ± 3.4 0.66 共同 66 24.2 ± 2.9 0.74 88 23.7 ± 3.5 0.66 异卵双生 分开 75 25.1 ± 3.0 0.15 143 24.9 ± 4.1 0.25 共同 89 24.6 ± 2.7 0.33 119 23.9 ± 3.5 0.27 *均值±1标准差。
BMI, Body mass index; DZ, dizygotic; MZ, monozygotic.BMI,身体质量指数;DZ,异卵双生;MZ,同卵双生。
Data from Stunkard AJ, Harris JR, Pedersen NL, et al: The body-mass index of twins who have been reared apart, NEJM 322:1483–1487, 1990.数据来源:Stunkard AJ, Harris JR, Pedersen NL, et al: The body-mass index of twins who have been reared apart, NEJM 322:1483–1487, 1990。
37/53
Complex Inheritance of Common Multifactorial Disorders 157 also a problem for twin studies.
Ch9 — Segment 37
Complex Inheritance of Common Multifactorial Disorders 157 also a problem for twin studies.常见多因素疾病的复杂遗传也是双生子研究的一个问题。
Many studies rely on asking one twin with a particular disease to recruit the other twin to participate in a study (volunteer-based ascertainment), rather than ascertaining them first as twins through a twin registry and only then examining their health status (population-based ascertainment).许多研究依赖于让患有特定疾病的一名双生子招募另一名双生子参与研究(基于志愿者的确认),而不是首先通过双生子登记册确认他们为双生子,然后才检查他们的健康状况(基于人群的确认)。
Volunteer-based ascertainment can give biased results because twins, particularly MZ twins who may be emotionally close, are more likely to volunteer if they are concordant than if they are not, which inflates the concordance rate.基于志愿者的确认可能产生有偏倚的结果,因为双生子,尤其是可能情感亲密的同卵双生子,如果患病一致则比不一致时更可能自愿参与,这会抬高一致率。
Similarly, because case-control studies of family history often rely, for practical reasons, on taking a history from the proband rather than examining all the relatives directly, there may be recall bias, in which a proband may be more likely to know of family members with the same or similar disease.类似地,由于家族史的病例对照研究出于实际原因常常依赖于从先证者获取病史,而非直接检查所有亲属,可能存在回忆偏倚,即先证者更可能了解患有相同或相似疾病的家庭成员。
Such biases will inflate the level of familial aggregation.此类偏倚会抬高家族聚集程度。
Other difficulties arise in measuring and interpreting heritability.在测量和解释遗传度时会出现其他困难。
The same trait may yield different measurements of heritability in different populations because of different allele frequencies or diverse environmental conditions.同一性状在不同人群中可能产生不同的遗传度测量值,因为等位基因频率不同或环境条件多样。
For example, heritability measurements of height would be lower when measured in a population with widespread famine that stunts growth in childhood as compared to the same population after food becomes plentiful.例如,与同一人群在食物充足后相比,在普遍存在饥荒导致儿童期生长迟缓的人群中测量身高遗传度会较低。
Heritability of a trait should therefore not be thought of as an intrinsic, universally applicable measure of “how genetic” the trait is, because it depends on the population and environment in which the estimate is being made.因此,不应将性状的遗传度视为该性状“遗传程度”的固有、普遍适用的度量,因为它取决于进行估计时的人群和环境。
Although heritability estimates are still made in genetic research, most geneticists consider them to be only preliminary but convenient estimates of the role of genetic variation in phenotypic variation.尽管遗传学研究仍在进行遗传度估计,但大多数遗传学家认为它们只是对遗传变异在表型变异中作用的初步但便捷的估计。
Potential Genetic or Epigenetic Differences Despite the evident power of twin studies, one must caution against thinking of such studies as perfectly controlled experiments that compare individuals who share either half or all of their genetic variation and are exposed either to the same or to different environments.潜在的遗传或表观遗传差异 尽管双生子研究具有明显的优势,但必须警惕不要将这类研究视为完美对照的实验,即比较共享一半或全部遗传变异并且暴露于相同或不同环境的个体。
Studies of MZ twins assume the twins are genetically identical.对同卵双生子的研究假定双生子在遗传上是相同的。
Although this is mostly true, genotype and gene expression patterns may come to differ between MZ twins because of genetic or epigenetic changes that occur after the cleavage event that produced the MZ twin embryos.尽管这基本正确,但同卵双生子之间的基因型和基因表达模式可能因产生同卵双生子胚胎的卵裂事件后发生的遗传或表观遗传变化而出现差异。
For example, genotype may differ due to somatic rearrangements and/or rare somatic mutations that occur after the cleavage event (see Chapter 3).例如,基因型差异可能因卵裂事件后发生的体细胞重排和/或罕见体细胞突变所致(见第3章)。
Epigenetic changes may occur in response to environmental or stochastic factors, thus leading to differences in gene expression between MZ twins.表观遗传变化可能因环境或随机因素而发生,从而导致同卵双生子之间基因表达的差异。
(Female MZ twins have an additional source of variability because of the stochastic nature of X inactivation patterns in various tissues, as presented in Chapter 6.) Other Limitations Another problem may arise when assuming that the environmental exposure of MZ and DZ twins has been held constant when they are reared together but not when twins are reared apart.(雌性同卵双生子由于X失活模式在不同组织中的随机性而具有额外的变异来源,如第6章所述。) 其他局限性 另一个问题可能出现于假定同卵双生子和异卵双生子的环境暴露在共同抚养时保持恒定,但在分开抚养时则不然。
Environmental exposures, including even intrauterine environment, may vary for twins reared in the same family.环境暴露,甚至包括宫内环境,对于在同一家庭抚养的双生子可能有所不同。
For example, MZ twins frequently share a placenta, and there may be a disparity between the twins in blood supply, intrauterine development, and birthweight.例如,同卵双生子常共享一个胎盘,双生子之间在血液供应、宫内发育和出生体重上可能存在差异。
For late-onset diseases, such as neurodegenerative disease of late adulthood, the assumption that MZ and DZ twins are exposed to similar environments throughout their adult lives becomes less and less valid, and thus a difference in concordance provides less strong evidence for genetic factors in disease causation.对于晚发疾病,例如成年晚期的神经退行性疾病,同卵双生子和异卵双生子在整个成年期暴露于相似环境的假设越来越不成立,因此一致率的差异为疾病病因中的遗传因素提供的证据力度较弱。
Conversely, one assumes that by determining disease concordance in MZ twins reared apart, one is measuring the effect of different environments on the same genotype.相反,人们假设通过
However, the environment of twins reared apart may actually not be as different as one might suppose.[TL:missing]
Thus no twin study is a perfectly controlled assessment of genetic versus environmental influence.[TL:missing]
Finally, caution is necessary when generalizing from twin studies.[TL:missing]
The most extreme situation would be when the phenotype being studied is only sometimes genetic in origin; that is, nongenetic phenocopies may exist.[TL:missing]
If genotype alone causes the disease in half the pairs of twins (MZ twin concordance of 100%) in your sample and a nongenetic phenocopy affects only one twin of the other half of twin pairs in your sample (MZ twin concordance of 0%), twin studies will show an intermediate level of 50% concordance that really applies to neither form of the disease.[TL:missing]
Genome-wide Association Studies Lei Sun and Wei Deng Complex diseases and traits are influenced by a combination of genetic and environmental risk factors.[TL:missing]
In the last decade, genome-wide association studies (GWAS) have been particularly fruitful in identifying disease susceptibility loci and biologic pathways, elucidating the genetic architecture of many complex diseases.[TL:missing]
The concept of GWAS is elegantly simple: examining each individual genetic variant – usually a biallelic locus (which is typically either a single nucleotide polymorphism [SNP] or an insertion deletion variant or indel) – for evidence of association with a disease phenotype or quantitative trait.[TL:missing]
Individual variant results can point to specific regions of the genome that are associated with a disease or trait.[TL:missing]
Sometimes the association signal implicates a protein-altering variant, which has a simple interpretation in terms of biologic and functional consequences, but most of the time the association signal falls among the noncoding regions of the genome.[TL:missing]
Many approaches have been developed to map associated genetic variants[TL:missing]
38/53
to their effector genes, but for complex disease the most sophisticated approaches only guess the correct gene a little …
Ch9 — Segment 38
to their effector genes, but for complex disease the most sophisticated approaches only guess the correct gene a little over half of the time.对于复杂疾病,即使是最精密的解析方法,也只有在略过半的次数中才能正确推断出效应基因。
These individual association results, over the entire genome, reveal the landscape of phenotype-genotype associations, where the association evidence is often summarized pictorially by the famous Manhattan plot of –log 10 p-values .这些覆盖全基因组的单个关联结果揭示了表型-基因型关联的整体图景,其中关联证据常通过著名的曼哈顿图(以-log10 p值为纵轴)进行图形化总结。
More recently, some have proposed a Brisbane plot enumerating the number of independently associated variants within a specified window size (useful for GWAS with many findings), or a Miami plot showing results from two similar GWASs above and below the horizontal axis (i. e., a reflection) (Figs. 9. 5 and 9. 6).最近,有人提出了布里斯班图(Brisbane plot),用于列举指定窗口大小内独立关联变异体的数量(适用于有多项发现的GWAS);或者迈阿密图(Miami plot),以水平轴为镜像轴,上下分别展示两个相似GWAS的结果(见图9.5和图9.6)。
Polygenic Risk Scores Lei Sun and Wei Deng In addition to discovery of novel association signals, GWAS have enabled the estimation of genetic effects associated with millions of individual SNPs across a range of complex diseases.多基因风险评分(Lei Sun 与 Wei Deng)。除了发现新的关联信号外,GWAS还使得能够估计一系列复杂疾病中数百万个单个SNP所关联的遗传效应。
The magnitude of these effects for each genetic variant is typically small compared to established clinical risk factors.每个遗传变异体的效应幅度通常小于已确立的临床风险因素。
As the higher risk allele at each SNP locus only moderately increases the risk of disease, it is natural to construct a composite score that can potentially capture the overall genetic burden across the genome.由于每个SNP位点的较高风险等位基因仅适度增加疾病风险,构建一个能够潜在捕捉全基因组整体遗传负担的综合评分是自然的思路。
This is indeed the idea behind the PRS: a weighted sum of the numbers of the risk alleles of all associated SNPs, where the weights are derived from effect size estimates from GWAS.这确实是PRS背后的理念:所有关联SNP的风险等位基因数量的加权和,其中权重来自GWAS的效应量估计。
The earlier polygenic risk research focused on association, and one of the most prominent early examples of PRS was in schizophrenia.早期的多基因风险研究侧重于关联,其中最突出的早期PRS实例之一是在精神分裂症中。
There, a large number of SNPs individually below the genome-wide detection threshold (of p-value &lt;5 × 10−8, which is the threshold typically used to allow for millions of tests for common variants distributed throughout the genome) was used to construct a PRS highly associated with schizophrenia status.在那项研究中,大量单个SNP虽低于全基因组检测阈值(p值<5×10⁻⁸,这是为允许对分布于全基因组的数百万个常见变异进行检验而通常采用的阈值),但仍被用于构建与精神分裂症状态高度关联的PRS。
More recent PRS research has focused on the utility of PRS for prediction of individuals at risk of disease, on the ability of PRS to help refine diagnoses by separating different types of cases, and on the ability of PRS to aid in the selection of optimal treatment regimes.更近期的PRS研究聚焦于PRS在预测疾病风险个体方面的效用、PRS通过区分不同类型病例来帮助精细化诊断的能力,以及PRS辅助选择最优治疗方案的能力。
For example, a higher PRS is capable of identifying individuals with a higher disease risk: Individuals with PRS values in the top 1% of the distribution have at least a threefold increase in risk of developing coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, or breast cancer.例如,较高的PRS能够识别出疾病风险较高的个体:分布于前1%的PRS值个体患冠状动脉疾病、心房颤动、2型糖尿病、炎症性肠病或乳腺癌的风险至少增加三倍。
Indeed, PRS of this magnitude could have many clinical utilities: They might inform screening strategies, motivate health-related behavior changes, and potentially help identify treatment targets for precision medicine.确实,如此量级的PRS可能具有多种临床用途:它们可为筛查策略提供依据、促进健康相关行为改变,并可能有助于确定精准医学的治疗靶点。
PRS for breast cancer was the first to be implemented in the Can Risk/BOADICEA risk algorithm, which was used clinically to predict individuals at higher risk of breast cancer, and became available for prediction of coronary artery disease by at least one private company in 2022 (Color).乳腺癌的PRS率先被纳入CanRisk/BOADICEA风险算法,该算法在临床上用于预测乳腺癌高风险个体,并在2022年至少有一家私营公司(Color)将其用于预测冠状动脉疾病。
In the case of diagnostic refinement, PRS was found to distinguish between type 1 and type 2 diabetes, which had previously been assigned primarily based on age of onset.在诊断精细化方面,PRS被发现能够区分1型与2型糖尿病,而此前两者主要依据发病年龄进行归类。
The construction of a good PRS, predictive of the risk of a disease, requires a systematic approach to (1) powerfully identify associated SNPs, (2) accurately estimate their genetic effects on the disease, and (3) ensure the PRS is applied to the appropriate people (e. g., whose age, ancestry, and clinical risk profiles are similar to those in whom the genetic effects were originally estimated).构建一个能够预测疾病风险的优质PRS需要系统性的方法:(1) 强效识别关联SNP;(2) 准确估计这些SNP对疾病的遗传效应;(3) 确保PRS应用于合适的人群(例如,其年龄、遗传背景和临床风险特征与最初估计遗传效应时的人群相似)。
The ideal PRS construction should include all disease-associated SNPs and no others; a realistic PRS typically not only misses some of the truly associated SNPs but includes false positives due to limited power of GWAS.理想的PRS构建应包含所有疾病相关SNP而不包含其他SNP;而实际的PRS通常不仅遗漏部分真正关联的SNP,还因GWAS效力有限而包含假阳性结果。
The low power of GWAS is a direct result of moderate-to-weak effects of individual SNPs on a complex disease with polygenic inheritance.GWAS效力低是由于在多基因遗传的复杂疾病中,单个SNP的效应为中至弱度所致。
The estimation of effect can be challenging, as it is often biased toward discovery – otherwise known as the winner’s curse – where the estimated genetic effect of an SNP is inflated relative to the true value.效应估计可能具有挑战性,因为其常偏向于发现——即所谓的“赢家诅咒”——即SNP的遗传效应估计值相对于真实值被夸大。
Finally, to avoid data dredging (also known as p-hacking), for both (1) and (2) above, a (discovery) sample of individuals is used, which is independent of the (target) sample of individuals where the PRS is evaluated for prediction accuracy.最后,为避免数据挖掘(又称p值操纵),上述(1)和(2)均使用一个(发现)样本,该样本独立于用于评估PRS预测准确性的(目标)样本。
While statistical methodology in capturing the polygenic risk associated with diseases has progressed significantly, for many complex diseases and quantitative traits the amount of phenotypic variance explained by PRS derived from GWASs is modest.尽管捕捉疾病相关多基因风险的统计方法学取得了显著进展,但对于许多复杂疾病和数量性状,由GWAS推导的PRS所解释的表型方差量仍然有限。
Thus for most common diseases we are still evaluating the value of PRS in the clinic and in other settings where they might influence the health and wellbeing of individuals.因此,对于大多数常见疾病,我们仍在评估PRS在临床及其他可能影响个体健康与福祉的环境中的价值。
With the availability of biobank-scale studies and more sophisticated PRS methods, we expect incremental improvement in the predictiveness of PRS, particularly for diseases with few clinical predictors.随着生物银行规模研究的开展以及更精密的PRS方法的出现,我们预期PRS的预测能力将逐步提升,尤其对于缺乏临床预测因素的疾病。
A major methodologic issue in the construction of a PRS is population heterogeneity, also known as the transportability issue.PRS构建中的一个主要方法学问题是人群异质性,也称为可迁移性问题。
Because genetic effect sizes may vary according to the ancestral background of the study population, the PRS weights derived in one population might not translate well into another.由于遗传效应量可能随研究人群的祖先背景而变化,因此在一个群体中推导出的PRS权重可能无法很好地应用于另一个群体。
Another major gap is the lack of methods to include SNPs from the X chromosome, thus missing 5% of the haploid genome.另一个主要缺口是缺乏纳入X染色体SNP的方法,从而遗漏了单倍体基因组的5%。
Finally, the genetic risk for a disease, predicted using PRS, has a theoretical upper bound that depends on the population-specific heritability of the disease.最后,利用PRS预测的疾病遗传风险具有一个理论上限,该上限取决于该疾病的人群特异性遗传度。
This is not to say, however, that heritability limits the absolute disease risk for an individual.然而,这并不是说遗传度限制了个体的绝对疾病风险。
Consider the well-known example of familial breast cancer, whereby individuals with pathogenic variants in the BRCA1 gene can have a substantially increased risk of early-onset cancer, even though these relatively highly penetrant variants only account for a small number of breast cancer cases overall and only explain a small proportion of cancer heritability in the populations where they occur.以家族性乳腺癌这一知名实例为例:携带BRCA1基因致病性变异的个体患早发性癌症的风险显著升高,尽管这些相对高外显率的变异在总体上仅占乳腺癌病例的一小部分,且在其发生人群中仅解释了癌症遗传度的一小部分。
39/53
Complex Inheritance of Common Multifactorial Disorders 159 70 60 50 40 30 IL1RL1 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 2…
Ch9 — Segment 39
Complex Inheritance of Common Multifactorial Disorders 159 70 60 50 40 30 IL1RL1 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 20 -log 10(p-value) 16 12 8 4 0 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 The association analysis compares genotypes at millions of genetic variants between asthma cases, defined on the basis of diagnosis codes in their individual medical records, and controls.常见多因素疾病的复杂遗传 159 70 60 50 40 30 IL1RL1 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 20 -log 10(p-value) 16 12 8 4 0 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 关联分析比较了哮喘病例(根据其个人病历中的诊断代码定义)与对照之间数百万个遗传变异位的基因型。
The results of each comparison are summarized in a p-value which is plotted in the graph.每次比较的结果以p值汇总,并在图中绘制。
The top few association signals have been labeled with the name of the nearest gene.前几个关联信号已用最近基因的名称进行标注。
(The figure is based on the UK Biobank association analyses reported in Taliun D, Harris DN, Kessler MD, et al: Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program, Nature 590: 290–299, 2021. https:// doi. org/10. 1038/s 41586-021-03205-y.) Meanwhile, the clinical use of PRS is not without caveats.该图基于Taliun D, Harris DN, Kessler MD等人在《自然》杂志590: 290–299, 2021中报告的英国生物银行关联分析(https:// doi. org/10. 1038/s 41586-021-03205-y),同时,多基因风险评分(PRS)的临床应用并非没有注意事项。
A main controversy is the potential health disparity resulting from the fact that most large GWAS and PRS have been conducted and tested only in large samples of European ancestry.一个主要争议是潜在的健康差异,这是由于大多数大型GWAS和PRS仅在欧洲血统的大样本中进行和测试所致。
Second, the interpretation of a PRS score (e. g., relative vs. absolute risk) can sometimes be a barrier in clinical settings, requiring continued education of clinicians to help them to better disclose results to families.其次,PRS评分(例如相对风险与绝对风险)的解释有时可能成为临床环境中的障碍,需要持续教育临床医生,帮助他们更好地向家庭披露结果。
Third, it remains an open question how best to combine PRS and modifiable clinical risk factors to inform disease prevention.第三,如何最好地结合PRS和可改变的临床风险因素来指导疾病预防仍然是一个未解决的问题。
Finally, standardized reporting of PRS results will facilitate knowledge translation.最后,PRS结果的标准化报告将促进知识转化。
EXAMPLES OF COMMON MULTIFACTORIAL DISEASES WITH A GENETIC CONTRIBUTION In this section and the next, we turn to considering examples of several common conditions that illustrate general concepts of multifactorial disorders and their complex inheritance, as summarized in the accompanying box (see 1)..具有遗传贡献的常见多因素疾病示例 在本节和下一节中,我们将考虑几个常见疾病的例子,这些例子阐述了多因素疾病及其复杂遗传的一般概念,如随附框中所总结(见1)。
CHARACTERISTICS OF INHERITANCE OF COMPLEX DISEASES Genetic variation contributes to diseases with complex inheritance, but these diseases are not single-gene disorders and do not demonstrate a simple mendelian pattern of inheritance.复杂疾病遗传的特征 遗传变异导致具有复杂遗传的疾病,但这些疾病不是单基因疾病,也不表现出简单的孟德尔遗传模式。
Diseases with complex inheritance often demonstrate familial aggregation because close relatives of an affected individual are likely to also share disease-predisposing alleles.具有复杂遗传的疾病通常表现出家族聚集性,因为患病个体的近亲也可能共享易感等位基因。
Diseases with complex inheritance are more common among the close relatives of a proband and become less common in relatives who are less closely related and therefore share fewer predisposing alleles.具有复杂遗传的疾病在先证者的近亲中更常见,而在关系较远(因此共享较少易感等位基因)的亲属中则较少见。
Greater concordance for disease is expected among monozygotic versus dizygotic twins.同卵双胞胎比异卵双胞胎预期有更高的一致性。
However, pairs of relatives who share disease-­ predisposing genotypes at relevant loci may still be discordant for phenotype (show lack of penetrance) because of the crucial role of nongenetic factors in disease causation.然而,在相关位点共享易感基因型的亲属对可能仍然在表型上不一致(表现出外显率不全),因为非遗传因素在疾病成因中起着关键作用。
The most extreme examples of lack of penetrance despite identical genotypes are discordant monozygotic twins.尽管基因型相同但外显率不全的最极端例子是不一致的同卵双胞胎。
40/53
IL1RL1 0 0 8 16 24 32 40 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 4 8 12 16 20 30 40 50 60 70 D2HGDH T…
Ch9 — Segment 40
IL1RL1 0 0 8 16 24 32 40 1 3 5 7 9 11 13 15 17 19 21 X 2 4 6 8 10 12 14 16 18 20 22 4 8 12 16 20 30 40 50 60 70 D2HGDH TSLP HLA-DQB1 IL33 GATA3 SMAD3 -log 10(p-value) -log 10(p-value) The association analysis on the top panel compares genotypes at millions of genetic variants between asthma cases, defined on the basis of diagnosis codes in their individual medical records, and controls.上方面板的关联分析比较了哮喘病例(根据个体医疗记录中的诊断代码定义)与对照组之间数百万个遗传变异位点的基因型。
The analysis on the bottom panel compares genotypes at the same variants between individuals with nasal polyps and controls.下方面板的分析比较了鼻息肉患者与对照组之间相同变异位点的基因型。
Note the many shared signals between the two analyses.注意两次分析之间存在许多共享的信号。
The results of each comparison are summarized in a p-value, which is plotted in the graph.每次比较的结果以p值汇总,并绘制在图中。
The top few association signals have been labeled with the name of the nearest gene in the asthma portion of the plot.图中哮喘部分的前几个显著关联信号已标记最邻近基因的名称。
(The figure is based on the UK biobank association analyses reported in Taliun D, Harris DN, Kessler MD, et al: Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program, Nature 590: 290–299, 2021. 41586-021-03205-y.)(此图基于英国生物样本库关联分析报告,该报告发表于Taliun D, Harris DN, Kessler MD等:Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program, Nature 590: 290–299, 2021. 41586-021-03205-y。)
41/53
Complex Inheritance of Common Multifactorial Disorders 161 3-m individuals.
Ch9 — Segment 41
Complex Inheritance of Common Multifactorial Disorders 161 3-m individuals.常见多因素疾病的复杂遗传 161 300万个体。
Each dot shows one of the 12,111 independent genome-wide significant variants associated with height.每个点显示与身高相关的12,111个独立的全基因组显著变异中的一个。
Density was calculated as the number of other independent associated variants within 100 kb.密度计算为100 kb内其他独立相关变异的数量。
(Figure from height GWAS in Yengo L, Vedantam S, Marouli E, et al: A saturated map of common genetic variants associated with human height, Nature 610(7933):704–712, 2022. https:doi. org/10. 1038/ s 41586-022-05275-y.) Multifactorial Congenital Malformations Many common congenital malformations, occurring as isolated defects and not as part of a syndrome, are multifactorial and demonstrate complex inheritance ( Among these, congenital heart malformations are some of the most common and serve to illustrate the current state of understanding of other categories of congenital malformation.(图来自Yengo L, Vedantam S, Marouli E, 等:与人类身高相关的常见遗传变异饱和图谱,Nature 610(7933):704–712, 2022. https://doi.org/10.1038/s41586-022-05275-y.)多因素先天性畸形 许多常见的先天性畸形,作为孤立缺陷发生而非综合征的一部分,是多因素的并表现出复杂遗传(在这些畸形中,先天性心脏畸形是最常见的一些,并有助于说明对其他类别先天性畸形当前的理解状态。
Congenital heart defects (CHDs) occur at a frequency of ~4 to 8 per 1000 births.先天性心脏缺陷(CHDs)的发生频率约为每1000例出生4至8例。
They are a heterogeneous group, caused in some cases by single-gene or chromosomal mechanisms and in others by exposure to teratogens, such as rubella infection or maternal diabetes.它们是一组异质性疾病,部分病例由单基因或染色体机制引起,其他病例由暴露于致畸物(如风疹感染或母亲糖尿病)引起。
The cause is usually unknown, however, and the majority of cases are believed to be multifactorial in origin.然而,病因通常未知,且大多数病例被认为源于多因素。
There are many types of CHDs, with different population incidences and empirical risks.CHD有多种类型,具有不同的人群发病率和经验风险。
It is known that when heart defects recur in a family, however, the affected children do not necessarily have exactly the same anatomic defect but instead show recurrence of lesions that are similar with regard to developmental mechanisms (see Chapter 14).然而,已知当心脏缺陷在一个家族中复发时,受累儿童不一定具有完全相同的解剖缺陷,而是显示出在发育机制方面相似的病变复发(见第14章)。
By using developmental mechanisms as a classification scheme, five main groups of CHDs can be distinguished: Flow lesions Defects in cell migration 4–1. 7 Cleft palate 0. 4 Congenital dislocation of hip 2* Congenital heart defects 4–8 Ventricular septal defect 1. 7 Patent ductus arteriosus 0. 5 Atrial septal defect 1. 0 Aortic stenosis 0. 5 Neural tube defects 2–10 Spina bifida and anencephaly Variable Pyloric stenosis 1,† 5* *Per 1000 males. †Per 1000 females.通过使用发育机制作为分类方案,可以区分出五大类CHD:血流病变 细胞迁移缺陷 4-1.7 腭裂 0.4 先天性髋关节脱位 2* 先天性心脏缺陷 4-8 室间隔缺损 1.7 动脉导管未闭 0.5 房间隔缺损 1.0 主动脉瓣狭窄 0.5 神经管缺陷 2-10 脊柱裂和无脑畸形 可变 幽门狭窄 1,† 5* *每1000名男性。†每1000名女性。
Data from Carter CO: Genetics of common single malformations, Br Med Bull 32:21–26, 1976; Nora JJ: Multifactorial inheritance hypothesis for the etiology of congenital heart diseases: The genetic environmental interaction, Circulation 38:604–617, 1968; Lin AE, Garver KL: Genetic counseling for congenital heart defects, J Pediatr 113:1105–1109, 1988.数据来自Carter CO: Genetics of common single malformations, Br Med Bull 32:21–26, 1976; Nora JJ: Multifactorial inheritance hypothesis for the etiology of congenital heart diseases: The genetic environmental interaction, Circulation 38:604–617, 1968; Lin AE, Garver KL: Genetic counseling for congenital heart defects, J Pediatr 113:1105–1109, 1988.
42/53
Defects in cell death Abnormalities in extracellular matrix Defects in targeted growth The subtype of congenital heart m…
Ch9 — Segment 42
Defects in cell death Abnormalities in extracellular matrix Defects in targeted growth The subtype of congenital heart malformations known as flow lesions illustrates the familial aggregation and elevated risk for recurrence in relatives of an affected individual, all characteristic of a complex trait ( Flow lesions, which constitute ~50% of all CHDs, include hypoplastic left heart syndrome, coarctation of the aorta, atrial septal defect of the secundum type, pulmonary valve stenosis, a common type of ventricular septal defect, and other forms .细胞死亡缺陷 细胞外基质异常 靶向生长缺陷 被称为血流病变的先天性心脏畸形亚型展示了家族聚集性和患者亲属复发风险升高,这些都是复杂性状的特征(血流病变约占所有先天性心脏病的50%,包括左心发育不全综合征、主动脉缩窄、继发孔型房间隔缺损、肺动脉瓣狭窄、一种常见类型的室间隔缺损及其他形式)。
Up to 25% of individuals with flow lesions, particularly tetralogy of Fallot, may have the deletion of chromosome region 22q11 seen in the velocardiofacial syndrome (see Chapter 6).高达25%的血流病变患者,尤其是法洛四联症患者,可能存在22q11染色体区域的缺失,这种缺失见于腭心面综合征(见第6章)。
Certain isolated CHDs are inherited as multifactorial traits.某些孤立性先天性心脏病作为多因素性状遗传。
Until more is known, the figures shown in There is, however, a rapid falloff in risk (to levels not much higher than the population risk) in second- and third-degree relatives of index patients with flow lesions.在了解更多之前,所示数据表明,对于血流病变先证者的二级和三级亲属,风险迅速下降(至略高于人群风险的水平)。
Similarly, relatives of index patients with types of CHDs other than flow lesions can be offered reassurance that their risk is not much greater than that of the general population.同样,对于非血流病变类型的先天性心脏病先证者的亲属,可以告知其风险并不比普通人群高很多,从而让他们放心。
For further reassurance, many CHDs can now be assessed prenatally by ultrasonography (see Chapter 18). 17 4. 3 25 Patent ductus arteriosus 0. 083 3. 2 38 Atrial septal defect 0. 066 3. 2 48 Aortic stenosis 0. 044 2. 6 59 Normal Atrial septal defect Coarctation of the aorta Tetralogy of Fallot Patent ductus arteriosus (PDA) Hypoplastic left heart AO PA LA RA RV LV AO PA LA LV RA RV RV RA LA PA AO LV AO PA LA RA RV LV AO PA PDA LA RA RV LV AO PA LA LV RV RA PDA Blood on the left side of the circulation is shown in red, on the right side in blue.为进一步安心,许多先天性心脏病现在可以通过超声检查进行产前评估(见第18章)。
Abnormal admixture of oxygenated and deoxygenated blood is purple.含氧血和脱氧血的异常混合呈紫色。
AO, Aorta; LA, left atrium; LV, left ventricle; PA, pulmonary artery; RA, right atrium; RV, right ventricle.AO,主动脉;LA,左心房;LV,左心室;PA,肺动脉;RA,右心房;RV,右心室。
Neuropsychiatric Disorders Mental illnesses are some of the most common and perplexing of human diseases, affecting 4% of the human population worldwide.神经精神疾病 精神疾病是人类最常见且最令人困惑的疾病之一,影响全球约4%的人口。
As of 2020, the annual cost in medical care and social services exceeds $200 billion in the United States alone.截至2020年,仅在美国,每年的医疗和社会服务费用就超过2000亿美元。
Among the most severe of the mental illnesses are schizophrenia and bipolar disease (manic-depressive illness).最严重的精神疾病包括精神分裂症和双相情感障碍(躁郁症)。
Schizophrenia affects 1% of the world’s population.精神分裂症影响全球1%的人口。
It is a devastating psychiatric illness, with onset commonly in late adolescence or young adulthood, and is characterized by abnormalities in thought, emotion, and social relationships, often associated with delusional thinking and disordered mood.这是一种破坏性的精神疾病,通常在青少年晚期或成年早期发病,其特征是思维、情感和社交关系异常,常伴有妄想思维和情绪紊乱。
A genetic contribution遗传因素的作用
43/53
Complex Inheritance of Common Multifactorial Disorders 163 5 Sibling 8–14 11 Nephew or niece 1–4 2. 5 Uncle or aunt 2 2 …
Ch9 — Segment 43
Complex Inheritance of Common Multifactorial Disorders 163 5 Sibling 8–14 11 Nephew or niece 1–4 2. 5 Uncle or aunt 2 2 First cousin 2–6 4 Grandchild 2–8 5 to schizophrenia is supported by both twin and family aggregation studies.常见多因素疾病的复杂遗传 163 5 兄弟姐妹 8-14 11 侄子或侄女 1-4 2.5 叔叔或阿姨 2 2 堂表亲 2-6 4 孙子孙女 2-8 5 精神分裂症的遗传基础得到双胞胎研究和家族聚集研究的支持。
MZ concordance in schizophrenia is estimated to be 40% to 60%; DZ concordance is 10% to 16%.同卵双生子精神分裂症的一致率估计为40%至60%;异卵双生子的一致率为10%至16%。
The recurrence risk ratio is elevated in first- and second-degree relatives of individuals with schizophrenia ( Although there is considerable evidence of a genetic contribution to schizophrenia, only a subset of the genes and alleles that predispose to the disease has been identified to date.精神分裂症患者的一级和二级亲属中再发风险比升高(尽管有大量证据表明遗传对精神分裂症有贡献,但迄今为止仅鉴定出部分易感基因和等位基因)。
A major exception is the small percentage (&lt;2%) of all schizophrenia that is found in individuals with interstitial deletions of particular chromosomes, such as the 22q11 deletion responsible for the velocardiofacial syndrome.一个主要例外是,在所有精神分裂症中,有一小部分(<2%)患者存在特定染色体的间质性缺失,例如导致腭心面综合征的22q11缺失。
It is estimated that 25% of individuals with 22q11 deletions develop schizophrenia, even in the absence of many or most of the other physical signs of the syndrome.据估计,22q11缺失的个体中有25%会发展为精神分裂症,即使缺乏该综合征的许多或大部分其他体征。
The mechanism by which a deletion of 3 Mb of DNA on 22q11 causes mental illness in individuals with this syndrome is unknown.22q11上3 Mb DNA缺失导致该综合征患者精神疾病的机制尚不清楚。
Chromosomal microarrays have been used to scan the entire genome for other deletions and duplications, many too small to be detectable by standard cytogenetic approaches, as introduced in Chapter 5.染色体微阵列已用于扫描整个基因组以寻找其他缺失和重复,其中许多太小而无法通过标准细胞遗传学方法检测,如第五章所述。
These studies have revealed numerous deletions and duplications (copy number variants) throughout the genome in both normal individuals and individuals with a variety of psychiatric and neurodevelopmental disorders (see Chapter 6).这些研究揭示了正常个体以及患有各种精神疾病和神经发育障碍的个体(见第六章)基因组中广泛存在缺失和重复(拷贝数变异)。
In particular, small (1–1. 5 Mb) interstitial deletions at 1q21. 1, 15q11. 2, and 15q13. 3 have been implicated repeatedly in a small fraction of individuals with schizophrenia.特别地,1q21.1、15q11.2和15q13.3上的小片段(1-1.5 Mb)间质性缺失在少数精神分裂症患者中反复被证实。
For the vast majority of people with schizophrenia, however, genetic lesions are not known, and counseling therefore relies on empirical risk figures (see Bipolar disease is predominantly a mood disorder in which episodes of mood elevation, grandiosity, high-risk dangerous behavior, and inflated self-esteem (mania) alternate with periods of depression, decreased interest in what are normally pleasurable activities, feelings of worthlessness, and suicidal thinking.然而,对于绝大多数精神分裂症患者,遗传损伤尚不清楚,因此遗传咨询依赖于经验风险数据(参见双相情感障碍主要表现为情绪障碍,其特征是情绪高涨、夸大、高风险危险行为、自尊膨胀(躁狂)与抑郁期、对通常愉悦活动兴趣减退、无价值感及自杀想法交替出现)。
The prevalence of bipolar disease is 0. 8%, approximately equal to that of schizophrenia, with a similar age at onset.双相情感障碍的患病率为0.8%,与精神分裂症大致相等,发病年龄相似。
The ­seriousness of this condition is underscored by the high rate of suicide in affected individuals.该疾病的严重性体现在患者的高自杀率上。
A genetic contribution to bipolar disease is strongly supported by twin and family aggregation studies.双胞胎研究和家族聚集研究强烈支持遗传对双相情感障碍的贡献。
MZ twin concordance is 40% to 60%; DZ twin concordance is 4% to 8%.同卵双生子一致率为40%至60%;异卵双生子一致率为4%至8%。
Disease risk is also elevated in relatives of affected individuals ( One striking aspect of bipolar disease in families is that the condition has variable expressivity; some members of the same family demonstrate classic bipolar illness, others have depression alone (unipolar disorder), and others carry a diagnosis of a psychiatric syndrome that involves both thought and mood (schizoaffective disorder).患者亲属的疾病风险也升高(双相情感障碍在家族中的一个显著特点是表现度可变;同一家庭中部分成员表现为典型双相障碍,其他成员仅有抑郁(单相障碍),还有一些被诊断为兼有思维和情绪障碍的精神病综合征(分裂情感性障碍))。
Even less is known about genes and alleles that predispose to bipolar disease than is known for schizophrenia; in particular, although an increase in de novo deletions or duplications has been identified in bipolar psychosis, recurrent copy number variants involving particular regions of the genome have not been identified.目前对双相情感障碍易感基因和等位基因的了解甚至少于精神分裂症;特别是,尽管在双相精神病中发现了新生缺失或重复的增加,但尚未鉴定出涉及特定基因组区域的重现拷贝数变异。
Counseling therefore typically relies on empirical risk figures (see Coronary Artery Disease Coronary artery disease (CAD) kills ~500,000 individuals in the United States yearly and is one of the most frequent causes of morbidity and mortality in the developed world.因此,遗传咨询通常依赖于经验风险数据(参见冠状动脉疾病 冠状动脉疾病(CAD)每年在美国导致约50万人死亡,是发达国家发病率和死亡率最常见的原因之一。
CAD due to atherosclerosis is the major cause of the nearly 1. 5 million cases of myocardial infarction (MI) and the more than 200,000 deaths from acute MI occurring annually.由动脉粥样硬化引起的CAD是每年近150万例心肌梗死(MI)和超过20万例急性MI死亡的主要原因。
In the aggregate, CAD costs more than $143 billion in health care expenses alone each year in the United States, not including lost productivity.总体而言,在美国,仅医疗保健费用一项,CAD每年就耗费超过1430亿美元,不包括生产力损失。
For unknown reasons, males are at higher risk for CAD both in the general population and within affected families.原因不明,无论是一般人群还是患病家族中,男性患CAD的风险均更高。
Family studies have repeatedly supported a role for heredity in CAD, particularly when it occurs in relatively young individuals.家族研究反复证实了遗传因素在CAD中的作用,尤其是当该病发生在相对年轻个体中时。
The pattern of increased risk suggests that when the proband is female or young there is likely to be a greater genetic contribution to MI in the family, thereby increasing the risk for disease in the proband’s relatives.风险增加的模式表明,当先证者为女性或年轻时,家族中MI的遗传贡献可能更大,从而增加先证者亲属的患病风险。
For example, the recurrence risk (5-fold increased risk in female relatives of a male proband.例如,再发风险(男性先证者的女性亲属风险增加5倍)。
44/53
A few mendelian disorders leading to CAD are known.
Ch9 — Segment 44
A few mendelian disorders leading to CAD are known.一些导致冠心病的孟德尔遗传病已知。
Familial hypercholesterolemia (Case 16), an autosomal dominant defect of the low-density lipoprotein (LDL) receptor discussed in Chapter 13, is one of the most common of these but accounts for only ~5% of survivors of MI.家族性高胆固醇血症(病例16)是一种常染色体显性遗传的低密度脂蛋白(LDL)受体缺陷,已在第13章讨论,是这些疾病中最常见的之一,但仅占心肌梗死幸存者的约5%。
Most cases of CAD show multifactorial inheritance, with both nongenetic and genetic predisposing factors.大多数冠心病病例表现出多因素遗传,同时具有非遗传性和遗传性易感因素。
There are many stages in the evolution of atherosclerotic lesions in the coronary artery.冠状动脉粥样硬化病变的演变有多个阶段。
What begins as a fatty streak in the intima of the artery evolves into a fibrous plaque containing smooth muscle, lipid, and fibrous tissue.最初为动脉内膜的脂纹,演变为包含平滑肌、脂质和纤维组织的纤维斑块。
These intimal plaques become vascular and may bleed, ulcerate, and calcify, thereby causing severe vessel narrowing as well as providing fertile ground for thrombosis, resulting in sudden, complete occlusion and MI.这些内膜斑块变得血管化,可能出血、溃疡和钙化,从而导致严重的血管狭窄,并为血栓形成提供有利条件,最终引起突然的完全闭塞和心肌梗死。
Given the many stages in the evolution of atherosclerotic lesions in the coronary artery, it is not surprising that many genetic differences affecting the various pathologic processes involved could predispose to or protect from CAD .鉴于冠状动脉粥样硬化病变演变的多个阶段,许多影响相关病理过程的遗传差异可能易感或保护免于冠心病,这并不令人意外。
Additional risk factors for CAD include other disorders that are themselves multifactorial with genetic components, such as hypertension, 5-fold in female first-degree relatives Female 7-fold in male first-degree relatives Female &lt;55 yr 11. 4-fold in male first-degree relatives Two male relatives &lt;55 yr 13-fold in first-degree relatives *Relative to the risk in the general population.冠心病的其他危险因素包括其他本身具有多因素性和遗传成分的疾病,如高血压,女性一级亲属风险增加5倍,男性一级亲属风险增加7倍,女性年龄<55岁的一级亲属风险增加11.4倍,两名男性年龄<55岁的一级亲属风险增加13倍(*相对于一般人群的风险)。
CAD, Coronary artery disease.CAD,即冠状动脉疾病。
Data from Silberberg JS: Risk associated with various definitions of family history of coronary heart disease, Am J Epidemiol 147:1133–1139, 1998. 39 6–8-fold 0. 26 3-fold Female 0. 44 15-fold 0. 14 2. 6-fold *Early myocardial infarction defined as age &lt;55 years in males, age &lt;65 years in females. †Relative to the risk in the general population.数据来自Silberberg JS:不同冠心病家族史定义相关的风险,Am J Epidemiol 147:1133–1139, 1998。39 6–8倍 0.26 3倍 女性0.44 15倍 0.14 2.6倍 *早期心肌梗死定义为男性年龄<55岁,女性年龄<65岁。†相对于一般人群的风险。
DZ, Dizygotic; MZ, monozygotic.DZ,双卵双生;MZ,单卵双生。
Data from Marenberg ME: Genetic susceptibility to death from coronary heart disease in a study of twins, NEJM 330:1041–1046, 1994.数据来自Marenberg ME:双生子研究中冠心病死亡的遗传易感性,NEJM 330:1041–1046, 1994。
Normal Fatty streaks White blood cells White blood cells Platelets and fibrin Thrombus Calcium Red blood cells Lipid-rich plaque Inflammation and calcification Scar development with calcification Scar Early Lipid rich Internal rupture Calcified shell Calcified plaque Vulnerable Rupture Thrombus Myocardial infarction Obstruction Genetic and environmental factors operating at any or all of the steps in this pathway can contribute to the development of this complex, common disease. ) When the proband is young (&lt;55 years) and female, the risk for CAD is more than 11 times greater than that of the general population.正常脂纹、白细胞、血小板和纤维蛋白、血栓、钙、红细胞、富含脂质的斑块、炎症和钙化、瘢痕形成伴钙化、早期富含脂质、内部破裂、钙化外壳、钙化斑块、易损斑块、破裂、血栓、心肌梗死、阻塞——遗传和环境因素作用于该通路的任何或所有步骤均可促成这种复杂常见疾病的发生;当先证者年轻(<55岁)且为女性时,冠心病风险比一般人群高11倍以上。
Having multiple relatives affected at a young age increases risk substantially as well.有多个亲属在年轻时患病也会显著增加风险。
Twin studies also support a role for genetic variants in CAD (双生子研究也支持遗传变异在冠心病中的作用(
45/53
Complex Inheritance of Common Multifactorial Disorders 165 obesity, and diabetes mellitus.
Ch9 — Segment 45
Complex Inheritance of Common Multifactorial Disorders 165 obesity, and diabetes mellitus.常见多因素疾病的复杂遗传 165 肥胖症和糖尿病。
The metabolic and physiologic derangements represented by these disorders also contribute to enhancing the risk for CAD.这些疾病所代表的代谢和生理紊乱也增加了冠心病的风险。
Finally, diet, physical activity, systemic inflammation, and smoking are environmental factors that also play a major role in influencing the risk for CAD.最后,饮食、体力活动、全身性炎症和吸烟等环境因素也在影响冠心病风险中起主要作用。
Given all the different processes, metabolic derangements, and environmental factors that contribute to the development of CAD, it is easy to imagine that genetic susceptibility to CAD could be a complex multifactorial condition (see 2). or how a particular genotype and set of environmental influences interact to cause a disease or to determine the value of a particular physiologic measurement.考虑到所有导致冠心病发展的不同过程、代谢紊乱和环境因素,很容易想象冠心病的遗传易感性可能是一种复杂的多因素疾病(见2),或者一个特定的基因型与一组环境因素如何相互作用以导致疾病或确定某个生理测量值。
In most cases, all we can show is that there is some genetic contribution and estimate its magnitude.在大多数情况下,我们只能证明存在某种遗传贡献,并估计其大小。
These genetic contributions reflect the aggregate effects of many variants that might number in hundreds or thousands.这些遗传贡献反映了可能数以百计或千计的多态变异的累积效应。
There are, however, a few multifactorial diseases with complex inheritance for which we have begun to identify the genetic and, in some cases, environmental factors responsible for increasing disease susceptibility.然而,有少数具有复杂遗传的多因素疾病,我们已经开始识别其遗传因素,在某些情况下还包括环境因素,这些因素导致疾病易感性增加。
We give a few examples in the next part of this chapter, illustrating increasing levels of complexity.我们在本章下一部分给出几个例子,说明复杂性逐步增加的情况。
Modifier Genes in Mendelian Disorders As discussed in Chapter 7, allelic variation at a single locus can explain variation in the phenotype in many single-gene disorders.孟德尔遗传病中的修饰基因 如第7章所述,单个位点的等位基因变异可以解释许多单基因疾病中表型的变异。
However, even for well-characterized mendelian disorders known to be due to defects in a single gene, variation at other gene loci may impact some aspect of the phenotype, illustrating features of complex inheritance.然而,即使对于已知由单个基因缺陷导致的明确表征的孟德尔遗传病,其他基因位点的变异也可能影响表型的某些方面,展示了复杂遗传的特征。
In cystic fibrosis (CF) (Case 12), for example, whether an individual has pancreatic insufficiency requiring enzyme replacement can be explained largely by which genetic changes are present in the CFTR gene (see Chapter 13).例如,在囊性纤维化(CF)(病例12)中,个体是否存在需要酶替代治疗的胰腺功能不全,很大程度上可通过CFTR基因中存在哪些遗传变化来解释(见第13章)。
The correlation is imperfect, however, for other phenotypes.然而,对于其他表型,这种相关性并不完美。
For example, the variation in the degree of pulmonary disease seen in CF patients remains unexplained by allelic heterogeneity.例如,CF患者肺部疾病程度的变异仍无法通过等位基因异质性来解释。
It has been proposed that the genotype at other genetic loci could act as genetic modifiers, that is, genes whose alleles have an effect on the severity of pulmonary disease seen in CF.有学者提出,其他遗传位点的基因型可作为遗传修饰因子,即其等位基因影响CF患者肺部疾病严重程度的基因。
For example, reduction in forced expiratory volume after 1 second (FEV1), calculated as a percentage of the value expected for CF patients (a CF-specific FEV1 percent), is a quantitative trait commonly used to measure deterioration in pulmonary function in CF.例如,第一秒用力呼气容积(FEV1)的下降(计算为CF患者预期值的百分比,即CF特异性FEV1百分比)是常用于衡量CF患者肺功能恶化的数量性状。
A comparison of CF-specific FEV1 percent in affected MZ versus affected DZ twins provides an estimate of the heritability of the severity of lung disease in CF patients of ~50%.比较患病同卵双胞胎与患病异卵双胞胎的CF特异性FEV1百分比,可估计CF患者肺部疾病严重程度的遗传度约为50%。
This value is independent of the specific CFTR allele(s) (because both kinds of twins will have the same pathogenic CF variants).该值与特定的CFTR等位基因无关(因为两种双胞胎均具有相同的致病CF变异)。
Two loci harboring alleles responsible for modifying the severity of pulmonary disease in CF are known: MBL2, a gene that encodes a serum protein called mannose-binding lectin; and the TGFB1 locus encoding the cytokine transforming growth factor β (TGFβ).已知两个携带等位基因的位点负责调节CF肺部疾病的严重程度:MBL2(编码一种称为甘露糖结合凝集素的血清蛋白的基因)和TGFB1位点(编码细胞因子转化生长因子β)。
Mannose-binding lectin is a plasma protein in the innate immune system that binds to many pathogenic organisms and aids in their destruction by phagocytosis and complement activation.甘露糖结合凝集素是先天免疫系统中的一种血浆蛋白,与多种病原生物结合,并通过吞噬作用和补体激活帮助破坏它们。
A number of common alleles that result in reduced blood levels of the lectin exist at the MBL2 locus in European populations.[TL:missing]
Lower levels of mannose-binding lectin appear 2 GENES AND GENE PRODUCTS INVOLVED IN THE STEPWISE PROCESS OF CORONARY ARTERY DISEASE A large number of genes and gene products have been suggested and, in some cases, implicated in promoting one or more of the developmental stages of coronary artery disease.[TL:missing]
These include genes involved in the following: Serum lipid transport and metabolism – cholesterol, apolipoprotein E, apolipoprotein C-III, the low-density lipoprotein (LDL) receptor, and lipoprotein(a) – as well as total cholesterol level.[TL:missing]
Elevated LDL cholesterol level and triglyceride levels, both of which elevate the risk for coronary artery disease, are themselves quantitative traits with significant heritabilities.[TL:missing]
Vasoactivity, such as angiotensin-converting enzyme.[TL:missing]
Blood coagulation, platelet adhesion, and fibrinolysis, such as plasminogen activator inhibitor 1, and the platelet surface glycoproteins Ib and IIIa.[TL:missing]
Inflammatory and immune pathways.[TL:missing]
Arterial wall components.[TL:missing]
CAD is often an incidental finding in family histories of individuals with other genetic diseases.[TL:missing]
In view of the high recurrence risk, physicians and genetic counselors may need to consider whether first-degree relatives of people with CAD should be evaluated further and offered counseling and therapy, even when CAD is not the primary genetic problem for which the patient or relative has been referred.[TL:missing]
Such an evaluation is clearly indicated when the proband is young, particularly if the proband is female.[TL:missing]
EXAMPLES OF MULTIFACTORIAL TRAITS FOR WHICH SPECIFIC GENETIC AND ENVIRONMENTAL FACTORS ARE KNOWN Up to this point we have described some of the epidemiologic approaches involving family and twin studies that are used to assess the extent to which there may be a genetic contribution to a complex trait.[TL:missing]
It is important to realize, however, that studies of familial aggregation, disease concordance, or heritability do not specify how many loci there are, which loci and alleles are involved,[TL:missing]
46/53
associated with worse outcomes for CF lung disease, perhaps because low levels of lectin result in difficulties with con…
Ch9 — Segment 46
associated with worse outcomes for CF lung disease, perhaps because low levels of lectin result in difficulties with containing respiratory pathogens, particularly Pseudomonas.与CF肺病预后较差相关,可能是因为低水平的凝集素导致难以控制呼吸道病原体,尤其是假单胞菌。
Alleles at the TGFB1 locus that result in higher TGFβ production are also associated with worse outcome, perhaps because TGFβ promotes lung scarring and fibrosis after inflammation.TGFB1基因座上导致更高TGFβ产生的等位基因也与更差的预后相关,可能是因为TGFβ在炎症后促进肺部瘢痕形成和纤维化。
Thus both MBL2 and TGFB1 are modifier genes, variants at which – while they do not cause CF – can modify the clinical phenotype associated with disease-causing alleles at the CFTR locus.因此,MBL2和TGFB1均为修饰基因,其上的变异虽不导致CF,但可改变由CFTR基因座致病等位基因引起的临床表型。
Digenic Inheritance The next level of complexity is a disorder determined by the additive effect of the genotypes at two or more loci.双基因遗传 下一层次的复杂性是由两个或更多基因座的基因型相加效应所决定的疾病。
One clear example of such a disease phenotype has been found in a few families of patients with a form of retinal degeneration called retinitis pigmentosa (RP) .此类疾病表型的一个明确例子见于少数患有视网膜变性(称为视网膜色素变性,RP)的患者家族。
Affected individuals in these families are heterozygous for pathogenic alleles at two different loci (double heterozygotes).这些家族中的受累个体在两个不同基因座上均为致病等位基因的杂合子(双杂合子)。
One locus encodes the photoreceptor membrane protein peripherin and the other encodes a related photoreceptor membrane protein called Rom 1.一个基因座编码光感受器膜蛋白peripherin,另一个编码相关的光感受器膜蛋白Rom 1。
Heterozygotes for only one or the other of these mutations are unaffected.仅携带其中一种或另一种突变的杂合子不发病。
Thus the RP in this family is caused by the simplest form of multigenic inheritance, inheritance due to the effect of variant alleles at two loci, without any known environmental factors that greatly influence disease occurrence or severity.因此,该家族中的RP是由最简单的多基因遗传形式所致,即由两个基因座上的变异等位基因效应引起的遗传,且没有任何已知的显著影响疾病发生或严重程度的环境因素。
The proteins encoded by these two genes are likely to have overlapping physiologic function because they are both located in the stacks of membranous disks found in retinal photoreceptors.这两个基因编码的蛋白可能具有重叠的生理功能,因为它们都位于视网膜光感受器中的膜盘堆叠中。
It is the additive effect of having an abnormality in two proteins with overlapping function that produces disease.正是两个功能重叠的蛋白质均出现异常所产生的相加效应导致了疾病。
A multigenic model has also been proposed in a few families with Bardet-Biedl syndrome, a rare birth defect characterized by obesity, variable degrees of intellectual disability, retinal degeneration, polydactyly, and genitourinary malformations.在少数患有巴德-比德尔综合征(一种罕见的出生缺陷,特征为肥胖、不同程度的智力障碍、视网膜变性、多指(趾)畸形和泌尿生殖系统畸形)的家族中也提出了多基因模型。
Fourteen different genes have been found in which pathogenic variants cause the syndrome.已发现14个不同基因的致病性变异可导致该综合征。
Although inheritance is clearly autosomal recessive in most families, a few families appear to demonstrate digenic inheritance, in which the disease occurs only when an individual is homozygous or compound heterozygous for mutations at one of these 14 loci and is heterozygous for a variant at another of the loci.尽管在大多数家族中遗传方式明确为常染色体隐性遗传,但少数家族似乎表现出双基因遗传,即仅当个体在这14个基因座之一上纯合或复合杂合突变,并在另一基因座上为变异杂合子时,疾病才会发生。
Gene-Environment Interactions in Venous Thrombosis Another example of gene-gene interaction predisposing to disease is found in the group of conditions referred to as hypercoagulability states, in which venous or arterial clots form inappropriately and cause life-­threatening complications of thrombophilia (Case 46).静脉血栓形成中的基因-环境相互作用 基因-基因相互作用易致疾病的另一个例子见于一组称为高凝状态的疾病,其中静脉或动脉内形成不适当的血栓,并导致血栓形成的危及生命的并发症(病例46)。
With hypercoagulability, however, there is a third factor, an environmental influence that in the presence of the predisposing genetic factors increases the risk for disease even more.然而,对于高凝状态,存在第三个因素,即环境因素,它在存在易感遗传因素的情况下进一步增加疾病风险。
One such disorder is idiopathic cerebral vein thrombosis, a disease in which clots form in the venous system of the brain, causing catastrophic occlusion of cerebral veins in the absence of an inciting event such as infection or tumor.其中一种疾病是特发性脑静脉血栓形成,即脑静脉系统中形成血栓,在无感染或肿瘤等诱发事件的情况下导致脑静脉的灾难性闭塞。
It affects young adults, and although quite rare (&lt;1 per 100,000 in the population), it carries a high mortality rate (5–30%).该病影响年轻人,虽然相当罕见(人群中<1/100,000),但死亡率很高(5–30%)。
Three relatively common factors – two genetic and one environmental – that lead to abnormal coagulability of the clotting system are each known to individually increase the risk for cerebral vein thrombosis : 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/1 Genotype: peripherin: 1 or mut ROM1: 1 or mut 1/1 1/mut 1/mut 1/mut Dark blue symbols are affected individuals.导致凝血系统异常凝血的三个相对常见的因素——两个遗传因素和一个环境因素——各自已知可独立增加脑静脉血栓形成的风险:1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/mut 1/1 1/mut 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/mut 1/1 1/1 1/1 1/mut 1/mut 1/mut 1/1 基因型:peripherin: 1 或 mut ROM1: 1 或 mut 1/1 1/mut 1/mut 1/mut 深蓝色符号为受累个体。
Each individual’s genotypes at the peripherin locus (first line) and ROM1 locus (second line) are written below each symbol.每个个体在peripherin基因座(第一行)和ROM1基因座(第二行)上的基因型写在每个符号下方。
The normal allele is 1; the allele carrying a mutation is mut.正常等位基因为1;携带突变的等位基因为mut。
Light blue symbols are unaffected, despite carrying a pathogenic variant in one or the other gene.浅蓝色符号为未受累个体,尽管携带一个或另一个基因的致病性变异。
(Redrawn from Kajiwara K, Berson EL, Dryja TP: Digenic retinitis pigmentosa due to pathogenic variants at the unlinked peripherin/RDS and ROM1 loci, Science 264:1604–1608, 1994.)(重新绘制自Kajiwara K, Berson EL, Dryja TP: Digenic retinitis pigmentosa due to pathogenic variants at the unlinked peripherin/RDS and ROM1 loci, Science 264:1604–1608, 1994.)
47/53
Complex Inheritance of Common Multifactorial Disorders 167 A missense variant in the gene for the clotting factor, facto…
Ch9 — Segment 47
Complex Inheritance of Common Multifactorial Disorders 167 A missense variant in the gene for the clotting factor, factor V A variant in the 3′ untranslated region (UTR) of the gene for the clotting factor prothrombin The use of oral contraceptives A common allele of factor V, factor V Leiden (FVL) (Case 46), in which arginine is replaced by glutamine at position 506 (Arg 506Gln), has a frequency of ~2. 5% in populations of European origin but is rarer in other population groups.常见多因素疾病的复杂遗传 167 凝血因子V基因的一个错义变体、凝血因子凝血酶原基因3′非翻译区的一个变体、口服避孕药的使用,以及V因子常见等位基因V因子Leiden(FVL)(病例46),其中第506位精氨酸被谷氨酰胺取代(Arg506Gln),在欧裔人群中频率约为2.5%,但在其他群体中较为罕见。
This alteration affects a cleavage site used to degrade factor V, thereby making the protein more stable and able to exert its procoagulant effect for a longer duration.这一改变影响用于降解V因子的一个切割位点,从而使该蛋白更稳定,并能更持久地发挥其促凝血作用。
Heterozygous carriers of FVL, ~5% of the European population, have a risk for cerebral vein thrombosis that, although still quite low, is 7-fold higher than that in the general population; homozygotes have a risk that is 80-fold higher.FVL的杂合子携带者约占欧洲人口的5%,其脑静脉血栓形成的风险虽仍然很低,但比一般人群高7倍;纯合子的风险则高80倍。
The second genetic risk factor, a pathogenic variant in the prothrombin gene, changes a G to an A at position 20210 in the 3′ UTR of the gene (prothrombin g. 20210 G&gt;A).第二个遗传风险因素为凝血酶原基因的一个致病性变体,该变体在基因的3′UTR中第20210位将G变为A(凝血酶原g.20210 G>A)。
Approximately 2. 4% of individuals of European ancestry are heterozygotes, but it is rare in other groups.约2.4%的欧裔个体为杂合子,但在其他群体中罕见。
This change appears to increase the level of prothrombin mRNA, resulting in increased translation and elevated levels of the protein.这一改变似乎增加了凝血酶原mRNA的水平,导致翻译增加及蛋白水平升高。
Being heterozygous for the prothrombin 20210 G&gt;A allele raises the risk for cerebral vein thrombosis three- to sixfold.凝血酶原20210 G>A等位基因的杂合状态使脑静脉血栓形成的风险增加三到六倍。
The use of oral contraceptives containing synthetic estrogen increases the risk for thrombosis 14- to 22-fold, independent of genotype at the factor V and prothrombin loci, probably by increasing the levels of many clotting factors in the blood.含有合成雌激素的口服避孕药的使用,独立于V因子和凝血酶原基因座的基因型,可能通过增加血液中多种凝血因子的水平,使血栓形成的风险增加14至22倍。
Although using oral contraceptives and being heterozygous for FVL cause only a modest increase in risk compared with either factor alone, oral contraceptive use in a heterozygote for prothrombin 20210 G&gt;A raises the relative risk for cerebral vein thrombosis 30- to 150-fold!尽管单独使用口服避孕药或单独为FVL杂合子仅带来轻微的风险增加,但在凝血酶原20210 G>A杂合子中使用口服避孕药,可使脑静脉血栓形成的相对风险提高30至150倍!
There is also interest in the role of FVL and prothrombin 20210 G&gt;A alleles in deep venous thrombosis (DVT) of the lower extremities, a condition that occurs in ~1 in 1000 individuals per year, far more common than idiopathic cerebral venous thrombosis.FVL和凝血酶原20210 G>A等位基因在下肢深静脉血栓(DVT)中的作用也备受关注,DVT每年约每1000人中发生1例,远比特发性脑静脉血栓形成常见。
Mortality due to DVT (primarily due to pulmonary embolus) can be up to 10%, depending on age and the presence of other medical conditions.DVT(主要由肺栓塞引起)的死亡率可达10%,具体取决于年龄及是否存在其他医学状况。
Many environmental factors are known to increase the risk for DVT and include trauma, surgery (particularly orthopedic surgery), malignant disease, prolonged periods of immobility, oral contraceptive use, and advanced age.已知许多环境因素可增加DVT风险,包括创伤、手术(尤其是骨科手术)、恶性肿瘤、长期不活动、口服避孕药使用及高龄。
The FVL allele increases the relative risk for a first episode of DVT 7-fold in heterozygotes; heterozygotes who use oral contraceptives see their risk increased 30-fold compared with controls.FVL等位基因使杂合子首次发生DVT的相对风险增加7倍;使用口服避孕药的杂合子其风险相比对照增加30倍。
Heterozygotes for prothrombin 20210 G&gt;A also have an increase in their relative risk for DVT of two- to threefold.凝血酶原20210 G>A杂合子的DVT相对风险也增加两到三倍。
Notably, double heterozygotes for FVL and prothrombin 20210 G&gt;A have a relative increased risk of 20-fold – a risk approaching a few percent of the population.值得注意的是,FVL和凝血酶原20210 G>A的双重杂合子其相对风险增加20倍——这一风险接近人群的百分之几。
Thus each of these three factors, two genetic and one environmental, on its own increases the risk for an abnormal hypercoagulable state; having two or all three of these factors at the same time raises the risk even more, to the point that thrombophilia screening programs for selected populations may be indicated in the future.因此,这三个因素(两个遗传性、一个环境性)各自独立地增加异常高凝状态的风险;同时存在其中两个或全部三个因素时,风险进一步升高,甚至可能有必要在未来对特定人群实施易栓症筛查计划。
Multiple Coding and Noncoding Elements in Hirschsprung Disease A more complicated set of interacting genetic factors has been described in the pathogenesis of a developmental abnormality of the enteric nervous system in the gut known as Hirschsprung disease (HSCR).先天性巨结肠症中的多个编码和非编码元件 在一种称为先天性巨结肠症(HSCR)的肠道肠神经系统发育异常的发病机制中,描述了一组更复杂的相互作用遗传因素。
In HSCR, there is complete absence of some or all of the intrinsic ganglion cells in the myenteric and submucosal plexuses of the colon.在HSCR中,结肠肌间神经丛和黏膜下神经丛的部分或全部内在神经节细胞完全缺失。
An aganglionic colon is incapable of peristalsis, resulting in severe constipation, symptoms of intestinal obstruction, and massive dilatation of the colon (megacolon) proximal to the aganglionic segment.无神经节细胞的结肠无法蠕动,导致严重便秘、肠梗阻症状以及无神经节段近端结肠的显著扩张(巨结肠)。
The disorder affects ~1 in 5000 newborns of European ancestry but is twice as common among Asian infants.该病影响约每5000名欧裔新生儿中的1例,但在亚洲婴儿中发病率是其两倍。
HSCR occurs as an isolated birth defect 70% of the time, as part of a chromosomal syndrome 12% of the time, and as one element of a broad constellation of congenital abnormalities in the remainder of cases.HSCR有70%的病例作为孤立性出生缺陷发生,12%作为染色体综合征的一部分,其余病例则作为广泛先天性异常组合中的一个要素。
Among individuals with HSCR as an isolated birth defect, 80% have only a single, short aganglionic segment of colon at the level of the rectum (hence, HSCR-S), whereas 20% Factor Xa Factor Va Factor V Factor X (↑OC) Thrombin Prothrombin (↑OC) Fibrin Clot Fibrinogen Intrinsic pathway Extrinsic pathway Once factor X is activated, through either the intrinsic or extrinsic pathway, activated factor V promotes the production of the coagulant protein thrombin from prothrombin, which in turn cleaves fibrinogen to generate fibrin required for clot formation.在HSCR作为孤立性出生缺陷的个体中,80%仅有单个、短段的直肠水平无神经节结肠(因此称为HSCR-S),而20% 因子Xa 因子Va 因子V 因子X(↑OC)凝血酶 凝血酶原(↑OC)纤维蛋白凝块 纤维蛋白原 内源性途径 外源性途径 一旦因子X通过内源性或外源性途径被激活,活化的因子V促进凝血酶原生成凝血蛋白凝血酶,进而切割纤维蛋白原以产生凝块形成所需的纤维蛋白。
Oral contraceptives (OC) increase blood levels of prothrombin and factor X as well as a number of other coagulation factors.口服避孕药增加血液中凝血酶原、因子X以及许多其他凝血因子的水平。
The hypercoagulable state can be explained as a synergistic interaction of genetic and environmental factors that increase the levels of factor V, prothrombin, factor X, and others to promote clotting.高凝状态可解释为增加因子V、凝血酶原、因子X等水平以促进凝血的遗传因素与环境因素的协同相互作用。
Activated forms of coagulation proteins are indicated by the letter a.凝血蛋白的活化形式以字母a表示。
Solid arrows are pathways; dashed arrows are stimulators.实线箭头为通路;虚线箭头为刺激物。
48/53
have aganglionosis of a long segment of colon, the entire colon or, occasionally, the entire colon plus the ileum (hence…
Ch9 — Segment 48
have aganglionosis of a long segment of colon, the entire colon or, occasionally, the entire colon plus the ileum (hence, HSCR-L).具有长段结肠、全结肠或偶尔全结肠加回肠的无神经节细胞症(因此称为HSCR-L)。
Familial HSCR-L is often characterized by patterns of inheritance that suggest dominant or recessive inheritance, but consistently with reduced penetrance.家族性HSCR-L通常表现为提示显性或隐性遗传的遗传模式,但始终伴有外显率降低。
HSCR-L is most commonly caused by loss-of-function missense or nonsense mutations in the RET gene, which encodes RET, a receptor tyrosine kinase.HSCR-L最常见由RET基因中的功能丧失性错义或无义突变引起,该基因编码受体酪氨酸激酶RET。
A small minority of families have pathogenic variants in genes encoding ligands that bind to RET, but with even lower penetrance than those families with RET variants.少数家族在编码与RET结合的配体的基因中存在致病变异,但其外显率低于携带RET变异的家族。
HSCR-S is the more common type of HSCR and has many of the characteristics of a disorder with complex genetics.HSCR-S是更常见的HSCR类型,具有复杂遗传疾病的多项特征。
The relative risk ratio for sibs, λs, is very high (~200), but MZ twins do not show perfect concordance, and families do not show any obvious mendelian inheritance pattern for the disorder.同胞的相对风险比λs非常高(约200),但同卵双胞胎未显示完全一致性,且家族未显示该疾病的任何明显孟德尔遗传模式。
When pairs of siblings concordant for HSCR-S were analyzed genome wide to see which loci and which sets of alleles at these loci each sib had in common with an affected brother or sister, alleles at three loci (including RET) were found to be significantly shared, suggesting gene-gene interactions and/or multigenic inheritance; indeed, most of the concordant sibpairs were found to share alleles at all three loci.当对HSCR-S一致的同对同胞进行全基因组分析,以查看每个同胞与患病兄弟姐妹共有的位点及这些位点上的等位基因集时,发现三个位点(包括RET)的等位基因显著共享,提示存在基因-基因相互作用和/或多基因遗传;事实上,大多数一致的同对同胞被发现共享所有三个位点的等位基因。
Although the non-RET loci have yet to be identified, (sometimes referred to as insulin dependent [IDDM]) and type 2 (T2D) (sometimes referred to as non–insulin dependent [NIDDM]), representing ~10% As with most complex traits, genetic susceptibility can be explained by common variants with relatively small impacts on individual risk, and a spectrum towards very rare alleles with a substantial impact on risk.尽管非RET位点尚未确定,(有时称为胰岛素依赖型[IDDM])和2型(T2D)(有时称为非胰岛素依赖型[NIDDM]),约占10%。与大多数复杂性状一样,遗传易感性可由对个体风险影响相对较小的常见变异以及从常见到对风险有实质性影响的非常罕见等位基因的谱系来解释。
HSCR shows phenotypic severity differences - whereas syndromic HSCR and TCA/L-HSCR are more rare and severe, whereas sporadic S-HSCR is more common.HSCR表现出表型严重程度差异——综合征性HSCR和TCA/L-HSCR较为罕见且严重,而散发性S-HSCR则更为常见。
(From Karim A, Tang CS, Tam PK.(摘自Karim A, Tang CS, Tam PK。
The Emerging Genetic Landscape of Hirschsprung Disease and Its Potential Clinical Applications.先天性巨结肠症的新兴遗传景观及其潜在临床应用。
Front Pediatr 2021;9:638093.)Front Pediatr 2021;9:638093。)
49/53
Complex Inheritance of Common Multifactorial Disorders 169 and 88% of all cases, respectively.
Ch9 — Segment 49
Complex Inheritance of Common Multifactorial Disorders 169 and 88% of all cases, respectively.常见多因素疾病的复杂遗传 分别占所有病例的169%和88%。
Familial aggregation is seen in both types of diabetes, but in any given family usually only T1D or T2D is present.两种类型糖尿病均可见家族聚集性,但在任何特定家庭中,通常仅存在T1D或T2D。
They differ in typical onset age, MZ twin concordance, and association with particular genetic variants at particular loci.它们在典型发病年龄、同卵双生子一致性以及与特定位点特定遗传变异的相关性方面存在差异。
Here, we focus on T1D to illustrate the major features of complex inheritance in diabetes.在此,我们重点关注T1D,以阐述糖尿病复杂遗传的主要特征。
T1D has an incidence in the population of European ancestry of ~2 per 1000 (0. 2%), but this is lower in African and Asian ancestry populations.T1D在欧洲血统人群中的发病率约为2/1000(0.2%),但在非洲和亚洲血统人群中较低。
It usually manifests in childhood or adolescence.它通常在儿童期或青春期显现。
It results from autoimmune destruction of the β cells of the pancreas, which normally produce insulin.它是由胰腺β细胞的自身免疫破坏所致,这些β细胞通常产生胰岛素。
A large majority of children who will go on to have T1D develop multiple autoantibodies early in childhood against a variety of endogenous proteins, including insulin, well before they develop overt disease.大多数将发展为T1D的儿童在儿童早期就产生针对多种内源性蛋白质(包括胰岛素)的多种自身抗体
There is strong evidence for genetic factors in T1D: concordance among MZ twins is ~40%, which far exceeds the 5% concordance in DZ twins.[TL:missing]
The lifetime risk for T1D in siblings of an affected proband is ~7%, resulting in an estimated λs of ≈35.[TL:missing]
However, the earlier the age of onset of the T1D in the proband, the greater is λs.[TL:missing]
The Major Histocompatibility Complex The major genetic factor in T1D is the major histocompatibility complex (MHC) locus, which spans some 3 Mb on chromosome 6 and is the most highly polymorphic locus in the human genome, with over 200 known genes (many involved in immune functions) and well over 2000 alleles known in populations around the globe .[TL:missing]
On the basis of structural and functional differences, two major subclasses, the class I and class II genes, correspond to the human leukocyte antigen (HLA) genes, originally discovered by virtue of their importance in tissue transplantation between unrelated individuals.[TL:missing]
The HLA class I (HLA-A, HLA-B, HLA-C) and class II (HLA-DR, HLA-DQ, HLA-DP) genes encode cell surface proteins that play a critical role in the presentation of antigen to lymphocytes, which cannot recognize and respond to an antigen unless it is complexed with an HLA molecule on the surface of an antigen-presenting cell.[TL:missing]
Within the MHC, the HLA class I and class II genes are by far the most highly polymorphic loci .[TL:missing]
The original studies showing an association between T1D and alleles designated as HLA-DR3 and HLA-DR4 relied on a serologic method in use at that time for distinguishing between different HLA alleles, one that was based on immunologic reactions in a test tube.[TL:missing]
This method has long been superseded by direct determination of the DNA sequence of different alleles, and sequencing of the MHC in a large number of individuals has revealed that the serologically determined “alleles” associated with T1D are not single alleles at all (see 3).[TL:missing]
Both DR3 and DR4 can be subdivided into a dozen or more alleles located at a locus now termed HLA-DRB1.[TL:missing]
The set of HLA alleles at the different class I and class II loci on a given chromosome together form a haplotype.[TL:missing]
Within any one ancestral group, some HLA alleles and haplotypes are found commonly; others are rare or never seen.[TL:missing]
The differences in the distribution and frequency of the alleles and haplotypes within the MHC are the result of complex genetic, environmental, and historical factors at play in each of the different populations.[TL:missing]
The extreme levels of genetic variation at HLA loci and their resulting haplotypes have been extraordinarily useful for identifying associations of particular variants with specific diseases (see Chapter 11), many of which (as one might predict) are autoimmune disorders, associated with an abnormal immune response apparently directed against one or more self-antigens resulting from polymorphism in immune response genes.[TL:missing]
Furthermore, it is now clear that the association between certain DRB1 alleles and T1D is due, in part, to alleles at two other class II loci, DQA1 and DQB1, located ~80 kb away from DRB1, that form a particular combination of alleles with each other (i. e., a haplotype) that is typically inherited as a unit (due to linkage disequilibrium; see Chapter 11).[TL:missing]
DQA1 and DQB1 encode the α and β chains of the class II DQ protein.[TL:missing]
Certain combinations of alleles at these three loci form a haplotype that increases the risk for T1D more than 11-fold over that for the general population, whereas other combinations of alleles reduce the risk 50-fold.[TL:missing]
The DQB1*0303 allele contained in this protective haplotype results in the amino acid aspartic acid at position 57 of the DQB1 product, whereas other amino acids at this position (alanine, valine, or serine) confer susceptibility.[TL:missing]
In fact, ~90% of individuals with T1D are homozygous for DQB1 alleles that do not encode aspartic acid at position 57.[TL:missing]
It is likely that differences in antigen binding, determined by which amino acid is at position 3 HUMAN ANTIGEN ALLELES AND HAPLOTYPES The human leukocyte antigen (HLA) system can be confusing at first because the nomenclature used to define and describe different HLA alleles has undergone a fundamental change with the advent of widespread DNA sequencing of the major histocompatibility complex (MHC).[TL:missing]
According to the older system of HLA nomenclature, the different alleles were distinguished from one another serologically.[TL:missing]
However, as the genes responsible for encoding the class I and class II MHC chains were identified and sequenced , single HLA alleles initially defined serologically were shown to consist of multiple alleles defined by different DNA sequence variants even within the same serological allele.[TL:missing]
The 100 serologic specificities at HLA-A, B, C, DR, DQ, and DP loci now comprise more than 1300 alleles defined at the DNA sequence level![TL:missing]
For example, what used to be a single B27 allele defined serologically is now referred to as HLA-B*2701, HLA-B*2702, and so on, based on DNA-based genotyping.[TL:missing]
50/53
Gene annotations Simple nucleotide polymorphisms Strand db SNP&gt;1% Base pairs 30,000,000 30,500,000 31,000,000 31,500,…
Ch9 — Segment 50
Gene annotations Simple nucleotide polymorphisms Strand db SNP&gt;1% Base pairs 30,000,000 30,500,000 31,000,000 31,500,000 32,000,000 32,500,000 33,000,000 6 ZFP57 ZNRD1 AS1 RNF39 HLA-F HLA-G HLA-A HCG9 ZNRD1 PPP1R11 TRIM40 TRIM15 TRIM39 RPP21 HLA-E TRIM31 TRIM10 TRIM26 HCG17 HCG18 PRR3 ABCF1 MRPS188 ATAT1 GNL1 PPP1R10 DHX16 NRM MDC1 IER3 TIGD1L SFTA2 HCG21 HLA-C HLA-B CDSN PSORS1C2 CCHCR1 POU5F1 PSORS1C3 MICB HCP5 MICA HCG27 TUBB HCG20 DPCR1 MUC21 HCG22 PSORS1C1 TCF19 DDR1 GTF2H4 VARS2 HCG23 HLA-DRA HLA-DQA1 HLA-DQA2 PSM89 BRD2 HLA-DPB1 HCG24 HLA-DPA1 HLA-DOA HLA-DMB HLA-DMA C6orf 10 BTNL2 HLA-DR65 HLA-DR61 HLA-DQB1 HLA-DQ62 HLA-DOB TAP2 PSM68 TAP1 PPT2 EGFL8 RNF5 HSPA1A HSPA1B C2 CFB STK19 C4A C4B CVP21A2 AJF1 APOM CSNK2B LY6G6B LY6G6F LY6G6D MSH5 NFKBIL1 TNF LTA LST1 BAT1 ATP6V1G2 LTB NCR3 BAT3 BAT4 BAT5 LY6G6C DDAH2 CLIC1 VARS LSM2 HSPA1L NEU1 SLC44A4 EHMT2 RDBP DOM3Z TNX5 ATF68 PRRT1 AGPAT1 AGER PBX2 NOTCH4 The classic MHC is shown on the short arm of chromosome 6, comprising the class I region (yellow) and class II region (blue), both enriched in human leukocyte antigen (HLA) genes.[TL:failed]
Sequence-level variation is shown for single nucleotide polymorphisms (SNPs) found with at least 1% frequency.[TL:failed]
Remarkably high levels of genetic variation are seen in regions containing the classic HLA genes where variation is enriched in coding exons involved in defining the antigen-binding cleft.[TL:failed]
Other genes (pink) in the MHC region show lower levels of genetic variation. db SNP, Minor allele frequency in the SNP database. ) 2 MZ twin 40 Sibling 7 Sibling with no DR haplotypes in common 1 Sibling with 1 DR haplotype in common 5 Sibling with 2 DR haplotypes in common 17* Child 4 Child of affected mother 3 Child of affected father 5 *20–25% for particular shared haplotypes.[TL:failed]
MZ, Monozygotic. 57, contribute directly to the autoimmune response that destroys the insulin-producing cells of the pancreas.[TL:failed]
Other loci and alleles in the MHC, however, are also important, as can be seen from the fact that some individuals with T1D have an aspartic acid at this position.[TL:failed]
Genes Other Than Class II Major Histocompatibility Complex Loci in Type 1 Diabetes The MHC haplotype alone accounts for only a portion of the genetic contribution to the risk for T1D in siblings of a proband.[TL:failed]
Family studies in T1D ( Thus there must be other genes elsewhere in the genome that contribute to the development of T1D (assuming that MZ twins and sibs have similar environmental exposures).[TL:failed]
Indeed, genetic association studies (to be described in Chapter 11) indicate that variation at more than 50 different loci around the genome can increase susceptibility to T1D, although most have very small effects on increasing disease susceptibility.[TL:failed]
51/53
Complex Inheritance of Common Multifactorial Disorders 171 It is important to stress, however, that genetic factors alon…
Ch9 — Segment 51
Complex Inheritance of Common Multifactorial Disorders 171 It is important to stress, however, that genetic factors alone do not cause T1D because the MZ twin concordance rate is only ~40%, not 100%.常见多因素疾病的复杂遗传 171 然而,必须强调,仅遗传因素并不导致T1D,因为单卵双生子一致性率仅为约40%,而非100%。
Until a more complete picture develops of the genetic and nongenetic factors that cause T1D, risk counseling using HLA haplotyping must remain empirical (see Alzheimer Disease Alzheimer disease (AD) (Case 4) is a fatal neurodegenerative disease that affects 1% to 2% of the US population.在形成导致T1D的遗传与非遗传因素的更完整图景之前,使用HLA单倍型进行风险咨询仍需基于经验(参见阿尔茨海默病 阿尔茨海默病(AD)(病例4)是一种致命性神经退行性疾病,影响美国1%至2%的人口。
It is the most common cause of dementia in older adults and is responsible for more than half of all cases of dementia.它是老年人痴呆最常见的原因,占所有痴呆病例的一半以上。
As with other dementias, patients experience a chronic, progressive loss of memory and other cognitive functions, associated with loss of certain types of cortical neurons.与其他痴呆一样,患者经历慢性、进行性的记忆及其他认知功能丧失,伴有特定类型皮质神经元的缺失。
Age, sex, and family history are the most significant risk factors for AD.年龄、性别和家族史是AD最重要的风险因素。
Once a person reaches 65 years of age, the risk for any dementia, and AD in particular, increases substantially with age and female sex ( AD can be diagnosed definitively only postmortem, on the basis of neuropathologic findings of characteristic protein aggregates (β-amyloid plaques and neurofibrillary tangles; see Chapter 13).一旦人达到65岁,任何痴呆(尤其是AD)的风险随年龄和女性性别显著增加(AD只能在死后根据特征性蛋白聚集物(β-淀粉样斑块和神经原纤维缠结;见第13章)的神经病理学发现来明确诊断)。
The most important constituent of the plaques is a small (39–42 amino acid) peptide, Aβ, derived from cleavage of a normal neuronal protein, the amyloid protein precursor.斑块最重要的成分是一种小肽(39-42个氨基酸)Aβ,来源于正常神经元蛋白——淀粉样蛋白前体的裂解。
The secondary structure of Aβ gives the plaques the staining characteristics of amyloid proteins.Aβ的二级结构赋予斑块淀粉样蛋白的染色特性。
In addition to three rare autosomal dominant forms of the disease (see Chapter 13), in which disease onset is in the third to fifth decade, there is a common form of AD with onset after the age of 60 years (late onset).除了三种罕见的常染色体显性遗传型(见第13章),其发病在第三至第五个十年,还有常见的AD形式,发病在60岁以后(晚发型)。
This form has no obvious mendelian inheritance pattern but shows familial aggregation and an elevated relative risk ratio (λs = ≈4) typical of disorders with complex inheritance.该形式无明显孟德尔遗传模式,但表现出家族聚集性和升高的相对风险比(λs ≈4),符合复杂遗传疾病的特征。
Twin studies have been inconsistent but suggest MZ concordance of ~50% and DZ concordance of ~18%.双生子研究结果不一致,但提示单卵一致性约50%,双卵一致性约18%。
The ε4 Allele of Apolipoprotein E The major locus with alleles found to be significantly associated with common late-onset AD is APOE, which encodes apolipoprotein E.载脂蛋白E的ε4等位基因 与常见晚发型AD显著关联的主要位点是APOE,其编码载脂蛋白E。
Apolipoprotein E is a protein component of the LDL particle and is involved in clearing LDL through an interaction with high-affinity receptors in the liver.载脂蛋白E是低密度脂蛋白颗粒的一种蛋白质成分,通过与肝脏中高亲和力受体的相互作用参与低密度脂蛋白的清除。
Apolipoprotein E is also a constituent of amyloid plaques in AD and is known to bind the Aβ peptide.载脂蛋白E也是AD中淀粉样斑块的组成成分,已知与Aβ肽结合。
The APOE gene has three alleles, ε2, ε3, and ε4, due to substitutions of arginine for two different cysteine residues in the protein (see Chapter 13).APOE基因有三个等位基因ε2、ε3和ε4,由蛋白质中两个不同半胱氨酸残基被精氨酸替代所致(见第13章)。
When the genotypes at the APOE locus were analyzed in individuals with AD and controls, a genotype with at least one ε4 allele was found two to three times more frequently among patients compared with controls in both the general US and Japanese populations ( Even more striking is that the risk for AD appears to increase further if both APOE alleles are ε4, through an effect on the age at onset of AD; individuals with two ε4 alleles have an earlier onset of disease than those with only one.当在AD患者和对照中分析APOE位点的基因型时,在普通美国和日本人群中,患者中至少有一个ε4等位基因的基因型频率是对照的两到三倍(更显著的是,若两个APOE等位基因均为ε4,AD风险似乎进一步增加,通过影响AD发病年龄;携带两个ε4等位基因的个体比仅携带一个者发病更早)。
In a study of people with AD and unaffected controls, the age at which AD developed in the affected individuals was earliest for ε4/ε4 homozygotes, next for ε4/ε3 heterozygotes, and significantly less for the other genotypes .在一项针对AD患者和未受影响对照的研究中,受影响个体发生AD的年龄最早为ε4/ε4纯合子,其次为ε4/ε3杂合子,其他基因型显著较晚。
In the population in general, the risk for developing AD by age 80 is approaching 10%.在一般人群中,到80岁时发展为AD的风险接近10%。
The ε4 allele is clearly a predisposing factor that increases the risk for development of AD by shifting the age at onset to an earlier age, such that ε3/ε4 heterozygotes have a 40% risk for developing the disease, and ε4/ε4 have a 60% risk by age 85.ε4等位基因显然是一种易感因素,通过将发病年龄前移而增加AD发生风险,使ε3/ε4杂合子到85岁时有40%的风险,ε4/ε4有60%的风险。
Despite this increased risk, other genetic and environmental factors must be important because a significant proportion of ε3/ε4 and ε4/ε4 individuals live to extreme old age with no evidence of AD.尽管风险增加,其他遗传和环境因素必然重要,因为相当比例的ε3/ε4和ε4/ε4个体活到极高龄且无AD证据。
There are also reports of association between the presence of the ε4 allele and neurodegenerative disease after traumatic head injury (as seen in professional boxers, football players, and soldiers who have suffered blast injuries), indicating that at least one environmental factor, brain trauma, can interact with the ε4 allele in the pathogenesis of AD. 3 10. 9 Female 12 19 65–100 yr Male 25 32. 8 Female 28. 1 45 AD, Alzheimer disease.也有报道称ε4等位基因的存在与创伤性头部损伤后的神经退行性疾病相关(如职业拳击手、足球运动员和遭受爆炸伤的士兵中所见),表明至少一种环境因素——脑创伤——可在AD发病机制中与ε4等位基因相互作用。3 10.9 女性 12 19 65–100岁 男性 25 32.8 女性 28.1 45 AD,阿尔茨海默病。
Data from Seshadri S, Wolf PA, Beiser A, et al: Lifetime risk of dementia and Alzheimer’s disease.数据来自Seshadri S, Wolf PA, Beiser A等:痴呆和阿尔茨海默病的终生风险。
The impact of mortality on risk estimates in the Framingham Study, Neurology 49:1498–1504, 1997. 64 0. 31 0. 47 0. 17 ε3/ε3; ε2/ε3; or ε2/ε2 0. 36 0. 69 0. 53 0. 83 *Frequency of genotypes with and without the ε4 allele among Alzheimer disease (AD) patients and controls from the United States and Japan.死亡率对Framingham研究中风险估计的影响,《神经学》49:1498–1504, 1997. 64 0.31 0.47 0.17 ε3/ε3; ε2/ε3; 或 ε2/ε2 0.36 0.69 0.53 0.83 *美国与日本阿尔茨海默病(AD)患者和对照中带和不带ε4等位基因的基因型频率。
52/53
The ε4 variant of APOE represents a prime example of a predisposing allele: It predisposes to a complex trait in a power…
Ch9 — Segment 52
The ε4 variant of APOE represents a prime example of a predisposing allele: It predisposes to a complex trait in a powerful way but does not predestine any individual carrying the allele to the disease.APOE的ε4变异是易感等位基因的一个典型例子:它以强大的方式易感复杂性状,但并不会注定任何携带该等位基因的个体患病。
Additional genes as well as environmental effects are also clearly involved; although several of these appear to have a significant effect, most remain to be identified.其他基因以及环境影响也明显参与其中;尽管其中一些似乎具有显著效应,但大多数仍有待鉴定。
In general, testing of asymptomatic people for the APOE ε4 allele remains controversial because knowing that one is a heterozygote or homozygote for the ε4 allele does not mean one will develop AD, and there are no preventative interventions to reduce disease risk in more susceptible individuals (see Chapter 19).总的来说,对无症状人群进行APOE ε4等位基因检测仍存在争议,因为知道自己是ε4等位基因的杂合子或纯合子并不意味着一定会患上阿尔茨海默病,并且目前尚无预防性干预措施可降低易感个体的疾病风险(见第19章)。
THE CHALLENGE OF MULTIFACTORIAL DISEASE WITH COMPLEX INHERITANCE The greatest challenge facing medical genetics and genomic medicine going forward is unraveling the complex interactions between the variants at multiple loci and the relevant environmental factors that underlie the susceptibility to common multifactorial disease.具有复杂遗传方式的多因素疾病的挑战 医学遗传学和基因组医学未来面临的最大挑战是解析多个位点变异与相关环境因素之间的复杂相互作用,这些相互作用构成了常见多因素疾病易感性的基础。
This area of research is the central focus of the field of population-based genetic epidemiology (to be discussed more fully in Chapter 10).这一研究领域是基于人群的遗传流行病学领域的核心焦点(将在第10章更全面地讨论)。
The field is developing rapidly, and it is clear that the genetic contribution to many more complex diseases in humans will be elucidated in the coming years.该领域发展迅速,显然在未来几年内,人类更多复杂疾病的遗传贡献将被阐明。
Such understanding will, in time, allow the development of novel preventive and therapeutic measures for the common disorders that cause such significant morbidity and mortality in the population.这种理解最终将有助于为那些在人群中导致显著发病率和死亡率的常见疾病开发新的预防和治疗措施。
GENERAL REFERENCES Chakravarti A, Clark AG, Mootha VK: Distilling pathophysiology from complex disease genetics, Cell 155:21–26, 2013.一般参考文献 Chakravarti A, Clark AG, Mootha VK: Distilling pathophysiology from complex disease genetics, Cell 155:21–26, 2013.
Rimoin DL, Pyeritz RE, Korf BR: Emery and Rimoin's essential medical genetics, Waltham, MA, 2020, Academic Press (Elsevier).Rimoin DL, Pyeritz RE, Korf BR: Emery and Rimoin's essential medical genetics, Waltham, MA, 2020, Academic Press (Elsevier).
Scott W, Ritchie M: Genetic analysis of complex disease, ed 3, Hoboken, NJ, 2022, John Wiley and Sons.Scott W, Ritchie M: Genetic analysis of complex disease, ed 3, Hoboken, NJ, 2022, John Wiley and Sons.
REFERENCES FOR SPECIFIC TOPICS Baylis RA, Smith NL, Klarin D, et al: Epidemiology and genetics of venous thromboembolism and chronic venous disease, Circ Res 128:1988–2002, 2021.特定主题参考文献 Baylis RA, Smith NL, Klarin D, et al: Epidemiology and genetics of venous thromboembolism and chronic venous disease, Circ Res 128:1988–2002, 2021.
Bellenguez C, Küçükali F, Jansen IE, et al: New insights into the genetic etiology of Alzheimer’s disease and related dementias, Nat Genet 54:412–436, 2022.Bellenguez C, Küçükali F, Jansen IE, et al: New insights into the genetic etiology of Alzheimer’s disease and related dementias, Nat Genet 54:412–436, 2022.
Bigdeli TB, Fanous AH, Li Y, et al: Genome-wide association studies of schizophrenia and bipolar disorder in a diverse cohort of US veterans, Schizophr Bull 47(2):517–529, 2021.Bigdeli TB, Fanous AH, Li Y, et al: Genome-wide association studies of schizophrenia and bipolar disorder in a diverse cohort of US veterans, Schizophr Bull 47(2):517–529, 2021.
Grant SFA, Wells AD, Rich SS: Next steps in the identification of gene targets for type 1 diabetes, Diabetologia 63:2260–2269, 2020.Grant SFA, Wells AD, Rich SS: Next steps in the identification of gene targets for type 1 diabetes, Diabetologia 63:2260–2269, 2020.
Karim A, Tang CS, Tam PK: The emerging genetic landscape of Hirschsprung disease and its potential clinical applications, Front Pediatr 9:638093, 2021.Karim A, Tang CS, Tam PK: The emerging genetic landscape of Hirschsprung disease and its potential clinical applications, Front Pediatr 9:638093, 2021.
Khera AV, Chaffin M, Aragam KG, et al: Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations, Nat Genet 50:1219–1224, 2018. https:// www. nature. com/articles/s 41588-018-0183-z Matzaraki M, Kumar V, Wijmenga C, et al: The MHC locus and genetic susceptibility to autoimmune and infectious diseases, Genome Bio 18:76, 2017.Khera AV, Chaffin M, Aragam KG, et al: Genome-wide polygenic scores for common diseases identify individuals with risk equivalent to monogenic mutations, Nat Genet 50:1219–1224, 2018. https:// www. nature. com/articles/s 41588-018-0183-z Matzaraki M, Kumar V, Wijmenga C, et al: The MHC locus and genetic susceptibility to autoimmune and infectious diseases, Genome Bio 18:76, 2017.
Shoaib M, Ye Q, Iglay Reger H, et al: Evaluation of polygenic risk scores to differentiate between type 1 and type 2 diabetes, Genet Epidemiol. 2023.Shoaib M, Ye Q, Iglay Reger H, et al: Evaluation of polygenic risk scores to differentiate between type 1 and type 2 diabetes, Genet Epidemiol. 2023.
Uffelmann E, Huang QQ, Munung NS, et al: Genome-wide association studies, Nat Rev Methods Primers 1:59, 2021.Uffelmann E, Huang QQ, Munung NS, et al: Genome-wide association studies, Nat Rev Methods Primers 1:59, 2021.
Wu P, Gifford A, Meng X, et al: Mapping ICD-10 and ICD-10-CM codes to phecodes: Workflow development and initial evaluation, JMIR Med Inform 7(4):e 14325, 2019.Wu P, Gifford A, Meng X, et al: Mapping ICD-10 and ICD-10-CM codes to phecodes: Workflow development and initial evaluation, JMIR Med Inform 7(4):e 14325, 2019.
Risk for Alzheimer disease 70 60 50 40 30 20 10 0 50 55 60 65 70 75 80 85 Age Population male Population female ε3/ε3 male ε3/ε3 female ε3/ε4 male ε3/ε4 female ε4/ε4 male ε4/ε4 female At one extreme is the ε4/ε4 homozygote, who has a ≈40% chance of remaining free of the disease by the age of 85 years, whereas an ε3/ε3 homozygote has ≈70% to ≈90% chance of remaining disease free at the age of 85 years, depending on the sex.阿尔茨海默病风险 70 60 50 40 30 20 10 0 50 55 60 65 70 75 80 85 年龄 一般人群男性 一般人群女性 ε3/ε3男性 ε3/ε3女性 ε3/ε4男性 ε3/ε4女性 ε4/ε4男性 ε4/ε4女性 极端情况是ε4/ε4纯合子,其在85岁时保持无病的概率约为40%,而ε3/ε3纯合子根据性别不同,在85岁时保持无病的概率约为70%至约90%。
General population risk is also shown for comparison. study, J Geriatr Psychiatry Neurol 18:250– 255, 2005.)同时显示一般人群风险以作比较。
53/53
Complex Inheritance of Common Multifactorial Disorders 173 PROBLEMS 1.
Ch9 — Segment 53
Complex Inheritance of Common Multifactorial Disorders 173 PROBLEMS 1.常见多因素疾病的复杂遗传 173 问题1.
For a specific disease, the concordance rate in monozygotic (MZ) twins is 80% and the concordance rate in dizygotic (DZ) twins is 20%.对于一种特定疾病,单卵双生子(MZ)的一致率为80%,双卵双生子(DZ)的一致率为20%。
What does this tell us about whether there are genetic or environmental contributions to susceptibility for this disease?关于该疾病易感性中是否存在遗传或环境因素的贡献,这一结果说明了什么?
Match the genetic mapping technique that would be most cost-effective to find genetic predictors in each scenario: a.将以下各场景中寻找遗传预测因子最具成本效益的遗传定位技术进行匹配:a.
Searching for the diseasecausing variant in a large pedigree with 4-5 related individuals who have a rare disease i.在一个包含4-5名患有罕见疾病的相关个体的大型家系中寻找致病变异 i.
Genome-wide association study using array-based genotyping and imputation b.使用基于芯片的基因分型和推断的全基因组关联研究 b.
Searching for common variants that increase susceptibility to a disease in a large case-control sample with well-studied ancestry ii.在具有良好研究背景的大型病例-对照样本中寻找增加疾病易感性的常见变异 ii.
Exome sequencing with linkage analysis c.外显子组测序结合连锁分析 c.
Searching for specific effector genes for a large number of diseases in a biobank sample iii.在生物库样本中寻找大量疾病的具体效应基因 iii.
Exome or genome sequencing with genebased burden testing 3.外显子组或基因组测序结合基于基因的负荷检验 3.
One of the most important challenges in human genetics is keeping up with the rapidly evolving literature about each trait.人类遗传学中最重要的挑战之一是跟上关于每种性状的快速发展的文献。
Pick a complex disease of your choice and find the largest genome wide association study published in the last 5 years and compare it to one published 5 years before that.选择一种你感兴趣的复杂疾病,找出在过去5年内发表的最大规模的全基因组关联研究,并将其与早5年发表的一项研究进行比较。
If you don’t have a favorite complex disease, consider atrial fibrillation as a potential example. a.如果你没有偏好的复杂疾病,可以考虑心房颤动作为潜在示例。a.
How many disease loci were identified in each of the two studies? b.这两项研究各自识别出了多少个疾病位点?b.
What was the strongest signal in each of the studies, defined by p-value?以p值定义,每项研究中最强的信号是什么?
Summarize disease allele frequency, odds ratio and nearest gene of all genomewide hits. c.总结所有全基因组显著位点的疾病等位基因频率、比值比和最邻近基因。c.
What was the strongest signal in each of the studies, defined by odds ratio?以比值比定义,每项研究中最强的信号是什么?
Summarize disease allele frequency, odds ratio and nearest gene. d.总结疾病等位基因频率、比值比和最邻近基因。d.
Did the two studies focus on similar questions?这两项研究是否聚焦于相似的问题?
What new questions did the new study consider?新研究考虑了哪些新问题?