Purpose:Diagnostic yield from exome and genome sequencing varies widely across studies. It remains unclear how much of this variation reflects patient-level factors (e.g., sex, clinical features, race/ethnicity, genetic ancestry) versus site-level practices such as sequencing modality or variant interpretation workflows. We aimed to quantify the contributions of these factors to diagnostic outcomes across five U.S. clinical sequencing sites. Methods:We performed a cross-sectional analysis of 3,008 prenatal, neonatal, and pediatric cases from the NHGRI Clinical Sequencing Evidence-Generating Research (CSER) consortium (2017-2023). Clinical indications spanned neurodevelopmental, neurological, immunological, metabolic, craniofacial, skeletal, cardiac, prenatal, and oncologic presentations. Genetic ancestry was inferred from sequencing data, and variants were interpreted using ACMG/AMP guidelines to classify DNA-based diagnoses. Generalized linear mixed models were used to estimate associations between diagnostic yield and fixed effects (sex, prenatal status, isolated cancer, number of clinical indications, sequencing modality, race/ethnicity, and genetic ancestry), while modeling study site as a random effect to quantify between-site variation. Results:The overall diagnostic yield was 19.0%. Multiple clinical indications (OR=1.47, 95% CI 1.20-1.80, p<0.001) were associated with higher diagnostic yield, and male sex (OR=0.80, 95% CI 0.66-0.96, p=0.017) and prenatal status (OR=0.63, 95% CI 0.44-0.90, p=0.012) were associated with lower yield. Sequencing modality, race/ethnicity, genetic ancestry, and isolated cancer were not statistically significantly associated with diagnostic outcomes.. A model without fixed effects attributed ~10% of variance in diagnostic yield to between-site differences. After adjusting for covariates, site-level variance decreased to 5.7%, indicating consistent variation across sites not explained by measured patient factors. Conclusion:Across five sites, patient-level clinical features influenced diagnostic yield, but substantial site-level variation remained even after adjustment. Differences in variant interpretation, or case-classification practices may contribute to this residual variability. Further efforts to increase consistency in exome- and genome-sequencing diagnostic workflows may help reduce inter-site differences.
Despite advances in exome and genome sequencing, many patients with suspected genetic disorders remain undiagnosed due to limitations in detecting complex structural variants. This study aimed to evaluate the diagnostic yield and clinical utility of Full-Genome Analysis (FGA), an integrated approach that combines short-read whole-genome sequencing (WGS), 10x Genomics linked-read sequencing, and Bionano optical genome mapping (OGM). Twenty-nine patients with unclear or inconclusive genetic diagnoses after standard testing were analyzed using an in-house FGA pipeline capable of simultaneously detecting single nucleotide variants (SNVs), copy number variants (CNVs), and structural variants (SVs). FGA established molecular diagnoses in 12 of 29 patients (41.4%), identifying nine pathogenic SNVs, three CNVs, and two complex SVs. Two CNVs were missed by chromosomal microarray, and both SVs were undetectable by short-read WES or WGS. Representative cases demonstrated that integrating OGM and linked-read sequencing improved detection of compound heterozygous variants and cryptic rearrangements that conventional methods failed to resolve. FGA substantially improved the diagnostic yield in patients with unresolved genetic disorders after conventional testing. Its ability to comprehensively detect small and large genomic variants within a single workflow highlights its potential as a next-generation diagnostic platform for rare disease evaluation.
Kidney dysfunction is a major cause of mortality, but its genetic architecture remains elusive. In this study, we conducted a multiancestry genome-wide association study in 2.2 million individuals and identified 1026 (97 previously unknown) independent loci. Ancestry-specific analysis indicated an attenuation of newly identified signals on common variants in European ancestry populations and the power of population diversity for further discoveries. We defined genotype effects on allele-specific gene expression and regulatory circuitries in more than 700 human kidneys and 237,000 cells. We found 1363 coding variants disrupting 782 genes, with 601 genes also targeted by regulatory variants and convergence in 161 genes. Integrating 32 types of genetic information, we present the “Kidney Disease Genetic Scorecard” for prioritizing potentially causal genes, cell types, and druggable targets for kidney disease.
BACKGROUND:Cornelia de Lange syndrome (CdLS) is a rare genetic disorder characterized by congenital multiple anomalies, developmental delay, and distinctive facial features. METHODS:We performed chromosomal microarray analysis (CMA), whole exome sequencing (WES), linked-read whole-genome sequencing (WGS) and optical genome mapping (OGM) to investigate an undiagnosed case of CdLS. RESULTS:A male patient presented clinical features consistent with CdLS, including a short nose, synophrys, small hands, hearing impairment, refractory complex partial seizures, and developmental delay. Amniocentesis at 28 gestational weeks and karyotyping revealed a presumably balanced translocation between chromosome 5 and chromosome 6. CMA and WES failed to identify copy number variants or a molecular diagnosis. Further analysis using WGS and OGM identified two translocation events on chromosome 5, resulting in three derivative chromosomes: 46,XY, der(2)t(2;5)(q32.3;p13.2),der(5)t(5;6)(p13.1;q12),der(6)t(2;6)(q32.3;q12)ins(6;5)(q12;p13.1p13.2). These rearrangements disrupted the NIPBL gene, a key gene with CdLS, splitting it across derivative chromosomes 2 and 6. Phasing studies revealed that these translocations originated from the paternal lineage. CONCLUSIONS:This case highlights the intricate genetic underpinnings of CdLS in this patient and underscores the diagnostic value of high-resolution genomic analyses in elucidating complex chromosomal rearrangements.
IntroductionChromosomal structural variations (SVs) play an important role in the formation of human cancers, including leukemias. However, many complex SVs cannot be identified by conventional tools, including karyotyping, fluorescence in situ hybridization, microarrays, and multiplex ligation-dependent probe amplification (MLPA).MethodsOptical genome mapping (OGM) and whole genome sequencing (WGS) were employed to analyze five leukemia samples with SVs detected by karyotyping, MLPA, and RNA sequencing (RNA-seq). OGM was performed using the Saphyr chip on a Bionano Saphyr system. Copy number variation and rare variant assembly analyses were performed with Bionano software v3.7. WGS was analyzed by the Manta program for SVs.ResultsThe leukemia samples had an average of 477 insertions, 457 deletions, and 32 inversions, which were significantly greater than those of the normal blood samples (p = 0.016, 0.028, and 0.028, respectively). In Case 1, OGM detected a sequential translocation between chromosomes 5, 8, 12, and 21 and ETV6::RUNX1 and BCAT1::BAALC gene fusions. Case 2 had two pathogenic SVs and a BCR::ABL1 fusion. Case 3 had one pathogenic SV and an IGH::DUSP22 fusion. Case 4 had two pathogenic SVs and a CBFB::MYH11 fusion. Case 5 had an STIL::TAL1 fusion. All breakpoint sequences were defined by WGS. An IGH::DUX4 fusion previously found by RNA-seq in Case 3 was not confirmed because DUX4, which has multiple pseudogenes, was refractory to OGM and WGS analyses.ConclusionOGM is a fundamental tool that complements G-banding analysis in identifying complex SVs in leukemia samples, and WGS effectively closes the gaps in OGM mapping.
Incorporating pharmacogenetics into clinical practice promises to improve therapeutic outcomes by optimizing drug selection and dosage based on genetic factors affecting drug response. A key advantage of PGx-guided therapy is to decrease the likelihood of adverse events. To evaluate the clinical impact of PGx risk variants, we performed a retrospective study using genetic and clinical data from the largest Han Chinese cohort, comprising 486,956 individuals, assembled by the Taiwan Precision Medicine Initiative. We found that nearly all participants carried at least one genetic variant that could affect drug response, with many carrying multiple risk variants. Here we show the detailed analyses of four gene-drug pairs, azathioprine (NUDT15/TPMT), clopidogrel (CYP2C19), statins (ABCG2/CYP2C9/SLCO1B1), and NSAIDs (CYP2C9), for which sufficient data exists for statistical power. While the results validate previous findings that PGx risk variants are significantly associated with drug-related adverse events or ineffectiveness, the excess risk of adverse events or lack of efficacy is small compared to that found in those without the PGx risk variants, and most patients with PGx variants do not suffer from adverse events. Our results point to the complexity of implementing PGx in clinical practice and the need for integrative approaches to optimize precision medicine.
Human height prediction based on genetic factors alone shows positive correlation, but predictors developed for one population perform less well when applied to population of different ancestries. In this study, we evaluated the utility of incorporating non-genetic factors in height predictors for the Han Chinese population in Taiwan. We analyzed data from 78,719 Taiwan Biobank (TWB) participants and 40,641 Taiwan Precision Medicine Initiative (TPMI) participants using genome-wide association study and multivariable linear regression least absolute shrinkage and selection operator (LASSO) methods to incorporate genetic and non-genetic factors for height prediction. Our findings establish that combining birth year (as a surrogate for nutritional status), age at measurement (to account for age-associated effects on height), and genetic profile data improves the accuracy of height prediction. This method enhances the correlation between predicted and actual height and significantly reduces the discrepancies between predicted and actual height in both males and females.
Predicting complex disease risks on the basis of individual genomic profiles is an advancing field in human genetics1,2. However, most genetic studies have focused on populations of European ancestry, creating a global imbalance in precision medicine and underscoring the need for genomic research in non-European groups3,4. The Taiwan Precision Medicine Initiative recruited more than half a million Taiwanese residents, providing a large dataset of genetic profiles and electronic medical record data for people with Han Chinese ancestry. Using extensive phenotypic data, we conducted comprehensive genomic analyses across the medical phenome with individuals genetically similar to Han Chinese reference populations. These analyses identified population-specific genetic risk variants and new findings for various complex traits. We developed polygenic risk scores, demonstrating strong predictive performance for conditions such as cardiometabolic diseases, autoimmune disorders, cancers and infectious diseases. We observed consistent findings in an independent dataset, Taiwan Biobank, and among people of East Asian ancestry in the UK Biobank and the All of Us Project. The identified genetic risks accounted for up to 10.3% of the overall health variation in the Taiwan Precision Medicine Initiative cohort. Our approach of characterizing the phenome-wide genomic landscape, developing population-specific risk-prediction models, assessing their performance and identifying the genetic effect on health serves as a model for similar studies in other diverse study populations.
Han Chinese people comprise nearly 20% of the global population but remain under-represented in genetic studies1,2, so there is an urgent need for large-scale cohorts to advance precision medicine. Here we present the Taiwan Precision Medicine Initiative (TPMI), established by Academia Sinica in collaboration with 16 major medical centres around Taiwan, which has recruited 565,390 participants who consent to provide DNA samples for genetic profiling and grant access to their electronic medical records (EMRs) for research. EMR access is both retrospective and prospective, allowing longitudinal studies. Genetic profiling is done with population-optimized arrays of single-nucleotide polymorphisms for people of Han Chinese ancestry, which enable genome-wide association3,4, phenome-wide association5,6 and polygenic risk score7,8 studies to be performed to evaluate common disease risk and pharmacogenetic response. Participants also agreed to be re-contacted for future research and receive personalized genetic risk profiles with health management recommendations. The TPMI has established the TPMI Data Access Platform, a central database and analysis platform that both safeguards the security of the data and facilitates academic research. As a large cohort of individuals with non-European ancestry that merges genetic profiles with EMR data and enables longitudinal follow-up, TPMI provides a unique resource that could be used to validate genetic risk prediction models, perform clinical trials of risk-based health management and inform health policies. Ultimately, the TPMI cohort will contribute to global genetic research and serve as a model for population-based precision medicine.
It has been suggested that diagnostic yield (DY) from Exome Sequencing (ES) may be lower among patients with non-European ancestries than those with European ancestry. We examined the association of DY with estimated continental/subcontinental genetic ancestry in a racially/ethnically diverse pediatric and prenatal clinical cohort. Cases (N = 845) with suspected genetic disorders underwent ES for diagnosis. Continental/subcontinental genetic ancestry proportions were estimated from the ES data. We compared the distribution of genetic ancestries in positive, negative, and inconclusive cases by Kolmogorov–Smirnov tests and linear associations of ancestry with DY by Cochran-Armitage trend tests. We observed no reduction in overall DY associated with any genetic ancestry (African, Native American, East Asian, European, Middle Eastern, South Asian). However, we observed a relative increase in proportion of autosomal recessive homozygous inheritance versus other inheritance patterns associated with Middle Eastern and South Asian ancestry, due to consanguinity. In this empirical study of ES for undiagnosed pediatric and prenatal genetic conditions, genetic ancestry was not associated with the likelihood of a positive diagnosis, supporting the equitable use of ES in diagnosis of previously undiagnosed but potentially Mendelian disorders across all ancestral populations.
AbstractDNA sequencing of patients with rare disorders has been highly successful in identifying “causal variants” for numerous conditions. However, there are many reports of healthy individuals who harbor these deleterious variants, leading to the concept of incomplete penetrance and doubt about the utility of genetic testing in clinical practice and population screening. As the deleterious variants are rare, the penetrance of these variants in the population is largely unknown. We analyzed the genetic and clinical data from 486,956 participants of the Taiwan Precision Medicine Initiative (TPMI) to determine the risk difference between those with and without deleterious variants. In all, we analyzed 292 disease-relevant variants and their clinical outcomes to assess their association. We found that only 15 variants show a risk difference exceeding 5% between those with or without the variants. In essence, 87.3% of deleterious variants exhibit minimal risk differences, suggesting a limited impact on the individual and population levels. Our analysis revealed increasing trends with age in six cardiovascular and degenerative diseases and bell-shaped trends in two cancers. Additionally, we identified three clinical outcomes exhibiting a dose-response relationship with the number of deleterious variants. Our findings show that large-scale testing of deleterious variants found in the literature is not warranted, except for those exhibiting large disease risk differences.
Predicting complex disease risks based on individual genomic profiles is an advancing field in human genetics. However, most genetic studies have focused on European populations, creating a global imbalance in precision medicine and underscoring the need for genomic research in non-European groups. The Taiwan Precision Medicine Initiative (TPMI) recruited over half a million Taiwanese residents, providing the largest datasets of genetic profiles and electronic medical records for the Han Chinese. Using extensive phenotypic data, we conducted the largest genomic analyses of Han Chinese across the medical phenome. These analyses identified population-specific genetic risk variants and novel findings on the genetic architecture of complex traits. We developed polygenic risk scores, demonstrating strong predictive performance for conditions such as cardiometabolic diseases, autoimmune disorders, cancers, and infectious diseases. We observed consistent findings in an independent dataset, Taiwan Biobank, and among East Asians in the UK Biobank and the All of Us Project. The identified genetic risks accounted for up to 9.1% of the disease burden variance in Taiwan. Our approach of characterizing the phenome-wide genomic landscape, developing population-specific risk prediction models, assessing their performance, and identifying the genetic impact on health, serves as a model for similar studies in other diverse study populations. ### Competing Interest Statement SDHH is a founder, shareholder, and serves on the Board of Directors of Genomic Prediction, Inc. (GP). EW is an employee and shareholder of GP. ### Funding Statement This study was funded in part by the Academia Sinica (40-05-GMM, AS-GC-110-MD02, and 236e-1100202 to P.-Y.K. and J.-Y.W.) and the National Development Fund, Executive Yuan (NSTC 111-3114-Y-001-001 to P.-Y.K.). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the Institutional Review Boards of Taipei Veterans General Hospital (2020-08-014A), National Taiwan University Hospital (201912110RINC), Tri-Service General Hospital (2-108-05-038), Chang Gung Memorial Hospital (201901731A3), Taipei Medical University Healthcare System (N202001037), Chung Shan Medical University Hospital (CS19035), Taichung Veterans General Hospital (SF19153A), Changhua Christian Hospital (190713), Kaohsiung Medical University Chung-Ho Memorial Hospital (KMUHIRB-SV(II)-20190059), Hualien Tzu Chi Hospital (IRB108-123-A),Far Eastern Memorial Hospital (110073-F), Ditmanson Medical Foundation Chia-Yi Christian Hospital (IRB2021128), Taipei City Hospital (TCHIRB-10912016), Koo Foundation Sun Yat-Sen Cancer Center (20190823A), Cathay General Hospital (CGH-P110041), Fu Jen Catholic University Hospital (FJUH109001) and Academia Sinica (AS-IRB01-18079), Taiwan. Written informed consent was obtained from the subjects in accordance with institutional requirements and the Declaration of Helsinki principles. All collected information was de-identified before statistical data analysis. The analysis with Health and Welfare Data Science Center (HWDC) was approved by Institutional Review Boards of Academia Sinica (AS-IRB-BM-23056). This research has been conducted using the UK Biobank Resource under UK Biobank Main Application 15326. Work with All of Us data was performed using the All of Us Researcher Workbench under the workspace, Duplicate of Prediction of Polygenic Traits. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The genotyping and electronic medical record (EMR) data analyzed in this study are from the Taiwan Precision Medicine Initiative (TPMI) with proper approval from the TPMI Data Access Committee. In compliance with the confidentiality laws governing genetic and health data in Taiwan, the de-identified TPMI data are kept in a secure server at the Academia Sinica and not released to the public.
Our study presents a 319-gene panel targeting inherited retinal dystrophy (IRD) genes. Through a multi-center retrospective cohort study, we validated the assay’s effectiveness and clinical utility and characterized the mutation spectrum of Taiwanese IRD patients. Between January 2018 and May 2022, 493 patients in 425 unrelated families, all initially suspected of having IRD without prior genetic diagnoses, underwent detailed ophthalmic and physical examinations (with extra-ocular features recorded) and genetic testing with our customized panel. Disease-causing variants were identified by segregation analysis and clinical interpretation, with validation via Sanger sequencing. We achieved a read depth of >200× for 94.2% of the targeted 1.2 Mb region. 68.5% (291/425) of the probands received molecular diagnoses, with 53.9% (229/425) resolved cases. Retinitis pigmentosa (RP) is the most prevalent initial clinical impression (64.2%), and 90.8% of the cohort have the five most prevalent phenotypes (RP, cone-rod syndrome, Usher’s syndrome, Leber’s congenital amaurosis, Bietti crystalline dystrophy). The most commonly mutated genes of probands that received molecular diagnosis are USH2A (13.7% of the cohort), EYS (11.3%), CYP4V2 (4.8%), ABCA4 (4.5%), RPGR (3.4%), and RP1 (3.1%), collectively accounted for 40.8% of diagnoses. We identify 87 unique unreported variants previously not associated with IRD and refine clinical diagnoses for 21 patients (7.22% of positive cases). We developed a customized gene panel and tested it on the largest Taiwanese cohort, showing that it provides excellent coverage for diverse IRD phenotypes.
ABSTRACTIncorporating pharmacogenetics into clinical practice promises to improve therapeutic outcome by choosing the medication and dosage optimized for a patient based on genetic factors that affect drug response1. One of the most promising benefits of PGx-guided therapy is the avoidance of adverse reactions2. To evaluate the clinical impact of PGx risk variants on adverse outcomes, we performed a retrospective study and analyzed the genetic and clinical data from the largest Han Chinese cohort assembled by the Taiwan Precision Medicine Initiative. We found that nearly all participants carried at least one genetic variant that could affect drug response, with many carrying multiple risk variants. Here we show that detailed analyses of four gene-drug pairs, for which sufficient data exist for statistical power, validate previous findings that PGx risk variants are significantly associated with drug-related adverse events or ineffectiveness. However, the excess risk of side effects or lack of efficacy is small compared to that found in those without the PGx risk variants, and most patients with PGx variants do not suffer from adverse events. Our results point to the need for identifying additional risk factors that cause adverse events in patients without PGx risk variants and factors that protect those with PGx risk variants from adverse events.
The Taiwan Precision Medicine Initiative (TPMI), a project initiated by the Academia Sinica in collaboration with 16 major medical centers around Taiwan, has recruited 565,390 participants who consented to provide DNA samples for genetic profiling and grant access to their electronic medical records (EMR) for studies to develop precision medicine. Access to the EMR is both retrospective and prospective, allowing researchers to conduct prospective studies over time. Genetic profiling is done with population-optimized SNP arrays for the Han Chinese populations that enable genetic analyses such as genome-wide association, phenome-wide association, and polygenic risk score studies to evaluate common disease risk and pharmacogenetic response. Furthermore, the TPMI participants agree to be contacted for future research opportunities related to their genetic risks and receive personalized genetic risk profiles with health management recommendations. TPMI has established the TPMI Data Access Platform (TDAP), a central database and analysis platform that both safeguards the security of the data and facilitates academic research. The TPMI is the largest non-European cohort that merges genetic profiles with EMR in the world. With a cohort that can be followed over time, it can be utilized to validate genetic risk prediction models, conduct clinical trials to show the efficacy of risk-based health management, and optimize health policies based on genetic risks. In this report, we describe the TPMI study design, the population and genetic characteristics of the TPMI cohort, and the power it provides to conduct crucial studies in developing precision medicine on a population and personal level. As Han Chinese represent almost 20% of the world's population, the results of TPMI studies will benefit >1.4 billion people around the world and serve as a model for developing population-based precision medicine. ### Competing Interest Statement The authors have declared no competing interest.
A 5-year-old female diagnosed with severe hemophilia B began experiencing frequent muscular and joint bleeds at 19 months old. Molecular studies, including Sanger sequencing, Giemsa banding, human androgen receptor (HUMARA) assay, array-based comparative genomic hybridization (aCGH), whole-exome sequencing (WES), and multiplex ligation-dependent probe amplification (MLPA), revealed a heterozygous factor IX (F9) intron 3 substitution (c.277+1G>T) inherited from her mother and a de novo heterozygous 441 kb deletion in the Xq28 region, which flanked intron 22 homologous regions 1 (int22h1) and 2 (int22h2). This rare genetic profile explains her severe phenotype and guides hereditary consultation for family planning.
The benefits of large-scale genetic studies for healthcare of the populations studied are well documented, but these genetic studies have traditionally ignored people from some parts of the world, such as South Asia. Here we describe whole genome sequence (WGS) data from 4806 individuals recruited from the healthcare delivery systems of Pakistan, India and Bangladesh, combined with WGS from 927 individuals from isolated South Asian populations. We characterize population structure in South Asia and describe a genotyping array (SARGAM) and imputation reference panel that are optimized for South Asian genomes. We find evidence for high rates of reproductive isolation, endogamy and consanguinity that vary across the subcontinent and that lead to levels of rare homozygotes that reach 100 times that seen in outbred populations. Founder effects increase the power to associate functional variants with disease processes and make South Asia a uniquely powerful place for population-scale genetic studies.
Background High sequence identity between segmental duplications (SDs) can facilitate copy number variants (CNVs) via non-allelic homologous recombination (NAHR). These CNVs are one of the fundamental causes of genomic disorders such as the 3q29 deletion syndrome (del3q29S). There are 21 protein-coding genes lost or gained as a result of such recurrent 1.6-Mbp deletions or duplications, respectively, in the 3q29 locus. While NAHR plays a role in CNV occurrence, the factors that increase the risk of NAHR at this particular locus are not well understood. Methods We employed an optical genome mapping technique to characterize the 3q29 locus in 161 unaffected individuals, 16 probands with del3q29S and their parents, and 2 probands with the 3q29 duplication syndrome (dup3q29S). Long-read sequencing-based haplotype resolved de novo assemblies from 44 unaffected individuals, and 1 trio was used for orthogonal validation of haplotypes and deletion breakpoints. Results In total, we discovered 34 haplotypes, of which 19 were novel haplotypes. Among these 19 novel haplotypes, 18 were detected in unaffected individuals, while 1 novel haplotype was detected on the parent-of-origin chromosome of a proband with the del3q29S. Phased assemblies from 44 unaffected individuals enabled the orthogonal validation of 20 haplotypes. In 89% (16/18) of the probands, breakpoints were confined to paralogous copies of a 20-kbp segment within the 3q29 SDs. In one del3q29S proband, the breakpoint was confined to a 374-bp region using long-read sequencing. Furthermore, we categorized del3q29S cases into three classes and dup3q29S cases into two classes based on breakpoints. Finally, we found no evidence of inversions in parent-of-origin chromosomes. Conclusions We have generated the most comprehensive haplotype map for the 3q29 locus using unaffected individuals, probands with del3q29S or dup3q29S, and available parents, and also determined the deletion breakpoint to be within a 374-bp region in one proband with del3q29S. These results should provide a better understanding of the underlying genetic architecture that contributes to the etiology of del3q29S and dup3q29S.