Biogeographical ancestry (BGA) prediction is a valuable Forensic DNA Intelligence (FDI) tool used to provide investigators with additional information about the donor of biological evidence. FDI is particularly useful in cases where routine DNA analysis has not led to an identification. Over the last two decades, substantial research effort has been directed towards the development of new panels of ancestry informative markers and analysis tools for BGA prediction. These efforts have supported the successful application of BGA prediction in casework across several international jurisdictions. This paper identifies critical components of operational BGA prediction methods to highlight considerations for method development, optimisation, validation and associated interpretation and to identify areas where further research is needed to enhance capabilities.
The successful application of Forensic Investigative Genetic Genealogy (FIGG) to the identification of unidentified human remains and perpetrators of serious crime has led to a growing interest in its use internationally, including Australia. Routinely, FIGG has relied on the generation of high-density single nucleotide polymorphism (SNP) profiles from forensic samples using whole genome array (WGA) (similar to 650,000 or more SNPs) or whole genome sequencing (WGS) (millions of SNPs) for DNA segment-based comparisons in commercially available genealogy databases. To date, this approach has required DNA of a quality and quantity that is often not compatible with forensic samples. Furthermore, it requires the management of large data sets that include SNPs of medical relevance. The ForenSeqTM Kintelligence kit, comprising of 10,230 SNPs including 9867 for kinship association, was designed to overcome these challenges using a targeted amplicon sequencing-based method developed for low DNA inputs, inhibited and/or degraded forensic samples. To assess the ability of the ForenSeq (TM) Kintelligence workflow to correctly predict biological relationships, a comparative study comprising of 12 individuals from a family (with varying degrees of relatedness from 1st to 6th degree relatives) was undertaken using ForenSeq (TM) Kintelligence and a WGA approach using the Illumina Global Screening Array-24 version 3.0 Beadchip. All expected 1st, 2nd, 3rd, 4th and 5th degree relationships were correctly predicted using ForenSeq (TM) Kintelligence, while the expected 6th degree relationships were not detected. Given the (often) limited availability of forensic samples, findings from this study will assist Australian Law enforcement and other agencies considering the use of FIGG, to determine if the ForenSeqTM Kintelligence is suitable for existing workflows and casework sample types considered for FIGG.
The successful application of Forensic Investigative Genetic Genealogy (FIGG) to the identification of unidentified human remains and perpetrators of serious crime has led to a growing interest in its use internationally, including Australia. Routinely, FIGG has relied on the generation of high-density single nucleotide polymorphism (SNP) profiles from forensic samples using whole genome array (WGA) (∼650,000 or more SNPs) or whole genome sequencing (WGS) (millions of SNPs) for DNA segment-based comparisons in commercially available genealogy databases. To date, this approach has required DNA of a quality and quantity that is often not compatible with forensic samples. Furthermore, it requires the management of large data sets that include SNPs of medical relevance. The ForenSeq™ Kintelligence kit, comprising of 10,230 SNPs including 9867 for kinship association, was designed to overcome these challenges using a targeted amplicon sequencing-based method developed for low DNA inputs, inhibited and/or degraded forensic samples. To assess the ability of the ForenSeq™ Kintelligence workflow to correctly predict biological relationships, a comparative study comprising of 12 individuals from a family (with varying degrees of relatedness from 1st to 6th degree relatives) was undertaken using ForenSeq™ Kintelligence and a WGA approach using the Illumina Global Screening Array-24 version 3.0 Beadchip. All expected 1st, 2nd, 3rd, 4th and 5th degree relationships were correctly predicted using ForenSeq™ Kintelligence, while the expected 6th degree relationships were not detected. Given the (often) limited availability of forensic samples, findings from this study will assist Australian Law enforcement and other agencies considering the use of FIGG, to determine if the ForenSeq™ Kintelligence is suitable for existing workflows and casework sample types considered for FIGG.
Biogeographical ancestry (BGA) inference can generate valuable investigative leads when STR-based identification is not immediately available. However, the absence of suitable Filipino ancestry-informative markers (AIMs) and reference population databases poses challenges in developing and establishing a BGA inference capability to complement current forensic DNA analysis methods. Complex genetic relationships exist among Filipino groups resulting from ancient and recent migrations, admixture, and isolation driven by socio-cultural, economic, and geographical factors. Assessments of existing forensic BGA assays indicate limited informativeness in local investigations within the Asia Pacific, including the Philippines. This review highlights the need to identify more suitable AIMs and establish a suitable population database from different groups for more informative forensic BGA inference in the Philippines. This paper reviews considerations relevant to the application of forensic BGA inference in the Philippines. The challenges associated with establishing representative population datasets, identifying ancestry-informative markers and capability implementation are highlighted.
DNA methylation plays essential roles in regulating physiological processes, from tissue and organ development to gene expression and aging processes and has emerged as a widely used biomarker for the identification of body fluids and age prediction. Currently, methylation markers are targeted independently at specific CpG sites, as part of a multiplexed assay, rather than through a unified assay. Methylation detection is also dependent on divergent methodologies, ranging from enzyme digestion and affinity enrichment to bisulfite treatment, alongside various technologies for high-throughput profiling, including microarray and sequencing. In this pilot study, we test the simultaneous identification of age-associated and body fluid-specific methylation markers using a single technology, nanopore adaptive sampling. This innovative approach enables the profiling of multiple CpG marker sites across entire gene regions from a single sample without the need for specialized DNA preparation or additional biochemical treatments. Our study demonstrates that adaptive sampling achieves sufficient coverage in regions of interest to accurately determine the methylation status, shows a robust consistency with whole-genome bisulfite sequencing data, and corroborates known CpG markers of age and body fluids. Our work also resulted in the identification of new sites strongly correlated with age, suggesting new possible age methylation markers. This study lays the groundwork for the systematic development of nanopore-based methodologies in both age prediction and body fluid identification, highlighting the feasibility and potential of nanopore adaptive sampling while acknowledging the need for further validation and expansion in future research.
The successful application of forensic genetic genealogy (FGG) to identify Jane and John Doe cases in the United States has raised the prospect of using the technique in Australia to assist in the reconciliation of unidentified human remains (UHRs) with long term missing persons. A study was conducted to explore the feasibility of FGG using whole genome array (WGA) data from both pristine control samples as well as compromised casework samples, with the view to explore how DNA quantity and quality impacted on the ability to generate search results when compared to a genetic genealogy database, such as GEDmatch. From this study, several insights were gained as to the impact DNA quantity and degradation had on the percentage of SNPs genotyped and heterozygote/homozygote ratio - which are critical for successful matching outcomes. It was noted in this study (using a control sample) that successful matching occurred when genotyping errors were 5% or less. Two UHR cases were matched to kits on GEDmatch PRO, which provided investigative leads for identification purposes. The effectiveness of the FGG approach to match casework samples (as well as volunteer samples used in the study) is indicative of the usage of 'direct-to-consumer' (DTC) genetic testing by Australians. Given the (often) limited availability of casework samples, findings from this study will assist Australian agencies considering the use of FGG, to determine if WGA is a suitable method for their application.
As an emerging technology, Rapid DNA has demonstrated its utility for law enforcement in the provision of DNA profiling data at the point of arrest, often not requiring analyst review of the profiles generated. Recently, efforts have centred on the evaluation of Rapid DNA (without analyst review) and modified Rapid DNA (requiring review by a trained analyst) for application to crime scene samples. In a broader forensic context, however, another application for Rapid DNA is its use to process post-mortem samples to assist with the identification of deceased persons; and while gaps in our knowledge remain as to how Rapid DNA instruments perform with these sample types (often compromised with regards to their yield and quality of DNA), they have been successfully deployed in the field to assist in the identification of disaster victims (as exemplified during the 2018 Californian wildfire). This review aims to provide the current research landscape for the forensic application of Rapid DNA as an emerging technology from a Disaster Victim Identification perspective.
The advancement in DNA sequencing technologies and its application to forensic analysis has enabled an expansion of forensic capabilities. This chapter reviews the evolution of first-generation sequencing to second-and third-generation technologies, chemistries, platforms, and forensic considerations.
Forensic DNA Phenotyping (FDP) is an established but evolving field of DNA testing. It provides intelligence regarding the appearance (externally visible characteristics), biogeographical ancestry and age of an unknown donor and, although not necessarily a requirement for its casework application, has been previously used as a method of last resort in New South Wales (NSW) Police Force investigations. FDP can further assist law enforcement agencies by re-prioritising an existing pool of suspects or generating a new pool of suspects. In recent years, this capability has become ubiquitous with a wide range of service providers offering their expertise to law enforcement and the public. With the increase in the number of providers offering FDP and its potential to direct and target law enforcement resources, a thorough assessment of the applicability of these services was undertaken. Six service providers of FDP were assessed for suitability for NSW Police Force casework based on prediction accuracy, clarity of reporting, limitations of testing, cost and turnaround times. From these assessment criteria, a service provider for the prediction of biogeographical ancestry, hair and eye colour was deemed suitable for use in NSW Police Force casework. Importantly, the study highlighted the need for standardisation of terminology and reporting in this evolving field, and the requirement for interpretation by biologists with specialist expertise to translate the scientific data to intelligence for police investigators.
DNA methylation plays a fundamental role in the control of gene expression and genome integrity. Although there are multiple tools that enable its detection from Nanopore sequencing, their accuracy remains largely unknown. Here, we present a systematic benchmarking of tools for the detection of CpG methylation from Nanopore sequencing using individual reads, control mixtures of methylated and unmethylated reads, and bisulfite sequencing. We found that tools have a tradeoff between false positives and false negatives and present a high dispersion with respect to the expected methylation frequency values. We described various strategies to improve the accuracy of these tools, including a consensus approach, METEORE ( https://github.com/comprna/METEORE ), based on the combination of the predictions from two or more tools that shows improved accuracy over individual tools. Snakemake pipelines are also provided for reproducibility and to enable the systematic application of our analyses to other datasets.
DNA intelligence, and particularly the inference of biogeographical ancestry (BGA) is increasing in interest, and relevance within the forensic genetics community. The majority of current MPS-based forensic ancestry-informative assays focus on the differentiation of major global populations. The recently published MAPlex (Multiplex for the Asia Pacific) panel contains 144 SNPs and 20 microhaplotypes and aims to improve the differentiation of populations in the Asia Pacific region. This study reports the first forensic evaluation of the MAPlex panel using AmpliSeq technology and Ion S5 sequencing. This study reports on the overall performance of MAPlex including the assay's sequence coverage distribution and stability, baseline noise and description of problematic SNPs. Dilution series, artificially degraded and mixed DNA samples were also analysed to evaluate the sensitivity of the panel with challenging or compromised forensic samples. As the first panel to combine biallelic SNPs, multiple-allele SNPs and microhaplotypes, the MAPlex assay demonstrated an enhanced capacity for mixture detection, not easily performed with common binary SNPs. This performance evaluation indicates that MAPlex is a robust, stable and highly sensitive assay that is applicable to forensic casework for the prediction of BGA.
Forensic genetic genealogy, a technique leveraging new DNA capabilities and public genetic databases to identify suspects, raises specific considerations in a law enforcement context. Use of this technique requires consideration of its scientific and technical limitations, including the composition of current online datasets, and consideration of its scientific validity. Additionally, forensic genetic genealogy needs to be considered in the relevant legal context to determine the best way in which to make use of its potential to generate investigative leads while minimising its impact on individual privacy. This article presents these issues from an Australian perspective, with the observations and conclusions likely to be applicable to other jurisdictions.
Current forensic ancestry-informative panels are limited in their ability to differentiate populations in the Asia-Pacific region. MAPlex (Multiplex for the Asia-Pacific), a massively parallel sequencing (MPS) assay, was developed to improve differentiation of East Asian, South Asian and Near Oceanian populations found in the extensive cross-continental Asian region that shows complex patterns of admixture at its margins. This study reports the development of MAPlex; the selection of SNPs in combination with microhaplotype markers; assay design considerations for reducing the lengths of microhaplotypes while preserving their ancestry-informativeness; adoption of new population-informative multiple-allele SNPs; compilation of South Asian-informative SNPs suitable for forensic AIMs panels; and the compilation of extensive reference and test population genotypes from online whole-genome-sequence data for MAPlex markers. STRUCTURE genetic clustering software was used to gauge the ability of MAPlex to differentiate a broad set of populations from South and East Asia, the West Pacific regions of Near Oceania, as well as the other globally distributed population groups. Preliminary assessment of MAPlex indicates enhanced South Asian differentiation with increased divergence between West Eurasian, South Asian and East Asian populations, compared to previous forensic SNP panels of comparable scale. In addition, MAPlex shows efficient differentiation of Middle Eastern individuals from Europeans. MAPlex is the first forensic AIM assay to combine binary and multiple-allele SNPs with microhaplotypes, adding the potential to detect and analyze mixed source forensic DNA.
AbstractThe ability to provide accurate DNA-based forensic intelligence requires analysis of multiple DNA markers to predict the biogeographical ancestry (BGA) and externally visible characteristics (EVCs) of the donor of biological evidence. Massively parallel sequencing (MPS) enables the analysis of hundreds of DNA markers in multiple samples simultaneously, increasing the value of the intelligence provided to forensic investigators while reducing the depletion of evidential material resulting from multiple analyses. The Precision ID Ancestry Panel (formerly the HID Ion AmpliSeq™ Ancestry Panel) (Thermo Fisher Scientific) (TFS)) consists of 165 autosomal SNPs selected to infer BGA. Forensic validation criteria were applied to 95 samples using this panel to assess sensitivity (1 ng-15 pg), reproducibility (inter- and intra-run variability) and effects of compromised and forensic casework type samples (artificially degraded and inhibited, mixed source and aged blood and bone samples). BGA prediction accuracy was assessed using samples from individuals who self-declared their ancestry as being from single populations of origin (n = 36) or from multiple populations of origin (n = 14). Sequencing was conducted on Ion 318™ chips (TFS) on the Ion PGM™ System (TFS). HID SNP Genotyper v4.3.1 software (TFS) was used to perform BGA predictions based on admixture proportions (continental level) and likelihood estimates (sub-population level). BGA prediction was accurate at DNA template amounts of 125pg and 30pg using 21 and 25 PCR cycles respectively. HID SNP Genotyper continental level BGA assignments were concordant with BGAs for self-declared East Asian, African, European and South Asian individuals. Compromised, mixed source and admixed samples, in addition to sub-population level prediction, requires more extensive analysis.
Massively parallel sequencing (MPS) of identity informative single-nucleotide polymorphisms (IISNPs) enables hundreds of forensically relevant markers to be analysed simultaneously. Generating DNA sequence data enables more detailed analysis including identification of sequence variations between individuals. The GeneRead DNAseq 140 IISNP MPS panel (QIAGEN) has been evaluated on both the MiSeq (Illumina) and Ion PGM™ (Applied Biosystems) MPS platforms using the GeneRead DNAseq Targeted Panels V2 library preparation workflow (QIAGEN). The aims of this study were to (1) determine if the GeneRead DNAseq panel is effective for identity testing by assessing deviation from Hardy-Weinberg (HWE) and pairwise linkage equilibrium (LE); (2) sequence samples with the GeneRead DNAseq panel on the Ion PGM™ using the QIAGEN workflow and assess specificity, sensitivity and accuracy; (3) assess the efficacy of adding biological samples directly to the GeneRead DNAseq PCR, without prior DNA extraction; and (4) assess the effect of varying coverage and allele frequency thresholds on genotype concordance. Analyses of the 140 SNPs for HWE and LE using Fisher's exact tests and the sequential Bonferroni correction revealed that one SNP was out of HWE in the Japanese population and five SNP combinations were commonly out of LE in 13 of 14 populations. The panel was sensitive down to 0.3125 ng of DNA input. A direct-to-PCR approach (without DNA extraction) produced highly concordant genotypes. The setting of appropriate allele frequency thresholds is more effective for reducing erroneous genotypes than coverage thresholds.
Single nucleotide polymorphisms (SNPs) have been widely used in forensics for prediction of identity, biogeographical ancestry (BGA) and externally visible characteristics (EVCs). Single base extension (SBE) assays, most notably SNaPshot® (Thermo Fisher Scientific), are commonly used for forensic SNP genotyping as they can be employed on standard instrumentation in forensic laboratories (e.g. capillary electrophoresis). High resolution melt (HRM) analysis is an alternative method and is a simple, fast, single tube assay for low throughput SNP typing. This study compares HRM and SNaPshot®. HRM produced reproducible and concordant genotypes at 500 pg, however, difficulties were encountered when genotyping SNPs with high GC content in flanking regions and differentiating variants of symmetrical SNPs. SNaPshot® was reproducible at 100 pg and is less dependent on SNP choice. HRM has a shorter processing time in comparison to SNaPshot®, avoids post PCR contamination risk and has potential as a screening tool for many forensic applications.
Short tandem repeats are the gold standard for human identification but are not informative for forensic DNA phenotyping (FDP). Single-nucleotide polymorphisms (SNPs) as genetic markers can be applied to both identification and FDP. The concept of DNA intelligence emerged with the potential for SNPs to infer biogeographical ancestry (BGA) and externally visible characteristics (EVCs), which together enable the FDP process. For more than a decade, the SNaPshot® technique has been utilised to analyse identity and FDP-associated SNPs in forensic DNA analysis. SNaPshot is a single-base extension (SBE) assay with capillary electrophoresis as its detection system. This multiplexing technique offers the advantage of easy integration into operational forensic laboratories without the requirement for any additional equipment. Further, the SNP panels from SNaPshot® assays can be incorporated into customised panels for massively parallel sequencing (MPS). Many SNaPshot® assays are available for identity, BGA and EVC profiling with examples including the well-known SNPforID 52-plex identity assay, the SNPforID 34-plex BGA assay and the HIrisPlex EVC assay. This review lists the major forensically relevant SNaPshot® assays for human DNA SNP analysis and can be used as a guide for selecting the appropriate assay for specific identity and FDP applications.
Forensic DNA‐based intelligence, or forensic DNA phenotyping, utilises SNPs to infer the biogeographical ancestry and externally visible characteristics of the donor of evidential material. SNaPshot® is a commonly employed forensic SNP genotyping technique, which is limited to multiplexes of 30–40 SNPs in a single reaction and prone to PCR contamination. Massively parallel sequencing has the ability to genotype hundreds of SNPs in multiple samples simultaneously by employing an oligonucleotide sample barcoding strategy. This study of the Illumina MiSeq massively parallel sequencing platform analysed 136 unique SNPs in 48 samples from SNaPshot PCR amplicons generated by five established forensic DNA phenotyping assays comprising the SNPforID 52‐plex, SNPforID 34‐plex, Eurasiaplex, Pacifiplex and IrisPlex. Approximately 3 GB of sequence data were generated from two MiSeq flow cells and profiles were obtained from just 0.25 ng of DNA. Compared with SNaPshot, an average 98% genotyping concordance was achieved. Our customised approach was successful in attaining SNP profiles from extremely degraded, inhibited, and compromised casework samples. Heterozygote imbalance and sequence coverage in negative controls highlight the need to establish baseline sequence coverage thresholds and refine allele frequency thresholds. This study demonstrates the potential of the MiSeq for forensic SNP analysis.
The analysis of human population variation is an area of considerable interest in the forensic, medical genetics and anthropological fields. Several forensic single nucleotide polymorphism (SNP) assays provide ancestry-informative genotypes in sensitive tests designed to work with limited DNA samples, including a 34-SNP multiplex differentiating African, European and East Asian ancestries. Although assays capable of differentiating Oceanian ancestry at a global scale have become available, this study describes markers compiled specifically for differentiation of Oceanian populations. A sensitive multiplex assay, termed Pacifiplex, was developed and optimized in a small-scale test applicable to forensic analyses. The Pacifiplex assay comprises 29 ancestry-informative marker SNPs (AIM-SNPs) selected to complement the 34-plex test, that in a combined set distinguish Africans, Europeans, East Asians and Oceanians. Nine Pacific region study populations were genotyped with both SNP assays, then compared to four reference population groups from the HGDP-CEPH human diversity panel. STRUCTURE analyses estimated population cluster membership proportions that aligned with the patterns of variation suggested for each study population's currently inferred demographic histories. Aboriginal Taiwanese and Philippine samples indicated high East Asian ancestry components, Papua New Guinean and Aboriginal Australians samples were predominantly Oceanian, while other populations displayed cluster patterns explained by the distribution of divergence amongst Melanesians, Polynesians and Micronesians. Genotype data from Pacifiplex and 34-plex tests is particularly well suited to analysis of Australian Aboriginal populations and when combined with Y and mitochondrial DNA variation will provide a powerful set of markers for ancestry inference applied to modern Australian demographic profiles. On a broader geographic scale, Pacifiplex adds highly informative data for inferring the ancestry of individuals from Oceanian populations. The sensitivity of Pacifiplex enabled successful genotyping of population samples from 50-year-old serum samples obtained from several Oceanian regions that would otherwise be unlikely to produce useful population data. This indicates tests primarily developed for forensic ancestry analysis also provide an important contribution to studies of populations where useful samples are in limited supply.