The development of multicellular organisms relies on a symphony of spatiotemporally coordinated signals that selectively regulate gene expression. In particular, G protein-coupled receptors (GPCRs), the largest superfamily of transmembrane receptors, play a pivotal role in transducing extracellular signals into physiological outcomes. Notably, neurotransmitter GPCRs, classically associated with neuronal tissue communication, are increasingly emerging as regulators of pattern formation and morphogenesis. However, how these receptors coordinate such morphogenetic processes remains poorly understood. To address this gap, we developed and employed a coupled, machine-learning-based analytical pipeline, MAPPER 2.0, that fuses quantitative and qualitative analyses of Drosophila melanogaster wing phenotypes to robustly identify both severe and more subtle phenotypes generated by RNAi expression. We phenotypically characterized the impact of RNAi-based inhibition of the 111 GPCRs and the G-protein subunits in Drosophila, a genetic model system for investigating conserved protein and gene regulatory pathways. Severe morphological phenotypes resulted from RNAi-mediated knockdown targeting several G-proteins and neuropeptide and neurotransmitter GPCRs, with seven knockdowns exhibiting greater than 80% penetrance. Beyond these strong qualitative hits, MAPPER 2.0 revealed a broader class of more subtle phenotypes, including quantitative differences in wing size, compartmental organization, and vein patterning. Quantitative reverse transcription polymerase chain reaction and meta-analysis of RNA expression data validated that positive hits are expressed in the wing disc. Overall, MAPPER 2.0 provides a phenotypic platform for drug testing and mechanism discovery in GPCR-implicated human diseases, ranging from cancers to neurological conditions.
The G protein alpha subunit, Gαq, transduces extracellular signals from G-protein-coupled receptors (GPCRs) into the cell, playing essential roles in developmental processes such as organ size control, wound healing, and disease. Hyperactivating mutations in the Gαq are associated with Sturge-Weber syndrome and uveal melanoma, and thus, it serves as an important candidate drug target. However, the downstream mechanisms of Gαq signaling remain poorly defined, creating a bottleneck for designing more effective and targeted therapeutics. Here, we used Drosophila melanogaster wing discs to investigate the cellular and transcriptional consequences of Gαq dysregulation in a model epithelial system. We found that overexpression of Gαq in the wing discs reduces adult wing size and induces systemic developmental delay. Additional notable phenotypes include decreased apoptosis and reduced proliferation. Transcriptomic profiling reveals that the JAK/STAT signaling pathway is specifically upregulated in Gαq overexpression, but not in Gαq knockdown. Furthermore, perturbing Gαq impacts the cytoskeleton, confirmed by altered localization of phosphorylated Myosin II. Gαq overexpression in the wing disc upregulates stress-response pathways and triggers secretion of Drosophila insulin-like peptide 8 (Dilp8), a hormone that coordinates growth and developmental timing. Functional experiments confirmed that IP₃ receptor (IP₃R)-dependent calcium signaling mediates this delay and that the delay is rescued by the knockdown of Dilp8. In sum, Gαq acts as a critical regulator of epithelial growth and developmental timing via Ca2+-dependent Dilp8 signaling. These findings establish mechanistic links between GPCR signaling, tissue regeneration, and systemic developmental coordination, with broader implications for understanding Gαq-related pathologies in humans. This study explores how a protein called Gαq helps organs grow to the optimal size and shape during development, using fruit flies as a model. Gαq is part of a signaling system that controls how cells communicate and respond to their environment. We found that Gαq helps produce waves of calcium activity in developing wing tissue. When we altered Gαq levels during larval development, the adult wings became smaller. This was due to fewer cells dividing and, unexpectedly, fewer cells dying. These effects may relate to how Gαq functions in human diseases like cancer, though more research is needed. Gαq also slowed overall development. This delay was linked to the release of a signal called Dilp8, which tells the body to slow down growth so tissues can catch up. We showed that blocking Dilp8, or interfering with calcium signaling, could restore normal development speed. This means Gαq plays a role in managing developmental timing through a hormone system that coordinates growth across the body. Further genetic analysis revealed that Gαq activates several important pathways involved in immunity, growth, and cell structure. It also affects how cells connect physically and multiply, which are crucial for shaping tissues. In summary, Gαq is a key regulator of growth and timing during development. By influencing both local cell behavior and whole-body signals, it ensures that organs form correctly and in sync with the rest of the organism.
Mechanosensitive Piezo channels regulate cell division, cell extrusion, and cell death. However, systems-level functions of Piezo in regulating organogenesis remain poorly understood. Here, we demonstrate that Piezo controls epithelial cell topology to ensure precise organ growth by integrating live-imaging experiments with pharmacological and genetic perturbations and computational modeling. Notably, the knockout or knockdown of Piezo increases bilateral asymmetry in wing size. Piezo's multifaceted functions can be deconstructed as either autonomous or non-autonomous based on a comparison between tissue-compartment-level perturbations or between genetic perturbation populations at the whole-tissue level. A computational model that posits cell proliferation and apoptosis regulation through modulation of the cutoff tension required for Piezo channel activation explains key cell and tissue phenotypes arising from perturbations of Piezo expression levels. Our findings demonstrate that Piezo promotes robustness in regulating epithelial topology and is necessary for precise organ size control.
Phlebotomine sand flies are of global significance as important vectors of human disease, transmitting bacterial, viral, and protozoan pathogens, including the kinetoplastid parasites of the genus Leishmania, the causative agents of devastating diseases collectively termed leishmaniasis. More than 40 pathogenic Leishmania species are transmitted to humans by approximately 35 sand fly species in 98 countries with hundreds of millions of people at risk around the world. No approved efficacious vaccine exists for leishmaniasis and available therapeutic drugs are either toxic and/or expensive, or the parasites are becoming resistant to the more recently developed drugs. Therefore, sand fly and/or reservoir control are currently the most effective strategies to break transmission. To better understand the biology of sand flies, including the mechanisms involved in their vectorial capacity, insecticide resistance, and population structures we sequenced the genomes of two geographically widespread and important sand fly vector species: Phlebotomus papatasi, a vector of Leishmania parasites that cause cutaneous leishmaniasis, (distributed in Europe, the Middle East and North Africa) and Lutzomyia longipalpis, a vector of Leishmania parasites that cause visceral leishmaniasis (distributed across Central and South America). We categorized and curated genes involved in processes important to their roles as disease vectors, including chemosensation, blood feeding, circadian rhythm, immunity, and detoxification, as well as mobile genetic elements. We also defined gene orthology and observed micro-synteny among the genomes. Finally, we present the genetic diversity and population structure of these species in their respective geographical areas. These genomes will be a foundation on which to base future efforts to prevent vector-borne transmission of Leishmania parasites.
Background G proteins mediate cell responses to various ligands and play key roles in organ development. Dysregulation of G-proteins or Ca 2+ signaling impacts many human diseases and results in birth defects. However, the downstream effectors of specific G proteins in developmental regulatory networks are still poorly understood. Methods We employed the Gal4/UAS binary system to inhibit or overexpress Gαq in the wing disc, followed by phenotypic analysis. Immunohistochemistry and next-gen RNA sequencing identified the downstream effectors and the signaling cascades affected by the disruption of Gαq homeostasis. Results Here, we characterized how the G protein subunit Gαq tunes the size and shape of the wing in the larval and adult stages of development. Downregulation of Gαq in the wing disc reduced wing growth and delayed larval development. Gαq overexpression is sufficient to promote global Ca 2+ waves in the wing disc with a concomitant reduction in the Drosophila final wing size and a delay in pupariation. The reduced wing size phenotype is further enhanced when downregulating downstream components of the core Ca 2+ signaling toolkit, suggesting that downstream Ca 2+ signaling partially ameliorates the reduction in wing size. In contrast, Gαq -mediated pupariation delay is rescued by inhibition of IP 3 R, a key regulator of Ca 2+ signaling. This suggests that Gαq regulates developmental phenotypes through both Ca 2+ -dependent and Ca 2+ -independent mechanisms. RNA seq analysis shows that disruption of Gαq homeostasis affects nuclear hormone receptors, JAK/STAT pathway, and immune response genes. Notably, disruption of Gαq homeostasis increases expression levels of Dilp8, a key regulator of growth and pupariation timing. Conclusion Gαq activity contributes to cell size regulation and wing metamorphosis. Disruption to Gαq homeostasis in the peripheral wing disc organ delays larval development through ecdysone signaling inhibition. Overall, Gαq signaling mediates key modules of organ size regulation and epithelial homeostasis through the dual action of Ca 2+ -dependent and independent mechanisms.
The development of multicellular organisms relies on a symphony of spatiotemporally coordinated signals that regulate gene expression. G protein-coupled receptors (GPCRs) are the largest group of transmembrane receptors that play a pivotal role in transducing extracellular signals into physiological outcomes. Emerging research has implicated neurotransmitter GPCRs, classically associated with communication in neuronal tissues, as regulators of pattern formation and morphogenesis. However, how these receptors interact amongst themselves and signaling pathways to regulate organogenesis is still poorly understood. To address this gap, we performed a systematic RNA interference (RNAi)-based screening of 111 GPCRs along with 8 G α , 3 G β , and 2 G γ protein subunits in Drosophila melanogaster . We performed a coupled, machine learning-based quantitative and qualitative analysis to identify both severe and more subtle phenotypes. Of the genes screened, 25 demonstrated at least 60% penetrance of severe phenotypes with several of the most severe phenotypes resulting from the knockdown of neuropeptide and neurotransmitter GPCRs that were not known previously to regulate epithelial morphogenesis. Phenotypes observed in positive hits mimic phenotypic manifestations of diseases caused by dysregulation of orthologous human genes. Quantitative reverse transcription polymerase chain reaction and meta-analysis of RNA expression validated positive hits. Overall, the combined qualitative and quantitative characterization of GPCRs and G proteins identifies an extensive set of GPCRs involved in regulating epithelial morphogenesis and relevant to the study of a broad range of human diseases. ### Competing Interest Statement The authors have declared no competing interest.
Organ development relies on a symphony of signals that are coordinated spatiotemporally. Emerging research has implicated neurotransmitter receptors, belonging to the GPCR superfamily, as possible regulators of pattern formation and morphogenesis. However, how these receptors interact amongst themselves and with other complex signaling pathways to regulate biophysical properties of cells during organogenesis is still unknown. To address this knowledge gap, we performed a systematic RNAi-based screen of GPCRs and G-proteins, combined with deep learning-based bio-image analysis tools, to characterize cell autonomous roles of these receptors in a biophysical model system of epithelial morphogenesis, the Drosophila wing imaginal disc. A major class of receptors, whose loss in the wing disc resulted in severe morphogenetic defects, belonged to the Class A neuropeptide receptor family. RT-qPCR and fluorescent labeled gene expression reporters confirmed the predicted, but previously unreported, expression of the identified neurotransmitter receptors in wing imaginal disc cells. Further, live imaging and immunohistochemistry assays of the organ system were used to elucidate further the mechanisms of 5-HT1B in morphogen signaling and biomechanics behind organ development. We found that 5-HT1B promotes growth and cell division vein patterning through dysregulation of multiple key signaling pathways, including Mitogen-Activated Protein Kinase, Wingless/Wnt and Decapentaplegic/Bone Morphogenetic Protein signaling pathways. In-vivo imaging demonstrate that the 5-HT1B is necessary for spontaneous calcium signaling transients during in vivo development. This work elucidates the molecular mechanisms of how multiple neural GPCRs regulate epithelial physiological processes during morphogenesis.
Background New sequencing technologies have lowered financial barriers to whole genome sequencing, but resulting assemblies are often fragmented and far from ‘finished’. Updating multi-scaffold drafts to chromosome-level status can be achieved through experimental mapping or re-sequencing efforts. Avoiding the costs associated with such approaches, comparative genomic analysis of gene order conservation (synteny) to predict scaffold neighbours (adjacencies) offers a potentially useful complementary method for improving draft assemblies. Results We evaluated and employed 3 gene synteny-based methods applied to 21 Anopheles mosquito assemblies to produce consensus sets of scaffold adjacencies. For subsets of the assemblies, we integrated these with additional supporting data to confirm and complement the synteny-based adjacencies: 6 with physical mapping data that anchor scaffolds to chromosome locations, 13 with paired-end RNA sequencing (RNAseq) data, and 3 with new assemblies based on re-scaffolding or long-read data. Our combined analyses produced 20 new superscaffolded assemblies with improved contiguities: 7 for which assignments of non-anchored scaffolds to chromosome arms span more than 75% of the assemblies, and a further 7 with chromosome anchoring including an 88% anchored Anopheles arabiensis assembly and, respectively, 73% and 84% anchored assemblies with comprehensively updated cytogenetic photomaps for Anopheles funestus and Anopheles stephensi . Conclusions Experimental data from probe mapping, RNAseq, or long-read technologies, where available, all contribute to successful upgrading of draft assemblies. Our evaluations show that gene synteny-based computational methods represent a valuable alternative or complementary approach. Our improved Anopheles reference assemblies highlight the utility of applying comparative genomics approaches to improve community genomic resources.
Ticks transmit more pathogens to humans and animals than any other arthropod. We describe the 2.1 Gbp nuclear genome of the tick, Ixodes scapularis (Say), which vectors pathogens that cause Lyme disease, human granulocytic anaplasmosis, babesiosis and other diseases. The large genome reflects accumulation of repetitive DNA, new lineages of retro-transposons, and gene architecture patterns resembling ancient metazoans rather than pancrustaceans. Annotation of scaffolds representing ∼57% of the genome, reveals 20,486 protein-coding genes and expansions of gene families associated with tick–host interactions. We report insights from genome analyses into parasitic processes unique to ticks, including host ‘questing’, prolonged feeding, cuticle synthesis, blood meal concentration, novel methods of haemoglobin digestion, haem detoxification, vitellogenesis and prolonged off-host survival. We identify proteins associated with the agent of human granulocytic anaplasmosis, an emerging disease, and the encephalitis-causing Langat virus, and a population structure correlated to life-history traits and transmission of the Lyme disease agent.
BACKGROUND:Southern house mosquito Culex quinquefasciatus belongs to the C. pipiens cryptic species complex, with global distribution and unclear taxonomy. Mosquitoes of the complex can transmit human and animal pathogens, such as filarial worm, West Nile virus and avian malarial Plasmodium. Physical gene mapping is crucial to understanding genome organization, function, and systematic relationships of cryptic species, and is a basis for developing new vector control strategies. However, physical mapping was not established previously for Culex due to the lack of well-structured polytene chromosomes.METHODS:Inbreeding was used to diminish inversion polymorphism and asynapsis of chromosomal homologs. Identification of larvae of the same developmental stage using the shape of imaginal discs allowed achievement of uniformity in chromosomal banding pattern. This together with high-resolution phase-contrast photography enabled the development of a cytogenetic map. Fluorescent in situ hybridization was used for gene mapping.RESULTS:A detailed cytogenetic map of C. quinquefasciatus polytene chromosomes was produced. Landmarks for chromosome recognition and cytological boundaries for two inversions were identified. Locations of 23 genes belonging to 16 genomic supercontigs, and 2 cDNA were established. Six supercontigs were oriented and one was found putatively misassembled. The cytogenetic map was linked to the previously developed genetic linkage groups by corresponding positions of 2 genetic markers and 10 supercontigs carrying genetic markers. Polytene chromosomes were numbered according to the genetic linkage groups.CONCLUSIONS:This study developed a new standard cytogenetic photomap of the polytene chromosomes for C. quinquefasciatus and was applied for the fine-scale physical mapping. It allowed us to infer chromosomal position of 1333 of annotated genes belonging to 16 genomic supercontigs and find orientation of 6 of these supercontigs; the new cytogenetic and previously developed genetic linkage maps were integrated based on 12 matches. The map will further assist in finding chromosomal position of the medically important and other genes, contributing into improvement of the genome assembly. Better assembled C. quinquefasciatus genome can serve as a reference for studying other vector species of C. pipiens complex and will help to resolve their taxonomic relationships. This, in turn, will contribute into future development of vector and disease control strategies.
Background Anopheles stephensi is the key vector of malaria throughout the Indian subcontinent and Middle East and an emerging model for molecular and genetic studies of mosquito-parasite interactions. The type form of the species is responsible for the majority of urban malaria transmission across its range. Results Here, we report the genome sequence and annotation of the Indian strain of the type form of An. stephensi . The 221 Mb genome assembly represents more than 92% of the entire genome and was produced using a combination of 454, Illumina, and PacBio sequencing. Physical mapping assigned 62% of the genome onto chromosomes, enabling chromosome-based analysis. Comparisons between An. stephensi and An. gambiae reveal that the rate of gene order reshuffling on the X chromosome was three times higher than that on the autosomes. An. stephensi has more heterochromatin in pericentric regions but less repetitive DNA in chromosome arms than An. gambiae . We also identify a number of Y-chromosome contigs and BACs. Interspersed repeats constitute 7.1% of the assembled genome while LTR retrotransposons alone comprise more than 49% of the Y contigs. RNA-seq analyses provide new insights into mosquito innate immunity, development, and sexual dimorphism. Conclusions The genome analysis described in this manuscript provides a resource and platform for fundamental and translational research into a major urban malaria vector. Chromosome-based investigations provide unique perspectives on Anopheles chromosome evolution. RNA-seq analysis and studies of immunity genes offer new insights into mosquito biology and mosquito-parasite interactions.
Variation in vectorial capacity for human malaria among Anopheles mosquito species is determined by many factors, including behavior, immunity, and life history. To investigate the genomic basis of vectorial capacity and explore new avenues for vector control, we sequenced the genomes of 16 anopheline mosquito species from diverse locations spanning ~100 million years of evolution. Comparative analyses show faster rates of gene gain and loss, elevated gene shuffling on the X chromosome, and more intron losses, relative to Drosophila. Some determinants of vectorial capacity, such as chemosensory genes, do not show elevated turnover but instead diversify through protein-sequence changes. This dynamism of anopheline genes and genomes may contribute to their flexible capacity to take advantage of new ecological niches, including adapting to humans as primary hosts.
Ty3/gypsy elements represent one of the most abundant and diverse LTR-retrotransposon (LTRr) groups in the Anopheles gambiae genome, but their evolutionary dynamics have not been explored in detail. Here, we conduct an in silico analysis of the distribution and abundance of the full complement of 1045 copies in the updated AgamP3 assembly. Chromosomal distribution of Ty3/gypsy elements is inversely related to arm length, with densities being greatest on the X, and greater on the short versus long arms of both autosomes. Taking into account the different heterochromatic and euchromatic compartments of the genome, our data suggest that the relative abundance of Ty3/gypsy LTRrs along each chromosome arm is determined mainly by the different proportions of heterochromatin, particularly pericentric heterochromatin, relative to total arm length. Additionally, the breakpoint regions of chromosomal inversion 2La appears to be a haven for LTRrs. These elements are underrepresented more than 7-fold in euchromatin, where 33% of the Ty3/gypsy copies are associated with genes. The euchromatin on chromosome 3R shows a faster turnover rate of Ty3/gypsy elements, characterized by a deficit of proviral sequences and the lowest average sequence divergence of any autosomal region analyzed in this study. This probably reflects a principal role of purifying selection against insertion for the preservation of longer conserved syntenyc blocks with adaptive importance located in 3R. Although some Ty3/gypsy LTRrs show evidence of recent activity, an important fraction are inactive remnants of relatively ancient insertions apparently subject to genetic drift. Consistent with these computational predictions, an analysis of the occupancy rate of putatively older insertions in natural populations suggested that the degenerate copies have been fixed across the species range in this mosquito, and also are shared with the sibling species Anopheles arabiensis.
Background Transposable elements (TEs) are mobile sequences found in nearly all eukaryotic genomes. They have the ability to move and replicate within a genome, often influencing genome evolution and gene expression. The identification of TEs is an important part of every genome project. The number of sequenced genomes is rapidly rising, and the need to identify TEs within them is also growing. The ability to do this automatically and effectively in a manner similar to the methods used for genes is of increasing importance. There exist many difficulties in identifying TEs, including their tendency to degrade over time and that many do not adhere to a conserved structure. In this work, we describe a homology-based approach for the automatic identification of high-quality consensus TEs, aimed for use in the analysis of newly sequenced genomes. Results We describe a homology-based approach for the automatic identification of TEs in genomes. Our modular approach is dependent on a thorough and high-quality library of representative TEs. The implementation of the approach, named TESeeker, is BLAST-based, but also makes use of the CAP3 assembly program and the ClustalW2 multiple sequence alignment tool, as well as numerous BioPerl scripts. We apply our approach to newly sequenced genomes and successfully identify consensus TEs that are up to 99% identical to manually annotated TEs. Conclusions While TEs are known to be a major force in the evolution of genomes, the automatic identification of TEs in genomes is far from mature. In particular, there is a lack of automated homology-based approaches that produce high-quality TEs. Our approach is able to generate high-quality consensus TE sequences automatically, requiring the user to only provide a few basic parameters. This approach is intentionally modular, allowing researchers to use components separately or iteratively. Our approach is most effective for TEs with intact reading frames. The implementation, TESeeker, is available for download as a virtual appliance, while the library of representative TEs is available as a separate download.
Culex quinquefasciatus (the southern house mosquito) is an important mosquito vector of viruses such as West Nile virus and St. Louis encephalitis virus, as well as of nematodes that cause lymphatic filariasis. C. quinquefasciatus is one species within the Culex pipiens species complex and can be found throughout tropical and temperate climates of the world. The ability of C. quinquefasciatus to take blood meals from birds, livestock, and humans contributes to its ability to vector pathogens between species. Here, we describe the genomic sequence of C. quinquefasciatus : Its repertoire of 18,883 protein-coding genes is 22% larger than that of Aedes aegypti and 52% larger than that of Anopheles gambiae with multiple gene-family expansions, including olfactory and gustatory receptors, salivary gland genes, and genes associated with xenobiotic detoxification.
Physical mapping is a useful approach for studying genome organization and evolution as well as for genome sequence assembly. The availability of polytene chromosomes in malaria mosquitoes provides a unique opportunity to develop high-resolution physical maps. We report a 0.6-Mb-resolution physical map consisting of 422 DNA markers hybridized to 379 chromosomal sites of the Anopheles stephensi polytene chromosomes. This makes An. stephensi second only to Anopheles gambiae in density of a physical map among malaria mosquitoes. Three hundred sixty-three (363) probes hybridized to single chromosomal sites, whereas 59 clones yielded multiple signals. This physical map provided a suitable basis for comparative genomics, which was used for determining inversion breakpoints, duplications, and origin of novel genes across species.
As an obligatory parasite of humans, the body louse (Pediculus humanus humanus) is an important vector for human diseases, including epidemic typhus, relapsing fever, and trench fever. Here, we present genome sequences of the body louse and its primary bacterial endosymbiont Candidatus Riesia pediculicola. The body louse has the smallest known insect genome, spanning 108 Mb. Despite its status as an obligate parasite, it retains a remarkably complete basal insect repertoire of 10,773 protein-coding genes and 57 microRNAs. Representing hemimetabolous insects, the genome of the body louse thus provides a reference for studies of holometabolous insects. Compared with other insect genomes, the body louse genome contains significantly fewer genes associated with environmental sensing and response, including odorant and gustatory receptors and detoxifying enzymes. The unique architecture of the 18 minicircular mitochondrial chromosomes of the body louse may be linked to the loss of the gene encoding the mitochondrial single-stranded DNA binding protein. The genome of the obligatory louse endosymbiont Candidatus Riesia pediculicola encodes less than 600 genes on a short, linear chromosome and a circular plasmid. The plasmid harbors a unique arrangement of genes required for the synthesis of pantothenate, an essential vitamin deficient in the louse diet. The human body louse, its primary endosymbiont, and the bacterial pathogens that it vectors all possess genomes reduced in size compared with their free-living close relatives. Thus, the body louse genome project offers unique information and tools to use in advancing understanding of coevolution among vectors, symbionts, and pathogens.
BACKGROUND:The genome of Anopheles gambiae, the major vector of malaria, was sequenced and assembled in 2002. This initial genome assembly and analysis made available to the scientific community was complicated by the presence of assembly issues, such as scaffolds with no chromosomal location, no sequence data for the Y chromosome, haplotype polymorphisms resulting in two different genome assemblies in limited regions and contaminating bacterial DNA.RESULTS:Polytene chromosome in situ hybridization with cDNA clones was used to place 15 unmapped scaffolds (sizes totaling 5.34 Mbp) in the pericentromeric regions of the chromosomes and oriented a further 9 scaffolds. Additional analysis by in situ hybridization of bacterial artificial chromosome (BAC) clones placed 1.32 Mbp (5 scaffolds) in the physical gaps between scaffolds on euchromatic parts of the chromosomes. The Y chromosome sequence information (0.18 Mbp) remains highly incomplete and fragmented among 55 short scaffolds. Analysis of BAC end sequences showed that 22 inter-scaffold gaps were spanned by BAC clones. Unmapped scaffolds were also aligned to the chromosome assemblies in silico, identifying regions totaling 8.18 Mbp (144 scaffolds) that are probably represented in the genome project by two alternative assemblies. An additional 3.53 Mbp of alternative assembly was identified within mapped scaffolds. Scaffolds comprising 1.97 Mbp (679 small scaffolds) were identified as probably derived from contaminating bacterial DNA. In total, about 33% of previously unmapped sequences were placed on the chromosomes.CONCLUSION:This study has used new approaches to improve the physical map and assembly of the A. gambiae genome.