Novel genes are essential for evolutionary innovations and differ substantially even between closely related species. Recently, multiple studies across many taxa showed that some novel genes arise de novo, that is, from previously noncoding DNA. To characterize the underlying mutations that allowed de novo gene emergence and their order of occurrence, homologous regions must be detected within noncoding sequences in closely related sister genomes. So far, most studies do not detect noncoding homologs of de novo genes because of incomplete assemblies and annotations, and long evolutionary distances separating genomes. Here, we overcome these issues by searching for de novo expressed open reading frames (neORFs), the not-yet fixed precursors of de novo genes that emerged within a single species. We sequenced and assembled genomes with long-read technology and the corresponding transcriptomes from inbred lines ofDrosophila melanogaster, derived from seven geographically diverse populations. We found line-specific neORFs in abundance but few neORFs shared by lines, suggesting a rapid turnover. Gain and loss of transcription is more frequent than the creation of ORFs, for example, by forming new start and stop codons. Consequently, the gain of ORFs becomes rate limiting and is frequently the initial step in neORFs emergence. Furthermore, transposable elements (TEs) are major drivers for intragenomic duplications of neORFs, yet TE insertions are less important for the emergence of neORFs. However, highly mutable genomic regions around TEs provide new features that enable gene birth. In conclusion, neORFs have a high birth-death rate, are rapidly purged, but surviving neORFs spread neutrally through populations and within genomes.
Novel genes are essential for evolutionary innovations and differ substantially even between closely related species. Recently, multiple studies across many taxa have suggested that some novel genes arise de novo, i.e. from previously non-coding DNA. In order to characterise the underlying mutations that allowed de novo gene emergence and their order of occurrence, homologous regions must be detected within non-coding sequences in closely related sister genomes. So far, most studies do not detect non-coding homologs of de novo genes due to inconsistent data and long evolutionary distances separating genomes. Here we overcome these issues by searching for proto-genes, the not-yet fixed precursors of de novo genes that emerged within a single species. We sequenced and assembled genomes with long-read technology and the corresponding transcriptomes from inbred lines of Drosophila melanogaster, derived from seven geographically diverse populations. We found line-specific proto-genes in abundance but few proto- genes shared by lines, suggesting a rapid turnover. Gain and loss of transcription is more frequent than the creation of Open Reading Frames (ORFs), e.g. by forming new START- and STOP-codons. Consequently, the gain of ORFs becomes rate limiting and is frequently the initial step in proto-gene emergence. Furthermore, Transposable Elements (TEs) are major drivers for intra genomic duplications of proto-genes, yet TE insertions are less important for the emergence of proto-genes. However, highly mutable genomic regions around TEs provide new features that enable gene birth. In conclusion, proto- genes have a high birth-death rate, are rapidly purged, but surviving proto-genes spread neutrally through populations and within genomes.
scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.
customersupport@researchsolutions.com
10624 S. Eastern Ave., Ste. A-614
Henderson, NV 89052, USA
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Copyright © 2024 scite LLC. All rights reserved.
Made with 💙 for researchers
Part of the Research Solutions Family.