Gene length and the number of exons are constrained by biological limiting factors
the verdict
INSUFFICIENT LEANING
refutedsupported
the weight of evidence
3 sources for · 0 against
Peer-reviewed studies indicate that selective pressures and functional requirements influence gene structural features such as exon number and length, but comprehensive biological limiting factors are only partially covered.
The genus Brassica contains a diverse group of important vegetables and oilseed crops. Genome sequencing has been completed for the six species (B. rapa, B. oleracea, B. nigra, B. carinata, B. napus, and B. juncea) in U’s triangle model. The purpose of the study is to investigate whether positively and negatively selected genes (PSGs and NSGs) affect gene feature and function differentiation of Brassica tetraploids in their evolution and domestication. A total of 9,701 PSGs were found in the A, B and C subgenomes of the three tetraploids, of which, a higher number of PSGs were identified in the C subgenome as comparing to the A and B subgenomes. The PSGs of the three tetraploids had more tandem duplicated genes, higher single copy, lower multi-copy, shorter exon length and fewer exon number than the NSGs, suggesting that the selective modes affected the gene feature of Brassica tetraploids. The PSGs of all the three tetraploids enriched in a few common KEGG pathways relating to environmental adaption (such as Phenylpropanoid biosynthesis, Riboflavin metabolism, Isoflavonoid biosynthesis, Plant-pathogen interaction and Tropane, piperidine and pyridine alkaloid biosynthesis) and reproduction (Homologous recombination). Whereas, the NSGs of the three tetraploids significantly enriched in dozens of biologic processes and pathways without clear relationships with evolution. Moreover, the PSGs of B. carinata were found specifically enriched in lipid biosynthesis and metabolism which possibly contributed to the domestication of B. carinata as an oil crop. Our data suggest that selective modes affected the gene feature of Brassica tetraploids, and PSGs contributed in not only the evolution but also the domestication of Brassica tetraploids.
The purpose of the study is to investigate whether positively and negatively selected genes (PSGs and NSGs) affect gene feature and function differentiation of Brassica tetraploids in their evolution and domestication. A total of 9,701 PSGs were found in the A, B and C subgenomes of the three tetraploids, of which, a higher number of PSGs were identified in the C subgenome as comparing to the A and B subgenomes. The PSGs of the three tetraploids had more tandem duplicated genes, higher single copy, lower multi-copy, shorter exon length and fewer exon number than the NSGs, suggesting that the selective modes affected the gene feature of Brassica tetraploids.
To differentiate these genes, the Ka/Ks ratio was calculated for each gene, and Ka/Ks > 1, Ka/Ks = 1, and Ka/Ks < 1 were used as the indicators for positive selection, neutral evolution and negative selection (from small to large, the same number with PSGs) during gene sequence divergence, respectively. The physical location distribution of PSGs and NSGs were visualized by ‘CMplot’ R package (the widows size was 1 Mb).
A total of 100,000 runs was carried out for the two algorithms, resulted in 50,000 intersections for each. The threshold of the right tail in the 99% confidence interval was calculated for each algorithm (50,000 intersections) as the mean gene number of corresponding 50,000 intersections plus 2.33*S.D (mean+2.33*S.D., 4.68). ( Rosner, 2015 ). Finally, N com was compared with the threshold of the right tails calculated for the two algorithms (4.68). If N com falls outside of the 99% confidence intervals, the PSGs in gene set 1 will be considered to play a positive role in the domestication of the given biologic process. Other descriptive statistics were performed using Excel 2021.
These results indicate that the A. thaliana genes possibly underwent a higher evolutionary rate and experienced more selective pressure than the Brassica genes. 3.2. Gene feature analyses of PSGs and NSGs A total of 9,701 PSGs were found in all combinations, with an average of 1,617 pairs in each combination. A higher number of PSGs were identified in the combination of C subgenome (average 2,028 pairs) than in A and B subgenomes (average 1373 and 1451 pairs) ( Table 2 ). This indicates that more genes of B. oleracea (CC) were positively selected during the evolution of B. napus (AACC) and B. carinata (BBCC).
oleracea genome; BjuA, A subgenomes of B. juncea ; BnaA, A subgenomes of B. napus ; BjuB, B subgenomes of B. juncea ; BcaB, B subgenomes of B. carinata ; BnaC, C subgenomes of B. napus ; BcaC, C subgenomes of B. carinata . 2 : OGPs, Orthologous gene pairs. 3 : Ka, The rates of nonsynonymous substitution; Ks, The rates of synonymous substitution. To understand whether and how selective modes affect gene features, the tandem duplication, the copy number, the exon length and number of these PSGs and NSGs were characterized. It was found that PSGs had a significantly more tandem duplicated genes (average 5.47% vs . 1.68%), higher single copy (average 99.43% vs .
65.01%), lower multi-copy (average 0.28% vs . 17.50%), shorter exon length (average 1051 bp vs . 1798 bp) and fewer exon number (average 5.1 vs . 7.0) than the NSGs ( Table 3 ), revealing obvious varied gene features between PSGs and NSGs. This finding indicated that selective modes may serve as an alternative indicator for gene compactness. Table 3 Gene feature analysis between PSGs and NSGs. Variables 1 PSGs NSGs P-value (t-test) B. napus B. juncea B. carinata B. napus B. juncea B.
carinata Single copies 99.62% 99.22% 99.46% 61.56% 67.68% 65.79% 4.53E-05 Double copies 0.38% 0.67% 0.54% 26.06% 21.92% 25.37% 4.88E-05 > Triple copies 0% 0.11% 0% 12.38% 10.4% 8.84% 5.11E-04 Exon length (bp) 954 1040 1158 1303 1424 2667 3.89E-17 Exon number 4.8 4.9 5.7 6.6 7.3 7.0 4.39E-18 TDGs 4.91% 6.34% 5.16% 1.47% 1.53% 2.03% 6.49E-03 1 TDGs: Tandem duplicated genes. 3.3. The divergence time of Brassica species To comprehensively and systematically estimate the divergence time of the Brassica species, we investigated the distribution of the Ks in nine combinations. The Ks peak between A. thaliana and diploid Brassica species ( B. rapa , B. oleracea and B.
The SPL (SQUAMOSA promoter binding protein-like) gene family is one of the plant-specific transcription factor families and controls a considerable number of biological functions, including floral development, phytohormone signaling, and toxin resistance. However, the evolutionary patterns and driving forces of SPL genes in the Oryza genus are still not well-characterized. In this study, we investigated a total of 105 SPL genes from six AA genome Oryza representative species (O. barthii, O. glumipatula, O. nivara, O. rufipogon, O. glaberrima, and O. sativa). Phylogenetic and motif analyses indicated that SPL proteins could be divided into two distinct lineages (I and II), and further studies showed lineage II consisted of three clades (IIA, IIB, and IIC). We found that clade I had comparable structural features with clade IIA, whereas genes in clade IIC displayed intrinsic differences, such as lower exon numbers and the presence of miR156 regulation elements. Nineteen orthologous groups of OsSPLs in Oryza were also identified, and most exons within those genes maintained constant length, whereas length of intron changed relatively. All groups were constrained by stronger purifying selection and diversified continually including alterative gene number, intron length, and miR156 regulation. Subsequently, cis-acting element analyses revealed the potential role of SPLs in wild rice, which might participate in light-responsive, phytohormone response, and plant growth and developme
Abstract The SPL ( SQUAMOSA promoter binding protein-like) gene family is one of the plant-specific transcription factor families and controls a considerable number of biological functions, including floral development, phytohormone signaling, and toxin resistance. However, the evolutionary patterns and driving forces of SPL genes in the Oryza genus are still not well-characterized. In this study, we investigated a total of 105 SPL genes from six AA genome Oryza representative species ( O. barthii , O. glumipatula, O. nivara, O. rufipogon, O. glaberrima , and O. sativa ).
Phylogenetic and motif analyses indicated that SPL proteins could be divided into two distinct lineages (I and II), and further studies showed lineage II consisted of three clades (IIA, IIB, and IIC). We found that clade I had comparable structural features with clade IIA, whereas genes in clade IIC displayed intrinsic differences, such as lower exon numbers and the presence of miR156 regulation elements. Nineteen orthologous groups of OsSPL s in Oryza were also identified, and most exons within those genes maintained constant length, whereas length of intron changed relatively.
All groups were constrained by stronger purifying selection and diversified continually including alterative gene number, intron length, and miR156 regulation. Subsequently, cis -acting element analyses revealed the potential role of SPL s in wild rice, which might participate in light-responsive, phytohormone response, and plant growth and development. Our results shed light on that different evolutionary rates and duplication events might result in divergent evolutionary patterns in each lineage of SPL genes, providing a guide in exploring diverse function in the rice gene family among six closely related Oryza species.
Gene Structure, Motif, and Homology Analysis Exon/intron site and length data were extracted based on six respective genome annotation GFF files from Ensembl Plants ( Bolser et al., 2017 ). The software MEME Suite 5.0.2 ( Bailey et al., 2009 ) was employed to identify conserved motifs with the maximum number 20. To predict putative functions of identified motifs, the consensus sequences were subjected to search against the Interpro database ( Mitchell et al., 2018 ). The phylogenetic tree combined with motif arrangement was drawn by EvolView v2 ( He et al., 2016 ), and exon/intron structures were shown by TBtools v0.665 proportionally ( Chen et al., 2018 ).
Thus, intron/exon structures of each orthologous group were generated based on genome sequences and corresponding coding sequences ( Figure 3A ). In clade I and IIA, SPL s contained more than 10 exons, while genes in IIB only harbored three exons; according to this, proteins of clade I and IIA had the long C terminus with more than 700 aa residues ( Supplementary Figure S3 and Table S1 ). Despite that exon copy number was constant in clade IIB or clade IIA, the first two introns elongated or shortened in distinct groups.
In clade IIC, we found all groups harbored four or fewer exons and divided them into four subclades IIC-1 to -4 based on phylogenetic and ohnologous relationships ( Figure 3B ). Most genes in one group containing a similar structure (exon/intron number and length) even belonged to different species. There was one exception in group 4, where the last exon of cultivated rice SPL s ( OsSPL4 and OglaSPL4 ) possessed a shorter length than other wild rice SPL s. In general, introns exhibited significant length change, whereas most exon maintained a relatively constant length during the course of the Oryza evolution.
Colorful stars display paralogous pairs in each species. (B) Orthologous groups in clade IIC. (C) Exons–introns and untranslated regions (UTRs) of SPL genes. Green boxes indicate exons; black lines indicate introns. The length of protein can be estimated using the scale at the bottom. To investigate the selective pressure of SPL s, we calculated the number of non-synonymous substitutions per non-synonymous site (Ka), synonymous (Ks), and Ka/Ks ratio in each ortholog group. A statistically significant Ka/Ks ratio, equal to 1.0, meant neutral or absence of evolution. Whereas lower than or greater than that represent purifying selection and positive selection, separately.
(C) Exons–introns and orthologous groups in clade I based on a phylogenetic tree. Colorful stars display paralogous pairs in each species; yellow boxes indicate exons; and black lines indicate introns. The
Steady-state levels of mRNA in cells theoretically depend on the rate and efficiency of transcription and posttranscriptional processing, on mRNA stability, on transcriptional interference from other genes, and on poorly defined long-range chromatin effects. Although each of these cellular processes has been studied in detail for a few genes, it is not possible to predict expression levels by simply examining gene sequences. In this report, we have used a bioinformatics approach to identify critical factors that influence expression levels. To simplify the problem, we have limited our analysis to the collection of genes expressed in all tissues, because such genes provide a unique opportunity to distinguish the role of general genomic features that constrain gene expression from the effect of tissue-specific factors. Using correlation and regression techniques, we have investigated the dependence between expression level and morphological parameters (distance to neighbors, gene, mRNA or 3'-UTR length, number of exons, etc.) that can be directly related to transcription, posttranscriptional processing, mRNA stability, or transcriptional interference. We found that, on a genome-wide scale, highly expressed genes are significantly farther from their closest neighboring genes, are smaller, contain a moderate number of exons, and produce shorter mRNAs with shorter 3'-UTRs. This confirms that transcriptional and posttranscriptional processes are highly interrelated and implies that transcriptional interference plays a role in determining steady-state levels of mRNA in cells.
To simplify the problem, we have limited our analysis to the collection of genes expressed in all tissues, because such genes provide a unique opportunity to distinguish the role of general genomic features that constrain gene expression from the effect of tissue-specific factors.Using correlation and regression techniques, we have investigated the dependence between expression level and morphological parameters (distance to neighbors, gene, mRNA or 3′-UTR length, number of exons, etc.) that can be directly related to transcription, posttranscriptional processing, mRNA stability, or transcriptional interference.We found that, on a genome-wide scale, highly expressed genes are significantly farther from their closest neighboring genes, are smaller, contain a moderate number of exons, and produce shorter mRNAs with shorter 3′-UTRs.This confirms that transcriptional and posttranscriptional processes are highly interrelated and implies that transcriptional interference plays a role in determining steady-state levels of mRNA in cells.
To investigate these effects, and whether or not they account for substantial portions of the effect of gene length, we considered the number of exons as a
( A - C, left panels) Scatter plots of ln(mRNA/cell) versus ln(exon number), ln(mRNA length), and ln(3′-UTR length), respectively. The red curves represent lowess smooths, and the black curves lowess smooths from random permutations. mRNA expression has strong negative associations with all these parameters. ( A - C, right panels) Added variable plots for ln(mRNA/cell) versus ln(gene length) after correcting for ln(exon number), ln(mRNA length), and ln(3′-UTR length), respectively. Red and black curves are again lowess smooths on original data and random permutations; (res) residual.
The right panels of Figure 3, A,B,C , contain the added variable plots (see Methods) for expression on gene length after correction for exon number, mRNA length, and 3′-UTR length, respectively. In all cases, the correlation coefficients, and the lowess smooths (red curves) in comparison to lowess smooths from random permutations (black curves), establish a weakened but still significant negative association with respect to the original plot ( Fig. 2 , inset).
4D ) reveals a stronger positive relationship (the correlation increases to 0.163, with a p -value of 0.001, and the upward pattern in the lowess smooth is more marked). We then examined the relationship between copies of mRNA per cell and exon number, gene length, mRNA length, and 3′-UTR length separately for the urban and rural classes ( Fig. 5A,B,C,D ). The lowess smooths show how the association between mRNA expression and gene length, mRNA length, 3′-UTR length, and exon number is substantially weakened in the urban but not in the rural class.
Scatter plots of ln(mRNA/cell) versus ln(gene length), ln(mRNA length) ln(3′-UTR length), and ln(exon number), with lowess smooths for rural and urban classes (blue and green curves), and overall (red curves). Correlations by class are in the upper right corner. The relationships between expression and each of the parameters are considerably weakened for urban genes, indicating that transcriptional interference might be dominant over the other morphological parameters influencing gene expression. (corr) Correlation coefficient. p -values are in parentheses.
Our results support a strong negative association between mRNA expression and mRNA stability proxied by mRNA and 3′-UTR length, and imply a complex role for splicing, with a pattern that is negative at large exon numbers, but positive at low exon numbers. This might be explained by recent findings indicating that splicing factors have a stimulatory effect on transcriptional elongation (Fong et al. 2001). For genes containing few exons, this elongation stimulus might overcome the negative effect of abortive splicing events.
For example, the quantity on the vertical axis of Figure 3A , right panel, res[ln(mRNA/cell) ln(exon number)], represents vertical distances between points and the red curve in Figure 3A , left panel. The quantity on the horizontal axis of Figure 3A , right panel, res[ln(gene length) ln(exon number)], represents vertical distances from an analogous lowess for ln(gene length) on ln(exon number) (data not shown). Thus, when computing a lowess on the points of the AVP (red curve in Fig.
Everything we examined (3)
This check searched the claim as stated. It did not run a separate search for evidence against it.