TUESDAY, SEPTEMBER 22, 2026|No. 16002
Genetics Research · Molecular Biology

Study Suggests Bacterial Gene Positioning Influenced by Expression Levels and Growth Rate

New research indicates that the placement of genes on a bacterium's circular chromosome is likely shaped by natural selection acting on both average gene expression and how expression changes with growth rate.

Microscopic view of bacterial DNA replication and gene expression.
Microscopic view of bacterial DNA replication and gene expression.
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

Abstract

In bacteria with circular chromosomes, genes near the replication origin ( oriC) are replicated earlier and consequently attain higher copy numbers than genes near the terminus ( ter), particularly during rapid growth. Hence, mutations altering a gene’s chromosomal position can affect its expression level and be subject to selection. Two non–mutually exclusive hypotheses regarding the target of this selection have been proposed. The mean expression hypothesis (MEH) posits that the target is a gene’s average expression across environments, whereas the growth-dependent expression hypothesis (GEH) proposes that the target is the growth dependence of gene expression, quantified by the expression slope—the change in expression level per unit change in growth rate. To test these hypotheses and assess their relative support, we analyze eight multi-environment protein expression datasets from three bacterial species, as well as Escherichia coli promoter strengths measured in two environments. Consistent with both MEH and GEH, we observe a significant decrease in both mean expression and expression slope from oriC to ter in six and four of the eight datasets, respectively. In regression models predicting gene position, the relative contributions of the two hypotheses differ across species. However, even when combined, the two hypotheses explain only a small fraction of chromosomal gene positioning, in part because the replication-dose effect is incompletely offset by compensatory evolution of individual promoter strengths. The positional gradients in mean and growth-dependent expression are disproportionately contributed by genes involved in translation and transcription. We conclude that patterns of bacterial chromosomal gene positioning are consistent with moderate effects of selection on both mean and growth-dependent expression.

Author summary

Bacterial gene positioning affects gene expression because genes near the origin of chromosomal replication are replicated earlier than those near the terminus and thus have higher gene copy numbers, especially during rapid population growth. Consequently, mutations altering the chromosomal position of a gene influence its expression so could be targeted by natural selection. Two hypotheses propose that the selection target is the mean expression level of a gene across environments and the growth dependence of expression, respectively. This study analyzed multi-environment protein expression datasets from three bacterial species to assess the two hypotheses. We found evidence for both hypotheses, but also discovered that, even when combined, the two hypotheses explain only a small fraction of chromosomal gene positioning, in part because the replication-dose effect is incompletely offset by compensatory evolution of individual promoter strengths. We conclude that patterns of bacterial chromosomal gene positioning are consistent with moderate effects of selection on both mean and growth-dependent expression, but more bacteria should be studied in the future to test the generality of our findings.

Figures

Fig 5

Fig 1

Table 1

Fig 2

Fig 3

Fig 4

Fig 5

Fig 1

Citation: Yuan R, Zhang J (2026) Bacterial chromosomal gene positioning is likely shaped by selection on both mean and growth-dependent expression. PLoS Genet 22(9): e1012316.

https://doi.org/10.1371/journal.pgen.1012316

Editor: Louis-Marie Bobay, North Carolina State University, UNITED STATES OF AMERICA

Received: May 23, 2026; Accepted: September 11, 2026; Published: September 21, 2026

Copyright: © 2026 Yuan, Zhang. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: Relevant data and code can be obtained from https://github.com/ruiqiy/Bacterial-gene-positioning.

Funding: This work was supported by the US National Institutes of Health research grant R35GM139484 to J.Z. The funder played no role in the study design, data collection and analysis, decision to publish, and preparation of the manuscript.

Competing interests: The authors have declared that no competing interests exist.

Introduction

Most bacterial genomes consist of a single circular chromosome [ 1]. During replication of the circular genome, two replication forks form at the chromosome replication origin ( oriC) and proceed towards the terminus ( ter). The two replichores produced in this process are typically symmetric and of the same size [ 24]. There can be more than two replication forks proceeding simultaneously in a bacterial cell during fast growth because a new round of replication can initiate before the previous one completes. The number of simultaneous replication rounds ( R) can be estimated by the ratio of the time for replication forks to travel from oriC to ter ( C) to the time interval between two successive cell divisions (i.e., doubling time G), given that exactly one round of replication is needed per cell cycle [ 1]. Because growth rate μ equals ln2/ G, R = C/ G = /ln2 [ 5].

Genes near oriC attain higher copy numbers than genes near ter during replication because the former are always replicated earlier than the latter. This replication-associated gene dosage disparity increases with growth rate, because a gene near oriC has a copy number that is 2_R_ = e times that of a gene near ter [ 4]. The relative dosage of gene X, D X, is defined as the ratio of the average copy number of X to the average copy number of a hypothetical gene at the midpoint between oriC and ter. Based on the Cooper-Helmstetter model of bacterial chromosome replication [ 6, 7], it can be shown that

(1)

Here, x is the gene position normalized to the range of [0, 1], which is the distance from the gene’s midpoint to the midpoint of oriC divided by one half of the chromosome length.

A mutation altering a gene’s chromosomal position influences its dosage, which in turn influences its mRNA or protein expression level. Indeed, experimentally inserting a transgene near oriC resulted in a higher transgene expression than inserting it near ter in both Salmonella typhimurium and Escherichia coli [ 810]. Furthermore, empirical data have shown that expression levels of artificially translocated genes are approximately proportional to the gene dose predicted by the Cooper-Helmstetter model [ 8, 9, 11]. Hence,

(2)

Here, is the relative expression level of gene X (i.e., expression level of X divided by that when it is located at the midpoint between oriC and ter). Although C is often assumed to be a constant at rapid growth [ 5, 7], studies have reported that C increases as μ decreases during slow growth [ 1214]. An empirical formula relating C with μ was obtained for E. coli by fitting experimental data [ 14]:

(3)

Plugging Eq 3 into Eq 2 yields:

(4)

Eq 4 shows that, given , declines with x. Furthermore, as rises, increases more when x is close to 0 than when it is close to 0.5 and decreases more when x is close to 1 than when it is close to 0.5. That is, gene position is expected to influence both the level of gene expression and its growth dependence. Hereinafter, we define the replication-dose effect as the effect of chromosomal replication-associated gene dose change on gene expression. Note that the replication-dose effect increases with the growth rate. We refer to this growth-dependent property of the replication-dose effect as the growth-dependent replication-dose effect.

That altering the chromosomal position of a gene influences its expression prompted the hypothesis that chromosomal gene positioning is subject to expression-related selection [ 5, 1517]. In theory, gene positioning can be subject to purifying and positive selection. First, upon the optimization of gene expression, gene position should generally be subject to purifying selection and conserved. The observation that gene position is conserved across bacterial species [ 16] despite high rates of genome rearrangement [ 4, 18] provides empirical evidence for purifying selection on gene position. Second, if gene expression is suboptimal, mutations relocating the gene could be advantageous. Empirical evidence for such adaptive gene relocation would support positive selection acting on gene positioning. However, such evidence is scarce. In Lenski’s Long-Term Evolution Experiment of 12 replicate E. coli populations, a large inversion in the same genomic region was found in three populations [ 19]. Although the repeated acceptance of this inversion could have been driven by positive selection on gene positioning, it could also be explained by positive selection for a switch of genes from the lagging to the leading strand that could prevent head-on collisions between DNA and RNA polymerases that are particularly deleterious [ 20].

To explore potential positive selection on gene positioning at the chromosomal scale, let us consider the following scenario. Genes are initially randomly thrown to a chromosome. So, their expression levels, which are jointly determined by their chromosomal positions and transcriptional regulations, are suboptimal. Optimal gene expression levels may be achieved by (i) beneficial mutations changing gene positions and/or (ii) beneficial mutations changing transcriptional regulations, especially promoter strengths. If only (i) occurs, the outcome would be a trend for the optimal gene expression level to decrease from genes near oriC to genes near ter to take advantage of the replication-dose effect. If only (ii) occurs, the outcome would be a trend for the promoter strength to increase from genes near oriC to genes near ter, to offset the replication-dose effect. If both (i) and (ii) occur, both outcomes may be observed but with weakened trends. A test of these predictions can help infer adaptive gene positioning.

Because Eq (4) indicates that gene position simultaneously influences gene expression and its growth dependence, two hypotheses propose different expression components as the target of selection and make distinct predictions about gene positioning ( Fig 1). The mean expression hypothesis (MEH) considers the expression level of a gene as the selection target regardless of the growth rate and predicts that selection pushes genes with higher optimal expression levels (i.e., higher ) toward oriC [ 17]. For example, in E. coli, Bacillus subtilis, and Streptomyces, a significant negative correlation exists between gene position x and expression level, with no association between gene function and x [ 21]. Under the assumption that the observed expression levels reflect optimal levels, these observations support MEH. Although MEH does not explicitly consider environmental variation, it predicts that genes with higher mean expression across environments () lie closer to oriC.

thumbnail

Download:

Fig 1. The mean expression hypothesis (MEH) and growth-dependent expression hypothesis (GEH).

MEH predicts that selection favors placing genes from oriC to ter in order of decreasing mean expression (), while GEH predicts that selection favors placing genes from oriC to ter in order of decreasing growth dependence in expression ( k).

https://doi.org/10.1371/journal.pgen.1012316.g001

By contrast, the growth-dependent expression hypothesis (GEH) emphasizes the growth dependence in expression as a target of selection [ 16]. One can regress on the scaled growth rate μ scaled, which equals 0.619 μ (see Eq 4), to obtain the expression slope k as a measure of the extent to which a gene’s expression is growth dependent. GEH predicts that selection pushes genes with larger positive optimal k toward oriC and genes with larger negative optimal k toward ter. An underlying assumption of GEH is that each gene has a single optimal k regardless of the specific environments concerned; otherwise, the k-based selection would vary in direction and strength depending on the environment and it would be difficult for GEH to predict gene position patterns.

Evidence supporting GEH is available from comparative genomics and experimental gene translocation. For instance, previous work reported that translation- and transcription-related genes (i.e., those encoding RNA polymerases, ribosomal RNAs, ribosomal proteins, and certain tRNAs), but not other highly expressed genes, are significantly enriched near oriC in hundreds of bacterial species [ 5]. Consistent with GEH, translocating near- oriC ribosomal protein genes to a position distal to oriC reduced cell fitness under fast-growth but not slow-growth conditions [ 22]. Furthermore, fitness at fast growth recovered when two copies, but not one copy, of the ribosomal protein genes were translocated to positions distal to oriC [ 22]. This latter finding, however, is consistent with both MEH and GEH. Hu et al. recently analyzed three proteomic and ribosome profiling datasets and reported higher k for genes with 0 0.25 h-1 and growth rate in the faster-growing condition is more than 1.5 times that in the slower-growing condition) in each dataset. For each condition pair, we tested whether genes in a chromosomal region ( oriC or ter) have significantly different expressions between the two conditions (two-sided paired Wilcoxon signed-rank test; S2 Fig). GEH predicts that oriC genes have higher expression in the faster-growth condition than in the slower-growth condition while the reverse is true for ter genes. We found the results differ by condition pair and species. For many condition pairs in E. coli datasets, the observed expression differences between conditions are consistent with the GEH predictions. Nonetheless, the results are mixed in B. subtilis and V. natriegens datasets, with some condition pairs supporting GEH and others against GEH. These results are robust to the choice of growth-rate cutoffs used to define growth-divergent condition pairs ( S3 and S4 Figs).

However, the above tests are not ideal because they do not explicitly compute the expression slope k and test GEH’s prediction of a negative correlation between k and x. To this end, we used multiple conditions in each dataset to estimate k. In four of the eight datasets, we found a significant, negative (rank and linear) correlation between k and x ( Figs 2, S5), supporting GEH. In the remaining datasets, the rank correlation was not significant ( Fig 2), although the linear correlation was significantly positive in Zhu E. coli dataset ( S5 Fig). Because most significant results support GEH, we conclude that there is overall support for GEH in the datasets studied.

The explanatory power of MEH and GEH

A key assumption of GEH is that each gene has a single optimal k across all environments. We tested this assumption by calculating the standard deviation of and the sign consistency of across the five datasets of E. coli and across the two datasets of B. subtilis, respectively. We converted k to , because k is a slope, which can naturally be viewed as an angle ranging from -90° to 90°. We found that the sign of is inconsistent between the two B. subtilis datasets for 33% of genes and across the five E. coli datasets for 56% of genes ( S6 Fig). The standard deviation of is on average 38° across the B. subtilis datasets and 37° across the E. coli datasets. Due to this high variance in , we compiled E. coli and B. subtilis consensus datasets from which we estimated inverse-variance weighted mean and k (see Materials and methods). Using these estimates, we assessed the relative contributions of and k in explaining the variation in gene position. Specifically, we fitted a beta regression model predicting gene position using z-score of and z-score of k (Model 1), based on the E. coli consensus data, B. subtilis consensus data, and Zhu V. natriegens data, respectively. MEH predicts a negative regression coefficient for while GEH predicts a negative regression coefficient for k. The hypothesis with the more negative regression coefficient contributes more to chromosomal gene positioning.

In the beta regression model fitted for the E. coli consensus data, both and k have significantly negative regression coefficients ( Fig 3, S1 Table), supporting both MEH and GEH. Additionally, k has a significantly larger effect size than , suggesting that GEH explains gene position better than MEH in E. coli. For B. subtilis, both and k have negative, albeit non-significant, regression coefficients. Effect sizes of and k in the model are also similar. Hence, neither GEH nor MEH is significantly supported in B. subtilis. For V. natriegens, has a significantly negative regression coefficient while k does not have a significant coefficient. The difference in effect size between and k is significant. Hence, MEH but not GEH is significantly supported in V. natriegens. Together, the results from the three species indicate that the relative importance of the two hypotheses varies with the species.

thumbnail

Download:

Fig 3. Beta regression models contrasting support for GEH and MEH.

Each circle shows a regression coefficient for mean expression () or growth dependence in expression ( k) in a beta regression model, with the error bar indicating its 95% confidence interval. Stars or N. S. below error bars indicate whether the regression coefficient is significantly different from 0 or not (Wald test). Stars or N. S. on brackets indicate whether the two regression coefficients in the model are significantly different or not (Wald test). N. S., not significant. ", P < 0.05. "", P < 0.01. """, P < 0.001.

https://doi.org/10.1371/journal.pgen.1012316.g003

To estimate the total explanatory power of and k for gene position, we calculated the squared Pearson correlation coefficient ( r 2) between observed and predicted positions as a measure of goodness-of-fit for the beta regression model. We found that r 2 ranged from 0.007 to 0.039 in the three species ( S1 Table), suggesting that, even combined, and k have limited explanatory power for gene position.

Expression adaptation through changes in gene position and promoter strength

While the above sections have demonstrated that gene position patterns are consistent with expression-related selection, a complementary question is whether the mean expression and growth-dependent expression have been optimized through changes in both gene position and promoter strength. To address this question, we analyzed previously quantified strengths of all candidate promoters in E. coli [ 33]. Specifically, Urtecho et al. sheared the E. coli genome into 200- to 300-nucleotide fragments with 8.5 × coverage. They then built a library where DNA fragments were placed upstream of the superfolder green fluorescent protein (sfGFP) gene including a barcode, inserted the library into a defined intergenic location in the E. coli genome, and performed targeted amplicon sequencing of the barcoded sfGFP transcripts to quantify the RNA levels of barcodes normalized to their DNA abundances in LB medium and M9 medium, respectively. For each nucleotide site in the genome, they calculated the median ratio of barcode RNA abundance to barcode DNA abundance across all fragments covering the site, which was considered the promoter activity of the site. They defined a candidate promoter as a genomic region where the promoter activity is continuously higher than an empirical threshold for at least 60 nucleotides. In the present study, we mapped the candidate promoters to genes they regulate and used the highest expression level within a promoter sequence as a proxy for the promoter strength (see Materials and methods).

For each E. coli promoter, we computed its arithmetic mean of ln-transformed promoter strength in LB and M9 as a measure of its mean strength. We found a significant positive correlation between gene position x and mean promoter strength (Pearson’s r = 0.038, P = 0.041; Fig 4a). By contrast, there is a significant negative correlation between gene position x and observed in the E. coli consensus data (Pearson’s r = -0.058, P = 0.018; Fig 4a). The negative correlation between x and suggests that the replication-dose effect is used for mean expression optimization, whereas the positive correlation between x and mean promoter strength suggests that the optimization of individual promoter strengths has an overall effect of partially offsetting the replication-dose effect.

thumbnail

Download:

Fig 4. Expression and promoter strengths of E. coli genes.

Each dot represents a gene. Expression is based on the E. coli consensus data. (a) Mean expression level (black) and mean promoter strength (blue) in LB and M9 media for each gene. (b) Growth dependence in expression (black) and promoter strength (orange) in LB and M9 for each gene. A line shows the linear regression of the dots of the same color as the line, with the regression equation and P-value of the null hypothesis of zero Pearson’s correlation coefficient also indicated in that color. Growth dependence in promoter strength is measured by the difference in ln(promoter strength) between LB and M9. For visual clarity, y-axis range for each variable is restricted to its 1st to 99th percentile but all data points are included in statistical analyses.

https://doi.org/10.1371/journal.pgen.1012316.g004

For each E. coli promoter, we also computed the difference in ln-transformed promoter strength between LB and M9 media as a measure of growth dependence in promoter strength. We found no significant correlation between gene position ( x) and growth dependence in promoter strength (Pearson’s r = -0.009, P = 0.63; Fig 4b). By contrast, gene position x and observed k are significantly negatively correlated in the E. coli consensus data (Pearson’s r = -0.193, P = 9.7 × 10-16; Fig 4b). This negative correlation suggests that the growth-dependent replication-dose effect is used for optimizing growth-dependent expression. The lack of a significant correlation between gene position and growth dependence in promoter strength suggests that the optimization of the growth-dependent strengths of individual promoters likely neither reinforces nor offsets the growth-dependent replication-dose effect.

Bacterial genomes contain core and accessory genes, which are respectively shared and unshared across different genotypes of a species. Because core genes are more stably retained in genomes than accessory genes, we tested whether the relative contributions of the replication-dose effect and promoter strength to mean-expression optimization differ between core and accessory genes. Specifically, we constructed interaction models predicting mean promoter strength or by gene position x, pangenome class C (core genes coded as 0 and accessory genes coded as 1), and the interaction between x and C (see Materials and methods). We detected a significant positive interaction effect on in the E. coli consensus data, but no significant interaction effect on in the B. subtilis consensus data or Zhu V. natriegens dataset ( S7 Fig). This means that the effect of gene position x on is more negative for core genes than for accessory genes in E. coli, supporting the conjecture of a stronger replication-dose effect in the optimization of mean expression of core genes than accessory genes. Furthermore, we detected a significant positive interaction effect on the mean promoter strength in the E. coli consensus data, meaning that the effect of x on the mean promoter strength is more positive for accessory than core genes in E. coli, consistent with the conjecture of a more important role of promoter strength evolution in the mean expression optimization of accessory genes than that of core genes. We fitted similar models for k and growth dependence in promoter strength to test whether optimization of growth dependence in expression or growth dependence in promoter strength differed for core and accessory genes in mechanism but found no significant interaction effect in any dataset ( S7 Fig).

Disproportional contribution of translation and transcription genes to positional gradients in mean and growth-dependent expression

Our analyses have thus far shown an overall pattern of decreasing and k from oriC to ter across genes on the chromosome. We next tested whether these trends are similar for all genes or are especially prominent for translation and transcription (TT) genes, defined as ribosomal protein genes and RNA polymerase genes hereinafter. We first built ordinary least squares models predicting gene position x by or k. The regression coefficients estimate positional gradients in or k. We then examined changes in the regression coefficients when TT genes were excluded and compared the results with those when an equal number of randomly picked genes were excluded (see Materials and methods). In the E. coli consensus data, B. subtilis consensus data, and Zhu V. natriegens dataset, changes of regression coefficients are consistently larger by the exclusion of TT genes than by the exclusion of the same number of random genes ( P 0.011; Fig 5), meaning that TT genes are more important than random genes in driving the positional gradients in both and k. Moreover, our results show that, relative to the contribution of random genes, the contribution of TT genes to the positional gradient is greater in the case of than in the case of k (see Z TT in Fig 5).

thumbnail

Download:

Fig 5. Disproportionally larger contribution of translation and transcription (TT) genes to positional gradients in mean expression () and growth dependence in expression ( k).

Each graph shows the distribution (blue bars) of the change in regression coefficient caused by the removal of the same number of genes as TT genes from all genes ( n = 10,000 replicates of random gene exclusions). The red dashed line indicates , the change in regression coefficient caused by the removal of TT genes from all genes. One-sided empirical P-value indicates the fraction of random exclusions yielding equal or greater changes in the positional regression coefficient than . measures TT genes’ contribution to the positional gradient relative to that of the same number of random genes.

https://doi.org/10.1371/journal.pgen.1012316.g005

Because most TT genes are core genes in the datasets analyzed, we tested whether the previously observed difference between core genes and accessory genes in the mechanism of mean expression optimization remains when TT genes are excluded. The interaction effects decrease and become non-significant for both and mean promoter strength when TT genes are excluded ( S7 Fig). Thus, the observed difference between core genes and accessory genes in mean expression optimization is largely attributable to the enrichment of TT genes in core genes.

Discussion

Because the position of a gene on a bacterial chromosome influences its mean expression as well as growth-dependent expression, MEH and GEH respectively posit that selection on mean and growth-dependent expression shapes gene positioning. We tested these two hypotheses using multi-environment protein expression and growth rate data and made the following observations. First, in six of the eight datasets analyzed, significantly negatively correlates with x across genes, whereas the reverse is true in none of the eight datasets. Similarly, in four of the eight datasets, k significantly negatively correlates with x across genes, whereas the reverse is true in only one of the eight datasets. Hence, both MEH and GEH are supported, although the level of support varies among species and datasets. Second, we fitted beta regression models to quantify and partition the contributions of MEH and GEH to chromosomal gene positioning. We found that k has a significantly larger effect than in E. coli but the reverse is true in V. natriegens. This difference may be biological, but it could also be caused by a difference in the number of conditions in each dataset (see below). Third, the total explanatory power of and k for gene positioning is low, with the highest r 2 between predicted gene position and observed gene position being only 0.039 (in the E. coli consensus data). Fourth, analysis of E. coli promoter strengths suggests that the chromosome-wide pattern of mean expression optimization is realized by both gene relocation and compensatory promoter strength alteration, but the growth-dependent expression optimization appears to be driven by gene relocation only. Fifth, analysis after the removal of TT genes revealed that TT genes contribute to the positional gradients in and k more than other genes.

To test MEH and GEH empirically, we relied on the strong assumption that the observed and k of a gene respectively equal its optimal and k. This assumption is reasonable, because gene expression levels are predicted to be subjected to strong natural selection [ 34] and have been shown to become optimized relatively quickly by adaptive evolution [ 35]. Nonetheless, there are also findings that the observed gene expression levels of bacteria in laboratory environments are suboptimal [ 36, 37]. If the observed expression is the sum of the optimal expression and a random noise, the correlation and regression betwee

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →