(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . High-throughput conjugation reveals strain specific recombination patterns enabling precise trait mapping in Escherichia coli [1] ['Thibault Corneloup', 'Université Paris Cité', 'Université Sorbonne Paris Nord', 'Inserm', 'Iame', 'Paris', 'Cnrs', 'Institut Cochin', 'Juliette Bellengier', 'Ecology'] Date: 2026-02 Genetic exchange is a cornerstone of evolutionary biology and genomics, driving adaptation and enabling the identification of genetic determinants underlying phenotypic traits. In Escherichia coli, horizontal gene transfer via conjugation and transduction not only promotes diversification and adaptation but has also been instrumental in mapping genetic traits. However, the dynamics and variability of bacterial recombination remain poorly understood, particularly concerning the patterns of recombined DNA fragments. To elucidate these patterns and simultaneously develop a tool for trait mapping, we designed a high-throughput conjugation method to generate recombinant libraries. Recombination profiles were inferred through whole-genome sequencing of individual clones and populations after selection of a marker from the donor strain in the recipient. This analysis revealed an extraordinary range of recombined fragment sizes, spanning less than ten kilobases to over a megabase—a pattern that varied across the three tested strains. Mathematical modelling indicated that this diversity in recombined fragment size enables precise identification of selected loci following genetic crosses. Consistently, population sequencing pinpointed a selected marker at kilobase-scale accuracy, offering a robust tool for identifying subtle genetic determinants that could include point mutations in core genes. These findings challenge the conventional view that conjugation always transfers large fragments, suggesting that even short recombined segments, traditionally attributed to transduction, may originate from conjugation. Escherichia coli demonstrates a remarkable ability to adapt to diverse environments. This adaptability is in part connected to its ability to recombine incoming DNA in its chromosome, as illustrated by the vast gene content diversity observed in the species. However, the precise dynamics and patterns of bacterial recombination remain poorly understood. To address this knowledge gap, we developed a high-throughput method to observe genetic transfer in the laboratory based on conjugation. Surprisingly, we found that the size of transferred DNA fragments varied from less than 10 kilobases up to a megabase. These results challenge the long-held belief that conjugation primarily transfers only large DNA segments. This diversity in fragment sizes has important consequences. First as a genetic tool, it enables us to pinpoint the location of specific traits in the genome with gene resolution. Second and more importantly, short fragment recombination allows a specific selected loci to be transferred from one strain to another with minimum collateral material allowing an almost surgical precision in genome evolution. Finally, using three different stains as recipient, we observe a strain specific pattern of recombination suggesting a genetic control of that process that will require further investigation. In the present study, we leverage the MAE approach to investigate the recombination patterns and the heterogeneity of recombination fragments across E. coli strains, focusing on a specific loci. We developed an alternative experimental design relying on the classical E. coli conjugative plasmid F. Our findings revealed that conjugation under identical conditions yields fragment lengths spanning more than two orders of magnitude, with recombination patterns and fragment lengths being strain-specific. This heterogeneity, combined with pool sequencing, enabled precise identification of selected loci following crosses, a result corroborated by mathematical modelling. By shedding light on the extent of bacterial recombination, our approach provides an additional tool for functional genomics and highlights the potentially broader significance of conjugation in shaping E. coli genome evolution. Multiple experimental designs have been developed recently to study the heterogeneity of transfer in bacterial populations and/or to do trait mapping or simply as tools to modify genomes. For instance, Silbert et al [ 16 ] modified the broad host range plasmid RP4 to be integrated directed in the chromosome with a transposase, while Khetrapal et al. [ 17 ] integrated the origin of transfer of that plasmid through transposon mutagenesis while providing the plasmid transfer genes in trans. In the latter case, Mass Allelic exchange (MAE) was observed [ 17 ] and transfer of a virulent phenotype could be used to identify a 314kb region associated with that trait, which further reduced to a single gene via cloning of candidate operons. Despite these advances, identifying the genetic basis of phenotypic traits in bacteria remains challenging. In a genetically diverse species like E. coli, phenotypic diversity can stem from gene presence or absence, as exemplified by the HPI [ 3 ] or antibiotic resistance genes. Yet, over a decade of experimental evolution, coupled with whole-genome sequencing, revealed that single-point mutations in conserved genes or regulatory regions can drive significant adaptive changes [ 12 – 14 ]. Adaptation may even involve mutations in essential genes [ 15 ]. Classical tools like transposon mutagenesis, while invaluable for gene inactivation studies, often fail to pinpoint these causative alleles. QTL-style crosses could address this limitation, provided the resulting genetic linkage is minimal. Horizontal gene transfer in E. coli drives genomic diversification via two primary mechanisms. One involves phages erroneously packaging host DNA—a process known as transduction—while the other relies on plasmid-mediated conjugation. The relative importance of these mechanisms in nature remains contested. Comparative analyses of closely related genomes suggest that transduction predominates, with recombinant fragments spanning tens of kilobases, matching experimental observations [ 9 , 10 ]. By contrast, conjugation transfers fragments potentially exceeding hundreds of kilobases. Mastery of these two processes has underpinned genetic engineering for decades. Transduction, with its low efficiency, has been most useful for transferring selectable markers across genetic backgrounds. Conjugation, by contrast, has been a workhorse for genetic mapping [ 5 ] and, more recently, for assembling complex genome constructs and high-throughput marker transfers [ 11 ]. While the mechanics of eukaryotic recombination are well-studied, bacterial recombination remains more enigmatic. For example, conjugation involves double crossovers, resulting in varying integration lengths [ 6 , 7 ]. In contrast, another type of recombination such as natural transformation, may integrate single stranded DNA without requiring double crosses over. Complicating matters further, as incoming DNA is often a threat, bacteria have evolved core stress responses or accumulated accessory defence systems [ 8 ], including restriction-modification genes, to degrade incoming DNA. Such defences fragment donor DNA, with the pattern of fragmentation potentially depending on strain genomic content. To better investigate these issues, we have developed a model of conjugation to unravel how conjugation shapes genome evolution and enables functional genomics. In bacterial genomes, where reproduction is uncoupled from genetic exchange, these transfers remain no less crucial. Processes such as conjugation, transduction, and natural transformation drive both genome evolution and scientific discovery [ 2 ]. These mechanisms facilitate the spread of beneficial alleles or gene clusters, irrespective of their genomic background. Take the High Pathogenicity Island (HPI) in Escherichia coli: this genomic region, essential for efficient iron acquisition in diverse host environments, owes its dissemination within the species to homologous recombination following conjugation [ 3 ]. Such transfers have not only fostered bacterial adaptation but have also enabled HPI identification as a major virulence determinant using agnostic statistical approaches such as GWAS [ 4 ]. In addition, similarly to eukaryotic crosses, conjugation was instrumental in crafting the first bacterial genetic maps by tracking the inheritance of selectable markers [ 5 ]. Genetic exchanges occupy a pivotal position in genomics, serving two complementary roles. First, by severing the physical linkage between alleles during a population’s evolutionary journey, it ensures that alleles would rise or fall in frequency thanks to their own selective value rather than being affected by the selective impact of their chromosomal neighbours [ 1 ]. Second, being able to isolate the effect of a single allele from the surrounding genome is also central to geneticists’ ability to uncover the genetic underpinnings of phenotypic traits. From tracing genealogies to conducting quantitative trait loci (QTL) analyses and genome-wide association studies (GWAS), genetic exchange is crucial for pinpointing chromosomal regions linked to traits. 2. Results 2.1. Rationale for the design of a library of High-frequency Recombination (Hfr) variants For more than 70 years, it has been known that the transfer of chromosomal fragments through conjugation can occur at high frequency with a conjugative plasmid inserted in the chromosome [18]. Parts of the chromosome directly adjacent to the integrated conjugative plasmid, 5’ upstream of the origin of transfer (oriT), are transferred from the High-frequency recombination (Hfr) bacteria [19] (Fig 1A). As a consequence, conjugation from a single given Hfr to a recipient strain results in a biased transfer of genetic material proximal to the integration site. A major objective for exploiting conjugative recombination is to discover unknown loci imparting novel phenotypic traits in natural isolates of E. coli [20]. Since the position of such loci is unknown, a biased transfer proximal to a fixed position in the chromosome may limit their discovery. Therefore, to increase the success of discovery, we need to initiate transfer from multiple positions along the chromosome simultaneously (Fig 1B). To achieve an unbiased and uniform transfer of DNA from a donor strain, rather than using a single Hfr, we decided to create a library of Hfr donors, each having the conjugative plasmid integrated at a different position on the chromosome. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 1. Creation of single-locus Hfr donor versus a multi-locus Hfr donor. A) Chromosomal DNA transfer from a single locus Hfr strain to a recipient showing position-dependent bias in the transferred DNA fragment. The dark blue box and the green arrow represent the integrated conjugative plasmid and its origin of transfer, respectively. B) DNA transfer between a library of Hfr and a single recipient. Here, different fragments are transferred between the donor and the recipient. C) A landing pad, here a Kanamycin cassette, can be integrated at a random position in the genome with a transposon mutagenesis approach based on Mariner or Tn5 transposons. Homology between the cassette on the plasmid and the one on the chromosome promotes plasmid integration at different location generating a library of Hfr Donor. https://doi.org/10.1371/journal.pgen.1011636.g001 Conjugative plasmid integration occurs mostly through homologous recombination between plasmid and chromosome sequences that are highly similar [19]. The similarity between the plasmid and the chromosome is often the result of insertion sequences (IS), which can easily change their position (jump) within the genome [21]. The distribution of IS is variable between strains, and most IS elements are present in a few copies on the chromosome [22]. Therefore, IS offer only a limited, variable and biased number of insertion sites. We decided to use a conjugative plasmid that has been shortened to be IS-free: pNTM3 [11]. This suicide plasmid was designed to enhance its integration into the chromosome by replacing its native origin of replication with an R6K origin of replication (S1A Fig). This origin requires the presence of the pir gene to maintain the plasmid [23]. The pir gene is provided in trans in a permissive strain that allows replication. In strains lacking pir, the plasmid, encoding an antibiotic resistance gene, can be grown in media with the corresponding antibiotic only if the plasmid has integrated in the genome to maintain the resistance genes. The pNTM3 plasmid has been previously modified by the integration of certain chromosomal fragments to integrate at various sites on the chromosome [20]. However, our goal was to be able to integrate it at random positions in the chromosome of E. coli. To achieve this, we decided to introduce a genomic landing pad with homology to the pNTM3 plasmid. We took advantage of the well-characterized transposon mutagenesis technique (Fig 1C) [24]. Using the machinery of a Tn5 transposable element (or Mariner transposon see Methods), we can easily introduce a Kanamycin resistance cassette (KmR) at random positions of the E. coli chromosome. Although not all natural isolates are amenable to genetic manipulation, many strains can be modified using transposon mutagenesis. To use the Kanamycin landing pad for the chromosomal integration of the pNTM3 plasmid, we also introduced a Kanamycin resistance gene (KmR) into the plasmid near the origin of replication. We constructed the plasmid pNTM3_TetA-sacB_KmR (S1A Fig). Upon transferring this modified plasmid to a transposon mutagenesis library of variants (with the KmR gene randomly integrated at different positions on the genome), we expect to obtain a library of Hfr variants with this conjugative plasmid randomly integrated at different sites across the genome. 2.2. Heterogeneous integration rates of the conjugative machinery We first tested whether we could obtain efficient and consistent rates of pNTM3_TetA-sacB_KmR plasmid integration into the chromosome via homologous recombination between KmR genes by using only the lambda-Red recombinase. We used strains from the Keio collection [25]. Conveniently, this collection is composed of strains, each with KmR inserted to replace and inactivate non-essential E. coli genes (Fig 2A). We individually transferred the pNTM3_TetA-sacB_KmR plasmid into 14 strains from the Keio collection, as well as a negative control – BW25113 (Parent of the Keio collection, derivative of K-12) lacking KmR. In each strain, to allow recombineering, we inserted the plasmid pCREPE [26]. This plasmid contains a chloramphenicol resistance CmR, lambda-red recombinase genes as well as an origin of replication pSC101 [27]. After the transfer, we selected for Hfr recipients (Material and Methods), and measured the frequency of integration as the number of CFUs of Hfr variants (presenting KmR) relative to the number of CFUs of the pir + pNTM3 donor. As shown in Fig 2B, we observed a low frequency of integration in the negative control. This suggests residual integration independent of the presence of our landing pad. A considerable heterogeneity in plasmid integration rates was observed across genomic loci. Indeed, while we failed to have integration at some positions such as nfrA, we obtained integration rates of up to 10-5 Hfr CFUs per donor CFU in others positions such as ybeX (Fig 2B). The integration rates were, however low overall, below 10-8 per donor, and sometimes similar to those of the negative control, suggesting that the Hfr libraries produced will have biased integration sites. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 2. Position-dependent variability in Hfr construction. A) Schematic representation of the conjugation step to create an Hfr strain. The suicide plasmid pNTM3_TetA-sacB_KmR with a tetracyclin resistance (tetA) and KmR as the homologous marker for integration into the recipient genome, is replicated into a pir+ strain. This host is a diaminopimelic acid (DAP) auxotroph strain and harbours the pir gene to replicate the plasmid R6K origin of replication. After conjugation, pNTM3_TetA_sacB_KmR plasmid is integrated into the chromosomal KmR copy. Hfr cells, with the integrated plasmid, are selected on LB + tetracycline without DAP. B) Integration efficiency of the pNTM3_TetA_sacB_KmR into mutants of the Keio collection is measured as the log10 ratio of CFU per mL of trans-conjugant recombinants (selected on LB + tetracycline) to CFU per mL of donors (selected on LB + DAP + tetracycline). The genes’ names is on the x-axis correspond to the locus in the Keio mutant where the KmR gene is integrated. Boxplots represent the results from at least 3 independent replicates (depending on the locus studied, 3 to 26 datapoints were obtained). Boxplots display the median (horizontal line), interquartile range (box) with individual data points shown as dots. https://doi.org/10.1371/journal.pgen.1011636.g002 2.3. Inducing targeted integration with double-strand break To improve the integration of pNTM3_TetA-sacB_KmR into the chromosome, we increased the recombination efficiency by using Cas9-mediated targeted DNA double-strand break (DSB) coupled to Lambda-red recombination [28]. Cas9 endonuclease is directed by a guide RNA (gRNA) for targeted DNA double-stranded break (DSB). The gRNA contains a spacer region with a 20 bp homology to the target site, proximal to a 5’-NGG-3’ Protospacer Adjacent Motif or PAM. Cas9:gRNA interaction with the PAM and gRNA::DNA binding at the homologous target site induces the DNA DSB. In E. coli, efficient Cas9-mediated targeted DNA DSB is cytotoxic and results in cell death. Cells may acquire immunity to the targeted DNA DSB by replacing the PAM or the spacer sequence by using a repair template and phage-mediated homologous recombination. Therefore, Cas9-mediated DNA DSB can select for precise genome modifications coupled with PAM/spacer mutations introduced using phage-mediated recombination [28]. In addition, recent studies have shown that having a targeted DNA DSB can also improve the efficiency of recombination [29–32]. We induced a DNA DSB in the genomic Kanamycin landing pad using the Cas9. For that purpose, we tested three different gRNAs. The killing efficiency can vary between gRNA and the editing efficiency is correlated to the killing efficiency. Therefore, we first tested the gRNA killing efficiency of the three gRNAs in Keio strains with an integrated copy of the kanamycin resistance gene. All gRNAs had similar efficiencies, so we chose one (sgRNA1) and introduced it into plasmid pNTM3_TetA-sacB_KmR (S1A Fig). To rescue the cells from the targeted DNA DSB using recombination, we modified the corresponding PAM motif in the plasmid-encoded KmR gene by introducing a synonymous PAM mutation (SPM). Therefore, genomic integration of the plasmid would replace the chromosomal KmR gene with the KmR* gene with the SPM and provide immunity from the Cas9-mediated DSB to select for the plasmid integration. The resulting plasmid was named pNTM3_sgRNA. In each of our Keio strains and the negative control, we transformed the pCREPE plasmid modified to encode Cas9 endonuclease under a constitutive promoter and lambda-red recombination system expressed under a temperature-inducible promoter (at 42OC). The pCREPE plasmid has a temperature sensitive origin of replication, due to which it only replicates below 34OC (Fig 3A). We induced the system to facilitate recombination between the sequences flanking the cut site in the chromosome and the repair template present on the pNTM3_sgRNA plasmid (see methods) [33]. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 3. Cas9-Lambda-red recombination-mediated Hfr construction. A) Schematic representation of DSB-induced Lambda-red recombination showing the gRNA and repair template for Cas9-Lambda-red recombination-mediated Hfr construction in the recipient. B) Integration efficiency of the pNTM3_TetA-sacB_KmR into mutants of the Keio collection, measured as the log10 ratio of CFU per mL of recombinants (selected on LB + tetracycline) to CFU per mL of donors (selected on LB + DAP + tetracycline with the Cas9-Lambda-red approach (purple) across the 14 loci and the BW25113 control. For comparison, results obtained without an induced double-strand break (corresponding to Fig 2B) are given in (cyan). Boxplots display the median (horizontal line), interquartile range (box) with individual data points shown as dots for each of the strains. https://doi.org/10.1371/journal.pgen.1011636.g003 With the Cas9/recombineering-mediated insertion protocol, we achieved a much higher integration efficiency (S1B Fig). Additionally, we observed less heterogeneity between sites, with an integration rate greater than 10-5 for all but one locus (13 out of 14) (Fig 3B). Therefore, Cas9-mediated recombination significantly improved pNTM3 integration efficiency into the chromosome to produce Hfr strains. 2.4. Validating the ability of the constructed Hfr to transfer the galK marker To validate that the selected strains were functional Hfr bacteria, we tested their ability to transfer DNA to a recipient strain. To quantify the transfer’s efficiency, we conducted conjugation experiments between Hfr donor strains and a ΔgalK recipient. The galK gene encodes for the Galactokinase protein, which is required for growth on galactose as the sole carbon source. In the ΔgalK strain, the galK gene was inactivated by introducing multiple stop codons had previously been introduced to inactivate the galK gene to prevent any possibility of reversion by mutations. Consequently, in this strain, only recombination of a fully functional gene upon conjugative transfer from the Hfr donor could restore the galK function (Fig 4A) [26]. In addition, the recipient strain was also rendered resistant to spectinomycin by electroporating a small non conjugative plasmid containing the aadA gene. We then crossed each constructed Hfr with the galK- recipient and quantified the transfer efficiency as the number of galK+ specR transconjugants to the number of specR recipient cells. In the Hfr strains constructed using the different Keio mutants, the Hfr-oriT was integrated at variable distances from the galK gene. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 4. Testing Hfr donor libraries for transfer of a galK marker. A) A schematic representing crosses between Hfr galK+ donors and galK- spectinomycin recipients. Recombinants that have gained a functional galK allele are selected on M9 minimal media with galactose and spectinomycin. B) Recombination efficiency measured as the ratio of CFU per mL of recombinant (growing on M9 + galactose + spectinomycin) to CFU per mL of all recipients (selected on M9 + glucose + spectinomycin) for Hfr donors constructed by integration of the conjugative plasmid at different distances from the galK locus. Boxplots display the median (horizontal line), interquartile range (box) with individual data points shown as dots. Statistical comparisons were made using the Wilcoxon test, ns: non-significant, *: p-value ≤ 0.05, **: p-value ≤ 0.01.C) Recombination efficiency measured as the log10 ratio of CFU per mL of recombinant (selected on M9 + galactose + spectinomycin) to CFU per mL of all recipients (selected on M9 + glucose + spectinomycin) for an Hfr donor library crossed with galK- spectinomycin recipients in different genetic backgrounds E. coli K-12 BW25113 (control), REL606, HS, and 536. Means are represented by blue triangles and data points by yellow dots. https://doi.org/10.1371/journal.pgen.1011636.g004 Upon conjugation, we found that the strain bearing the oriT inserted at 100kb from galK could transfer the galK+ allele with an extremely high efficiency. Corroborating previous observations, we observed that Hfr-oriT integration at a larger distance resulted in lower transfer rates (from 939 kb on, the transfer rates were below detection limit), but the strain with an Hfr-oriT at 300 kb could still transfer the galK+ allele with 20% efficiency (Fig 4B). 2.5. Creating and testing a library Hfr After functional validation of our targeted Hfr constructs, we tested if we could create a library of Hfr. To construct the library, we used a transposon mutagenesis library in which the KmR landing pad is randomly inserted in the chromosome. We constructed the transposon libraries using laboratory strains BW25113 and REL606 [34]. We then used the afore mentioned Cas9/recombineering-mediated system to integrate the pNTM3_sgRNA plasmid across the different sites of integration of the transposon. Upon selection of TetR colonies growing on LB plates, we obtained 103 to 104 colonies that had integrated the pNTM3_sgRNA plasmid into the transposon that were pooled in Hfr libraries. To test if in these Hfr libraries, the integration of pNTM3_sgRNA conjugative plasmid occurred across the genome, we investigated the integration sites. We adapted a method for transposon sequencing to our libraries of Hfr donors to determine the precise locations of the landing pad in which pNTM3_sgRNA was integrated [35] focusing on library made on BW25113. This approach is based on the amplification of a short DNA sequence spanning the extremity of the transposon and the proximal chromosomal sequence (S2A Fig) where the integration has occurred. We then used sequencing to extract the insertion positions. Of note, as the landing pad is larger than 1kb, our approach is indeed identifying the insertion sites of the Kanamycin cassette that were sampled during the creation of the Hfr library. However, our previous experiments have shown the vast majority of integration of the plasmid (from 99 to 99.9%) occurs in the Kanamycin cassette, so the position of the sampled Kanamycin integration sites should be representative of the site of integrations of the plasmids. Overall, these analyses showed that the pNTM3_sgRNA plasmid was integrated at multiple loci across the chromosome. The genome architecture may lead to biases in the integration of the plasmid. A source of bias is differences in copy number variation between different regions of the genome (with higher copies closer to the origin of replication, compared to the terminus) due to differences between rate of chromosomal replication and cell division. In addition, organization of the E. coli chromosome into macrodomains may result in a biased accessibility of their DNA [36]. Additionally, the randomized integration doesn’t guarantee that the distribution of the positions of integration will be homogenous along the genome. For instance, integration inside essential genes will be lethal for the bacteria leading to some local biases of insertion. We therefore investigated the distribution of insertion sites along the chromosome using the genomic coverage derived from transposon sequencing data (the number of reads that match a specific insertion site) (S2B Fig). We compared the density of insertions, calculated on 250-kb bins, along the chromosome as the log2 fold-change in coverage compared to the coverage of the Ter macrodomain. There are two ways to analyse the data: we can focus on the insertion sites (red in S2B Fig), or the coverage of these insertion sites (blue on S2B Fig). Insertion sites refer to the identified positions of insertions and likely refer to independent insertions during the process of Hfr construction; we counted 1140 different insertion sites (blue on S2D Fig). Coverage of these sites corresponds to potential differences in the representation of the various Hfr strains in the library, but also integrates potential biases associated with sequencing [37]. As mentioned, higher copy number of the Ori macrodomain within the cells usually leads to higher coverage of that region relative to the Ter macrodomain in sequencing data. Analysing density profiles along the genomes with bins of 250kb, we observed that, there was a bias for both metrics: fewer integration sites were found in the Ter macrodomain compared to the rest of the genome (S2C Fig) and the coverage of insertions in the Ter macrodomain was lower. The bias was visible using both metrics, but was, as expected, more marked considering coverage of insertion sites (up to 10-fold, S2C Fig) than density of insertion sites (up to 3-fold, S2C Fig). This is expected given the afore mentioned Ori to Ter coverage bias observed in whole genome sequencing. The biases in the density of insertion sites suggest a biological bias during transposon mutagenesis and/or during pNTM3_sgRNA integration, that likely results from the lower number of copies and/or lower accessibility of the Ter macrodomain. Such a bias has been reported before, notably by Khetrapal et al [17]. Despite this bias, integrations of the conjugation machinery were spread along the chromosome and our library could therefore be used to transfer any locus in the genome to a recipient strain. To test if this library could transfer DNA to various genetic backgrounds, we mated the donor K12 BW25113 Hfr library (constructed twice independently) with 4 recipients with different genetic backgrounds: the laboratory strains BW25113 and REL606, the natural isolate HS [38], and the pathogenic strain 536 [39]. While the first three correspond to strains from phylogroup A of E. coli, the latter from B2 phylogroup is more genetically distant. All four strains were made galK- and resistant to spectinomycin via the presence of a plasmid. After the mating with the Hfr library, we plated on minimal media with galactose and spectinomycin to select only for transconjugants that recovered a galK+ allele and get rid of donors and recipients. We estimated the recombination efficiency as the ratio of number of recombinant colonies (selected by plating minimal media with galactose and spectinomycin) to the number of recipient colonies (estimated by plating on non-selective minimal media with glucose and spectinomycin). The frequency of recombinants varied largely depending on the recipient’s genetic background (Fig 4C). In BW25113 and REL606 backgrounds, the recombination efficiency was up to 5–10%, This efficiency reduced to ~0.5% in HS giving only a few colonies of recombinants and it goes as low as 0.00015% when 536 was the recipient. This difference is correlated with the genomic divergence between K12 and the recipients. Additionally, this effect might be amplified for the strain 536 compared to HS or REL606, as we found that 536 is expressing toxins that might kill the donor strains during the conjugation assay (S3 Fig) and harbours also a candidate restriction system of type 2 that could degrade the incoming DNA [40]. 2.6. Analysis of the genome of the recombinants Since we initiated the transfers from multiple loci, we expected to observe different patterns of recombination along the genome (Fig 5A). For the various crosses we made, we sampled 19–20 clones for sequencing (only 8 in the case of 536 due to low efficiency) and determined the parts of their genomes that came from the donor or the recipient strain. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 5. Genome sequencing of recombinant clones. A) A schematic representing the creation of recombinants from a conjugation between an Hfr library used as donor and a recipient. B) Identification of the recombined region of a recombinant. Frequency of donor alleles along the genome of a REL606 recombinant. The red line shows the galK locus position. The broken line shows the delineation of the recombined fragment using the changepoint package. D, E) Recombined fragments (yellow fraction) across the genomes for recipient REL606 (C), HS (D) and 536 (E). Red line indicates the galK loci position. F) Distribution of recombination fragment length (log10 scale), colours correspond to the different recipients (REL606, dark blue, HS, light blue, 536 light pink). y-axis represents the count of recombinant genomes, the bins height corresponds to the number of a recombined fragment in a certain length interval. G) Boxplot showing the size of recombination fragments for each recipient (log10 scale), for fragments including the selected marker (upper box) and the ones not including it (lower box, lighter colour). Wilcoxon-test was used for statistical analysis. ns: non-significant, *: p value ≤ 0.05, **: p value ≤ 0.01. H) Distribution of the total length recombined (log10) per genome of recombinants (REL606, dark blue, HS, light blue, 536 light pink). https://doi.org/10.1371/journal.pgen.1011636.g005 To identify the recombined tracks, we used two independent methods that produced almost identical results. In the first approach, we focused on read mapping. Sequencing output, consisting of 150-bp reads, was mapped competitively against the donor and recipient genomes. Three outcomes were possible. (i) Reads could match both genomes equally well, in which case they were not informative. (ii) Reads could match only one of the two genomes, thereby providing information on accessory genes that were retained or transferred. Since our interest here is in the consequences of recombination on the core genome, we did not use these reads. (iii) Reads could match both genomes but with a preference for one over the other, owing to the presence of alleles specific to one of the strains. The positions of reads matching preferentially to the donor genome were then used to delineate genomic fragments that originated from the donor in a given recombinant. To do this, we applied a Hidden Markov Model to define transitions between two states: donor-derived and recipient-derived. This method is relatively robust to sequencing errors because it considers the entire read for mapping, and in most cases a 150-bp stretch of the core genome contains more than one difference between the donor and recipient genomes. As an alternative approach, we relied on allele frequencies at polymorphic sites between the donor and recipient core genomes. We first identified all polymorphic sites and then mapped reads to compute the frequency of donor and recipient alleles along the genome. A step function was then used to detect transitions between regions enriched for donor or recipient alleles. This method offers the highest possible resolution but can be more sensitive to sequencing errors. Comparison of the two methods yielded nearly identical (S4 Fig) and robust detection (S5 Fig) of recombined fragment and estimation of fragment length and revealed as shown in Fig 5B–E, that all genomes but one exhibited a clear recombined region around the selected site. The patterns of recombination were however quite different between recombinants and between recipient strains (Fig 5C, 5D and 5E). First, 22 recombinants (1/8 for 536, 7/18 REL606 or 14/20 for HS) showed a single fragment of recombination that overlapped the galK but the length of this fragment was very variable: from a few kilobases (1.5 kb or 7.5 kb), up to a full megabase with a fragment as long as 1.103 Mb. Transfer size can therefore span about 3 orders of magnitude. Second, the 23 other recombined strains harboured not only a recombined fragment around the locus of selection galK, but also several other recombined regions in the neighbourhood. They contained between 2 and 16 fragments including the one containing galK. This is compatible with an incoming DNA fragment being cut and recombined into pieces [6]. This pattern was most drastic with strain 536 as recipient with 7.125 fragments per recombinants compared to 1.65 for HS and 2.22 for REL606. We also noticed that a few recombinants revealed some distant distinct sites of recombination (5/19 for REL, and 3/8 in 536). Finally, a few strains had not resolved the recombination in the colony that was sequenced as the recombined regions appeared to have both overlapping donor and recipient-specific sequences. Overall, the distribution of the recombined fragments varied across strains (Fig 5C–G). The median size of recombined fragments was 198,195 bp, 126,214 bp and 21,353 bp for the recipients REL606, HS and 536 respectively (Fig 5F and 5G); when considering only fragments that included the galK locus, their lengths were significantly greater across all strains (327,529, 156,831.8 and 82,130 bp median size in REL606, HS and 536 respectively) (Fig 5G). We tested the lengths’ differences between the galK-containing fragments across recipient strains and found that there was a significant difference between REL606 recombinants compared to HS and 536 recombinants (Fig 5G). When we looked at the total length of recombined material, the difference between strains remained with the mean sizes being 152kb, 208kb and 440kb for 536, HS and REL606 respectively (Fig 5H). 2.7. Expected power to detect a selected site Sequencing of individual clones confirmed the occurrence of recombination and provided critical insights into the length, position, and number of recombined fragments. Beyond this descriptive analysis, the experiment can be viewed as a genetic cross coupled with selection, which enables the detection of loci under selection in a QTL-like framework. However, unlike the genetic exchanges typical of eukaryotic systems used in QTL approaches, the recombined fragment lengths here are highly variable. This raises the question of how such variability impacts the ability to pinpoint loci under selection. To explore this generically, we employed mathematical modelling. We analyzed the statistical properties of recombined fragments conserved among a pool of selected recombinants near the allele under selection. The genomic region where donor DNA is consistently found across all selected recombinants is likely to contain the selected allele—here, the functional galK allele. Recombined fragments were modelled as independent, identically distributed intervals with positions uniformly distributed across the genome. This model demonstrated that for a large number (n) of selected recombinants, the length of donor DNA shared by all recombinants scales with the harmonic mean of the length of the recombined fragment encompassing the selected loci divided by the number of recombinants. Specifically, the expected size of the region likely to harbour the selected locus, denoted as, , is given by , where represents the harmonic mean of the recombined fragment lengths that include the selected loci (see Proposition 4.1 in S1 File). This result is significant because the harmonic mean is particularly sensitive to smaller values, meaning that the presence of short fragments—just a few kilobases in length—can substantially reduce the size of the shared recombined region across recombinants. Consequently, the model suggests that the observed heterogeneity in recombined fragment lengths may, counterintuitively, enhance the precision of locus detection under selection. 2.8. Analysis of the pool of recombinants To test our ability to detect with precision the selected loci, we decided to use another approach based on sequencing a pool of recombinants. Using this methodology, we can analyse the whole population and look at the density of donor-specific reads along the genome to visualize patterns of recombination at the pool scale. Because we are enforcing a strong selection at galK loci, we expected a clear and strong signal of donor DNA integrated into the recipient. For each of the recipients’ backgrounds, we identified a peak of donor-specific sequences around the galK loci (Fig 6A, 6B and 6C). The delineation of the target region was even more precise using the ratio of donor to recipient-specific reads along bins of 1 kb along the genome. The maximum of that ratio landed exactly in the galK gene for all crosses (Fig 6A, 6B and 6C). The pattern of decay of the donor DNA presence differed between the recipient strains (Fig 6D) with a decay less marked in REL606 than in HS and 536. This pattern is consistent with the distribution of recombination length detected as longer recombination tracks should lead to a slower decay. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 6. Pooled Genome sequencing. A, B, C) Log10 of the per kb coverage of the genome with reads preferentially mapping to the recipient genome (in A dark blue for REL606, in B blue for HS, in C pink for 536) or the donor one (yellow). Red lines show the galK loci position. D) Log10-ratio of donor-to-recipient coverage along the genome relative to galK position for the three recipient genotypes (dark blue for REL606, blue for HS, red for 536.E) Log10-ratio of donor-to-recipient coverage along the genome for REL606 pool with the Terminus Ter (light blue) and Ori (green) macrodomain positions shown. F) Log10-ratio of donor-to-recipient coverage along the genome for three different ∆matP REL606 mutans (purple, light blue and median blue) and REL606 (dark blue). The red line shows the galK loci position with the Terminus Ter (light blue) and Ori (green) macrodomain positions represented by colored rectangles. https://doi.org/10.1371/journal.pgen.1011636.g006 For all recipients, we observed a steeper reduction of donor DNA towards the Ter macrodomain (Fig 6E). This pattern could result from the integration bias of the conjugative plasmid in our library. However, we also know that the Ter macrodomain is structured and made less accessible through the interaction between the MatP protein and the matS sites that are frequent in that macrodomain [36]. To test if that structuration of the Ter Macrodomain could reduce the efficiency of recombination and contribute to the observed pattern, we used a ∆matP mutant as a recipient. We observed no difference between ∆matP and the wild type (Fig 6F) suggesting that recombination was not affected by macrodomain structural organization and that the biases resulted most likely from our skewed distribution of conjugative plasmid integration sites. [END] --- [1] Url: https://journals.plos.org/plosgenetics/article?id=10.1371/journal.pgen.1011636 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/