(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . Unknotting RNA: A method to resolve computational artifacts [1] ['Simón Poblete', 'Facultadde Ingeniería', 'Arquitectura Y Diseño', 'Universidad San Sebastián', 'Santiago', 'Centro Basal Ciencia', 'Vida', 'Mikolaj Mlynarczyk', 'Institute Of Computing Science', 'Poznan University Of Technology'] Date: 2025-04 RNA 3D structure prediction often encounters entanglements, computational artifacts that complicate structural models, resulting in their exclusion from further studies despite the potentially accurate prediction of regions outside the entanglement. This study presents a protocol aimed at resolving such issues in RNA models while preserving the overall 3D fold and structural integrity. By employing the SPQR coarse-grained model and short Molecular Dynamics simulations, the protocol imposes energy terms that enable selective modifications to disentangle structures without causing significant distortions. The method was validated on 195 entangled RNA models from CASP15 and RNA-Puzzles, successfully resolving over 70% of interlaces and approximately 40% of lassos, with minimal impact on the original geometry but notable improvement in ClashScore. The efficiency of untangling conformations that are unequivocally classified as artifacts is 81%. Certain cases, particularly those involving dense packing of atoms or complex secondary structures, posed challenges that limited the efficiency of the method. In this paper, we present quantitative results from the application of the protocol and discuss examples of both successfully disentangled and unresolved structures. We show a viable approach for refining models previously deemed unsuitable due to topological artifacts. Most algorithms for 3D RNA structure prediction, including a recent version of AlphaFold, in the pool of solutions return models with entanglements, which turn out to be artifacts of computational modeling. Excluding the entanglement region, these models can reflect the native conformation quite well; however, due to the presence of the artifact, they are not considered promising and are therefore discarded from further study. In this paper, we present the first method that enables the systematic classification and removal of RNA entanglements while preserving the model’s consistency and its global 3D fold. Tests on the in silico models submitted to RNA-Puzzles and CASP15 RNA show that the system can automatically solve a significant number of issues, including 81% of entanglements considered incorrect conformations. Funding: SP was supported by the Fondecyt Regular project No. 1231071 and Centro Ciencia & Vida, FB210008, Financiamiento Basal para Centros Científicos y Tecnológicos de Excelencia de ANID ( https://anid.cl ). MM and MS were supported by the statutory funds of Poznan University of Technology ( https://www.put.poznan.pl/en ) and the Institute of Bioorganic Chemistry, Polish Academy of Sciences ( https://www.ibch.poznan.pl/en.html ). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Data Availability: RNA models predicted in CASP15 are available at https://predictioncenter.org/download_area/CASP15/predictions/RNA/ , while those of RNA-Puzzles at http://www.rnapuzzles.org/results/ . RNA 3D models with entanglements resolved using the SPQR-based protocol are available at doi: 10.5281/zenodo.13840004 . RNAspider is accessible at https://rnaspider.cs.put.poznan.pl/ and SPQR at doi: 10.5281/zenodo.14658435 and https://github.com/srnas/spqr . Copyright: © 2025 Poblete et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Recent studies have shown that all current methods for predicting 3D RNA structures tend to generate entangled models [ 30 , 31 ]. In response, we present the first systematic solution to this problem, combining the RNAspider tool for entanglement identification with SPQR (SPlit–and–conQueR), a coarse-grained model specifically designed for the prediction and refinement of RNA structure [ 32 , 33 ]. SPQR can effectively resolve unwanted structural entanglements by allowing the introduction of arbitrary energy terms. These terms can distort the geometry of selected nucleotides while preserving essential RNA interactions and conformations [ 33 – 35 ]. Furthermore, SPQR facilitates the straightforward incorporation of all-atom details [ 36 ]. By leveraging data from RNAspider, we can define precise repulsive or attractive interactions between specific nucleotides, minimizing entanglements with marginal impact on the overall geometry of the structure. However, we have identified instances where the disentanglement process cannot be applied without significantly distorting the secondary structure, or where the compactness of the molecule poses challenges to the correction procedure. In the paper, we present the untangling procedure and its results on example RNA 3D models. We also address cases of problematic entanglements in which the method proved ineffective and suggest directions for future research on resolving these challenges in RNA structure prediction and refinement. Prediction accuracy is typically evaluated in the context of a reference structure, meaning a molecule with a known experimental structure is used as a benchmark. The computer-predicted model is then compared to this reference structure, and similarity or distance measures are calculated. These metrics consider both the 3D structure parameters and the compatibility of secondary interactions; exemplary approaches applied for RNA, the focus of this paper, have been proposed in [ 19 – 23 ]. In contrast, the quality of a computational model can be assessed without referencing a specific target. This involves analyzing the geometry of the structure to ensure that its parameters (such as angles, bond lengths, interatomic distances, and planarity of bases) fall within acceptable ranges [ 24 ]. Programs like MAXIT [ 25 ] and MolProbity [ 26 ] facilitate this process by identifying and often correcting geometric abnormalities in the RNA 3D models. Furthermore, recent tools – KymoKnot [ 27 ], Topoly [ 28 ], and RNAspider [ 29 ] – analyze the topology of 3D structures, identifying topological anomalies such as interpenetrating fragments that form entanglements or complex topological knots, which are unlikely to occur in natural molecules. In recent years, computational modeling has emerged as a leading technique for elucidating the secondary and tertiary structures of biological molecules. Deep learning-driven modeling has already successfully replaced experimental methods in protein research [ 1 – 3 ]. For nucleic acids, predictive algorithms complement wet-lab experiments. Current computational tools effectively generate secondary structures of RNA and DNA from sequences, primarily reflecting short-range canonical interactions. However, non-canonical pairings and long-range contacts are often absent in output models [ 4 – 6 ]. Obtaining the three-dimensional structure of nucleic acids from sequences or base pairing information remains a challenge, as evidenced by genomic studies and blind prediction initiatives such as CASP and RNA-Puzzles [ 7 – 11 ]. Nonetheless, computer-based 2D and 3D structure prediction methods are frequently employed to create initial models of nucleic acid molecules, which are then refined and validated experimentally. In silico generated models sometimes verify experimental data and often serve as starting points for designing new RNA- or DNA-based diagnostics and therapeutics, as well as applications in bio- and nanotechnology [ 12 – 18 ]. Thus, the accuracy of structural modeling and the quality of nucleic acid predictions are becoming increasingly important. Once the structures were treated by the disentanglement protocol, their topology was analyzed again with RNAspider. We also examined the secondary structure of the entangled loops after the refinement, since the rearrangement of the structure can produce the disruption of their constitutive base pairs which might turn an entanglement undetectable. For the cases where this happened, we checked visually the presence or absence of the entanglement, and we reported them accordingly in the Supporting Information ( S1 Table for CASP15 and S2 Table for RNA-Puzzles). To resolve entanglements of structural elements, we applied the untangling protocol (see Appendix A in S1 Text ). The protocol involves a series of simulations utilizing both coarse-grained (CG) and atomistic representations [ 36 ], and follows three main steps ( Fig 3 ). RNAspider processes the tertiary RNA structure by first applying the RNApolis Annotator [ 37 , 38 ] to annotate base pairs. This step allows the system to derive the secondary structure of RNA. Based on the latter, RNAspider identifies structural elements within the input 3D structure, such as loops, dinucleotide steps, and single-stranded fragments. Next, the full-atom 3D structure is converted into a wireframe model. In this model, nodes represent selected backbone atoms (P, C4 ′ ) and the centroids of heavy atoms that form Watson-Crick-Franklin hydrogen bonds, while edges illustrate the RNA backbone and base pairings. Thus, each loop and each dinucleotide step is encircled by a closed polynomial chain. The interior of each circular structure element is covered with a mesh and is recursively triangulated. The subdivision continues until sufficiently small triangles are obtained, with the user able to set a stop criterion. The punctures are checked for each triangle using the Moeller-Trumbore algorithm [ 39 ]. This process allows RNAspider to identify intersection points where one structural element penetrates another. Once an intersection point is identified, the system determines which structure elements form the entanglement and whether it is an interlace or a lasso. To identify entanglements of structural elements, we used the RNAspider program with its default settings [ 29 ]. RNAspider analyzes the input 3D structure for entanglements, outputs the identified entanglements, and provides their locations. The system can distinguish between three types of interlaces (also called links; D&D, D&L, L&L) and six types of lasso entanglements (D(D), D(L), D(S), L(D), L(L), L(S)) ( Fig 2 ). Each entanglement involves two structural elements that are either circular (loops and dinucleotide steps) or linear (single strands). An interlace is formed by two circular elements, while a lasso-type entanglement involves a circular component that encircles another structural element, which can be of any type. In [ 31 ], we further distinguish between shallow and deep L(*) type lassos. Shallow lassos, occurring when the lassoed fragment threaded through the loop spans no more than five nucleotides, are considered non-pathological due to their potential for spontaneous disentanglement. In contrast, deep L(*) lassos and all D(*) type lassos are classified as artifacts. To test the disentanglement protocol, we downloaded RNA 3D models predicted in the CASP15 and RNA-Puzzles competitions, available in their online repositories as of January 2024. The CASP15 dataset ( https://predictioncenter.org/download_area/CASP15/predictions/RNA/ ) included 1,660 models generated in CASP15 for 12 RNA targets. The RNA-Puzzles dataset ( https://github.com/RNA-Puzzles ) contained 1,028 models targeting 22 RNA sequences in rounds I-IV of RNA-Puzzles. From this collection, we discarded redundant structures and blobs, focusing our analysis on the remaining models for the entanglements. Specifically, among the 122 entangled predictions from CASP15, 21 were redundant, 8 were counted as blobs, and 4 were discarded due to clear artifacts in their coordinates. Blobs generally had a ClashScore greater than 400, making their structures impossible to analyze. The exception was the R1126TS470_5 model, which exhibited an analyzable structure and appeared visibly reasonable despite the high number of clashes and its ClashScore of 416.7. In the remaining set of 89 CASP15 RNA models, 88 had up to 7 entanglements (with 51 of them having only a single entanglement) (see Fig A in S1 Text ). One outlier, the R1138TS239_5 model (720 nts), had a ClashScore of 86.1 and contained as many as 15 entanglements. For reference, based on our experience from CASP and RNA-Puzzles, a ClashScore is insignificant and considered acceptable if it is below 10. The RNA-Puzzles collection included 123 entangled structures, 12 of which were discarded as duplicates. Five additional structures were removed due to entanglement resulting from misplaced phosphates. Among the remaining 106 non-redundant models, each contained up to five entanglements, with the vast majority (88 models) exhibiting a single entanglement. The complete distribution of the number of entanglements in each dataset is shown in Fig A in S1 Text . Ultimately, 195 entangled structures were included in the test collection: 89 from CASP15 and 106 from RNA-Puzzles ( Fig 1 ). They contained a total of 329 structural element entanglements classified into 9 groups: 3 classes of interlaces (dinucleotide step-dinucleotide step, labeled D&D; dinucleotide step-loop, labeled D&L; and loop-loop, labeled L&L), 3 classes where a dinucleotide step forms a lasso (around another dinucleotide step, labeled D(D); around a loop, labeled D(L); or around a single strand, labeled D(S)), and 3 classes where a loop forms a lasso (around a dinucleotide step, labeled L(D); around another loop, labeled L(L); or around a single strand, labeled L(S)). Note that in the class labels, D stands for a dinucleotide step, L for a loop, and S for a single strand, & denotes interlace, and brackets are used to denote lasso-type entanglement. A list of models, along with the entanglement information, is provided in S1 Table (CASP15 subset) and S2 Table (RNA-Puzzles subset). Results and discussion We applied the disentanglement protocol to each of the 195 entangled structures from the benchmark set (the resulting structures are available at doi: 10.5281/zenodo.13840004). Table 1 presents the aggregate results of this experiment. The protocol successfully resolved approximately half (49%) of the entanglements in eligible RNA structures, with 72% of successful cases coming from CASP15 predictions and 28% from RNA-Puzzles. It was notably more effective for interlaces, resolving 77% of cases, compared to 40% for lassos across both datasets. However, it is important to note that lassos are more prevalent – occurring twice as often as interlaces in the CASP15 dataset and 13 times more frequently in the RNA-Puzzles dataset – generally harder to remove and, importantly, may not be artifacts. Artifacts definitively include all types of interlaces and D(*) lassos. Among the 99 such conformations identified in the dataset, 80 (81%) were successfully untangled. When analyzing results at the whole-structure level – acknowledging that individual RNA structures may contain multiple entanglements – 41% of problematic structures were completely resolved (i.e., all entanglements were removed), 14% were partially resolved (some entanglements were addressed while others remained), and 45% of structures remained unresolved. Among the unresolved cases (87 RNA models), 98% contained lasso-type entanglements, while only 2% involved solely interlaces. We evaluated the impact of disentangling on overall structure quality by analyzing changes in RMSD (Root Mean Square Deviation) [19], INF (Interaction Network Fidelity) [20], and ClashScore [43] after applying the protocol. S3 Table (for CASP15) and S4 Table (for RNAPuzzles) provide detailed values of these measures for all models in the benchmark set, while Table 2 summarizes the results of their analysis. RMSD (distance measure) and INF (similarity measure) were calculated relative to the reference structure both before and after disentanglement. When multiple reference models were available (as in some CASP15 targets), the first model was used. ClashScore, a reference-free measure of stereochemical quality, did not require comparison to the native structure. It showed significant improvement after refinement – 70 out of 108 RNA models had a lower post-refinement ClashScore (the largest decrease was 410.97), especially in larger RNA models. For 37 out of 108 models, the measure increased slightly (the largest increase was 25.57), mostly in cases where the pre-refinement ClashScore was already low, and the structures were densely packed in 3D space. Changes in RMSD and INF values were negligible, confirming that the untangling protocol preserves conformational integrity. RMSD differences ranged from –3.5Å to 5.9Å, while INF were between -0.12 and 0.04. The small reduction in INF was mostly attributed to the disentanglement of loops, which were stabilized by base pairs. Discrepancies in these values can be improved by running simulations with stronger restraints in the last steps of the refinement procedure, for which we report a range of reasonable values in Appendix A in S1 Text. We also tested the disentanglement protocol on selected blobs and found them generally difficult to relax during Step 1 of the procedure. Additionally, Step 2 presented challenges in disentangling due to insufficient space to accommodate the loops in the resulting structure. Consequently, we confirmed that the blobs were not suitable for untangling. Finally, we assessed the time required for disentanglement. These tests were conducted on two machines with different configurations: MacOS 12.6 (i5-8259U, 4 cores/8 threads) and Linux Ubuntu 20.04.1 (AMD Ryzen Threadripper 2950X, 16 cores/32 threads). For a medium-sized structure (363 nucleotides, 1 entanglement), the disentanglement simulations took approximately 2 minutes, while the backmapping relaxation step required an additional 5 minutes. The runtime for disentanglement varied based on the structure size and the number of entanglements, ranging from a few seconds for smaller structures (<100 nucleotides) to several minutes for larger ones (>700 nucleotides). The backmapping relaxation time depended solely on the structure size, ranging from 1 to 15 minutes. Lastly, the all-atom refinement step, performed with Gromacs 2021.2 using 8 threads and a single GPU, required 4 minutes for a 69-nucleotide structure and up to 12 minutes for a 720-nucleotide structure. Interlaces Interlaces typically arise from the superposition of fragments, leading to entangled geometry after eliminating the clashes. These include D&D, D&L, and L&D conformations, all of which are considered artifacts. Fig 4 presents examples of successfully resolved interlaces for each type. Our protocol effectively disentangles D&D interlaces, achieving a high success rate. Of the 30 D&D entanglements in RNA models from the benchmark set, 29 were successfully resolved (97%). In the single unresolved case, two large fragments of the structure overlapped, spanning six loops and creating five adjacent entanglements (Fig C in S1 Text). The attempt to simultaneously remove all these entanglements was unsuccessful. Similarly, for D&L entanglements, the disentanglement protocol proved highly efficient, resolving 18 out of 25 cases (72%). The permanent interlaces of this type are distributed across five RNA structures. In three of these (CASP15 models), the D&L interlaces coexist with at least four other entanglements. In the fourth unresolved structure (also from CASP15), numerous clashes indicated by a high ClashScore, lead to a compact geometry that leaves no space for proper disentanglement (cf S1 Table and S3 Table). In the fifth unresolved model, from RNA-Puzzles, the disentanglement protocol transformed the D&L interlace into an L(D) lasso. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 5. Reduction and removal of the R1138TS076_3 interlace L&L, (A) unreduced, (B) reduced and (C) reduced and untangled. For L(S) lasso from the R1107TS054_3 model: (D) unreduced, (E) reduced, (F) reduced and untangled. https://doi.org/10.1371/journal.pcbi.1012843.g005 For L&L interlaces, the disentanglement protocol is less efficient, successfully resolving 12 out of 22 cases (55%). Failures typically occur when the interlaces are adjacent to other entanglements or when the puncturing loop cannot be reduced to a shorter segment during preprocessing. When loops remain extended, the virtual sites used in refinement simulations become less localized around the entanglement, reducing the procedure’s effectiveness. The reduction of the loops is crucial for an efficient untangling simulations, and it can be significant as shown in Fig 5 for both lassos and interlaces. In both datasets, CASP15 and RNA-Puzzles, we observed instances where L&L interlaces degenerated into L(L) lassos. This outcome can be attributed to the specific geometry, where reducing the loops around the intersection point is not optimal. In such cases, the disentangling simulation displaces the loops in the wrong direction (Fig D in S1 Text). To prevent this, manual intervention in loop reduction, disregarding the intersection point, can be employed. [END] --- [1] Url: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1012843 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/