(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . DRfold2 is a deep learning-based tool that enables efficient and accurate RNA structure prediction [1] ['Yang Li', 'Cancer Science Institute Of Singapore', 'National University Of Singapore', 'Singapore', 'Chenjie Feng', 'School Of Science', 'Ningxia Medical University', 'Yinchuan', 'Xi Zhang', 'Center For Ai'] Date: 2026-02 RNA structures are essential for understanding their biological functions and developing RNA-targeted therapeutics. However, accurate RNA structure prediction from sequence remains a crucial challenge. We introduce DRfold2, a deep learning framework that integrates a novel pre-trained RNA Composite Language Model (RCLM) with a denoising structure module for end-to-end RNA structure prediction. Based solely on single sequence, DRfold2 achieves superior performance in both global topology and secondary structure predictions over other state-of-the-art approaches across multiple benchmark tests from diverse species. Detailed analyses reveal that the improvements primarily stem from the RCLM’s ability to capture co-evolutionary pattern and the effective denoising process, with a more than 100% increase in contact prediction precision compared to existing methods. Furthermore, DRfold2 demonstrates high complementarity with AlphaFold3, achieving statistically significant accuracy gains when integrated into our optimization framework. By uniquely combining composite language modeling, denoising-based end-to-end learning, and deep learning-guided post-optimization, DRfold2 establishes a distinct direction for advancing ab initio RNA structure prediction. Funding: This work was supported in part by the Ministry of Education, Singapore (T1251RES2309 and T2EP20125-0039 to YZ), the Agency for Science, Technology and Research (A*STAR), Singapore (IAF-PP H25J6a0034 to YZ), and Natural Science Foundation of Ningxia Province (2023AAC05036 to CF). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Data Availability: The datasets collected and used in this work are available at https://zhanggroup.org/DRfold2/data.zip and https://doi.org/10.5281/zenodo.18401691 . All underlying data are available via the Supporting information file ( S1 Data ). We show structures of 7QR3, 8FZA, and 8DP3 obtained by four-digit accession codes in the PDB repository ( https://www.rcsb.org/ ). All experimental RNA structures tested in this study are publicly available in the Protein Data Bank with the following accession codes: 7ELP, 7QR3, 7QR4, 7V9E, 7YR6, 7YR7, 8BU8, 8DP3, 8FZA, 8GXC, 8HB8, 8HZD, 8HZL, 8ITS, 8JHP, 8QO2, 8QO3, 8S95, 8SYK, 8T2P, 8TVZ, 8U5Z, 8UO6, 8V1H, 8VCI, 8VT5, 8XTR, 8XZL, 8Z8Q, 8Z8U, 9BCI, 9BUN, 9BZC, 9DE6, 9DXL, 9EOW, 9G7C, 9IS7, 9KPO. CASP16 results are available through the CASP prediction center ( https://predictioncenter.org/ ). The RNA sequence data for RCLM training was download from RNAcentral ( https://rnacentral.org/ ). The online server and standalone package of DRfold2 are freely available at https://zhanggroup.org/DRfold2/ and https://doi.org/10.5281/zenodo.18401691 . Copyright: © 2026 Li et al. This is an open access article distributed under the terms of the Creative Commons Attribution License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. In this study, we introduce DRfold2, a new deep learning-based framework for high-quality RNA structure prediction. The core novelty of DRfold2 consists of a pretrained RNA Composite Language Model (RCLM), which improves likelihood approximation and captures co-evolutionary signals from unsupervised RNA sequences more effectively than the previously used embeddings learned from structure prediction pipeline [ 21 ]. Next, the sequential and pairwise representations from RCLM are processed by an RNA Transformer-based pipeline, powered with a denoising module for end-to-end structure and geometry prediction. A post-processing protocol is then implemented for final RNA model selection and optimization, built on the end-to-end and geometry potentials. The large-scale benchmark results demonstrate that co-evolution features learned by RCLM from unsupervised data, when integrated with denoising structure learning, significantly enhance the accuracy of ab initio RNA structure prediction. Despite the progress, RNA structure prediction accuracy has not yet reached a so-called “AlphaFold moment” [ 23 ], partly due to limited solved structures compared to proteins. Language models have shown promise in protein structure prediction by capturing key features like co-evolution patterns and improving sequence representation from massive unsupervised sequence data [ 24 ]. Their ability to predict structures from single sequences is particularly valuable for RNAs, while multiple sequence alignment (MSA) construction [ 25 ] is more resource-intensive and the usefulness often varies depending on the depth of MSAs. Computer-based RNA structure prediction has evolved considerably over the past decades. Traditional template- and fragment-based approaches [ 9 – 12 ] construct structure models using known templates or fragments but are constrained by the limited availability of experimentally solved structures. Ab initio methods [ 13 – 15 ] attempt to predict structures from sequence alone; however, traditional physical or statistical force fields often lack the accuracy and specificity needed to fold RNAs with complex topologies. Recent advances in deep learning have introduced several innovative approaches. For example, DeepFoldRNA [ 16 ] and trRosettaRNA [ 17 ] utilize transformer-based models to predict local geometries and reconstruct coordinates from the predicted restraints. On the other hand, end-to-end models like RhoFold [ 18 ] and RoseTTAFoldNA [ 19 ] directly predict 3D coordinates, inspired by AlphaFold2 [ 20 ] in protein structure predictions. Our previous approach, DRfold [ 21 ], combines both end-to-end and geometry-based predictions through optimizing a hybrid potential function with a differentiable fold program [ 22 ]. Researchers have developed multiple approaches to study RNA structures, ranging from chemical probing [ 4 ] and proximity ligation methods [ 5 ] for detecting base pairs or other interactions, to structural biology techniques such as X-ray crystallography [ 6 ], NMR spectroscopy [ 7 ], and cryo-EM [ 8 ]. Despite higher resolution, experimental RNA 3D structure determination is often prohibitively expensive and sometimes infeasible. Consequently, computational approaches have faced growing demands for high-quality 3D structure modeling directly from sequences. Understanding the structure and function of RNA molecules has been a central focus of molecular biology and the pharmaceutical industry. RNAs, particularly non-coding RNAs, fold into specific structures that can be functional in various cellular processes, including gene regulation (e.g., transcription and translation), catalysis, cellular signaling, stress responses, and genome defense [ 1 ]. Due to their functional significance, RNAs have emerged as promising targets for small-molecule drug development, especially for human diseases that lack traditional protein targets [ 2 , 3 ]. With the exponential growth of RNA sequence data from high-throughput sequencing and the widening gap between known sequences and experimentally resolved RNA structures, determination of atomistic structures of RNAs from the primary sequences has become increasingly urgent. Results The overall DRfold2 pipeline for RNA structure prediction is illustrated in Fig 1A. Starting from an RNA sequence, DRfold2 first employs a pretrained RCLM to embed the query into sequence and pair representations. Here, RCLM achieves superior sequence pattern recognition through training on large-scale unsupervised sequence data using a composite likelihood [26] maximization approach (Fig 1B). These sequence and pair representations are processed by a set of RNA Transformer Blocks (Fig 1C), which produce necessary representations for structure folding. DRfold2 then generates RNA conformations through a Denoising RNA Structure Module (DRSM) in an end-to-end fashion (Fig 1D). The final RNA models are obtained through a post-processing CSOR protocol that is designed to select and refine conformation decoys generated from a pool of checkpoints (Fig 1E). PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 1. The pipeline of DRfold2 for end-to-end RNA structure prediction. (A) Overview of DRfold2 pipeline. (B) Details of training RCLM with masked negative composite log likelihood loss function. (C) Details of RNA Transformer Block. (D) Details of Denoising RNA Structure Module. (E) Detailed pipeline of CSOR Protocol as post-process to select and refine final RNA models. https://doi.org/10.1371/journal.pbio.3003659.g001 To objectively benchmark the performance of DRfold2, we constructed an independent test dataset containing 41 RNA structures with sequence length < 400 nts collected from (1) recent RNA-Puzzles targets [27], (2) CASP15 RNA targets [28,29], and (3) recently released RNA-only structures after Year 2024 (before March 25th 2025) from Protein Data Bank (PDB) [30]. Notably, large synthetic RNAs from CASP15 were excluded due to their deviation from naturally occurring RNAs, which serve as the primary focus for most functional assays and drug design efforts. To ensure rigorous evaluation, our training set only includes RNA structures released before 2024, while excluding any RNAs sharing more than 80% sequence identity with the test RNAs. Both the training and test datasets are made publicly available at https://zhanggroup.org/DRfold2/data.zip and https://doi.org/10.5281/zenodo.18401691. DRfold2 enables robust RNA structure prediction for targets with novel folds While DRfold2 shows highly competitive performance against existing top-performing methods such as AlphaFold3, it is also of interest to examine how DRfold2 performs on targets whose structural folds differ from those present in the training set. We utilized US-align [37] to filter out targets that have maximum normalized TM-score >0.45 against all training samples. This procedure yields a total of 25 novel-fold targets. We note that, while structurally homologous folds are excluded by this procedure, homologous sequences or closely related RNA families may still exist in the large-scale sequence corpus (RNAcentral, release 22) used for language model pretraining, which is unavoidable in self-supervised sequence modeling. However, no experimentally resolved structural information from these sequences is used during training, and the novel-fold analysis therefore does not imply reliance on trivial structural memorization. On these novel folds, DRfold2 achieves an average TM-score of 0.309, which is lower than its performance on the full benchmark set (0.350), indicating that prediction accuracy decreases for unseen structural topologies. Importantly, DRfold2 is trained to learn a mapping from sequence to structure, and for targets with distinct sequences, the model cannot rely on simple memorization of previously observed structures to produce correct predictions. The generally higher performance of DRfold2 on those known-fold targets over novel targets, actually demonstrates that the DRfold2 model can capture latent structural representations that go beyond sequence similarity. Notably, DRfold2 consistently outperforms AlphaFold3 on this subset of new-fold RNAs (0.309 versus 0.281), with a p-value of 5.29E−02 for one-tailed t test (S2 Fig). Furthermore, DRfold2 achieves a significantly lower RMSD (15.888 versus 18.832) with a p-value of 1.97E−02 (one-tailed t test), despite the limited number of samples. These results suggest that DRfold2 can better retain nontrivial generalization ability by extracting structural principles from sequence and does not simply memorize previously seen folds or sequences. DRfold2 enables robust RNA structure prediction with few homologs Next, we examine how the performance of DRfold2 varies across different levels of evolutionary information, as reflected by the MSA depth retrieved by AlphaFold3. In Fig 3F, we present this comparison together with AlphaFold3. For reference, we also include the performance of DRAF as well as that of AlphaFold3 without MSA input. In general, both DRfold2 and AlphaFold3 achieve better performance with the increase of numbers of homologous sequences, likely due to richer co-evolutionary constraints. When no MSA is available, AlphaFold3 achieves an average TM-score of only 0.333, which is moderately lower than its full-version performance (0.347). Although DRfold2 and AlphaFold3 show comparable overall accuracy, their behavior differs across MSA-depth groups. For targets without any detectable homologs (MSA depth = 0), DRfold2 achieves slightly higher accuracy (TM-score = 0.300) compared to AlphaFold3 (TM-score = 0.287). A similar trend is observed for low-homology targets (MSA depth 1–10), where DRfold2 outperforms AlphaFold3 with TM-scores of 0.442 versus 0.407. For targets with richer evolutionary information (for example, MSA depth 11–20), however, AlphaFold3 performs better (TM-score = 0.357 versus 0.311 for DRfold2). Nonetheless, the hybrid folding pipeline, DRAF, which integrates DRfold2 predictions with the AlphaFold3 potential, achieves the highest accuracy (TM-score = 0.368) in this group, demonstrating the complementarity of the two approaches. Taking together, these results further highlight the complementary strengths of DRfold2 and AlphaFold3. DRfold2 is trained primarily on single-sequence features enriched with co-evolutionary signals derived from RCLM. The hidden universal covariance learned by RCLM becomes more valuable, for example, when explicit homologs are insufficient or not detected. Consequently, DRfold2 can provide more reliable predictions for RNAs with few or even no homologs. DRfold2 produces physically realistic RNA structures In addition to achieving high global structure accuracy, DRfold2 also produces physically correct models, addressing a crucial problem in current RNA bioinformatics [38]. DRfold2 employs a hybrid refinement strategy that integrates Monte Carlo simulation with L-BFGS optimization to generate the final full-atomic models. We evaluated the geometric quality of the refined structures on the test set using three standard metrics, i.e., the root mean square (RMS) of bond length and bond angle deviations, and the clash score (S3 Table). The corresponding values for native structures are also provided as references. The refinement procedure substantially improves all three metrics, reducing them from 0.042, 5.854, and 42.823 to 0.012, 1.559, and 2.171, respectively. Notably, DRfold2’s refinement step yields structures with an average clash score even lower than that of the experimental structures, highlighting its strong physical regularization capability. S3 Table further reports the number of structures containing at least one entanglement [39], as identified by RNAspider, a webserver for detecting topological entanglements in RNA 3D structures [40]. Interestingly, although DRfold2 refinement is not explicitly designed to resolve such motifs, the refinement successfully removes one D&L interlace, in which a dinucleotide step (D) and a loop (L) are interwoven (S3 Fig). We further examined the local torsion-angle similarity by evaluating the Mean of Circular Quantities (MCQ), a torsion-based distance metric that assesses the local geometry defined by every consecutive four atoms, with a lower MCQ value indicating better local structure [41]. S4 Fig presents the MCQ comparison across all baseline methods and DRfold2. On average, DRfold2 achieves an MCQ of 23.8, significantly outperforming many control methods (RNAComposer, trRosettaRNA, and RhoFold). For example, RNAComposer, which assembles structures using a high-quality fragment library and thus can generally provide good local torsion geometry, still shows a higher MCQ of 26.1 (p-value = 1.06E−03, one-tailed t test). The MCQ value of DRFold2 is, however, comparable or slightly higher than the rest of the control methods (RoseTTAFoldNA, DeepFoldRNA, AlphaFold3), suggesting potential future improvements, for example, through the incorporation of torsion-based energy potential. In summary, DRfold2 provides not only accurate global folds but also physically well-behaved structures, ensuring that the predicted models are suitable for reliable downstream experimental design and computational analyses. DRfold2 networks enhance the accuracy of RNA secondary structure prediction Many RNA functions are closely associated with their secondary structure arrangements, which are primarily stabilized by hydrogen bonding and base-pair stacking interactions [42]. Fig 4A–4D analyzes the Interaction Network Fidelity (INF) indices, which are defined as the geometric mean of precision and recall of secondary structure predictions, i.e., . First, for Watson-Crick base pairs (INF_wc) and stacking interactions (INF_stack), while most methods (except for trRosettaRNA) demonstrate reasonable performance, DRfold2 achieved the INF values of 0.831 and 0.754, slightly lower than those of AlphaFold3 (0.849 and 0.762), but higher than all other control methods. For non-canonical interactions (INF_nwc), however, all methods, including DRfold2 and AlphaFold3, exhibit much lower INF values. Although DRfold2 and AlphaFold3 still achieves the top-2 INF_nwc, their low value (0.184 and 0.202, respectively) underscores that accurately modeling non-canonical base-pair interactions remains an urgent challenge in RNA structure prediction. Fig 4D presents a combined secondary structure accuracy across all three indices. DRfold2 achieves INF_all = 0.750, which is statistically comparable to that of AlphaFold3 (p-value = 2.92E−01, two-tailed t test), and 7.4% higher (p-value = 5.97E−04, one-tailed t test) than the third-best method, DeepFoldRNA (INF_all = 0.698). PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 4. Comparative analysis of RNA secondary structure prediction. (A–E) Box plots comparing eight RNA structure prediction methods across different RNA secondary structure performance metrics. The green point and the white horizontal line represent the mean and median values respectively. Note that the mean Deformation Indexes for trRosettaRNA and RhoFold are much higher than the median, due to their poor performance on a few targets where no correctly modeled base pairs are present in the predicted structures. (F) Cumulative distribution plot of deformation index for different methods. Underlying numerical data for this figure can be found in S1 Data (see sheets “S1_Data_Figure4_A-F”). https://doi.org/10.1371/journal.pbio.3003659.g004 While the INF indices assesses the conformational accuracy at the nucleotide pair level, the Deformation Index (DI), defined as RMSD/INF_all, has been introduced to provide a more balanced evaluation of 3D topological and 2D base-pairing patterns [43]. As shown in Fig 4E, DRfold2 achieves the lowest average DI of 20.21, which is 9.6% lower than the second-best method, AlphaFold3 (DI = 22.15), The corresponding p-value of 1.43E-01 in two-tailed t test indicates that this improvement is not statistically significant. However, the method that combines DRfold2 and AlphaFold3 potentials (DRAF) achieves the lowest DI of 20.15, which is significantly lower than that of AlphaFold3 (p-value = 2.14E−02, one-tailed t test). Fig 4F further presents the accumulative fraction distribution at different DI cut-offs, where DRfold2 outperforms all third-party control methods, achieving the highest normalized area under the curve (AUC = 0.520), which is 5.3% higher than the best control method, AlphaFold3 (AUC = 0.494). These results suggest that DRfold2 excels not only in predicting global structural topology but also in capturing detailed hydrogen-bonding and base-pair stacking patterns within RNA structures. It is notable that DRfold2 has achieved superior secondary structure predictions using sequence information alone, without incorporating any precomputed secondary structure predictions, as many control methods do. This success can be attributed to several key components of DRfold2. First, the RCLM effectively captures inter-residue co-evolutionary patterns, providing reliable embedding features for the deep learning model. Second, the Denoising Structure Module helps enhance the end-to-end learning of inter-nucleotide frames through supervised training, ensuring proper spatial relationships between nucleotides. Finally, the precise prediction of inter-nucleotide geometries through multi-task learning helps guide the refinement of base pairing conformations during the post-processing optimization stage. RCLM more effectively captures co-evolutionary patterns than existing language models One of the main innovations of the DRfold2 pipeline is the introduction of RCLM which is expected to provide a more precise approximation of the full likelihood probability by considering higher-order inter-nucleotide interactions (see Materials and methods). To examine the assumption, first, S5 Fig evaluates the basic nucleotide sequence recovery rate of RCLM against three control language model (LM) methods, i.e., RNA-FM, RiNALMo, and RNAErnie. Here, RNA-FM is a representative RNA sequence language model [44] trained on RNAcentral database [45] with conventional Masked Language Model Loss. RiNALMo was trained with 650M parameters on 36M non-coding RNA sequences from several database [46]. RNAErnie was trained through a specifically designed motif-level random masking strategy [47]. During the training, the composite likelihood of RCLM was calculated based on the sum of two probabilities for nucleotide-wise masked type and pair-wise joint predictions. In S5 Fig, we also examine the sequence recovery ability derived from the pair-wise joint probability. Here, we marginalize the predicted joint distribution into nucleotide-wise distribution for evaluation, denoted as RCLM-M. The data show an improved masked nucleotide recovery ability for RCLM, with a recovery rate of 45.82%, compared to RNA-FM (41.47%) and RNAErnie (42.87%). Interestingly, the marginalized prediction (RCLM-M) achieves slightly better performance, with a sequence recovery rate of 47.87%, marginally higher than the second-best model, RiNALMo. Notably, these evaluation metrics are marginalized across positions, representing only a subset of RCLM. Even with this partial probability, RCLM demonstrates consistently superior performance over RNA-FM. The result of RCLM-M suggests that RCLM also generates meaningful predictions for the pair-wise joint nucleotide distributions. In addition to sequence recovery rate, which is highly related to the masking and loss function strategy during the training and cannot fully capture the performance of the model, a more critical assessment is on the co-evolution patterns captured by the language models, which are directly relevant to 3D RNA structure predictions. To evaluate this capability, we employed a completely unsupervised method, Categorical Jacobian [48], to extract the RNA contact maps from the language models. Fig 5A–5C presents a comparison of top-N contact map precision between RCLM and control LMs across distance thresholds of 12, 16, and 20 Å, where RCLM consistently has dramatically higher precisions than all control language models. For example, when using a contact distance cutoff of 12 Å, the Top-L precision of RCLM (49.0%) is more than 100% higher than that achieved by RNA-FM (23.6%). RCLM also produces significantly higher contact prediction than the second-best LM model, RiNALMo, with p-values of 4.82E−09, 5.00E−09, and 1.04E−09 for thresholds of 12, 16, and 20 Å, respectively, in one-tailed t test. In Fig 5A–5D, we further examine the precisions of the Categorical Jacobian contact and secondary structure predictions based on marginalized pairwise prediction from RCLM (RCLM-M). While RCLM-M produces slightly lower precisions than RCLM, it still maintains a large performance margin compared to those control methods. In Fig 5D, we also compare the precision of unsupervised secondary structure prediction, by assuming that the Categorical Jacobian produces the base pairing score maps. RCLM, RCLM-M, and RiNALMo exhibit comparable performance, while RCLM once again achieves highest precision and shows a statistically significant advantage over the second-highest model, RiNALMo, in Top-L/2 precision, with p-value of 8.57E-03 in one-tailed t test. These results indicate that RCLM can more effectively learn high-quality co-evolution patterns from sequence data than existing language modeling approaches. Since Categorical Jacobian requires mutation of all positions into all possible nucleotides during the extraction process, the high contact precision by RCLM also suggests that the Composite Language Model has achieved a reliable level of sensitivity. It is important to note that, RCLM, containing 47M parameters, demonstrates superior performance in unsupervised contact precision tasks, with only 50% of the model size of RNA-FM (100M) and RNAErnie (87M), and even less than 10% of RiNALMo (650M). Such efficiency makes this form of language model training a promising approach for RNA sequence language model learning. Our next step is to upscale this methodology to explore the full potential of RCLM. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 5. Comparative analysis of RCLM and control RNA language models for RNA sequence learning. (A–C) Top-N precision curves at different contact thresholds (12, 16, and 20 Å) for RCLM, RCLM-M, and control language models. (D) Precision comparison of unsupervised RNA secondary prediction between RCLM, RCLM-M, and control language models. (E) 3D structure visualization of the Class I type III preQ1 riboswitch from E. coli (PDB ID: 8FZA). (F, G) Contact maps of 8FZA, with gray dots representing ground-truth contacts, red dots false predicted contacts, and blue dots true positive predictions. Upper-left triangle shows results of RiNALMo predictions, and lower-right triangle is that of RCLM (F) and RCLM-M (G) predictions. Underlying numerical data for this figure can be found in S1 Data (see sheets “S1_Data_Figure5_A-D” and “S1_Data_Figure5_F-G”). https://doi.org/10.1371/journal.pbio.3003659.g005 As an illustration, Fig 5E–5G showcase an example of a Class I type III preQ1 riboswitch from Escherichia coli (PDB ID: 8FZA) [49], which consists of 30 nucleotides as shown in Fig 5E. RCLM and RCLM-M achieve reasonable predictions on this target with Top-L contact precision of 56.67% and 63.33%, respectively, where the contacts are defined by inter-N atom distances <12 Å. The second-highest precision among the control methods is 23.33% from RiNALMo. The upper-left and lower-right panels of Fig 5F and 5G are the unsupervised Top-L predicted contacts by RiNALMo and RCLM/RCLM-M, respectively, where the ground-truth contacts are shown as gray dots. RiNALMo only predicts 7 out of 30 (23.33%) positive contacts, while RCLM and RCLM-M correctly produce 17 (56.67%) and 19 (63.33%) positive contacts when using single nucleotide prediction and marginalized pairwise prediction, respectively, both exceeding RiNALMo’s accuracy by more than 2-fold. The result of this example highlights again the efficiency of RCLM to learn co-evolution information that is crucial for producing higher-precision structural restraints and ab initio RNA models, especially for the cases with limited sequence and structural homologies. RCLM learns base pairing interactions in an unsupervised manner Although the model is optimized purely for likelihood-based sequence reconstruction, the coevolutionary statistics it captures are validated to correlate strongly with physical contacts (Fig 5). To further understand how RCLM learns these physical interactions, we conducted a nucleotide-masking experiment (in addition to the Categorical Jacobian analysis). Specifically, for each annotated base-pairing interaction, we masked one nucleotide and assessed how accurately RCLM could recover the masked residue from the remaining context, and the corresponding attention weight distributions. Representative examples are shown in S6 Fig for two targets whose PDB IDs are 8HZD (chain A) and 8Z8U (chain A), respectively. In 8HZD_A, masking positions 2 or 6 causes the attention mechanism to assign higher weights to their long-range canonical and non-canonical partners, indicating that RCLM learns these interactions by attending directly to the physically paired nucleotides. A similar behavior is observed in 8Z8U_A for positions 1 and 20. However, this physically grounded attention does not always generalize, especially for some non-canonical interactions. For instance, when masking position 8 in 8HZD_A or position 13 in 8Z8U_A, the model relies more on local neighboring nucleotides rather than attending to the true long-range interaction partner. Overall, for canonical Watson–Crick pairs, the recovery success rate was 57.1%, indicating that RCLM effectively captures canonical interacting patterns. However, for non-canonical interactions, the recovery rate dropped to 36.6%, suggesting that RCLM learns weaker signals for these interaction types (S7 Fig). Both canonical and non-canonical base pairs are recovered at higher accuracies above random expectation (~25%), indicating that RCLM has learned to recognize true physical pairing signals without any structural supervision. Such difference in accuracy may partly account for the higher performance of DRfold2 on canonical, compared to non-canonical, base-pair modeling (Fig 4B). We further performed a two-nucleotide masking experiment on annotated non-canonical pairs. In this setting, both nucleotides in a non-canonical interaction were masked, and we evaluated whether RCLM could correctly recover both nucleotide identities using the sequence-level prediction head and the pairwise prediction head (supervised by the pairwise likelihood term). Overall, the pairwise head achieved a slightly higher recovery rate (18.0%) compared with the sequence head (17.8%). The pairwise head outperformed the sequence head in 9 targets, while the sequence head outperformed the pairwise head in 5 targets. However, the difference is not statistically significant (p-value = 7.81E−01, two-tailed t test), indicating that the composite likelihood does not provide a significant advantage for non-canonical interactions. This observation, however, aligns with our expectation since all nucleotide pairs are treated equally without bias terms for non-canonical pairs during the RCLM training. In summary, these results confirm that, while RCLM provides strong support for canonical base-pair learning, its ability to model non-canonical interactions remains limited. Nevertheless, they also suggest potential directions for improvement, such as explicitly modeling non-canonical interactions in composite likelihood. Positive impact of Composite Language Model on RNA 3D structure prediction To directly examine the effectiveness and impact of different language models on 3D RNA structure prediction, we present in Fig 6A the TM-scores of identical end-to-end structural learning networks built on one-hot and RCLM representations. We also include RNA-FM as a reference model, which was previously shown to be helpful for RNA structure prediction [50]. It is observed that models using RCLM features achieve the highest mean and median TM-scores of 0.319 and 0.299, which are 6.0% and 10.7% higher than that with RNA-FM (0.301 and 0.270) and 7.8% and 12.4% higher than that of the one-hot representation (0.296 and 0.266), respectively. Fig 6B and 6C further list head-to-head TM-score comparisons of RCLM against the control models, where RCLM outperforms one-hot and RNA-FM in 30 and 25 cases, while underperforming in only 11 and 16 cases, respectively. PPT PowerPoint slide PNG larger image TIFF original image Download: Fig 6. Comparison of RNA structure prediction performance across different encoding methods and correlation analysis between unsupervised contact/secondary structure precision and TM-score/RMSD. (A) Violin plots of TM-scores by end-to-end RNA structure predictions based on different encoding strategies (One-hot, RNA-FM, and RCLM), with white horizontal lines indicating medians and green points showing means. (B, C) Head-to-head TM-score comparisons between RCLM and One-hot and RNA-FM encoding based models. (D) Correlation between TM-score and Top-L/2 secondary structure precision. (E) Correlation between TM-score and Top-L/2 Contact precision at 12Å threshold. (F) Correlation between RMSD and Top-L/2 Contact precision at 12Å threshold. Red lines indicate linear regression fits. Underlying numerical data for this figure can be found in S1 Data (see sheets “S1_Data_Figure6_A-C” and “S1_Data_Figure6_D-F”). https://doi.org/10.1371/journal.pbio.3003659.g006 To quantitatively interpret the contribution of language models to the 3D RNA structure modeling, Fig 6D shows the relationship between the TM-score of RCLM-based models and the Top-L/2 secondary structure precision from RCLM Categorical Jacobian, where a medium-high Pearson Correlation Coefficient (PCC = 0.453) is observed, suggesting that targets with a higher RCLM-based secondary structure precision generally have more accurate 3D structure prediction. In addition to base pairing precision, Fig 6E and 6F also examine the correlation between unsupervised contact map precision with TM-score and RMSD, respectively, which show slightly weaker but robust PCC values of 0.370 and −0.358. Collectively, these results suggest that the comprehensive representations provided by advanced language models, like RCLM, can enhance the ability to learn co-evolutionary patterns and spatial restraints, leading to more accurate 3D RNA structure modeling through the end-to-end network in DRfold2. Performance of DRfold2 in CASP16 An earlier version of DRfold2 participated in the 16th Critical Assessment of protein Structure Prediction (CASP16) experiment for RNA structure prediction under the group ID “dNAfold.” S4 Table summarizes the average TM-score and accumulated Z-scores for the Top 20 groups on the 20 RNA targets with known experimental structure and sequence lengths below 400 nucleotides. Although the CASP16 assessment contains both RNA and DNA targets, only monomer RNA targets, for which dNAfold submitted predictions, were included in this analysis. The results revealed that Vfold, which utilizes a template-based modeling approach integrated with AlphaFold3 [50], achieved the highest performance, with an average TM-score of 0.566 and Z-score of 19.71, outperforming other groups. The remaining top-ranked groups had comparable average TM-scores around 0.5, with DRfold2 securing 5th place in both TM-score (0.510) and Z-score (13.46), indicating its strong competitiveness among the state-of-the-art programs. Consistent with the benchmark results shown in Fig 3, DRfold2 slightly outperformed the AlphaFold3 server, which was ranked 7th in TM-score (0.501) and 16th in Z-score (9.78) (S4 Table). It is noteworthy that the AlphaFold3 server was released before the CASP16 experiment, allowing the participating groups to integrate AlphaFold3 predictions into their modeling pipelines, which many did [51]. However, we chose not to incorporate AlphaFold3 or any third-party predictions in our CASP16 submissions, aiming to more critically assess the DRfold2’s independent modeling capabilities. To further assess the uniqueness of the RNA structure predictions, we introduced an AlphaFold dissimilarity score, defined as 1 − TM M , where TM M is the maximum TM-score between the first submitted model by each group and the five models submitted by AlphaFold3 (Group ID: AF3-server). In S8 Fig, we plot the average TM-score on 20 RNA monomer targets versus the AlphaFold dissimilarity for all groups that submitted predictions for all these targets, where dNAfold exhibited the highest dissimilarity score among the six groups that have an average TM-score outperforming AlphaFold3. Moreover, dNAfold achieved the top rank among all groups with a dissimilarity score above 0.45. These results highlight the capability of DRfold2 for producing distinct and independent structure predictions. Nevertheless, given the complementary nature of DRfold2 and AlphaFold3 in both methodology and the benchmark results shown in Fig 3, we expect that appropriate integration of DRfold2 with AlphaFold3 or other top-performing models represents a promising direction for future improvements in DRfold2-based RNA structure prediction. [END] --- [1] Url: https://journals.plos.org/plosbiology/article?id=10.1371/journal.pbio.3003659 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/