(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . miRScore: A rapid and precise microRNA validation tool [1] ['Allison Vanek', 'Bioinformatics', 'Genomics Ph.D. Program', 'Huck Institutes Of The Life Sciences', 'The Pennsylvania State University', 'University Park', 'Pennsylvania', 'United States Of America', 'Department Of Biology', 'Sam Griffiths-Jones'] Date: 2026-02 MicroRNAs (miRNAs) are small non-protein-coding RNAs that regulate gene expression in many eukaryotes. Next-generation sequencing of small RNAs (small RNA-seq) is central to the discovery and annotation of miRNAs. Newly annotated miRNAs and their longer precursors encoded by MIRNA loci are typically submitted to databases such as the miRBase microRNA registry following the publication of a peer-reviewed study. However, genome-wide scans using small RNA-seq data often yield high rates of false-positive MIRNA annotations, highlighting the need for more robust validation methods. miRScore was developed as an independent and efficient tool for evaluating new MIRNA annotations using sRNA-seq data. miRScore combines structural and expression-based analyses to provide rapid and reliable validation of new MIRNA annotations. By providing users with detailed metrics and visualization, miRScore enhances the ability to assess confidence in MIRNA annotations. miRScore has the potential to advance the overall quality of MIRNA annotations by improving accuracy of new submissions to miRNA databases and serving as a resource for re-evaluating existing annotations. MicroRNAs (miRNAs) play a major role in gene regulation in most eukaryotic organisms. Genome-wide analysis of miRNAs and miRNA-encoding precursors (here, MIRNAs), can lead to numerous false positive annotations. Criteria for MIRNA annotation often use a combination of short RNA sequencing and predicted RNA secondary structural properties of precursors. However, implementation of these criteria varies. Here, we introduce a tool, miRScore, developed to standardize MIRNA validation using accepted annotation criteria. miRScore is intended to improve the quality of MIRNA annotation. miRScore takes as input one or more putative miRNA sequences, one or more corresponding precursor RNAs, and short RNA sequencing data. miRScore quickly and accurately evaluates each candidate in the context of the provided data and determines whether each annotation meets all criteria. These results can be used to determine high-confidence miRNAs for cataloging and downstream analysis. Funding: This work was supported by the National Science Foundation (2130884 to MJA and BCM; 2450802 to BCM), the Biotechnology and Biological Sciences Research Council (BB/W018438/1 to SGJ), and a seed grant from The Huck Institutes of the Life Sciences at Penn State (un-numbered award to MJA). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Introduction MicroRNAs (miRNAs) are a class of small, non-coding RNAs that regulate gene expression within eukaryotes. This regulation typically occurs when a miRNA, which is loaded into an RNA-induced silencing complex (RISC), imperfectly base pairs to a target messenger RNA (mRNA). The RISC then frequently acts as an endonuclease to cleave the mRNA or to otherwise inhibit its translation [1–4]. miRNA-directed regulation of mRNAs is crucial in various biological processes such as developmental timing [5–7], metabolism [8,9], and defensive pathways [10–13] in both plants and animals. Although miRNA biogenesis varies somewhat between animals and plants, the fundamental aspects of miRNA structure and function are conserved [14]. In both plants and animals, the precursors of miRNAs are generally transcribed by RNA polymerase II from an endogenous MIRNA gene. While many MIRNA primary transcripts are transcribed as independent genes from intergenic regions, some are processed from the introns of protein-coding mRNAs [4]. Transcription results in a long single-stranded RNA containing a hairpin, called the primary miRNA (pri-miRNA). The hairpin embedded within the primary transcript is then processed by sequential endonuclease activity (Drosha and Dicer in animals, or by a single Dicer-Like protein in plants) to release a miRNA duplex. The miRNA duplex is a double-stranded RNA, typically with a few mismatched and/or bulged nucleotides, which consists of the mature functional strand (miRNA) and passenger strand (miRNA*). The miRNA duplex is unwound, and a single-stranded mature miRNA is bound to an Argonaute protein to form the RISC. Most frequently a single strand from this duplex is incorporated into RISC and regulates mRNAs; in some cases both strands from the miRNA duplex become separately bound to different RISCs and have two distinct constellations of mRNA targets. For details of microRNA biogenesis, see [4,15,16]. Alignment of deep small RNA-sequencing (sRNA-seq) data to a reference genome is a common method for MIRNA annotation and quantification. Several tools such as ShortStack [17,18], miRador [19], miRDeep [20], and miRDeep-P2 [21], have been developed to annotate miRNAs and other small RNAs using sRNA-seq data. These tools typically work by aligning sRNA-seq data to a reference genome, followed by evaluation of potential miRNA-encoding loci (MIRNA). One way candidate MIRNAs are identified is by the distinctive alignment pattern of the miRNA/miRNA* duplex reads to the hairpin precursor. miRNA and miRNA* reads from sRNA-seq align to a single genomic strand, as their precursors are single-stranded transcripts. These reads align a short distance from each other, forming two distinct “stacks” of read coverage [17,22]. MIRNA primary transcripts are typically short-lived and hard to detect using sRNA-seq or regular mRNA-seq. Most sRNA-seq centered MIRNA identification tools thus annotate “hairpin” sequences that encompass the stem-loop region and some adjacent sequence of pre-determined length. The start and stop positions of these annotations do not necessarily correspond to the ends of the actual primary transcripts. The secondary structure of this putative hairpin precursor is then predicted. For true MIRNAs, the predicted secondary structure of the putative precursor RNA is an imperfect stem-loop. Furthermore, two stacks of aligned sRNA-seq reads from the miRNA and the miRNA* are found on opposite arms of the predicted stem-loop with a diagnostic two nucleotide 3’-overhang. Generally, the sequence with the most abundant set of reads is termed the ‘mature’ miRNA, while the sequence with less abundant reads is the ‘star’ sequence. Detection of reads from both arms of the miRNA duplex is required to confirm the predicted duplex [23–25]. Identification of candidate MIRNAs using sRNA-seq is therefore dependent on empirical evaluation of read alignment patterns in the context of the presumed precursor’s predicted RNA secondary structure. The identification of MIRNAs through deep sequencing data poses some challenges. One is the handling of multimapping reads, in which there are multiple best-scoring alignments for a single read. This occurs frequently with sRNA-seq data due to shorter read lengths and the fact that identical miRNAs can be encoded by paralogous loci [18]. Another challenge is distinguishing true MIRNAs from other sRNA classes such as short-interfering RNAs (siRNAs), which have their own unique alignment patterns and criteria [23,26]. Each MIRNA discovery tool employs distinct methods for handling these challenges, with varying degrees of performance for identification of novel MIRNAs in plants and animals [17,19–21]. The lack of uniform implementation of well-defined MIRNA criteria, coupled with the challenging nature of informatically distinguishing miRNAs from noise or other sRNA species, has led to diminishing confidence in the overall quality of existing MIRNA annotations [24,27–30]. There have been considerable efforts to define MIRNA criteria to improve the quality of annotations [23–25,31]. Some miRNA databases contain a significant number of false positive annotations [28,30,32]. miRBase for example relies on researchers and peer reviewers to assess the validity of miRNAs before submission, and has adopted methods of determining confidence in these community-based annotations [27,32]. The current release of miRBase (V. 22.1) contains over 48,000 mature miRNA sequences from 271 diverse species including animals, plants, and some protists [32]. MirGeneDB has taken a different approach, manually curating MIRNA annotations of metazoan species through structural, expression, and conservation analysis [25,31,33]. Whether by database curators or the research community, the assessment of novel miRNAs relies on a degree of manual inspection and evaluation. However, manual inspection of incoming annotations takes significant effort and currently lacks standardized implementation. While there are many de novo miRNA annotation tools and miRNA databases available, a secondary method to quickly analyze novel and annotated MIRNAs following genome-wide sRNA annotation is not available. Such a tool, to rapidly check new annotations, could be useful for database curators by removing the need for time-consuming manual inspection of new submissions. A standardized and quick method of automatically validating new MIRNA annotations would improve the quality of annotations published and subsequently submitted to online repositories. Retrospective application of such a method could also be used to flag and remove problematic entries in repositories such as miRBase. To address this need, we developed miRScore – a rapid and precise miRNA validation tool. miRScore can rapidly evaluate the annotation of both existing and novel miRNAs against specific sRNA-seq datasets using widely accepted MIRNA criteria in plants and animals. It offers a comprehensive evaluation of MIRNA loci, analyzing each criterion and producing visualizations of hairpin secondary structure and expression patterns. In this study, miRScore is described and tested using both annotated and novel MIRNAs from plants and animals. [END] --- [1] Url: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1013663 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/