(C) PLOS One This story was originally published by PLOS One and is unaltered. . . . . . . . . . . Training biologists in Unix command-line skills: From curriculum to interactive online tutorials [1] ['Lucie Khamvongsa-Charbonnier', 'Ifb-Core', 'Institut Français De Bioinformatique', 'Ifb', 'Cnrs', 'Inserm', 'Inrae', 'Cea', 'Villejuif', 'Robert Aboukhalil'] Date: 2026-04 As the generation of data in the life and health sciences expands rapidly, there is a growing need for professionals and students in these fields to master core bioinformatics skills, particularly those relating to Unix-like systems, most commonly used in bioinformatics. This paper introduces two key contributions to address this need: (1) A Unix curriculum for life scientists with little or no command-line experience, based on progressive Unix skill levels for bioinformatics and (2) An implementation of this curriculum into a series of interactive online tutorials deployed through Sandbox.bio—an open-source platform for learning bioinformatics that embeds a command line in the browser, which removes barriers related to software installation and access. We performed an overall evaluation of this teaching framework in different contexts. This inclusive, sustainable approach provides widespread access to essential bioinformatics skills for life science students and professionals alike. Funding: This work was partially supported by the French Institute of Bioinformatics (IFB), which was funded by the Future Investment Programme and subsidized by the National Research Agency under grant number ANR-11-INBS-0013. The IFB provided a salary for one of the authors, LK-C. The funder had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Introduction As data generation continues to accelerate, the life science and health communities are facing a growing need to master core bioinformatics skills. These communities encompass a diverse range of profiles from university undergraduates to professionals in research and medical institutions, eager to learn new analysis techniques [1–4]. Training such a large and diverse community raises significant challenges, especially regarding content design, trainer availability, and access to computational resources [5,6]. The International Society for Computational Biology Education committee has recently proposed with international experts a competency framework for bioinformatics including knowledge, skills, and attitudes (KSAs) for several bioinformatics career profiles [7]. In this context, Unix, Python, and R are commonly identified as core technical skills in data science in general and also in many bioinformatics training programs [8,9]. One notable example is the Software Carpentry initiative [10], which has been playing a key role in providing high-quality training and open material to the community (see: https://software-carpentry.org). However, this initiative largely depends on face-to-face training, which creates a bottleneck due to the need for a sufficient number of trained instructors. It also requires using a commercial cloud (Amazon Web Services) or installing several bioinformatics tools locally (using the command-line interface!). In recent years, self-paced learning has emerged as a key pillar of education, offering flexible, learner-centered approaches that allow individuals to progress at their own pace, manage their schedules, and revisit foundational concepts as needed [11]. This format is particularly well-suited for acquiring prerequisites ahead of in-person courses, managing large student cohorts, and supporting lifelong learning in professional environments. To meet this demand, several platforms—such as DataCamp (https://www.datacamp.com], Katacoda [now closed], and Killercoda (https://killercoda.com)—have contributed to democratizing self-paced learning in fields like data science and programming. However, these solutions remain proprietary, raising concerns about long-term availability (e.g., Katacoda was discontinued in 2022 [12]), and are generally not designed for bioinformatics. More broadly, many free and open educational resources have been developed to support training in computational biology. For instance, the Galaxy Training Network [13,14] is a prominent collaborative platform offering open-source tutorials tailored for both scientists and trainers across a wide range of topics. TeSS, the ELIXIR Training eSupport System, has a dedicated section to discover existing e-learning materials (https://tess.elixir-europe.org/elearning_materials). Additionally, an increasing number of bioinformatics tools now come with dedicated tutorials for educators and beginners, facilitated by the accessibility and collaborative features of platforms like GitHub and GitLab. Yet, they often assume access to sufficient computing infrastructure—whether local machines or remote clusters—and a minimum level of system configuration skills. These technical requirements pose a significant barrier for learners with limited prior experience, especially in institutional settings and countries where access to resources is uneven. In this context, WebAssembly (Wasm) technologies offer a transformative alternative by enabling programs to run directly within the user’s browser. This approach harnesses the user’s own computing power, eliminating the reliance on large centralized computational infrastructures. Moreover, software installation and local configuration are handled seamlessly, enabling effortless deployment and a smooth user experience. As most introductory tutorials make use of small datasets for demonstration purposes, a basic web browser serving basic files for a website (e.g., HTML, CSS, Javascript) together with Wasm-compiled programs should be generally sufficient for training. Wasm thus appears as a promising solution to meet the needs of online training for bioinformatics core skills where complex toolchains and system dependencies often hinder early learning. For instance, JupyterLite (https://jupyterlite.readthedocs.io/en/latest/), a Wasm implementation of JupyterLab, allows running a fully-fledged JupyterLab in the browser without the need of any external server. Relying on the WebAssembly technology, Sandbox.bio (https://sandbox.bio) represents a forward-looking step in bioinformatics education. At the initiative of one of the authors of this article (RA), it has now become an open-source online platform specifically designed for bioinformatics training. This platform leverages the use of the Wasm technology to provide self-paced guided tutorials for popular bioinformatics programs such as BLAST, SeqKit, or fastp. Sandbox.bio also provides virtual playgrounds where learners can experiment with Unix tools such as Grep, Awk, or Sed. Despite the consensus that Unix is a fundamental skill in bioinformatics, the specific commands and levels required in the field of bioinformatics are rarely defined beyond a generic level like “basic,” “intermediate,” or “advanced,” generally not attached to any precise operational skills. For example, does a biology student learning Unix need to know all or only some of the Unix commands related to data manipulation (like, grep, cut, count, awk…)? Does he/she need to learn to write a complete Bash script or rather only need to understand how a given script is working and be able to adapt it to his/her needs? It has become crucial to go beyond generic levels to truly unlock self-paced learning, and globally facilitate learning of Unix skills for the life sciences. In this work, we address two issues: Firstly, we designed a Unix curriculum targeting life scientists with little or no command line experience and defined progressive Unix skill levels in bioinformatics that go beyond the three generic basic, intermediate, and advanced levels. Secondly, we propose sustainable online tutorials to help biologists master these skills. The Unix curriculum details and organizes each individual Unix concept in coherent and progressive categories. This curriculum can be used for various purposes, such as self-assessing student level, designing a training curriculum based on learner profile, or specifying target skills to be acquired upon completion of the training. Secondly, we propose a series of interactive online tutorials based on the WebAssembly technology and hosted on the open-source sandbox.bio platform. This enables access without installation and supports a large number of trainees. These two resources can be used independently, yet they were designed to be complementary, with the tutorials directly covering the topics outlined in the skill levels. Designed for biologists, these resources enable the progressive learning of bioinformatics with biology-specific examples. Hosted by the sandbox.bio platform, they are inexpensive in terms of computing resources and can be deployed on a large scale without posing any beginner-related risk to computing infrastructures, such as breaking something on shared servers. [END] --- [1] Url: https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1014133 Published and (C) by PLOS One Content appears here under this condition or license: Creative Commons - Attribution BY 4.0. via Magical.Fish Gopher News Feeds: gopher://magical.fish/1/feeds/news/plosone/