[HN Gopher] Lord of the Io_uring (2020)
       ___________________________________________________________________
        
       Lord of the Io_uring (2020)
        
       Author : LAC-Tech
       Score  : 149 points
       Date   : 2025-01-06 07:11 UTC (15 hours ago)
        
 (HTM) web link (unixism.net)
 (TXT) w3m dump (unixism.net)
        
       | jauntywundrkind wrote:
       | Definitely one of the best pieces of documentation out there for
       | io_uring. But I'm not sure how much if at all it's been updated
       | since 2020 & Linux 5.5.
       | https://web.archive.org/web/20200527021134/https://unixism.n...
        
         | alecco wrote:
         | Yeah, it should have (2020)
         | 
         | Previous discussion
         | https://news.ycombinator.com/item?id=23132549
        
       | hinkley wrote:
       | But they were all of them deceived?
        
         | LAC-Tech wrote:
         | This made me laugh a lot. I can spot Tolkien's language from a
         | mile off!
        
           | ratherbefuddled wrote:
           | Technically Peter Jackson's :)
        
         | Cthulhu_ wrote:
         | And nine, nine async I/O programming APIs were gifted to the
         | race of Linux users, who above all else desire power.
        
           | api wrote:
           | But this next API, we'll get it right. Let's call it
           | io_uring2!
        
           | hinkley wrote:
           | I'm now imagining Torvalds riding a hell-hawk.
        
         | friend_Fernando wrote:
         | One io_uring to root them all.
        
       | mgaunard wrote:
       | a lot of the functionality was significantly improved in 6 and
       | isn't reflected there.
       | 
       | In practice io_uring can be used in many different ways, and it
       | can be challenging to find the most efficient one.
        
         | LAC-Tech wrote:
         | What are the big changes in 6? links welcome.
        
           | alecco wrote:
           | https://kernelnewbies.org/Linux_6.0#io_uring_features but
           | only mentions zero copy and https://lwn.net/Articles/879724/
           | 
           | also https://www.phoronix.com/news/Linux-6.0-IO-Block-
           | IO_uring
        
       | accelbred wrote:
       | I'd like to use io_uring, but as long as it bypasses seccomp it
       | should be disabled whenever seccomp is in use. As such, I use
       | epoll, and find it annoying when kernel APIs like ublk require
       | io_uring. The places I'd want to use ublk are inside sandboxes
       | using seccomp. Given that container runtimes, hardened kernels,
       | chromeos, etc., disable io_uring, using it means needing an epoll
       | fallback anyways, so might as well just use epoll and not
       | maintain two async backends for your application.
        
         | poincaredisk wrote:
         | Is there a specific io_uring opcode you would like disabled in
         | your sandboxes? It's not like io_uring is a complete seccomp
         | bypass, just another syscall that provides an alternative way
         | to do many things. I doubt you block "read" or "accept" in
         | docker, for example. You can't execute a sysctl or mount a
         | filesystem using io_uring, which are things that are actually
         | blocked in Docker by default.
         | 
         | edit: on the other hand, a good reason to disable uring in
         | containers is that it's infested with vulnerabilities. It's
         | new, complex, and does a whole lot of things - all of which
         | make serious security bugs there quite common right now.
        
           | accelbred wrote:
           | Out of current ones, at a quick glance: connect, openat,
           | openat2, renameat, mkdirat, and bind. More importantly, I'd
           | like to block any opcode I haven't whitelisted, even when my
           | software runs on future kernels with more opcodes available.
           | 
           | Now that I think about it, how does io_uring interact with
           | landlock?
        
           | ibotty wrote:
           | It's not only potentially infested with vulnerabilities. It's
           | also not possible to filter io_uring using seccomp at all. So
           | if you allow io_uring, you allow all that is possible with
           | it.
        
           | JoshTriplett wrote:
           | > infested with vulnerabilities
           | 
           | Current io_uring is not particularly prone to
           | vulnerabilities. The original version of it had a design that
           | often led to them (a kernel thread doing operations on behalf
           | of the process and not always remembering to set the
           | appropriate privileges), but it no longer uses that design,
           | and the current design is much more resilient. Unfortunately,
           | the original design led to a reputation that it's still
           | trying to shake.
        
             | quotemstr wrote:
             | > Current io_uring is not particularly prone to
             | vulnerabilities
             | 
             | The tech industry: launch early! Develop in public! Many
             | eyes make all bugs shallow!
             | 
             | Also the tech industry: we will never forgive you for that
             | one segfault you had ten years ago.
        
         | samlightfoot wrote:
         | https://github.com/containerd/containerd/issues/9048
        
         | fulafel wrote:
         | Does this mean you shouldn't use it in containers?
         | 
         | edit: it does seem it is disabled there now:
         | https://github.com/containerd/containerd/pull/9320 (thanks to
         | sibling comment for an adjancent link)
        
         | JoshTriplett wrote:
         | ublk, specifically, is something I'd expect to be primarily
         | used in privileged contexts anyway, because the primary use of
         | the resulting block device is to mount it, which requires
         | privileges for most interesting filesystems. If you want an
         | unprivileged mechanism, you may be interested in the upcoming
         | uring-accelerated FUSE support.
         | 
         | For other uses, uring has a "restriction" mechanism that does
         | part of what you want. See REGISTER_RESTRICTIONS in the
         | documentation. Any process that's setting up its own seccomp
         | restrictions can also set up a uring with restrictions,
         | limiting the opcodes it can use.
         | 
         | That said, that mechanism would benefit from a way to apply
         | such restrictions to a process that isn't doing the setup
         | itself, such as when setting up seccomp restrictions on a
         | container or daemon. For instance, a way to set restrictions on
         | all rings created by child processes, or a way for seccomp to
         | enforce that any uring created has restrictions applied to it.
        
           | quotemstr wrote:
           | > For instance, a way to set restrictions on all rings
           | created by child processes, or a way for seccomp to enforce
           | that any uring created has restrictions applied to it.
           | 
           | SELinux or your favorite MAC is there to solve this exact
           | problem.
        
           | haberman wrote:
           | > you may be interested in the upcoming uring-accelerated
           | FUSE support.
           | 
           | Do you have a reference for this? What is the anticipated
           | timeframe?
        
             | JoshTriplett wrote:
             | https://lore.kernel.org/io-uring/20241209-fuse-uring-
             | for-6-1...
             | 
             | I don't know when it'll be merged, but it seems like it's
             | getting close to ready.
        
           | accelbred wrote:
           | The main problem I have with fuse is inotify not working. If
           | inotify just worked for fuse, I'd just use it. Ideally I
           | could just run the software in a mount namespace with a fuse
           | fs, but I need inotify.
           | 
           | I mainly was trying to use ublk to implement a sort of fuse
           | like thing with the kernel handling the fs and thus having
           | inotify support.
        
         | quotemstr wrote:
         | > find it annoying when kernel APIs like ublk require io_uring
         | 
         | Good. That's a forcing function for making io_uring work in
         | your environment.
         | 
         | > bypasses seccomp
         | 
         | Seccomp sucks.
         | 
         | We shouldn't be enforcing security by filtering system calls,
         | the set of which will grow forever, but instead by describing
         | access control rules on objects, e.g. with SELinux. If your
         | security policy is that your sandbox should be able to read
         | from some file but not write to it, you should do that with
         | real MAC, which applies to all operations , il_uring included.
         | You shouldn't just filter read(2) and write(2) in particular.
         | 
         | We shouldn't hold back evolution in systems interfaces because
         | some people are stuck on bad ways of doing things and won't
         | move.
        
           | accelbred wrote:
           | Since when can you use a MAC as an unprivileged user on an
           | arbitrary distro?
        
       | t00 wrote:
       | There are examples of cat and cp using io_uring. What are the
       | chances of having io_uring utilised by standard commands to
       | improve overall Linux performance? I presume GNU utils are not
       | Linux specific hence such commands are programmed for a generic
       | *nix.
       | 
       | Another one is I could not find a benchmark with io_uring - this
       | would confirm the benefit of going from epoll.
        
         | fweimer wrote:
         | GNU coreutils already has tons of Linux-specific code. But it
         | would be a bit of a kernel fail if io_uring were faster or
         | other preferable to copy_file_range for cp (at least for files
         | that do not have holes).
        
           | Sesse__ wrote:
           | Not at all; with io_uring, you can copy multiple files in
           | parallel (and in fewer syscalls), which is a huge win for
           | small files.
        
         | mahkoh wrote:
         | >Another one is I could not find a benchmark with io_uring -
         | this would confirm the benefit of going from epoll.
         | 
         | One of the advantages of io_uring, unrelated to performance, is
         | that it supports non-blocking operations on blocking file
         | descriptors.
         | 
         | Using io_uring is the only method I recall to bypass
         | https://gitlab.freedesktop.org/wayland/wayland/-/issues/296.
         | This issue deals with having to operate on untrusted file
         | descriptors where the blocking/non-blocking state of the file
         | descriptions might be manipulated by an adversary at any time.
        
           | lukeh wrote:
           | Also useful for things like SPI with only blocking user space
           | API.
        
       | samsquire wrote:
       | This document helped me learn the io_uring API.
       | 
       | You can use io_uring with epoll to monitor eventfd to wake up
       | your sleeping with io_uring wait for completions.
       | 
       | I have implemented a barrier and thread safe techniques that I am
       | trying to turn into a command line tool
       | 
       | My goal is that thread safe performant servers are easy to write.
       | 
       | I am using bloom filters for fast set intersection. I intend to
       | use Simd instructions with the bloom hashes.
        
       | Thaxll wrote:
       | Someone can comment on the security implications of sharing a
       | buffer between user space and kernel space?
        
         | fragmede wrote:
         | As you suspect, it's not awesome.
         | 
         | https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=io_uring
        
         | alexgartrell wrote:
         | Sharing a queue itself is not new
         | https://www.kernel.org/doc/html/v5.8/networking/packet_mmap....
         | and https://docs.kernel.org/next/userspace-
         | api/perf_ring_buffer.... are two examples.
         | 
         | Issues with io_uring security mostly stemmed from an old
         | architecture and just the fact that there's a ton of surface
         | area.
        
         | quotemstr wrote:
         | binder shares a buffer between kernel and user space on
         | billions of Android devices, and Android is by far the most
         | secure Linux distribution.
         | 
         | There's nothing wrong with the general concept.
        
       ___________________________________________________________________
       (page generated 2025-01-06 23:01 UTC)