[HN Gopher] Lord of the Io_uring (2020)
___________________________________________________________________
Lord of the Io_uring (2020)
Author : LAC-Tech
Score : 149 points
Date : 2025-01-06 07:11 UTC (15 hours ago)
(HTM) web link (unixism.net)
(TXT) w3m dump (unixism.net)
| jauntywundrkind wrote:
| Definitely one of the best pieces of documentation out there for
| io_uring. But I'm not sure how much if at all it's been updated
| since 2020 & Linux 5.5.
| https://web.archive.org/web/20200527021134/https://unixism.n...
| alecco wrote:
| Yeah, it should have (2020)
|
| Previous discussion
| https://news.ycombinator.com/item?id=23132549
| hinkley wrote:
| But they were all of them deceived?
| LAC-Tech wrote:
| This made me laugh a lot. I can spot Tolkien's language from a
| mile off!
| ratherbefuddled wrote:
| Technically Peter Jackson's :)
| Cthulhu_ wrote:
| And nine, nine async I/O programming APIs were gifted to the
| race of Linux users, who above all else desire power.
| api wrote:
| But this next API, we'll get it right. Let's call it
| io_uring2!
| hinkley wrote:
| I'm now imagining Torvalds riding a hell-hawk.
| friend_Fernando wrote:
| One io_uring to root them all.
| mgaunard wrote:
| a lot of the functionality was significantly improved in 6 and
| isn't reflected there.
|
| In practice io_uring can be used in many different ways, and it
| can be challenging to find the most efficient one.
| LAC-Tech wrote:
| What are the big changes in 6? links welcome.
| alecco wrote:
| https://kernelnewbies.org/Linux_6.0#io_uring_features but
| only mentions zero copy and https://lwn.net/Articles/879724/
|
| also https://www.phoronix.com/news/Linux-6.0-IO-Block-
| IO_uring
| accelbred wrote:
| I'd like to use io_uring, but as long as it bypasses seccomp it
| should be disabled whenever seccomp is in use. As such, I use
| epoll, and find it annoying when kernel APIs like ublk require
| io_uring. The places I'd want to use ublk are inside sandboxes
| using seccomp. Given that container runtimes, hardened kernels,
| chromeos, etc., disable io_uring, using it means needing an epoll
| fallback anyways, so might as well just use epoll and not
| maintain two async backends for your application.
| poincaredisk wrote:
| Is there a specific io_uring opcode you would like disabled in
| your sandboxes? It's not like io_uring is a complete seccomp
| bypass, just another syscall that provides an alternative way
| to do many things. I doubt you block "read" or "accept" in
| docker, for example. You can't execute a sysctl or mount a
| filesystem using io_uring, which are things that are actually
| blocked in Docker by default.
|
| edit: on the other hand, a good reason to disable uring in
| containers is that it's infested with vulnerabilities. It's
| new, complex, and does a whole lot of things - all of which
| make serious security bugs there quite common right now.
| accelbred wrote:
| Out of current ones, at a quick glance: connect, openat,
| openat2, renameat, mkdirat, and bind. More importantly, I'd
| like to block any opcode I haven't whitelisted, even when my
| software runs on future kernels with more opcodes available.
|
| Now that I think about it, how does io_uring interact with
| landlock?
| ibotty wrote:
| It's not only potentially infested with vulnerabilities. It's
| also not possible to filter io_uring using seccomp at all. So
| if you allow io_uring, you allow all that is possible with
| it.
| JoshTriplett wrote:
| > infested with vulnerabilities
|
| Current io_uring is not particularly prone to
| vulnerabilities. The original version of it had a design that
| often led to them (a kernel thread doing operations on behalf
| of the process and not always remembering to set the
| appropriate privileges), but it no longer uses that design,
| and the current design is much more resilient. Unfortunately,
| the original design led to a reputation that it's still
| trying to shake.
| quotemstr wrote:
| > Current io_uring is not particularly prone to
| vulnerabilities
|
| The tech industry: launch early! Develop in public! Many
| eyes make all bugs shallow!
|
| Also the tech industry: we will never forgive you for that
| one segfault you had ten years ago.
| samlightfoot wrote:
| https://github.com/containerd/containerd/issues/9048
| fulafel wrote:
| Does this mean you shouldn't use it in containers?
|
| edit: it does seem it is disabled there now:
| https://github.com/containerd/containerd/pull/9320 (thanks to
| sibling comment for an adjancent link)
| JoshTriplett wrote:
| ublk, specifically, is something I'd expect to be primarily
| used in privileged contexts anyway, because the primary use of
| the resulting block device is to mount it, which requires
| privileges for most interesting filesystems. If you want an
| unprivileged mechanism, you may be interested in the upcoming
| uring-accelerated FUSE support.
|
| For other uses, uring has a "restriction" mechanism that does
| part of what you want. See REGISTER_RESTRICTIONS in the
| documentation. Any process that's setting up its own seccomp
| restrictions can also set up a uring with restrictions,
| limiting the opcodes it can use.
|
| That said, that mechanism would benefit from a way to apply
| such restrictions to a process that isn't doing the setup
| itself, such as when setting up seccomp restrictions on a
| container or daemon. For instance, a way to set restrictions on
| all rings created by child processes, or a way for seccomp to
| enforce that any uring created has restrictions applied to it.
| quotemstr wrote:
| > For instance, a way to set restrictions on all rings
| created by child processes, or a way for seccomp to enforce
| that any uring created has restrictions applied to it.
|
| SELinux or your favorite MAC is there to solve this exact
| problem.
| haberman wrote:
| > you may be interested in the upcoming uring-accelerated
| FUSE support.
|
| Do you have a reference for this? What is the anticipated
| timeframe?
| JoshTriplett wrote:
| https://lore.kernel.org/io-uring/20241209-fuse-uring-
| for-6-1...
|
| I don't know when it'll be merged, but it seems like it's
| getting close to ready.
| accelbred wrote:
| The main problem I have with fuse is inotify not working. If
| inotify just worked for fuse, I'd just use it. Ideally I
| could just run the software in a mount namespace with a fuse
| fs, but I need inotify.
|
| I mainly was trying to use ublk to implement a sort of fuse
| like thing with the kernel handling the fs and thus having
| inotify support.
| quotemstr wrote:
| > find it annoying when kernel APIs like ublk require io_uring
|
| Good. That's a forcing function for making io_uring work in
| your environment.
|
| > bypasses seccomp
|
| Seccomp sucks.
|
| We shouldn't be enforcing security by filtering system calls,
| the set of which will grow forever, but instead by describing
| access control rules on objects, e.g. with SELinux. If your
| security policy is that your sandbox should be able to read
| from some file but not write to it, you should do that with
| real MAC, which applies to all operations , il_uring included.
| You shouldn't just filter read(2) and write(2) in particular.
|
| We shouldn't hold back evolution in systems interfaces because
| some people are stuck on bad ways of doing things and won't
| move.
| accelbred wrote:
| Since when can you use a MAC as an unprivileged user on an
| arbitrary distro?
| t00 wrote:
| There are examples of cat and cp using io_uring. What are the
| chances of having io_uring utilised by standard commands to
| improve overall Linux performance? I presume GNU utils are not
| Linux specific hence such commands are programmed for a generic
| *nix.
|
| Another one is I could not find a benchmark with io_uring - this
| would confirm the benefit of going from epoll.
| fweimer wrote:
| GNU coreutils already has tons of Linux-specific code. But it
| would be a bit of a kernel fail if io_uring were faster or
| other preferable to copy_file_range for cp (at least for files
| that do not have holes).
| Sesse__ wrote:
| Not at all; with io_uring, you can copy multiple files in
| parallel (and in fewer syscalls), which is a huge win for
| small files.
| mahkoh wrote:
| >Another one is I could not find a benchmark with io_uring -
| this would confirm the benefit of going from epoll.
|
| One of the advantages of io_uring, unrelated to performance, is
| that it supports non-blocking operations on blocking file
| descriptors.
|
| Using io_uring is the only method I recall to bypass
| https://gitlab.freedesktop.org/wayland/wayland/-/issues/296.
| This issue deals with having to operate on untrusted file
| descriptors where the blocking/non-blocking state of the file
| descriptions might be manipulated by an adversary at any time.
| lukeh wrote:
| Also useful for things like SPI with only blocking user space
| API.
| samsquire wrote:
| This document helped me learn the io_uring API.
|
| You can use io_uring with epoll to monitor eventfd to wake up
| your sleeping with io_uring wait for completions.
|
| I have implemented a barrier and thread safe techniques that I am
| trying to turn into a command line tool
|
| My goal is that thread safe performant servers are easy to write.
|
| I am using bloom filters for fast set intersection. I intend to
| use Simd instructions with the bloom hashes.
| Thaxll wrote:
| Someone can comment on the security implications of sharing a
| buffer between user space and kernel space?
| fragmede wrote:
| As you suspect, it's not awesome.
|
| https://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=io_uring
| alexgartrell wrote:
| Sharing a queue itself is not new
| https://www.kernel.org/doc/html/v5.8/networking/packet_mmap....
| and https://docs.kernel.org/next/userspace-
| api/perf_ring_buffer.... are two examples.
|
| Issues with io_uring security mostly stemmed from an old
| architecture and just the fact that there's a ton of surface
| area.
| quotemstr wrote:
| binder shares a buffer between kernel and user space on
| billions of Android devices, and Android is by far the most
| secure Linux distribution.
|
| There's nothing wrong with the general concept.
___________________________________________________________________
(page generated 2025-01-06 23:01 UTC)