[HN Gopher] Things Unix can do atomically (2010)
       ___________________________________________________________________
        
       Things Unix can do atomically (2010)
        
       Author : onurkanbkrc
       Score  : 240 points
       Date   : 2026-02-06 05:29 UTC (17 hours ago)
        
 (HTM) web link (rcrowley.org)
 (TXT) w3m dump (rcrowley.org)
        
       | 0xbadcafebee wrote:
       | You can use `ln` atomicity for a simple, portable(ish) locking
       | system: https://gist.github.com/pwillis-
       | els/b01b22f1b967a228c31db3cf...
        
         | akoboldfrying wrote:
         | Really nice explanation of a useful pattern. I was surprised to
         | discover that even the famously broken NFS honours atomicity of
         | hardlink creation.
        
       | exac wrote:
       | Sorry, there is zero chance I will ever deploy new code by
       | changing a symlink to point to the new directory.
        
         | sholladay wrote:
         | Why? What do you prefer to do instead?
        
           | gib444 wrote:
           | Anything less than an entire new k8s cluster and switching
           | over is just amateur hour obviously
        
         | iberator wrote:
         | why? it works and its super clever. Simple command instead some
         | shit written in JS with docker trash
        
           | lloeki wrote:
           | Ah, the memories of capistrano, complete with zero-downtime
           | unicorn handover
           | 
           | https://github.com/capistrano/capistrano/
        
             | 10us wrote:
             | Still use php deployer each day and works with symlinks as
             | well. https://deployer.org/
        
         | alpb wrote:
         | Nobody's saying you should deploy code with this, but symlinks
         | are a very common filesystem locking method.
        
         | slopusila wrote:
         | that's how some phone OSes update the system (by having 2 read
         | only fs)
         | 
         | that's how Chrome updates itself, but without the symlink part
        
           | x4132 wrote:
           | not surprised about the chrome part, but pretty shocked at
           | the phone OS part. I know APFS migration was done in this
           | way, but wouldn't storage considerations for this be massive?
        
             | slopusila wrote:
             | what would be more massive would be phones not booting up
             | because of a botched update. this way you can just switch
             | back to the old partition
        
             | marmarama wrote:
             | Not really, because only the OS core is swapped in this
             | way. Apps and data live in their own partitions/subvolumes,
             | which are mutable and shared between OS versions.
             | 
             | The OS core is deployed as a single unit and is a few GB in
             | size, pretty small when internal storage is into the
             | hundreds of GB.
        
           | dizhn wrote:
           | No snapshotting at all? Thinking about it.. The filesystem
           | does not support it I suppose.
        
             | LiamPowell wrote:
             | Android does use snapshots:
             | https://source.android.com/docs/core/ota/virtual_ab
        
               | dizhn wrote:
               | Oh cool. I was a bit confused about not using snapshots
               | and relying on symlinks but it couldn't be so simple. I
               | guess it's just a simple userspace cow mount. https://sou
               | rce.android.com/docs/core/ota/virtual_ab#compress...
        
         | bandrami wrote:
         | Works pretty well for Nix
        
           | mananaysiempre wrote:
           | And for Stow[1] before it, and for its inspiration Depot[2]
           | before even that. It's an old idea.
           | 
           | [1] https://www.gnu.org/software/stow/
           | 
           | [2] http://ftp.gregor.com/download/dgregor/depot.pdf
        
             | bandrami wrote:
             | I really liked stow. My toy distro back in the day was
             | based on it.
        
           | atmosx wrote:
           | Worked pretty well in production systems, serving huge amount
           | of RPS (like ~5-10k/s) running on a LAMP stack monolith in
           | five different geographical regions.
           | 
           | Just git branch (one branch per region because of compliance
           | requirements) -> branch creates "tar.gz" with predefined name
           | -> automated system downloads the new "tar.gz", checks
           | release date, revision, etc. -> new symlink & php
           | (serverles!!!) graceful restart and ka-b00m.
           | 
           | Rollbacks worked by pointing back to the old dir & restart.
           | 
           | Worked like a charm :-)
        
         | gonzus wrote:
         | Then you are locking yourself out of a pretty much ironclad
         | (and extremely cost-effective) way of managing such things.
        
         | 1718627440 wrote:
         | Isn't that the standard way to do that? Why wouldn't you?
        
         | silisili wrote:
         | I don't do devops/sysadmin anymore, so this would have been
         | before the age of k8s for everything. But I once interviewed
         | for a company hiring specifically because their deployment
         | process lasted hours, and rollbacks even longer.
         | 
         | In the interview when they were describing this problem, I
         | asked why the didn't just put all of the new release in a new
         | dir, and use symlinks to roll forward and backwards as needed.
         | They kind of froze and looked at each other and all had the
         | same 'aha' moment. I ended up not being interested in taking
         | the job, but they still made sure to thank me for the idea
         | which I thought was nice.
         | 
         | Not that I'm a genius or anything, it's something I'd done
         | previously for years, and I'm sure I learned it from someone
         | else who'd been doing it for years. It's a very valid
         | deployment mechanism IMO, of course depending on your
         | architecture.
        
       | sega_sai wrote:
       | rename() is certainly the easiest to use for any sort of file-
       | system based synchronization.
        
         | compressedgas wrote:
         | As long as you don't run into or want freedom from possible
         | path races, for that you need the missing:
         | frenameat2(srcdirfd, srcfd, srcname, dstdirfd, dstfd, dstname)
        
       | MintPaw wrote:
       | Not much apparently, although I didn't know about changing
       | symlinks, that could be very useful.
        
       | ta8903 wrote:
       | Not technically related to atomicity, but I was looking for a way
       | to do arbitrary filesystem operations based on some condition
       | (like adding a file to a directory, and having some operation be
       | performed on it). The usual recommendation for this is to use
       | inotify/watchman, but something about it seems clunky to me. I
       | want to write a virtual filesystem, where you pass it a trigger
       | condition and a function, and it applies the function to all
       | files based on the trigger condition. Does something like this
       | exist?
        
         | Brian_K_White wrote:
         | incron
        
           | ta8903 wrote:
           | Thanks, I didn't find this when I was looking for a solution
           | for my problem. This is pretty much the exact solution for my
           | usecase, though for some reason inotify feels more
           | complicated than some kind of filesystem mount solution for
           | me.
        
         | direwolf20 wrote:
         | are you asking for if statements?
         | 
         | if(condition) {do the thing;}
        
           | ta8903 wrote:
           | I know this is trivial to do programmatically, but I was
           | looking for a way this will be handled by the filesystem. For
           | instance, if I have some processes generating log files, and
           | I have a script that converts them to html, I wanted the
           | script to be called every time a log file is updated, without
           | having a daemon running in the background to monitor the
           | directory, just some filesystem mount. This would have made
           | some deployments easier.
        
         | laz wrote:
         | Sounds half baked. What context does this function run in? Is
         | it an interpreted language or an executable that you provide?
         | 
         | Inotify is the way to shovel these events out of the kernel,
         | then userspace process rules apply. It's maybe not elegant from
         | your pov, but it's simple.
        
         | quesera wrote:
         | I've used FUSE for something similar.
         | 
         | There are sample "drivers" in easily-modified python that are
         | fast enough for casual use.
        
         | zbentley wrote:
         | The challenge with that approach is memory: trigger conditions,
         | if added irresponsibly, can result in unbounded memory and
         | (depending on implementation) potentially linear performance
         | degradation of filesystem operations as well. Unbounded kernel
         | memory growth leads to stability or security risks.
         | 
         | That tradeoff is at the root of why most notify APIs are either
         | approximate (events can be dropped) or rigidly bounded by
         | kernel settings that prevent truly arbitrary numbers of
         | watches. fanotify and some implementations of kqueue are better
         | at efficiently triggering large recursive watches, but that's
         | still just a mitigation on the underlying memory/performance
         | tradeoffs, not a full solution.
        
       | zzo38computer wrote:
       | Even though it can do some things atomically, it only does with
       | one file at a time, and race conditions are still possible
       | because it only does one operation at a time (even if you are
       | only need one file). Some of these are helpful anyways, such as
       | O_EXCL, but it is still only one thing at a time which can cause
       | problems in some cases.
       | 
       | What else it does not do is a transaction with multiple objects.
       | That is why, I would design a operating system, that you can do a
       | transaction with multiple objects.
        
         | ptx wrote:
         | Windows had APIs for this sort of thing added in Vista, but
         | they're now deprecating it "due to its complexity and various
         | nuances which developers need to consider":
         | 
         | https://learn.microsoft.com/en-us/windows/win32/fileio/about...
        
         | akoboldfrying wrote:
         | I don't follow, sorry. Are you saying that if we run:
         | mv a b         mv c d
         | 
         | We could observe a state where a and d exist? I would find such
         | "out of order execution" shocking.
         | 
         | If that's not what you're saying, could you give an example of
         | something you want to be able to do but can't?
        
           | jstimpfle wrote:
           | I don't think that's happening in practice, but 1) it may not
           | be specified and 2) What you say could well be the persisted
           | state after a machine crash or power loss. In particular if
           | those files live in different directories.
           | 
           | You can remedy 2) by doing fsync() on the parent directory in
           | between. I just asked ChatGPT which directory you need to
           | fsync. It says it's both, the source and the target
           | directory. Which "makes sense" and simplifies
           | implementations, but it means the rename operation is atomic
           | only at runtime, not if there's a crash in between. It think
           | you might end up with 0 or 2 entries after a crash if you're
           | unlucky.
           | 
           | If that's true, then for safety maybe one should never rename
           | across directories, but instead do a coordinated link(source,
           | target), fsync(target_dir), unlink(source), fsync(source_dir)
        
             | jstimpfle wrote:
             | why is this being downvoted? If there's something wrong,
             | explain?
        
           | devnonymous wrote:
           | I'm almost certain what the OP meant was if the commands were
           | run synchronously (ie: from 2 different shells or as `mv a b
           | &; mv c d`) yes there is a possibility that a and d exist
           | (eg: On a busy system where neither of the 2 commands can be
           | immediately scheduled and eventually the second one ends up
           | being scheduled before the first)
           | 
           | Or to go a level deeper, if you have 2 occurrences of
           | rename(2) from the stdlibc ...
           | 
           | rename('a', 'b'); rename('c', 'd');
           | 
           | ...and the compiler decides on out of order execution or
           | optimizing by scheduling on different cpus, you can get a and
           | d existing at the same time.
           | 
           | The reason it won't happen in the example you posted is the
           | shell ensures the atomicity (by not forking the second mv
           | until the wait() on the first returns)
        
             | isodude wrote:
             | nitpick, it should be `touch a c & mv a b & mv c d` as `&;`
             | returns `bash: syntax error near unexpected token `;'`. I
             | always find this oddly weird, but that would not be the
             | first pattern in BASH that is.
             | 
             | `inotifywait` actually sees them in order, but nothing
             | ensure that it's that way.                 $ inotifywait -m
             | /tmp       /tmp/ MOVED_FROM a       /tmp/ MOVED_TO b
             | /tmp/ MOVED_FROM c       /tmp/ MOVED_TO d
             | 
             | `stat` tells us that the timestamps are equal as well.
             | $ stat b d | grep '^Change'       Change: 2026-02-06
             | 12:22:55.394932841 +0100       Change: 2026-02-06
             | 12:22:55.394932841 +0100
             | 
             | However, speeding things up changes it a bit.
             | 
             | Given                 $ (         set -eo pipefail
             | for i in {1..10000}         do           printf '%d ' "$i"
             | touch a c           mv a b &           mv c d &
             | wait           rm b d         done       )       1 2 3 4 5
             | 6 .....
             | 
             | And with `inotifywait` I saw this when running it for a
             | while.                 $ inotifywait -m -e
             | MOVED_FROM,MOVED_TO /tmp > /tmp/output       cat
             | /tmp/output | xargs -l4 | sort | uniq -c       9104 /tmp/
             | MOVED_FROM a /tmp/ MOVED_TO b /tmp/ MOVED_FROM c /tmp/
             | MOVED_TO d       896 /tmp/ MOVED_FROM c /tmp/ MOVED_TO d
             | /tmp/ MOVED_FROM a /tmp/ MOVED_TO b
        
           | duped wrote:
           | All you need for this to occur is the window where both
           | renames occurs overlap. A system polling to check if a, b, c,
           | and d exist while the renames are happening might find all
           | four of them.
        
             | jstimpfle wrote:
             | Assuming that the two `mv` commands are run in sequence,
             | there shouldn't be any possibility for a and d to be
             | observed "at once" (i.e. first d and then afterwards still
             | a, by a single process).
        
           | zbentley wrote:
           | Depending on metadata cache behavior configuration, if the
           | system is powered off immediately after the first command,
           | then that could indeed happen I think.
           | 
           | As to whether it's technically possible for it to happen on a
           | system that stays on, I'm not sure, but it's certainly
           | vanishingly rare and likely requires very specific
           | circumstances--not just a random race condition.
        
             | LgWoodenBadger wrote:
             | Uhh, if the system powers off immediately after the first
             | command (mv a b), the second command (mv c d) would never
             | run. So where would d come from if the command that created
             | it never executed?
        
               | zbentley wrote:
               | Er, sorry: I meant: if the first command runs, the plug
               | is pulled, system starts again, second command runs.
        
               | lpribis wrote:
               | Sure, but splitting "atomic" operations across a reboot
               | is an interesting design choice. Surely upon reboot you
               | would re-try the first `mv a b` before doing other
               | things.
        
         | Orphis wrote:
         | In some cases, you can start by using the "at" functions
         | (openat...) to work on a directory tree. If you have your
         | logical "locking" done at the top-level of the tree, it might
         | be a fine option.
         | 
         | In some other cases, I've used a pattern where I used a symlink
         | to folders. The symlink is created, resolved or updated
         | atomically, and all I need is eventual consistency.
         | 
         | That last case was to manage several APT repository indices.
         | The indices were constantly updated to publish new testing or
         | unstable releases of software and machines in the fleet were
         | regularly fetching the repository index. The APT protocol and
         | structure being a bit "dumb" (for better or worse) requires you
         | to fetch files (many of them) in the reverse order they are
         | created, which leads to obvious issues like the signature is
         | updated only after the list of files is updated, or the list of
         | files is created only after the list of packages is created.
         | 
         | Long story short, each update would create a new folder that's
         | consistent, and a symlink points to the last created folder (to
         | atomically replace the folder as it was not possible to swap
         | them), and a small HTTP server would initiate a server side
         | session when the first file is fetched and only return files
         | from the same index list, and everything is eventually
         | consistent, and we never get APT complaining about having
         | signature or hash mismatches. The pivotal component was indeed
         | the atomicity of having a symlink to deal with it, as the Java
         | implementation didn't have access to a more modern "openat"
         | syscall, relative to a specific folder.
        
       | amstan wrote:
       | Missing (probably because of the date of the article): `mv
       | --exchange` aka renameat2+RENAME_EXCHANGE. It atomically swaps 2
       | file paths.
        
         | oguz-ismail2 wrote:
         | Title says Unix, renameat2 is Linux-only.
        
           | jasode wrote:
           | _> Title says Unix,_
           | 
           | You're misinterpreting the title. The author didn't intend
           | "Unix" to literally mean only the official _AT
           | &T/TheOpenGroup UNIX(r) System_ to the exclusion of Linux.
           | 
           | The first sentence of "UNIX-like" makes that clear : _> This
           | is a catalog of things UNIX-like/POSIX-compliant operating
           | systems can do atomically, _
           | 
           | Further down, he then mentions some Linux specifics : _>
           | fcntl(fd, F_GETLK, &lock), fcntl(fd, F_SETLK, &lock), and
           | fcntl(fd, F_SETLKW, &lock) . [...] There is a "mandatory
           | locking" mode but Linux's implementation is unreliable as
           | it's subject to a race condition._
        
             | monibious wrote:
             | But I also don't think the auther meant Things you can do
             | in Linux but not Unix
        
               | jasode wrote:
               | _> But I also don't think the auther meant Things you can
               | do in Linux but not Unix_
               | 
               | I wasn't claiming that. I just thought the ggp had a
               | useful comment about renameat2() which led to gp's
               | "correction" which wasn't 100% accurate.
               | 
               | IBM z/OS UNIX also has renameat2(). It doesn't have the
               | Linux specific flag RENAME_EXCHANGE.
               | 
               | https://www.ibm.com/docs/en/zos/3.1.0?topic=functions-
               | rename...
        
               | mghackerlady wrote:
               | pedantic but z/OS isn't a unix, it can just pretend to be
               | one enough for the open group to call it one. IBM has a
               | unix still, AIX.
        
               | skissane wrote:
               | In recent versions, z/OS has been copying lots of Linux-
               | specific APIs (e.g. unshare [0]) in order to support the
               | z/OS port of Kubernetes.
               | 
               | If Kubernetes starts using renameat2(RENAME_EXCHANGE),
               | they could very plausibly add it.
               | 
               | [0] https://www.ibm.com/docs/en/zos/3.2.0?topic=csd-
               | unshare-bpx1...
        
             | shawn_w wrote:
             | Bit rot alert: Linux doesn't even have mandatory file locks
             | these days.
             | 
             | Linux-specific open file description locks could be brought
             | up in a modern version of TFA though.
        
             | stephenr wrote:
             | Sounds like the key term then is probably this:
             | 
             | > POSIX-compliant
             | 
             | Which, FWIW, doesn't mean Linux. AFAIK there is _no_ Linux
             | distro that 's fully compliant, even before you worry about
             | the specifics of whether it's _certified_ as compliant.
        
               | rascul wrote:
               | EulerOS was certified UNIX some years ago.
        
               | stephenr wrote:
               | Huh, TIL. Thanks.
        
               | dietr1ch wrote:
               | AFAIK you don't even want to be POSIX-compliant unless
               | having a sticker means more to you than being reasonable.
               | Most projects knowingly steer away from compliance (and
               | certifying compliance is probably also expensive)
        
               | jasode wrote:
               | _> POSIX-compliant Which, FWIW, doesn't mean Linux. AFAIK
               | there is no Linux distro that's fully compliant_
               | 
               | I read author's use of "POSIX-compliant" as a _loose and
               | fuzzy family category_ rather than an exhaustive and
               | authoritative reference on 100% strict compliance.
               | Therefore, the author mentioning non-100%-compliant Linux
               | is ok.
               | 
               | There seems to be 2 different expectations and
               | interpretations of what the article is about.
               | 
               | - (1) article is attempting to be a strict _intersection_
               | of all Unix-like systems that conform to official UNIX
               | POSIX API. I didn 't think this was a reasonable
               | interpretation since we can't be sure the author actually
               | verified/tested other POSIX-like systems such as FreeBSD,
               | HP-UX, IBM AIX, etc.
               | 
               | - (2) article is a looser _union_ of operating systems
               | and can _also include idiosyncracies of certain systems
               | like Linux that the author is familiar with_ that don 't
               | apply to all other UNIX systems. I think some readers
               | don't realize that all the author's citations to man
               | pages point to _Linux_ specific urls at :
               | https://linux.die.net/man/
               | 
               | The ggp's (amstan) additional comment about
               | renameat2(,,,,RENAME_EXCHANGE) is useful info and is
               | consistent with interpretation (2).
               | 
               | If the author really didn't want Linux to be lumped in
               | with "POSIX-like", it seems he would avoid linux.die.net
               | and instead point to something more of a UNIX standard
               | such as: https://unix.org/apis.html
               | 
               | [0] Intersection vs Union: https://en.wikipedia.org/wiki/
               | Set_(mathematics)#Intersection
        
               | mionhe wrote:
               | The slash is read as "OR" in this case.
               | 
               | As in: Unix-like OR POSIX-compliant
               | 
               | In that light, it's probably fine to not nitpick over
               | certifications here.
        
             | pjmlp wrote:
             | Except POSIX doesn't specify some of them as happening
             | atomically.
             | 
             | Many people write UNIX/POSIX without ever reading what it
             | says.
        
             | bee_rider wrote:
             | They aren't misinterpreting the title, the title is
             | incorrect.
        
               | jasode wrote:
               | _> , the title is incorrect._
               | 
               | Differing philosophies of how to interpret titles.
               | Prescriptive vs Descriptive language.[0]
               | 
               | There can be different usages of the word _" Unix"_:
               | 
               | #1: Unix is a UNIX(tm) System V descendent. More emphasis
               | that the kernel needs to be UNIX. In this strict
               | definition, you get the common reminder that _" Linux is
               | not a Unix!"_
               | 
               | #2: "Unix" as a loose _generic_ term for a family of o /s
               | that looks/feels like Unix. This perspective includes
               | using an o/s that has userland Unix utilities like
               | cat/grep/awk. Sometimes deliberately styled as asterisk
               | _" *nix"_ or a suffix-qualifier _" Unix-like"_ but often
               | just written as a naked _" Unix"_.
               | 
               | A _Prescriptivist_ says the author 's title is
               | "incorrect". On the other hand, a _Descriptivist_ looks
               | at the whole content of the article -- notices the text
               | has a lot of Linux specific info such as
               | fcntl(,F_GETLEASE /F_SETLEASE), and every hyperlink to a
               | man page url points to https://linux.die.net/man/ , etc
               | -- and thus determines that the author is using
               | "Unix"(#2) in the looser way that can include some Linux
               | idiosyncrasies.
               | 
               | "Unix" instead of "*nix" as a generic term for Linux is
               | not uncommon. Another example article where the authors
               | use the so-called incorrect "Unix" in the title even
               | though it's mostly discussing Linux CUPS instead of
               | Solaris :
               | https://www.evilsocket.net/2024/09/26/Attacking-UNIX-
               | systems...
               | 
               | [0] https://en.wikipedia.org/wiki/Linguistic_prescription
        
         | rustybolt wrote:
         | I tried using this a while back and found it was not widely
         | available. You need coreutils version 9.1 or later for this,
         | many distros do not ship this.
         | 
         | I made https://github.com/rubenvannieuwpoort/atomic-exchange
         | for my usecase.
        
       | klempner wrote:
       | This document being from 2010 is, of course, missing the
       | C11/C++11 atomics that replaced the need for compiler intrinsics
       | or non portable inline asm when "operating on virtual memory".
       | 
       | With that said, at least for C and C++, the behavior of
       | (std::)atomic when dealing with interprocess interactions is
       | slightly outside the scope of the standard, but in practice (and
       | at least recommended by the C++ standard) (atomic_)is_lock_free()
       | atomics are generally usable between processes.
        
         | senderista wrote:
         | That's right, atomic operations work just fine for memory
         | shared between processes. I have worked on a commercial product
         | that used this everywhere.
        
       | Igrom wrote:
       | >fcntl(fd, F_GETLK, &lock), fcntl(fd, F_SETLK, &lock), and
       | fcntl(fd, F_SETLKW, &lock)
       | 
       | There's also `flock`, the CLI utility in util-linux, that allows
       | using flocks in shell scripts.
        
         | cachius wrote:
         | What are flocks in this context? Surely not a number of
         | sheep...
        
           | ncruces wrote:
           | File locks.
        
           | gbacon wrote:
           | https://man.openbsd.org/flock.2
           | 
           | https://man7.org/linux/man-pages/man2/flock.2.html
        
         | pjmlp wrote:
         | In UNIX/POSIX file locks are advisory, not enforced, it only
         | works if all processes play ball.
        
           | zbentley wrote:
           | Sure, but the discussion is around whether they're atomic,
           | not whether they're advisory.
        
         | zbentley wrote:
         | Aren't flock and POSIX locks backed by totally different
         | systems?
        
       | andrewstuart wrote:
       | Anywhere there is atomic capability you can build a queuing
       | application.
        
       | ncruces wrote:
       | I use several of these to implement alternative SQLite locking
       | protocols.
       | 
       | POSIX file locking semantics really are broken beyond repair:
       | https://news.ycombinator.com/item?id=46542247
        
       | pjmlp wrote:
       | Unless they can be guaranteed by the POSIX specification, they
       | are implementation specific and should not be relied upon for
       | portable code.
        
         | kccqzy wrote:
         | Which of these are not guaranteed by the POSIX specification?
         | It's been a while since I studied it, but if I recall correctly
         | the ones mentioned in the article are guaranteed.
        
       | jeffbee wrote:
       | I wonder why the author left out atomic writes with O_APPEND.
        
         | zbentley wrote:
         | Unsure. Aren't there filesystems which make O_APPEND less
         | durable than it's specified to be, which might be interpreted
         | to adversely affect atomicity? Could that be it?
        
         | ozgrakkurt wrote:
         | This requires O_SYNC and O_DIRECT afaik.
         | 
         | Even then it is only some file systems that guarantee it and
         | even then file size updating isn't atomic afaik.
         | 
         | Not so sure about file size update being atomic in this case
         | but fairly sure about the rest.
         | 
         | Matklad had some writing or video about this.
         | 
         | Also there is a tool called ALICE and authors of that tool have
         | a white paper about this subject.
         | 
         | Also there was a blog post about how badger database fixed some
         | issues around this problem.
        
           | jeffbee wrote:
           | I don't think any part of your post is right. Aside from NFS,
           | there should not be filesystems where this doesn't work. If
           | there are, those are just bugs. The flags you mentioned are
           | not required or relevant. Setting the fd offset to the end of
           | the file atomically is the entire purpose of O_APPEND.
        
             | ozgrakkurt wrote:
             | It depends on what you mean by atomic. If it is only
             | writing to page cache and you are writing a small amount
             | then yes?
             | 
             | If there is a failure like a crash or power outage etc.
             | then it doesn't work like that.
             | 
             | You might as well be pushing into an in-memory data
             | structure and writing to disk at program exit in terms of
             | reliability
        
               | jeffbee wrote:
               | You are projecting imaginary features onto O_APPEND and
               | then hypothesizing that your imaginary features might not
               | work.
               | 
               | POSIX says that for a file opened with O_APPEND "the file
               | offset shall be set to the end of the file prior to each
               | write." That's it. That's all it does.
        
       | KevinChasse wrote:
       | Nice catalog. One subtle thing I've found in building
       | deterministic, stateless systems is that atomic filesystem and
       | memory operations are the only way to safely compute or persist
       | secrets without locks. Combining rename/link/O_EXCL patterns with
       | ephemeral in-memory buffers ensures that sensitive data is never
       | partially written to disk, which reduces race conditions and
       | side-channel exposure in multi-process workflows.
        
       | nialv7 wrote:
       | The mmap/msync one is incorrect I believe? (Correct me if I am
       | wrong).
       | 
       | msync() sync content in memory back to _disk_. But multiple
       | processes mapping the same file always see the same content
       | (barring memory consistency, caching, etc.) already. Unless the
       | file is mapped with MAP_PRIVATE.
        
         | DSMan195276 wrote:
         | Yeah I agree that one isn't very clear, perhaps the idea is to
         | use `msync()` as a barrier to achieve consistent ordering of
         | the writes without having to handle that yourself with more
         | complex primitives. But then, they do mention some of those
         | primitives at the bottom of the article, so it's hard to say
         | what exactly the idea is.
        
         | icedchai wrote:
         | mmap/msync is behavior is also very platform specific. On some
         | systems (like AIX, at least older versions), even without
         | msync, memory mapped data is synced back to disk periodically.
         | 
         | I worked on a code base that was portable between Linux, AIX,
         | and some other Unix flavors. mmap/msync was a source of bugs.
         | Just imagine your system running for days, never syncing any
         | data to disk... then someone pulls the plug. Where'd my data
         | go? Even worse, it happened "in production" at a beta site.
         | Fortunately we had a way to recover data from a log.
        
       ___________________________________________________________________
       (page generated 2026-02-06 23:00 UTC)