Post B6R14ZS4WzGpXOkE8e by codebergstatus@social.anoxinon.de
(DIR) More posts by codebergstatus@social.anoxinon.de
(DIR) Post #B6R14Wiyfg0j4xdBZY by codebergstatus@social.anoxinon.de
1 likes, 0 repeats
You might've heard it by now: the Linux kernel is so insecure that AI can find many LPEs for them, so we'll be taking this evening to migrate to *BSD.In all seriousness, although we've mitigated the currently known LPEs via blocking modules, seccomp and sysctl, we still want to run with a kernel that has patched it. We will be doing that this evening. That means that each server will be taken down to boot into the new kernel, and services will be temporarily inaccessible.
(DIR) Post #B6R14XAyzYPcTojYDw by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
To be clear, we mean to boot into a patched Linux kernel version. We have no intentions to migrate to a BSD distribution.
(DIR) Post #B6R14XVBmQZjUUBgie by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
We'll be starting to reboot the first server. This server hosts Pages, Translate, Forgejo runner, e.V. forum and join.codeberg.org.
(DIR) Post #B6R14XsEOl0Udwy5dQ by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
This rebooted completed. We noticed that HAProxy didn't start on its own because it tried to read from Ceph before its mounted, we forget to add a mount dependency in the systemd service.Unfortunately all services that uses the Galera cluster takes more time to start. We noticed for a while that when a Galera node joins the cluster it does a full SST instead of a incremental sync. Which is painfully slow.
(DIR) Post #B6R14YI6qXhtwD4kyG by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
We'll be starting the reboot of the second server. This server only has services that are deployed with high availability, except the Woodpecker agent.
(DIR) Post #B6R14YlB6SxXOMfyHQ by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
We noticed that during this reboot, CephFS barely responded to any FS I/O requests. This caused 504 errors to be seen on codeberg.org during this reboot; that was not our intention sorry for that.The server we rebooted was the same that had "problems" where Ceph didn't allow it to go into maintenance mode: https://social.anoxinon.de/@codebergstatus/116523668707929156We now have a feeling the problem is still not completely fixed and there's still a big dependency on this server for Ceph to work.RT: https://social.anoxinon.de/ap/users/115651402898351746/statuses/116523668707929156
(DIR) Post #B6R14Z6RpNyOSKcxQu by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
We're starting the last reboot of the evening, the one that has the main Forgejo instance. We've moved this to another server temporarily, all data, queues and session were already stored high availability a while ago so this is luckily a trivial task. Unfortunately we forget to sync SSH host keys, that will stay downThis one has not been rebooted for 110(!) days. We'll also be updating the BIOS and iDRAC with this reboot, so this one will take a bit longer. But no services should be impacted.
(DIR) Post #B6R14ZS4WzGpXOkE8e by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
This reboot has now finalized, BIOS and iDRAC update was successful. We're now waiting on the Galera node to be fully synced to the cluster and then we'll move the main Forgejo instance back.
(DIR) Post #B6R14ZnhEaZGcSrUqO by codebergstatus@social.anoxinon.de
0 likes, 0 repeats
The main Forgejo instance is back, SSH also works again. We have no remaining maintenance for this evening, all services should be working again without any hiccups.If you feel or see any problem, please report it to https://codeberg.org/Codeberg/Community/issues