Posts by codebergstatus@social.anoxinon.de
 (DIR) Post #B5kNrtQ2vPQGqiRa4G by codebergstatus@social.anoxinon.de
       1 likes, 0 repeats
       
       We're investigating a downtime of our primary instance.
       
 (DIR) Post #B5xBkzGi6bRiZKUrmS by codebergstatus@social.anoxinon.de
       1 likes, 0 repeats
       
       Codeberg.org appears to be very slow right now, we are trying to resolve this as we speak.
       
 (DIR) Post #B5zzf4GS0ChiLkfMPY by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We'll be performing a reboot of all our servers this evening.This means that our services (CI, pages, translate) will be temporarily offline. We invested a lot of effort in the last few months to be able to move Forgejo between servers easily and this should see minimal downtime.
       
 (DIR) Post #B5zzf4Ud9U2x3jIgVs by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       Hi,It's taking longer than expected. Ceph made a surprise move on us at the last moment by telling us our network configuration is maybe no longer up to date after we replaced servers (not today, but a while ago). This is preventing us from deploying new monitors to keep quorum. We're working on fixing this.It's _simply_ preventing us from gracefully shutting down, no data is at risk.
       
 (DIR) Post #B5zzf4lI9XNFtP5zU0 by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We're uncertain about the timeframe of when we can shutdown this machine. Therefore, we're going to bring up services again for the time being, we only require them to be down just before bringing the machine down.
       
 (DIR) Post #B5zzf4ylLS9KZBOkTo by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       Ceph is happy, we're able to bring this machine down finally and are doing so now.
       
 (DIR) Post #B5zzf58gkY5b3y2fx2 by codebergstatus@social.anoxinon.de
       1 likes, 0 repeats
       
       The maintenance is completed and services are expected to recover soonish (still fighting with the Codeberg e. V. forum).We're postponing the reboot of the third machine.
       
 (DIR) Post #B6R14Wiyfg0j4xdBZY by codebergstatus@social.anoxinon.de
       1 likes, 0 repeats
       
       You might've heard it by now: the Linux kernel is so insecure that AI can find many LPEs for them, so we'll be taking this evening to migrate to *BSD.In all seriousness, although we've mitigated the currently known LPEs via blocking modules, seccomp and sysctl, we still want to run with a kernel that has patched it. We will be doing that this evening. That means that each server will be taken down to boot into the new kernel, and services will be temporarily inaccessible.
       
 (DIR) Post #B6R14XAyzYPcTojYDw by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       To be clear, we mean to boot into a patched Linux kernel version. We have no intentions to migrate to a BSD distribution.
       
 (DIR) Post #B6R14XVBmQZjUUBgie by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We'll be starting to reboot the first server. This server hosts Pages, Translate, Forgejo runner, e.V. forum and join.codeberg.org.
       
 (DIR) Post #B6R14XsEOl0Udwy5dQ by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       This rebooted completed. We noticed that HAProxy didn't start on its own because it tried to read from Ceph before its mounted, we forget to add a mount dependency in the systemd service.Unfortunately all services that uses the Galera cluster takes more time to start. We noticed for a while that when a Galera node joins the cluster it does a full SST instead of a incremental sync. Which is painfully slow.
       
 (DIR) Post #B6R14YI6qXhtwD4kyG by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We'll be starting the reboot of the second server. This server only has services that are deployed with high availability, except the Woodpecker agent.
       
 (DIR) Post #B6R14YlB6SxXOMfyHQ by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We noticed that during this reboot, CephFS barely responded to any FS I/O requests. This caused 504 errors to be seen on codeberg.org during this reboot; that was not our intention sorry for that.The server we rebooted was the same that had "problems" where Ceph didn't allow it to go into maintenance mode: https://social.anoxinon.de/@codebergstatus/116523668707929156We now have a feeling the problem is still not completely fixed and there's still a big dependency on this server for Ceph to work.RT: https://social.anoxinon.de/ap/users/115651402898351746/statuses/116523668707929156
       
 (DIR) Post #B6R14Z6RpNyOSKcxQu by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       We're starting the last reboot of the evening, the one that has the main Forgejo instance. We've moved this to another server temporarily, all data, queues and session were already stored high availability a while ago so this is luckily a trivial task. Unfortunately we forget to sync SSH host keys, that will stay downThis one has not been rebooted for 110(!) days. We'll also be updating the BIOS and iDRAC with this reboot, so this one will take a bit longer. But no services should be impacted.
       
 (DIR) Post #B6R14ZS4WzGpXOkE8e by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       This reboot has now finalized, BIOS and iDRAC update was successful. We're now waiting on the Galera node to be fully synced to the cluster and then we'll move the main Forgejo instance back.
       
 (DIR) Post #B6R14ZnhEaZGcSrUqO by codebergstatus@social.anoxinon.de
       0 likes, 0 repeats
       
       The main Forgejo instance is back, SSH also works again. We have no remaining maintenance for this evening, all services should be working again without any hiccups.If you feel or see any problem, please report it to https://codeberg.org/Codeberg/Community/issues
       
 (DIR) Post #B6hQ097WZd6fSbEJ8a by codebergstatus@social.anoxinon.de
       1 likes, 0 repeats
       
       You might have noticed that SSH connections are no longer being served at the moment.As part of our investigation into the most recent slowness of Codeberg (which usually results in 504s being given), we have temporarily stopped those. This confirms a suspicion of ours that SSH connections have become, for whatever reason, heavy on the CPU.We will resume serving SSH connections with a smaller queue to preserve the availability of Codeberg.
       
 (DIR) Post #B6j5m4RDZn2e7JN7x2 by codebergstatus@social.anoxinon.de
       2 likes, 0 repeats
       
       Good news: We have addressed recent SSH performance degradation by replacing linear parsing of the authorized_keys file with a custom AuthorizedKeys command.This is what GitHub and GitLab have been doing for years, and we have grown to a size where this has become necessary for us as well.The cause for the degradation was still abusive patterns, as connections without valid key take more resources (scanning the file to the end) than legitimate users (scanning is stopped after match).
       
 (DIR) Post #B6j5m4rnywJDRloMOO by codebergstatus@social.anoxinon.de
       2 likes, 0 repeats
       
       Bad news: Performance on Codeberg is still slow. We're seeing a lot of abusive cloning and crawling and have limited the ability to clone for several minutes to investigate the situation.We're still investigating the situation and taking respective countermeasures. We try to keep the impact on legitimate usage as small as possible.
       
 (DIR) Post #B7QP8Zm91wBWpZggxE by codebergstatus@social.anoxinon.de
       2 likes, 0 repeats
       
       We are struggling to keep Codebreg.org available for unauthenticated users due to massive abuse of expensive endpoints.Our current priority is keeping Codeberg.org responsive for authenticated users.