Post B5lb8LhJ306tkTecee by arclight@oldbytes.space
(DIR) More posts by arclight@oldbytes.space
(DIR) Post #B5kZwP7YCRRvmKi6Hg by arclight@oldbytes.space
0 likes, 0 repeats
RE: https://c.im/@cdarwin/116479704797697865Nothing "went rogue". AI didn't delete the firm's database and backups. A human operator built admin automation and ran it in production without adequate testing or backups.I'm sorry, but no: a human gave admin privileges to unverified tools and run them in a production environment.Own your work. You as sysadmin, developer, etc. are paid to perform a job with skill and diligence. Ultimately you are responsible for your professional work. If there was someone upstream responsible for V&V of the tool, ensuring users are trained, cautions and limitations of the tool are communicated, and confirming the tool is fit for its intended use, they bear a share of that responsibility.If you're the manager that forced worker to use an unreliable tool on production systems without putting it through proper V&V, without effective user training, use case development, or risk assessment, you bear a share of the responsibility.Repeating this for those in the back: AI does not launder away your job responsibilities.RT: https://c.im/users/cdarwin/statuses/116479704797697865
(DIR) Post #B5kZwPj7wjVVesI6r2 by snowyfox@deadinsi.de
0 likes, 1 repeats
.
(DIR) Post #B5lb8Ksc5TYpDFw8dk by arclight@oldbytes.space
0 likes, 0 repeats
Once upon a time - around 2007 or so, just before I left sysadminery to do risk assessment on legacy radioctive waste cleanup - I was upgrading a Blackboard LMS. Hardware load balancer in front of two unreliable web/app servers, Oracle db with RMAN backups, NFS file store backed by iSCSI SAN storage for user files. The web front-ends were intentionally provisioned with low disk because they didn't need it - content was on the NFS server or in the database and all the front-ends needed disk for was swap and log files from Apache and Tomcat. Those were religiously scraped because why would you expect the vendor to rotate their giant useless log files or record to a remote log host when they could just let crap accumulate everywhere until their system fell over?But I digress. I had automated log cleanup and database backup (with tested restores) and main IT managed the SAN backups. It was that brief Windows of low usage between semesters when we could run the vendor-supplied binaries to upgrade this expensive and cursed assemblage of Java, Perl, Oracle, and human misery. Read the documentation multiple times to understand the order of upgrade operations, clean and quiesce the system, take a few final backup snapshots and pull the trigger.The upgrade worked as intended, dutifully _moving_ files from the NFS mount of the large iSCSI drive to the web front-ends, filling the disk, then shitting the bed and falling over leaving the system in an unknown and unrecoverable state. As one does when you are Blackboard, the usurious vandal of LMS vendors.Surveying the flaming wreckage, I called Bb support to as for guidance. The support peon was impressed with my calm tone. I responded that being outwardly furious and losing my shit at them was unlikely to recover my system or improve any outcome I cared about. I did ask in my support ticket if this behavior from the update was documented and if there was any way I had missed a critical "do this to avoid incinerating prod" step in the upgrade process.A few hours later I got a response back from upper tier support that no, I had read and done everything correctly according to their documentation and that this whole fiasco could have been avoided by the use of an _intentionally undocumented_ option to the updater.Their words: _"intentionally undocumented"_.Why? Somebody might get confused by the explanation so it was omitted in the interest of ... clarity?I spent several harrowing hours waiting for the iSCSI restore to complete. I did my best to verify no user content was lost but to this day I don't know if we lost data.Deleting one symlink in the filesystem would have prevented this problem, a symlink that was required in a previous version of the code for the system to work properly (one actually described in the vendor documentation).We were a small university and did not have a full replica dev system to test the updater on. Why would we? We explicitly did not do development, we ran a vendor-supplied code in production. Dev systems were for developers which we were not. This wasn't a matter of a spinning up some virtuals in the cloud - we bought and managed real hardware and there was no way in hell we could justify doubling our hardware investment to have a test environment just to verify vendor supplied code worked as advertised. There wasn't a possibility of auditing the updater to detect that it would copy and delete the entirety of user-uploaded content, not without decompiling a big blob of Java. I exercised what diligence I could given the garbage state of the vendor's code and still got royally fucked over.I owned that. I informed my management chain of the situation and kept them updated with new developments and a revised ETA until the system was stabilized, recovered, and updated. That is how I practiced server ops for a decade (1998-2008). You did your dligence, said a prayer as you pushed the button, and you owned the outcome.No idea what current practice is. That was almost 20 years ago before devops, virtuals, clouds, and containers replaced real machines and dedicated sysadmins. I would like to believe that outlook and practice carried forward since then but I don't know - I left for greener, safer, more relaxing and fulfilling pastures helping package, transport, and store Cold War era uranium-metal-bearing radioactive sludge, moving it out of crumbling fuel pools at Hanford to interim storage elsewhere at Hanford. Then later projects for Dounreay, Sellafield, US commercial plants, Swedish interim waste storage at CLAB underneath Oskarshamn, fire PRA for plants in the US, Sweden, and Spain, and a whole lot of safety analysis code development and software QA. Somewhere in there I live-tweeted Fukushima melting and exploding. That's possibly the most direct act of nuclear safety I've performed with the goal of keeping people informed well enough to contextualize what was happening so they didn't panic and hurt themselves.
(DIR) Post #B5lb8LhJ306tkTecee by arclight@oldbytes.space
0 likes, 0 repeats
So if you wonder how I kept my shit together when Blackboard torched itself, it came from the perspective of nuclear safety analysis. Nobody was going to die, nobody was going to be hurt, the entire web service indistry wasn't going to be wiped out, and the worst thing that could happen is a few people would be inconvenienced for a few hours and there might be minor data loss of instructional material. Important in the local context but meaningless when compared with human safety and environmental protection.Skill and diligence are underrated. They may not save you when all hell breaks loose but more often than not they turn a disaster into a near miss. We ignore the deskilling and low/no-responsibility aspects of AI at our collective peril.
(DIR) Post #B5lb8LzjwSr6feHLO4 by arclight@oldbytes.space
0 likes, 0 repeats
Note that this isn't a "back in my day..." post as much as a description of how easy it is to get screwed over by _deterministic_ tools in an environment you mostly control. Modern systems have so many layers abstraction between an admin and the actual hardware, there are so many hidden points of failure and latent vulnerabilities that a) it's understandable that someone with tenuous employment (i.e. contractor held to a Jira ticket clearance rate or time window) would turn to a chatbot especially if chatbot use was now a KPI and a mandate, and b) if we couldn't avoid Blackboard-level disasters in an age of deterministic tools and more mentally tractable architectures, what makes anyone believe a chatbot can do any better today? We've painted ourselves into a corner chasing velocity and scale and layer upon layer of abstraction and the same people pushing all this are now pushing the magic expensive pirated text extrusion engine as a solution.Do you honestly need all that to build a stupid web page?"If I want to meet my KPIs and keep my pittance of health insurance I do."I can't fault that logic but seriously wonder why more tech workers haven't invested more of their energy into unions and guillotines. That supposes there's any discretionary energy and support networks left and that anything can survive the decades of libertarian propaganda. I suppose the silver lining in all this is with chatbots, you don't need to organize and recruit saboteurs to fight back; the whole system has evolved to tear itself apart on its own.
(DIR) Post #B5lb8MMmYnHrp73kIq by arclight@oldbytes.space
0 likes, 1 repeats
We really need to organize to take back control of our technology and our jobs and our profession. We are failing ourselves, the people who use and are affected by our systems, and the generation that is entering the field. Until we build more solidarity than a bucket full of lobsters, this BS is just going to keep happening.I'm lucky that I had my original profession to go back to when I left ops in 2008. Not everyone can or wants to work on literal radioactive sludge cleanup and ideally nobody should have to make that choice given how soulcrushingly awful dev & ops have become (OTOH it's a really easy choice if you've ever managed Blackboard. Beyond all probability, the company is worse than its software. As an example, radioactive sludge has never intentionally and actively made my life worse. Sludge mostly just sits there emitting energy and hydrogen.)Nobody should have to choose between using broken black-box automation that risks their code, systems, and skills and quitting their job. But people do every damn day - how do we break this cycle of abuse and incompetence?
(DIR) Post #B5lb8NZw3NP1aCDlRY by arclight@oldbytes.space
0 likes, 0 repeats
I will grudgingly accept "arson" as a viable solution. Not my first choice but I'm hard pressed to see an effective alternative.