[HN Gopher] Xee: A Modern XPath and XSLT Engine in Rust
       ___________________________________________________________________
        
       Xee: A Modern XPath and XSLT Engine in Rust
        
       Author : robin_reala
       Score  : 254 points
       Date   : 2025-03-28 06:48 UTC (16 hours ago)
        
 (HTM) web link (blog.startifact.com)
 (TXT) w3m dump (blog.startifact.com)
        
       | athanagor2 wrote:
       | The fact it could be compiled in WASM is a good thing, given the
       | Chrome team was considering removing libxml and XSLT support a
       | few years back. The reasons cited were mostly about security (and
       | share of users).
       | 
       | It's another proof that working on fundamental tools is a good
       | thing.
        
         | therealmarv wrote:
         | not WASM but there is also https://www.npmjs.com/package/saxon-
         | js
        
       | montroser wrote:
       | Fun fact: XSLT still enjoys broad support across all major
       | browsers: https://caniuse.com/?search=xslt
        
         | Telemakhos wrote:
         | This is true only of XSLT 1.0. The current standard is 3.0.
        
           | falcor84 wrote:
           | Oh, a shame. Is there any way to track browser version
           | adoption on caniuse, or any other site?
           | 
           | Also, is it up to browser implementations, or does WHATWG
           | expect browsers to stay at version XSLT 1?
        
             | Telemakhos wrote:
             | I think it's up to browser implementations, but JSON and
             | JavaScript stole much of XML's thunder in the browser
             | anyway, plus HTML5's relaxed tags won out over XHTML 4's
             | strictness (strictness was a benefit if you were actually
             | working with the data). There are still plenty of web-
             | available uses of XML, like RSS and RDF and podcasts/OPML,
             | but people are more likely to call xmlhttp.responseXML and
             | parse a new DOM than wrap their head around XSL templates.
             | 
             | The big place I've successfully used XSLT was in TEI, which
             | nobody outside digital humanities uses. Even then, the XSLT
             | processing is usually minimal, and Javascript is going to
             | do a lot of work that XSL could have done.
        
             | tannhaeuser wrote:
             | There's nothing to track here really. For better or worse,
             | browsers are stuck with 1999's XSLT 1.0, and it's a miracle
             | it's still part of native browser stacks given PDF
             | rendering has been implemented using JS for well over a
             | decade now.
             | 
             | XSLT 2 and 3 is a W3C standard written by the sole
             | commercial provider of an XSLT 2 or 3 processor, which is
             | problematic not only because it reduces W3C to a moniker
             | for pushing sales, but also because it undermines W3C's own
             | policy of at least two interworking implementations for a
             | spec to get "recommendation" status.
             | 
             | XSLT is of course a competent language for manipulating
             | XML. It would be a good fit if your processing requires
             | lots of XML literals/fragments to be copied into your
             | target document since XSLT is an XML language itself.
             | Though OTOH it uses XPath embedded in strings excessively,
             | thereby distrusting XML syntax for a core part of its
             | language itself, and coding XPath in XML attributes can be
             | awkward due to restrictive contextual encoding rules for
             | special characters and such.
             | 
             | XSLT can be a maintenance burden if used casually/rarely,
             | since revisiting XSLT requires substantial relearning and
             | time investment due to its somewhat idiosyncratic nature.
             | IDE support for discovery, refactoring, and test automation
             | etc. is lacking.
        
               | velcrovan wrote:
               | My favorite and only use of XSLT that still works pretty
               | well is to allow people to browse my RSS feed as if it
               | were a web page.
               | 
               | https://joeldueck.com/feed.atom
        
               | giantrobot wrote:
               | To this day I'm frustrated that developers (web devs
               | mostly) tossed XML aside for JSON and the requisite
               | JavaScript to replace relatively straight forward things
               | like converting structured data to something a browser
               | could display.
               | 
               | I bought into and still believe in the separation of data
               | and its presentation. This a place where XML/XSLT was
               | very awesome despite some poor ergonomics.
               | 
               | An RSS XML document could live at an endpoint and contain
               | in-line comments, extra data in separate namespaces, and
               | generally be really useful structured data for any user
               | agent or tool to ingest. An RSS reader or web spider
               | could process the data directly, an XSLT stylesheet could
               | let a web browser display a nice HTML/CSS output, and any
               | other tools could use the data as well. Even better any
               | user agent ingesting the XML could use in-built tools to
               | validate the document.
               | 
               | XSLT to convert an XML feed to pretty HTML is a great
               | example of the utility. Browsers have _fast_ built-in
               | conversion engines and the resulting HTML produced has
               | all the normal capabilities of HTML including CSS and
               | JavaScript. To the uninitiated: the XML feed just links
               | to an external XSL stylesheet, when a web browser fetches
               | the XML it grabs the stylesheet and transforms the XML to
               | an HTML (or XHTML) representation that 's then fed back
               | into the browser.
               | 
               | A feed reader will fetch the XML and process it directly
               | as RSS data and ignore the stylesheet. Some other user
               | agent could fetch the XML and ignore its linked
               | stylesheet but provide its own to process the RSS data.
               | Since the feed has a declared schema pretty much any
               | stylesheet written to understand that schema will work.
               | For instance you could turn an RSS feed into a PDF with
               | XSLT.
        
               | ajxs wrote:
               | That's an awesome idea! Very cool! I might just do that
               | for my own RSS feed, and credit you for the great idea.
        
               | int_19h wrote:
               | > XSLT 2 and 3 is a W3C standard written by the sole
               | commercial provider of an XSLT 2 or 3 processor,
               | 
               | I was following the W3C XSLT mailing list for quite some
               | time back when they were doing 3.x, and this does not
               | strike me as accurate.
        
         | eyelidlessness wrote:
         | I can't say this with certainty, but I have some reason to
         | suspect I might be partially to blame for this fun fact!
         | 
         | A couple years ago, I stumbled on a discussion considering
         | deprecation/removal of XSLT support in Chrome. At some point in
         | the discussion, they mentioned observing a notable uptick in
         | usage--enough of an uptick (from a baseline of approximately
         | zero) that they backed out.
         | 
         | The timing was closely correlated with work I'd done to adapt a
         | library, which originally used XSLT via native Node extensions,
         | to browser XSLT APIs. The project isn't especially "popular" in
         | the colloquial sense of the term, but it does have a
         | substantial niche user base. I'm not sure how much uptake the
         | browser adaptation of this library has had since, but some
         | quick napkin math suggested it was at least plausible that the
         | uptick in usage they saw might have been the _onslaught of
         | automated testing_ I used to validate the change while I was
         | working on it.
        
         | ajxs wrote:
         | Being interested in archaic technologies, I built a website
         | using XML/XSLT not that long ago. The site was an archive of a
         | band I was in, which made it fundamentally data oriented: We
         | recorded multiple albums, with different tracks, and a
         | different lineup of musicians each time. There's lots of
         | different databases I could built a static site generator
         | around, but what if the browser could render the page straight
         | from the data? That's what's cool about XML/XSLT. On paper, I
         | think it's actually a pretty nice idea: The browser starts by
         | loading the actual data, and then renders it into HTML
         | according to a specific stylesheet. Obviously the history of
         | browser tech forked in a different direction, but the idea
         | remains good. What if there was native browser support for
         | styling JSON into HTML?
        
       | mvc wrote:
       | Nice work. Xpath is a beast. Obvious why paligo would be
       | interested too. Must be a lot of commercial documentation out
       | there where the best representation they can get looks a bit
       | XMLish.
        
       | therealmarv wrote:
       | Great to see that somebody else creates a true open source XSLT 3
       | and XPATH 3 implementation!
       | 
       | I worked on projects which refused to use anything more modern
       | than XSLT & XPATH 1.0 because of lack of support in the non
       | Java/Net World (1.0 = tech from 1999). Kudos to Saxon though, it
       | was and is great but I wished there were more implementations of
       | XSLT 2.0 & XPATH 2.0 and beyond in the open source World... both
       | are so much more fun and easier to use in 2.0+ versions. For that
       | reason I've never touched XSLT 3.0 (because I stuck to Saxon B
       | 9.1 from 2009). I have no doubt it's a great spec but there
       | should be other ways than only Saxon HE to run it in an open
       | source way.
       | 
       | It's like we have an amazing modern spec but only one browser
       | engine to run it ;)
        
         | Finnucane wrote:
         | I've worked on archive projects with complex TEI xml files
         | (which is why when people say xml is bad and it should be all
         | json or whatever, I just LOL), and fortunately, my employer
         | will pay for me to have an editor (Oxygen) that includes the
         | enterprise version of Saxon and other goodies. An open-source
         | xml processing engine that wasn't decades out of date would be
         | a big deal in the digital humanities world.
        
           | sramsay wrote:
           | I don't think people realize just how important XML is in
           | this space (complex documentary editing, textual criticism,
           | scholarly full-text archives in the humanities). JSON cannot
           | be used for the kinds of tasks to which TEI is put. It's not
           | even an option.
           | 
           | Nothing could compel me to like XSLT. I admire certain
           | elements of its design, but in practice, it just seems
           | needlessly verbose. But I really love XPath, though.
        
             | miki123211 wrote:
             | XML is great for documents.
             | 
             | If your data is essentially a long piece of text, with
             | annotations associated with certain parts of that text,
             | this is where XML shines.
             | 
             | When you try to use XML to represent something like an
             | ecommerce order, financial transaction, instant message and
             | so on, this is where you start to see problems. Trying to
             | shove some extremely convoluted representation of text
             | ranges and their attributes into JSON is just as bad.
             | 
             | A good "rule of thumb" would be "does this document still
             | make sense if all the tags are stripped, and only the text
             | nodes remain?" If yes, choose XML, if not, choose JSON.
        
           | faassen wrote:
           | My hope is that we can get a little collective together that
           | is willing to invest in this tooling, either with time or
           | money. I didn't have much hope, but after seeing the positive
           | response today more than before.
        
         | smartmic wrote:
         | Well, it's not as if this is the first free alternative. Here
         | is a wonderful, incredibly powerful tool, not written in Java,
         | but in Free Pascal, which is probably too often underestimated:
         | Xidel[1]. Just have a look at the features and check its Github
         | page[2]. I've often been amazed at its capabilities and, apart
         | from web scraping, I mainly use it for XQuery executions - so
         | far the latest version 0.9.9 has also implemented XPath/XQuery
         | 3.1 perfectly for my requirements. Another insider tip is that
         | XPath/XQuery 3.1 can also be used to transform JSON wonderfully
         | - JSONiq is therefore obsolete.
         | 
         | [1] https://www.videlibri.de/xidel.html
         | 
         | [2] https://github.com/benibela/xidel
        
           | smartmic wrote:
           | Forget to add, for latest XQuery up to 4.0, there is also
           | BaseX [1] -- this time a Java program. It has a great GUI/IDE
           | for XQuery rapid prototyping.
           | 
           | [1] https://basex.org/basex/xquery/
        
       | stuaxo wrote:
       | eXcellent, it's good to see new work on XSLT, reviled bysome it's
       | actually great tech and useful in all sorts of places.
        
       | airstrike wrote:
       | What problems are {elegantly, neatly, best} solved by using XPath
       | and XSLT today that would make them reasonable choices over
       | alternatives?
        
         | mickeyp wrote:
         | XPath is a superb query language for XML (or anything that you
         | can structure as a DOM) --- it is also, with some obscure
         | exceptions, the _only_ query language with serious adoption, so
         | it 's an easy choice and readily available in XML tools. The
         | only caveat is there are various spec versions and most never
         | added support for newer versions.
         | 
         | Let's look at JSON by comparison. Hmm, let's see: JSONPath,
         | JMESPath, jq, jsonql, ...
        
           | trallnag wrote:
           | Recently discovered Jsonata thanks to AWS adding it to Step
           | Functions. Feel free to add it to your enumeration
        
           | never_inline wrote:
           | JQ is the most feature-rich of the bunch. It's defacto
           | standard and I usually just default to it because it offers
           | so much - assignment, various builtins such as base64
           | encoding.
           | 
           | The disadvantage is that it's not easily embeddable in your
           | own programs - so programs use JSONPath / Go templates often.
        
             | bbkane wrote:
             | I also don't think there's a specification written for the
             | jq query language, unlike https://jmespath.org/ , which as
             | you mentioned also has more client libraries.
             | 
             | I too am probably going to embed jmespath in my app.I need
             | it to allow users to fill CLI flags from config files, and
             | it'll replace my crappy homegrown version ( https://github.
             | com/bbkane/warg/blob/740663eeeb5e87c9225fb627... )
        
         | therealmarv wrote:
         | E.g. massive XML documents with complexity which you need to be
         | transformed into other structured XML. Or if you need to parse
         | complex XML. Some people hate XSLT, XPATH with a passion and
         | would rather write much more complex lxml code. It has a steep
         | learning curve but once you understand the fundamentals you can
         | transform XML more easily and especially predictable and
         | reliable than ever.
         | 
         | Another example: If you have very large XML you cannot fit even
         | into memory you can still stream process them with XSLT.
         | 
         | It makes you the master of XML transformations and fetching
         | information out of complex XML ;)
        
         | jeffbee wrote:
         | What alternatives exist for extracting structured data from the
         | web? I have several ETL pipelines that use htmltidy to turn tag
         | soup into something approximately valid and xmlstarlet to
         | transform it into tabular data.
        
         | jerf wrote:
         | XPath is a very nice language for querying over XML. Most
         | places pitch it as a "declarative" syntax, but as I am quite
         | skeptical of "declarative" as a concept, you can also look at
         | the vast majority of the XPath standard as a way to
         | imperatively drive a multicursor over an XML document, diving
         | in out and out nodes and extracting bits of text and such,
         | without having to write the equivalent code in your language to
         | do so, which will be inevitably quite a bit more verbose. When
         | you need it, it's really useful.
         | 
         | In my very opinionated opinion, XPath is about 99% of the value
         | of XSLT, and XSLT itself is a misfire. Embedding an XML
         | language in XML, rather than being an amazing value
         | proposition, is actually a huge and really annoying mistake, in
         | much the same way and for much the same reason as anyone who
         | has spent much time around shell scripting has found trying to
         | embed shell strings in shell strings (and, if the situation is
         | particularly dire, another third or fourth level of such
         | nesting) is quite unpleasant. Imagine trying to deal with bash,
         | except you have to first quote all the command lines as bash
         | strings like you're using bash -c, all the time. I think "XPath
         | + your favorite language" has all the power of XSLT and,
         | generally, better ergonomics and comprehensibility. Once you've
         | got the selection of nodes in hand, a general-purpose
         | programming language is a better way to deal with their
         | contents then what XSLT provides. Hence why it has always
         | languished.
        
           | akshayshah wrote:
           | To someone who hasn't worked much with XML, this seems like a
           | reasonable take!
           | 
           | For cases where a host system wants to execute user-defined
           | data transformations safely, XSLT seems like it might be
           | useful. When they mature, maybe WASM and WASI will fill the
           | same niche with better developer ergonomics?
        
           | therealmarv wrote:
           | Interesting take about XSLT. But I agree... XSLT could be
           | something much more simple (and non XML initself) and
           | combined with XPATH. It feels like a lot of boiler code to
           | write XSLT.
        
           | jrpelkonen wrote:
           | It's been a while since I've had to deal with XML, but I
           | remember finding it fairly convenient to restructure XML
           | documents with XSLT. Modifying the data in those documents,
           | much less so. I think there's a sweet spot.
        
           | int_19h wrote:
           | XQuery is the best of both worlds - you get almost all the
           | benefits of XSLT like e.g. the ability to define your own
           | functions, but with non-XML-based syntax that is a superset
           | of XPath.
           | 
           | Basically the only thing it's missing in XQuery vs XSLT is
           | template rules and their application; but IMO simple ones are
           | just as easy to write explicitly, and complex rulesets are
           | hard to reason about and maintain anyway.
        
         | never_inline wrote:
         | I have used it when using scraping some data from web pages
         | using scrapy framework. It's reliable way to extract something
         | from web pages compared to regex.
        
         | password4321 wrote:
         | XPATH+XSLT is SQL for XML, declarative selection and
         | transformation.
         | 
         | Using an XML library to iterate through an entire XML document
         | without XPATH is like looping through entire database tables
         | without a JOIN filter or a WHERE clause.
         | 
         | XSLT is the SELECT, transforming XML output with a new level of
         | crazy for recursion.
        
         | Devasta wrote:
         | I manage a team who build and maintain trading data reports for
         | brokers, we have everything generate in a fairly standard
         | format and customize to those brokers exact needs with XSLT.
         | Hundreds of reports, couldnt manage without it.
        
       | mickeyp wrote:
       | It's interesting to see the slow rehabilitation of XML and its
       | tooling now that there's a new generation of developers who have
       | not grown up in the shadow of XML's prime in the late 90s / early
       | 2000s, and who have not heard (or did not buy into) the anti-XML
       | crowd's ranting --- even though some of their criticisms were
       | legitimate.
       | 
       | I've always liked XML, and especially XPath, and even though
       | there were a large number of missteps in the heyday of XML, I
       | feel it has always been unfairly maligned. Look at all the people
       | who reinvent XML tooling but for JSON, but not nearly as well.
       | Luckily, people who value XML can still use it, provided the fit
       | is right. But it's nice to see the tides turning.
       | 
       | Most fashions really are cyclical.
        
         | mickeyp wrote:
         | Oh, and just to pile on to my own post:
         | 
         | If you like React's JSX; enjoy its strictures and clean,
         | readable "HTML"; then good news, you're writing XML (but
         | without namespacing).
        
           | Lammy wrote:
           | See also: ECMAScript for XML (E4X) https://ecma-
           | international.org/wp-content/uploads/ECMA-357_2...
        
         | Mountain_Skies wrote:
         | I made extensive use of XPath and XSL(T) back in their heyday
         | and in general was fine with them but the architect astronauts
         | who love showing off how clever they are with artificial
         | complexity had a tendency to make use of XML tech to complicate
         | things unnecessarily. Think that might be where many people's
         | dislike of it came from, especially those whose first exposure
         | wasn't learning through simple structures when XML was new but
         | were thrown into the type of morass that develops when a tech
         | is climbing the maturity curve.
        
         | linguae wrote:
         | It's the "slope of enlightenment" phase of the Gartner hype
         | cycle, where people are able to make sober assessments of
         | technologies without undue influence from hype or its backlash.
         | We're long past the days where XML is used for everything, even
         | when it's inappropriate, and we're also past the "trough of
         | disillusionment" phase where people sought alternatives to XML.
         | 
         | I think XML is good for expressing document formats and for
         | configuration settings. I prefer JSON for data serialization,
         | though.
        
           | 01HNNWZ0MV43FF wrote:
           | For phone users
           | https://en.m.wikipedia.org/wiki/Gartner_hype_cycle
        
         | j-pb wrote:
         | XML is still a huge mistake for most stuff. It's fine for
         | _documents_ but not as a data storage solution. Bloat,
         | ambiguities, virtually impossible to canonicalise.
         | 
         | XPath is cute, but if you don't mind bloat, text-only and lack
         | of ergonomics, anyways then Conjunctive Regular Path Queries
         | and RDF are miles ahead of XML as a data storage solution. (Not
         | serialised as XML please xD)
        
         | JTyQZSnP3cQGa8B wrote:
         | It's only a sample of one but I'm really unhappy with the
         | issues and limitations that JSON and YAML have, and I welcome
         | XML if it has good tools.
        
           | bluGill wrote:
           | That depends on what I'm doing. Most what what I'm doing is
           | simple and so xml is just way to complex for the task.
           | However when I need something complex xml can handle things
           | that the others cannot - at the expense of being really
           | complex to work with.
        
         | ctrlp wrote:
         | XML/XPath are very useful but I've definitely lived through
         | their abuses. Still _abusus non tollit usam_ and I 've had many
         | positive experiences with XPath especially. XmlStarlet has been
         | especially useful, also xmllint. I welcome more tooling like
         | this. The major downside to XML is the verbosity and cognitive
         | load. Tooling that manages that is a godsend.
        
         | kgwxd wrote:
         | XML, and other X[x] standards, are just horrible to read. On
         | top of that, XML was made 10x worse by wrapping things in SOAP
         | and the like over the wire, back in the day.
         | 
         | XSD, XPath, XSLT are all domains where I'd argue that
         | reading/reasoning about are way more important.
         | 
         | When troubleshooting an issue, I don't mind scanning XML for a
         | few data points so I can confirm what values are being
         | communicated, but when I need to figure out how/why a specific
         | value came to be, I don't want the logic spread throughout a
         | giant text file wrapped in attribute value strings, and other
         | non-debuggable "code". I'd rather it just be in a proper
         | programming language.
        
           | faassen wrote:
           | The specifications are certainly not easy to read, and I
           | wouldn't recommend them to learn about XML. But from the
           | perspective of someone implementing them they are quite
           | useful!
           | 
           | As someone who has used many programming languages and who
           | went through the process of implementing this one I have many
           | opinions about XPath and XSLT as programming languages. I
           | myself am more interested in implementing them for others who
           | value using them than using them myself. I do recognize there
           | is a sizeable community of people who do use these tools and
           | are passionate about them - and that's interesting to see and
           | more power to them!
        
         | da_chicken wrote:
         | My complaints about XML remain pretty much unchanged since 10
         | years ago.
         | 
         | - Not including self-closing tags, there should only be one
         | close tag: </>
         | 
         | - Elements are for data. Attributes are evil
         | 
         | - XPath indexing should be 0-based
         | 
         | - Documents without a schema should not make your tools panic
         | or complain
         | 
         | - An xml document shouldn't have to waste it's time telling you
         | it's an xml document in xml
         | 
         | I maintain that one of the reasons JSON got so popular so
         | quickly is because it does all of the above. The problem with
         | JSON is that you lose the benefits of having a schema to
         | validate against.
        
           | ebruchez wrote:
           | There have been proposals a long time ago, including by Tim
           | Bray, for an XML 2.0 that would remove some warts. But there
           | was no appetite in the industry to move forward.
        
           | bambax wrote:
           | > _Elements are for data. Attributes are evil_
           | 
           | This is like, your opinion, man... ;-) You can devise your
           | schema any way you want. Attributes are great, and they exist
           | in HTML in the form of datasets, which, as usual, are a
           | poorly-specified and ill-designed rethinking of XML
           | attributes
           | 
           | > _Documents without a schema should not make your tools
           | panic or complain_
           | 
           | They don't. You absolutely don't need a schema. If you
           | declare a schema, it should exist. If not, no problem?
        
             | da_chicken wrote:
             | No, the problem with attributes is that people consistently
             | misuse them. So many things about XML break down when you
             | make everything a self closing tag with 50 attributes. So
             | many programmers just seem to say, "oh, it's shorter text
             | so it must be inherently better" or "oh it's one-to-one so
             | I should strictly avoid anything resembling a heirarchy."
             | 
             | Like I think this guy is mostly correct in identifying bad
             | XML: https://www.devever.net/~hl/xml
             | 
             | Though I don't necessarily agree with the "data format"
             | framing. This idea that markup languages are not data
             | formats seems confused.
             | 
             | > They don't. You absolutely don't need a schema. If you
             | declare a schema, it should exist. If not, no problem?
             | 
             | I agree that they should not.
             | 
             | However, I have used many tools that puke when presented
             | with XML fragments or XML with no schema.
        
             | Mountain_Skies wrote:
             | Sometime attribute use goes too far such as when they
             | contain comma separated lists of items.
        
               | bambax wrote:
               | Sure, but that's not the fault of the format itself, is
               | it? You can write extremely long enumerations in any
               | natural language -- that's the author's fault.
        
           | Mountain_Skies wrote:
           | Microsoft seems to be especially obsessed with making as much
           | as possible into attributes. Makes me wonder if there is some
           | hidden historical reason for that like an especially powerful
           | evangelist inside the company that loved attributes during
           | the early days of adopting XML.
        
             | int_19h wrote:
             | Attributes are way shorter to write.
             | 
             | That said, these days most Microsoft XML dialects are
             | actually XAML-based, and in XAML attributes are basically
             | syntactic sugar - you can write:                 <Foo
             | Bar="123">
             | 
             | or                 <Foo>         <Foo.Bar>123</Foo.Bar>
             | </Foo>
             | 
             | (the dot in the syntax makes it possible for the XAML
             | parser to distinguish nested elements that represent
             | properties from nested elements that represent child
             | objects)
        
         | int_19h wrote:
         | Curiously, one of the driving forces behind renewed interest in
         | XML is that language models seem to handle large XML documents
         | better than JSON. I suspect this has something to do with it
         | being more redundant - e.g. closing tags including the element
         | name - making it easier for the model to keep track of
         | structure.
        
       | egh wrote:
       | Very cool! I recently wrote an XSLT 2 transpiler for js
       | (https://github.com/egh/xjslt) - it's nice to see some options
       | out there! Writing the xpath engine is probably the hard part (I
       | relied on fontoxpath). I'm going to be looking into what you have
       | done for inspiration!
        
       | vessenes wrote:
       | This, thirty years later, is the best pitch for XML I've read.
       | Essentially, it's a slow moving, standards-based approach to data
       | interoperability.
       | 
       | I hated it the minute I learned about it, because it missed
       | something I knew I cared about, but didn't have a word for in the
       | 90s - developer ergonomics. XML sucks shit for someone who wants
       | to think tersely and code by hand. Seriously, I hate it with a
       | fiery passion.
       | 
       | Happily to my mind the economics of easier-for-creators -> make
       | web browsers and rendering engines either just DEAL with weird
       | HTML, or else force people to use terse data specs like JSON won
       | out. And we have a better and more interesting internet because
       | of it.
       | 
       | However, I'm old enough now to appreciate there is a place for
       | very long-standing standards in the data and data transformation
       | space, and if the XML folks want to pick up that banner, I'm for
       | it. I guess another way to say it is that XML has always seemed
       | to be a data standard which is intended to be what _computers_
       | prefer, not _people_. I'm old enough to welcome both, finally.
        
         | tzcnt wrote:
         | Developer ergonomics is drastically underappreciated, even in
         | modern times. Since we're talking about textual data formats,
         | I'll go out on a limb here and say that I hate YAML. Double
         | checking exactly how many spaces are present on each line is
         | tedious. It manages to make a simple task like copy-pasting
         | something from a different file (at a different indentation
         | level) into an error-prone process. I'll take angle brackets
         | any day.
        
           | formerly_proven wrote:
           | Working with large YAML documents is incredibly annoying and
           | shows the benefit of closing tags.
        
             | 4ndrewl wrote:
             | It all went downhill after we stopped using .ini files
        
               | Locutus_ wrote:
               | Well....toml isn't that much more than .ini files
               | slightly brough up in feature support.
               | 
               | Again not great for bigger documents.
        
           | 01HNNWZ0MV43FF wrote:
           | JSON5 is a real sweet spot for me. Closing brackets, but I
           | don't have to type every tag twice. Comments and trailing
           | commas.
        
             | consteval wrote:
             | I find for deeply hierarchical data that XML is much easier
             | to read.
        
           | chuckadams wrote:
           | You haven't felt hate until you've counted spaces in your
           | Helm templates in order to know what value to put after
           | `nindent`. The punchline is that k8s doesn't even speak yaml,
           | the protocol is all json and it's the tooling that inflicts
           | yaml on us. I can live with yaml as a config format, but once
           | logic starts creeping in, give me anything else.
        
           | Pet_Ant wrote:
           | > Developer ergonomics is drastically underappreciated, even
           | in modern times.
           | 
           | When was the last time you had an editor that wouldn't just
           | auto close the current tag with "</" ? I mean it's a god-send
           | for knowing where you are at in large structure. You aren't
           | scrolling to the top to find which tag you are in.
        
         | tannhaeuser wrote:
         | > _XML has always seemed to be a data standard which is
         | intended to be what computers prefer, not people._
         | 
         | On one hand, you aren't wrong: XML has in fact been used for
         | machine-to-machine communication mostly. OTOH, XML was just
         | introduced as a subset of SGML doing away with the need of
         | vocabulary-specific markup declarations for mere parsing in
         | favor of always requiring explicit start- and end-element tags.
         | Whereas HTML is chock full of SGMLisms such as tag inference
         | (for example inferring paragraph ends on block elements), empty
         | ("self-closing") elements and enumerated ("boolean") attributes
         | driven by per-element declarations.
         | 
         | One can argue to death whether the web should work as a mere
         | document delivery network with rigid markup a la XML, or that
         | browsers should also directly support SGML authoring idioms
         | such as the above shortform mechanisms. SGML also has text
         | macros/shared fragments (entities) and even allows defining own
         | parsing tokens for markdown, math, CSV, or custom syntaxes.
         | HTML leans towards SGML in that its documentation portrays HTML
         | as an authoring language, but browsers are lacking even in
         | basic SGML features such as entities.
        
           | IgorPartola wrote:
           | That's a flame war that's been raging for decades for sure.
           | 
           | I do wonder what web application markup would look like today
           | if designed from scratch. It is kind of amazing that HTML and
           | CSS can be used for creating beautiful documents viewable on
           | pretty much any device with a screen AND also for creating
           | dynamic applications with pixel-perfect rendering, special
           | effects, integrations with the device's hardware, and even
           | external peripherals.
           | 
           | If there was ever scope creep in a project this would be it.
           | And given the recent discussion on here of curses based
           | interfaces it reminded me just how primitive other GUI
           | application layout tools can be while still achieving amazing
           | results. Even something like GTK does not need the intense
           | level of layout engine support and yet is somehow considered
           | richer in some ways and probably more performant for a lot of
           | stuff that's done with it.
           | 
           | So I am curious what web application development would look
           | like today if it wasn't for HTML being "good enough".
        
             | caspper69 wrote:
             | Had we had better process isolation in the mid-90s, I
             | assume web application development would mostly be Java
             | apps, with a mini-vm for each one (sort of a qubes like
             | environment).
             | 
             | We just couldn't keeps apps' hands out of the cookie jar
             | back then.
        
         | wongarsu wrote:
         | And not only does the XML format have bad developer ergonomics,
         | most XML parsers are equally terrible to use. There are many
         | things I like about XML: name spaces, schemas, XPath, to some
         | degree even XSLT. But the typical XML developer experience is
         | terrible on every layer
        
         | baq wrote:
         | XML is a big improvement over YAML.
         | 
         | There, I said it.
        
         | IshKebab wrote:
         | The main thing I hate about XML ( _apart_ from the tedious
         | syntax and terrible APIs - who thought SAX was a sane idea?) is
         | that the data model is wrong for 99% of use cases.
         | 
         | XML gives you an object soup where text objects can be anywhere
         | and data can be randomly stored in tags or attributes.
         | 
         | It just doesn't at all match the object model used by basically
         | all programming languages.
         | 
         | I think that's a big reason JSON is so successful. It's
         | literally the object model used by JavaScript. There's no weird
         | impedance mismatch between the data represented on disk and in
         | your program.
         | 
         | Then someone had to go and screw things up with YAML...
         | 
         | JSON5 is the way.
        
         | velcrovan wrote:
         | You should try using a LISP like Racket for XML. Because XML
         | can be expressed directly as S-expressions, XML and LISP go
         | together like peanut butter and jelly.
         | <greeting attr="val" href="#">Hello
         | <thing>world</thing><greeting>              (greeting ((attr
         | "val") (href "#")) "Hello " (thing "world"))
        
           | koito17 wrote:
           | In my experience, at least with Clojure, it's much more
           | convenient to serialize XML into a map-like structure. With
           | your example, the data structure would look like so.
           | {:tag     :greeting        :attrs   {:href "#" :attr "val"}
           | :content ["Hello" {:tag :thing :content ["world"]}]}
           | 
           | Some people use namespaced keywords (e.g. :xml/tag) to help
           | disambiguate keys in the map. This kind of data structure
           | tends to be more convenient than dealing with plain sexps or
           | so-called "Hiccup syntax". i.e.                 [:greeting
           | {:href "#" :attr "val"} "Hello" [:thing "world"]]
           | 
           | The above syntax is convenient to write, but it's tedious to
           | manipulate. For instance, one needs to dispatch on types to
           | determine whether an element at some index is an attribute
           | map or a child. By using the former data structure, one
           | simply looks up the :attrs or :content key. Additionally, the
           | map structure is easier to depth-first search; it's a one-
           | liner with the tree-seq function.
           | 
           | I've written a rudimentary EPUB parser in Clojure and found
           | it easier to work with zippers than any other data structure
           | to e.g. look for <rootfile> elements with a <container>
           | ancestor.
           | 
           | Zippers are available in most programming languages,
           | thankfully, so this advantage is not really unique to Clojure
           | (or another Lisp). However, I will agree that something like
           | sexps (or Hiccup) is more convenient than e.g. JSX, since you
           | are dealing with the native syntax of the language rather
           | than introducing a compilation step and non-standard syntax.
        
             | velcrovan wrote:
             | I have not looked into the use of zippers for this purpose,
             | but I will do so!
             | 
             | Racket has helper libraries like TxExpr
             | (https://docs.racket-lang.org/txexpr/index.html) that make
             | it pretty easy to manipulate S-expressions of this kind.
        
           | zoogeny wrote:
           | This looks like it loses the distinction between attributes
           | and nested tags?
           | 
           | As in, I don't see a difference between `(attr "val")` which
           | expresses an attribute key/value pair and `(thing "world")`
           | which expresses a tag/content relationship. Even if I thought
           | the rule might be "if the first element of the list is a list
           | itself then it should be interpreted as a set of attribute
           | key value pairs" then I would still be ambiguous with:
           | (foo (bar "baz") "content")
           | 
           | which could serialize to either:                   <foo
           | bar="baz">content</foo>
           | 
           | or:                   <foo><bar>baz</bar>content</foo>
           | 
           | In fact, this ambiguity between attributes and children has
           | always been one of the head scratching things for me about
           | XML. Well, the thing I've always disliked the most is
           | namespaces but that is another matter.
        
             | shawn_w wrote:
             | There's no ambiguity. The first element is a symbol that's
             | the name of a tag. If the second element is a list of two
             | element symbol + string lists, it's the attributes. If it's
             | one of the other recognized types, it's part of the
             | contents of the tag.
             | 
             | See a grammar for the representation at
             | https://docs.racket-
             | lang.org/xml/index.html#%28def._%28%28li...
             | 
             | Most Scheme tools for working with XML use a different
             | layout where a list starting with the symbol @ indicates
             | attributes. See https://en.wikipedia.org/wiki/SXML for it.
        
         | weinzierl wrote:
         | _" This, thirty years later, is the best pitch for XML I've
         | read."_
         | 
         | I wish someone would write _" XML - The Good Parts"_.
         | 
         | Others might argue that this is JSON but I'd disagree:
         | 
         | - No comments is a non-starter
         | 
         | - No proper integers
         | 
         | - No date format
         | 
         | - Schema validation is a primitive toy compared what we had for
         | XML
         | 
         | - Lack of allowed trailing commas
         | 
         | YAML ain't better. I hated whitespace handling in XML, it's a
         | miracle how YAML could make it even worse.
         | 
         | XML is from era long past and I certainly don't want to go back
         | there, but it had its good parts and I feel we have not really
         | learned a lot from its mistakes.
         | 
         | In the end maybe it is just that developer ergonomics is
         | largely a matter of taste and no language will ever please
         | everyone.
        
           | jlarocco wrote:
           | It's funny to hear people in the comments here talk about XML
           | in the past tense.
           | 
           | I know it's passe in the web dev world, but in my work we
           | still work with XML all the time. We even have work in our
           | queue to add support for new data sources built on XML
           | (specifically QIF https://qifstandards.org/).
           | 
           | It's fine with me... I've come to like XML. It's nice to have
           | a standard, easy way to do seschemas, validators, processors,
           | queries, etc. It can be overdone and it's not for every use
           | case, but it's pretty good at what it does.
        
         | bambax wrote:
         | I used to do a lot of XSLT coding, by hand, in text editors
         | that weren't proper IDEs, and frankly it wasn't very hard to
         | do.
         | 
         | There's something very zen-like with this language; you put a
         | document in a kind of sieve and out comes a "better" document.
         | It cannot fail; it can be wrong, full of errors, of course
         | (although if you're validating the result against a schema it
         | cannot be very wrong); but it will almost never explode in your
         | face.
         | 
         | And then XSLT work kind of disappeared; I miss it a lot.
        
         | iamthepieman wrote:
         | >XML has always seemed to be a data standard which is intended
         | to be what computers prefer, not people
         | 
         | Interesting take, but I'm always a little hesitant to accept
         | any anthropomorphizing of computer systems.
         | 
         | Isn't it always about what we can reason and extrapolate about
         | what the computer is doing? Obviously computers have no
         | preference so it seems like you're really saying
         | 
         | "XML is a poor abstraction for what it's trying to accomplish"
         | or something like that.
         | 
         | Before jQuery, chrome, and web 2.0, I was building xslt driven
         | web pages that transformed XML in an early nosql doc store into
         | html and it worked quite beautifully and allowed us to skip a
         | lot of schema work that we definitely were ready or
         | knowledgeable enough to do.
         | 
         | EDIT: It was the perfect abstraction and tool for that job.
         | However the application was very niche and I've never found a
         | person or team who did anything similar (and never had the
         | opportunity to do anything similar myself again)
        
           | madkangas wrote:
           | Re xslt based web applications - a team at my employer did
           | the same circa 2004. It worked beautifully except for one
           | issue: inefficiency. The qps that the app could serve was
           | laughable because each page request went through the xslt
           | engine more than once. No amount of tuning could fix this
           | design flaw, and the project was killed.
           | 
           | Names withheld to protect the guilty. :)
        
         | wiremine wrote:
         | > developer ergonomics
         | 
         | That was a huge reason JSON took over.
         | 
         | Another reason was the overall XML ecosystem grew unwieldy and
         | difficult to navigate: XPath, XSLT, SOAP, WSDL, Xpointer,
         | XLink, SOAP, XForms... They all made sense in their own way,
         | but it was difficult to master them all. That complexity, plus
         | the poor ergonomics, is what paved the way for JSON to become
         | preferred.
        
           | tonyedgecombe wrote:
           | I quite liked it when it first came out, I'd been dealing
           | with a ton of bespoke formats up until then. Pretty much
           | every one was ambiguous and painful to deal with. It was a
           | step forward being able to push people towards a standard for
           | document transfer.
           | 
           | I suspect it was SOAP and WSDL that killed it for a lot of
           | people though. That was a typical example of a technical
           | solution looking for a problem and complete overkill for most
           | people.
           | 
           | The whole namespace thing was probably a step too far as
           | well.
        
         | kibwen wrote:
         | _> XML sucks shit for someone who wants to think tersely and
         | code by hand. Seriously, I hate it with a fiery passion._
         | 
         | At the risk of glibly missing the main point of your comment,
         | take a look at KDL. Unlike JSON/TOML/YAML, it features XML-
         | style node semantics. Unlike XML, it's intended to be human-
         | readable and writeable by hand. It has specifications for both
         | a query language and a schema language as well as
         | implementations in a bunch of languages. https://kdl.dev/
        
       | samsk wrote:
       | Nice ! I've a scrapper using XPath/XSLT extensively and 90% of
       | the XPath selectors work like for years without a change. With
       | CSS selectors I've had more problems...
        
         | ebruchez wrote:
         | CSS selectors have spent the last few decades reinventing
         | XPath. XPath introduced right from the beginning the notion of
         | axes, which allow you to navigate down, up, preceding,
         | following, etc. as makes sense. XPath also always had
         | predicates, even in version 1.0. CSS just recently started
         | supporting :has() and :is(), in particular. Eventually, CSS
         | selectors will match XPath's query abilities, although with
         | worse syntax.
        
           | bambax wrote:
           | > _CSS selectors have spent the last few decades reinventing
           | XPath_
           | 
           | YES! This is so true! And ridiculous! It's a mystery why we
           | didn't simply reuse XPath for selectors... it's all in
           | there!!
        
             | masklinn wrote:
             | > It's a mystery why we didn't simply reuse XPath for
             | selectors... it's all in there!!
             | 
             | It's not really a mystery:
             | 
             | > CSS was first proposed by Hakon Wium Lie on 10 October
             | 1994. [...] discussions on public mailing lists and inside
             | World Wide Web Consortium resulted in the first W3C CSS
             | Recommendation (CSS1) being released in 1996
             | 
             | > XPath 1.0 was published in 1999
             | 
             | CSS2 was released before XPath 1.0.
        
           | samsk wrote:
           | The problem with CSS selectors (at least in scrapers) is also
           | that they change relatively often, compared to (html)
           | document structure, thats why XPath last longer. But you are
           | right, CSS selectors compared to 20 years old XPath are realy
           | worse.
        
           | masklinn wrote:
           | On the other hand:
           | 
           | - XPath literally didn't exist when CSS selectors were
           | introduced
           | 
           | - XPath's flexibility makes it a lot more challenging to
           | implement efficiently, even more so when there are thousands
           | of rules which need to be dynamically reevaluated at each
           | document update
           | 
           | - XPath is lacking conveniences dedicated to HTML semantics,
           | and handrolling them in xpath 1.0 was absolutely heinous (go
           | try and implement a class predicate in xpath 1.0 without
           | extensions)
        
       | trympet wrote:
       | I recently had the pleasure of using XSLT after never having seen
       | it before. I used it to transform a huge 130K line XML manifest
       | with MAPI property metadata into C# source code. It was so
       | simple, readable, and intuitive to use.
        
       | mattrighetti wrote:
       | I will definitely try this out!
       | 
       | I have a service that extracts <meta> tags in webpages and to do
       | that I'm currently using (and need) three different dependencies:
       | html5ever, markup5ever_rcdom, markup5ever. I don't like those to
       | be honest, the documentation is quite bad and it was difficult to
       | understand how I should have used the libraries to achieve such a
       | simple task.
       | 
       | XPath on the other hand makes this extremely easy in comparison,
       | I wonder how this will perform compared to my current solution.
        
         | faassen wrote:
         | Thanks!
         | 
         | Unfortunately at this point there's no HTML parser frontend for
         | Xee (and its underlying library Xot) yet (HTML 5 parser
         | serialization is supported at least in code). It shouldn't be
         | too hard to add at least HTML 5 support using something like
         | html5ever.
        
       | infogulch wrote:
       | There are many humongous XML sources. E.g. the Wikipedia archive
       | is 42GB of uncompressed text. Holding a fully parsed
       | representation of it in memory would take even more, perhaps even
       | >100GB which immediately puts this size of document out of reach.
       | 
       | The obvious solution is streaming, but streaming appears to not
       | be supported, though is listed under Challenging Future Ideas:
       | https://github.com/Paligo/xee/blob/main/ideas.md
       | 
       | How hard is it to implement XML/XSLT/XPATH streaming?
        
         | wongarsu wrote:
         | 100GB doesn't sound _that_ out of reach. It 's expensive in a
         | laptop, but in a desktop that's about $300 of RAM and our
         | supported by many consumer mainboards. Hetzner will rent me a
         | dedicated server with that amount of ram for $61/month.
         | 
         | If the payloads in question are in that range, the time spent
         | to support streaming doesn't feel justified compared to just
         | using a machine with more memory. Maybe reducing the size of
         | the parsed representation would be worth it though, since that
         | benefits nearly every use case
        
           | infogulch wrote:
           | I just pulled the 100GB number out of nowhere, I have no idea
           | how much overhead parsed xml consumes, it could be less or it
           | could be more than 2.5x (it probably depends on the specific
           | document in question).
           | 
           | In any case _I_ don 't have $1500 to blow on a new computer
           | with 100GB of ram in the unsubstantiated hope that it happens
           | to fit, just so I can play with the Wikipedia data dump. And
           | I don't think that's a reasonable floor for every person that
           | wants to mess with big xml files.
        
             | hyhjtgh wrote:
             | Xml and textual formats in general are ill suited to such
             | large documents. Step 1 should really be to convert and/or
             | split the file into smaller parts.
        
             | philipkglass wrote:
             | In the case of Wikipedia dumps there is an easy work-
             | around. The XML dump starts with a small header and
             | "<siteinfo>" section. Then it's just millions of "page"
             | documents for the Wiki pages.
             | 
             | You can read the document as a streaming text source and
             | split it into chunks based on matching pairs of "<page>"
             | and "</page>" with a simple state machine. Then you can
             | stream those single-page documents to an XML parser without
             | worrying about document size. This of course doesn't apply
             | in the general case where you are processing arbitrary huge
             | XML documents.
             | 
             | I have processed Wikipedia many times with less than 8 GB
             | of RAM.
        
         | 01HNNWZ0MV43FF wrote:
         | Is that all in one big document?
        
           | magicalhippo wrote:
           | We regularly parse ~1GB XML documents at work, and got
           | laughed at by someone I know who worked with bulk invoices
           | when I called it a large XML file.
           | 
           | Not sure how common 100GB files are but I can certainly image
           | that being the norm in certain niches.
        
         | dleeftink wrote:
         | StackExchange also, not necessarily streamable but records are
         | newline delimited which makes it easier to sample (at least the
         | last time I worked with the Data Dump).
        
         | faassen wrote:
         | Anything could be supported with sufficient effort, but
         | streaming hasn't been my priority so far and I haven't explored
         | it in detail. I want to get XSLT 3.0 working properly first.
         | 
         | There's a potential alternative to streaming, though - succinct
         | storage of XML in memory:
         | 
         | https://blog.startifact.com/posts/succinct/
         | 
         | I've built a succinct XML library named Xoz (not integrated
         | into Xee yet):
         | 
         | https://github.com/Paligo/xoz
         | 
         | The parsed in memory overhead goes down to 20% of the original
         | XML text in my small experiments.
         | 
         | There's a lot of questions on how this functions in the real
         | world, but this library also has very interesting properties
         | like "jump to the descendant with this tag without going
         | through intermediaries".
        
           | infogulch wrote:
           | 0.2x of the original size would certainly make big documents
           | more accessible. I've heard of succinct storage, but not in
           | the context of xml before, thanks for sharing!
        
             | faassen wrote:
             | I myself actually had no idea succinct data structures
             | existed until last December , but then I found a paper that
             | used them in the context of XML. Just to be clear: it's
             | 120% of the original size; as it stands this library still
             | uses more memory than the original document, just not a
             | _lot_ of overhead. Normal tree libraries, even if the tree
             | is immutable, take a parent pointer, and a first child
             | pointer and next and previous sibling pointers per node.
             | Even though some nodes can be stored more compactly it does
             | add up.
             | 
             | I suspect with the right FM-Index Xoz might be able to
             | store huge documents in a smaller size than the original,
             | but that's an experiment for the future.
        
               | lambda wrote:
               | Would you be able to parse it in a streaming fashion and
               | just store the structure of the document in memory, with
               | just offsets for all of the string locations, and then
               | re-read those from disk as needed?
               | 
               | With modern SSDs and disk cache, that's likely enough to
               | be plenty performant without having to store the whole
               | document in memory at once.
        
           | bambax wrote:
           | > _I want to get XSLT 3.0 working properly first_
           | 
           | May I ask why? I used to do a lot of XSLT in 2007-2012 and
           | stuck with XSLT 2.0. I don't know what's in 3.0 as I've never
           | actually tried it but I never felt there was some feature
           | missing from 2.0 that prevented me to do something.
           | 
           | As for streaming, an intermediary step would be the ability
           | to cut up a big XML file in smaller ones. A big XML document
           | is almost always the concatenation of smaller files (that's
           | certainly the case for Wikipedia for example). If one can
           | output smaller files, transform each of them, and then
           | reconstruct the initial big file without ever loading it in
           | full in memory, that should cover a huge proportion of
           | "streaming" needs.
        
             | faassen wrote:
             | XSLT has been a goal of this project from the start, as my
             | customer uses it. XSLT 3.0 simply as that's the latest
             | specification. What tooling do you use for XSLT 2.0?
        
               | bambax wrote:
               | Saxon's free version, which IIRC only implemented 2.0.
        
         | jerf wrote:
         | "How hard is it to implement XML/XSLT/XPATH streaming?"
         | 
         | It's actually quite annoying on the general case. It is
         | completely possible to write an XPath expression that says to
         | match a super early tag on an arbitrarily-distant further tag.
         | 
         | In another post in this thread I mention how I think it's
         | better to think of it as a multicursor, and this is part of
         | why. XPath doesn't limit itself to just "descending", you can
         | freely cursor up and down in the document as you write your
         | expression.
         | 
         | So it is easy to write expressions where you literally can't
         | match the first tag, or be sure you shouldn't return it, until
         | the whole document has been parsed.
        
           | riedel wrote:
           | I think from a grammar side, XPath had made some decisions
           | that make it really hard to generally implement it
           | efficiently. About 10 years ago I was looking into binary XML
           | systems and compiling stuff down for embedded systems
           | realizing that it is really hard to e.g. create efficient
           | transducers (in/out pushdown automata) for XSLT due to
           | complexity of XPath.
        
       | jchw wrote:
       | I wonder if this could perhaps some day be used in Wine, for the
       | MSXML implementations. Maybe not, since those implementations
       | need to be bug-compatible where applications depend on said bugs;
       | but the current implementation(s) are also not fantastic. I
       | believe it is still using libxml2.
       | 
       | (Aside: A long time ago, I had written an alternate XPath 1.1
       | implementation for Wine during GSoC, but rather shamefully, I
       | never actually got it merged. Life became very hectic for me
       | during that time period and I never really looped back to it.
       | Still feel pretty bad about it all these years later.)
        
       | infogulch wrote:
       | > I was at XML Prague, an XML conference
       | 
       | There's an _XML conference_?!
        
         | bartschuller wrote:
         | There are at least 5:
         | 
         | https://www.xmlprague.cz/ https://www.balisage.net/
         | https://declarative.amsterdam/ https://markupuk.org/
         | https://xmlsummerschool.org/
        
       | dev_l1x_be wrote:
       | XML is the what OOP is for programming languages. Often
       | overcomplicated, hard to follow, full of footguns.
        
       | notfed wrote:
       | Does it preserve whitespace? Something that I always found
       | asinine about XSLT is that it wipes out whitespace when
       | transforming. Imagine you have thousands of corporate XML files
       | in source control, and you want to tranform them all, performing
       | some simple mutation. XSLT claims to be fit for this job, but in
       | practice your diff is going to be full of unintentional
       | whitespace mangling.
        
         | ebruchez wrote:
         | XSLT will perform the transformations that you instruct it to
         | do. It does not wipe out whitespace just on its own. Do you
         | mean that you'd like facilities to nicely reindent the output?
        
       | 1shooner wrote:
       | I miss the declarative purity of XSLT as an HTML templating
       | layer. I'd love to know if there is a similar system for more
       | popular/current web stack.
        
       | yxhuvud wrote:
       | I hope this will be packaged into shared libraries at some point
       | so that languages that isn't rust will get access to it.
        
       | threecheese wrote:
       | This is great, I've been looking for performant and safe XML
       | processing to replace IBM stuff (websphere/datapower) that we
       | really only keep around for hw accelerated payload processing. At
       | our scale, lxml and others + BYO gateway tech has a similar run
       | cost even considering IBM licensing. I hate running their crap,
       | which requires k8s at a version that's some hair-thin slice above
       | the minimum supported EKS version, it's almost like they want us
       | to live in 24/7 fear of being OOS.
        
       | ianand wrote:
       | Fun fact: A decade ago the designer of HAML and Sass created a
       | modern alternative to XSLT.
       | https://en.wikipedia.org/wiki/Tritium_(programming_language)
        
       | smitty1e wrote:
       | > XML is now niche technology, but it's a bigger niche than you
       | might think, and it's not going to go away any time soon.
       | 
       | When you consider that .docx, .pptx, and .xlsx files are zipped
       | XML archives, "niche" seems a misnomer.
        
       | blacklion wrote:
       | Does XSLT still used in a new projects? I have impression, that
       | it was not popular even when XML was.
       | 
       | For example, apache HTTPD never has official module to serve XML
       | via XSLT transformation.
       | 
       | And XSL:FO looks even more obscure.
        
         | kondro wrote:
         | There are a lot of APIs out there that are still XML-based,
         | especially from enterprise suppliers.
         | 
         | Equifax and Experian's APis immediately come to mind as
         | documents that generate complex results that people often want
         | to turn into some type of visual representation with XSLT.
        
         | int_19h wrote:
         | XSL:FO is dead for all practical purposes.
         | 
         | XSLT was not popular for its _original intended application_ -
         | which is to say, serving XML data from web servers and
         | translating it to HTML (or XSL:FO, or ...) on the client as
         | needed. However, it was used plenty for XML processing outside
         | of that particular niche.
         | 
         | New projects these days rarely have to process complicated XML
         | to begin with. But when you do, I'd say XSLT (or perhaps better
         | yet, XQuery) is a very useful tool to have in your toolbox.
        
       | riedel wrote:
       | Love to see stuff outside the Java space since I really like
       | thedoing stuff in XSLT. Question: Does this work on a textual XML
       | representation or can you plug in different XML readers? I have
       | had really great fun in the past using http://www.ananas.org/xi/
       | transforming arbitrarily for formated files using XSLT. Also it
       | is today really important that XML Reader has error correction
       | capabilities, since lots of tools don't write well-formed XML,
       | which often is a showstopper for employing to transforms from my
       | experience.
        
       | tracnar wrote:
       | Nice! I tried using XQuery (superset of XPath 3) for a while
       | through the BaseX implementation. It's pretty nice, but you have
       | to face XML problems like namespaces, document order, attributes
       | vs nodes, you don't know if you can have 0, 1 or more nodes, etc.
       | Something I wish was more readily available would be to run XPath
       | against JSON, yaml, etc. It's a nicer language than say jq, but
       | its ties to XML sometimes make it hard to transfer.
       | 
       | Another pain point with XML is the lack of inline schema, so the
       | languages around like XPath have to work with arbitrary
       | structures unlike say JSON where you at least have basic
       | primitives like map/dict, numbers, bool, etc
        
       | shadowtree wrote:
       | Throwback shoutout to Steve Muench and his genius method of
       | grouping elements in XSLT 1.0.
       | 
       | So good it has its own Wikipedia page!
       | 
       | https://en.wikipedia.org/wiki/XSLT/Muenchian_grouping
       | 
       | I mean, talk about hacker cred.
        
       ___________________________________________________________________
       (page generated 2025-03-28 23:00 UTC)