[HN Gopher] Tips on adding JSON output to your command line util...
       ___________________________________________________________________
        
       Tips on adding JSON output to your command line utility. (2021)
        
       Author : fanf2
       Score  : 77 points
       Date   : 2024-04-20 16:42 UTC (6 hours ago)
        
 (HTM) web link (blog.kellybrazil.com)
 (TXT) w3m dump (blog.kellybrazil.com)
        
       | enriquto wrote:
       | But why? Just for people to gron it back into usable form?
       | 
       | Unix tools work best with line-based data. Using some
       | "structured" monstrosity like xml or json forces all other tools
       | to deal with this particular format, thus breaking the
       | orthogonality between programs.
        
         | simonw wrote:
         | Newline-delimited JSON is so much more useful to me than weird
         | Unix line-formatted output that I have to parse with pattern
         | matching or regular expressions.
         | 
         | It's basically a line-based format that can represent all of
         | the JSON types and includes support for nested data structures
         | where necessary. What's not to like?
        
           | qwertox wrote:
           | It's also supported by simdjson [0] (which has a lot of
           | language bindings [1]):
           | 
           | > Multithreaded processing of gigantic Newline-Delimited JSON
           | (ndjson) and related formats at 3.5 GB/s
           | 
           | [0] https://simdjson.org/
           | 
           | [1] https://github.com/simdjson/simdjson?tab=readme-ov-
           | file#bind...
        
           | bsdetector wrote:
           | JSON is not immediately usable by and is cumbersome to parse
           | correctly in a shell.
           | 
           | A simple line-based shell variable name=value format works
           | unreasonably well. For example:                   # ls
           | --shell-var ./thefile         dir="/home/user" file="thefile"
           | size=1234 ...         # eval $(ls --shell-var ./thefile);
           | echo $size         1234
           | 
           | If this had been in shells and cmdline tools since the
           | beginning it would have saved so much work, and the security
           | problems could have been dealt with by an eval that only set
           | variables, adding a prefix/scope to variables, and so on.
           | 
           | Unfortunately it's too late for this and today you'll be
           | using a pipeline to make the json output shell friendly or
           | use some substring hacks that probably work most of the time.
        
             | simonw wrote:
             | I use jq - which ChatGPT knows inside out, so I can
             | generally get exactly what I want from it with a single
             | prompt.
        
             | kitd wrote:
             | That's not your original request though, to use line-based
             | data. It seems you're determined not to use jq but if
             | anything, json output | jq is more the unix way than piping
             | everything through shell vars.
        
               | mlhpdx wrote:
               | And, the author isn't suggesting only having JSON output,
               | but adding it as an option for those of use that would
               | make use of it. The plain text should remain as well (and
               | has to or many, many things would break).
               | 
               | On a separate point, I find the JSON much easier to
               | reason about. The wall of text output doesn't work for my
               | brain - I just can't see it all. Structuring/nesting with
               | clear delineations makes it far easier for me to grok.
        
               | bsdetector wrote:
               | > That's not your original request though, to use line-
               | based data.
               | 
               | It wasn't my request and OP (not me) said "line-based
               | data" is best. The comment I replied to said "Newline-
               | delimited JSON ... a line-based format".
               | 
               | If the only objection you have is "but that's line-
               | based!" then you're in a completely different
               | conversation.
               | 
               | > if anything, json output | jq is more the unix way than
               | piping everything through shell vars.
               | 
               | The unix way is line-based. The comment I replied to is
               | talking about line-based output. Line-based output is the
               | only structure for data universal to unix cmdline tools -
               | even tab/space isn't universal; sending structured non-
               | line-delimited data to a program to unpack it is the
               | least unix-like way to do it.
               | 
               | Also there's no pipe in the shell-variable output scheme
               | I described, whereas "json | jq" is a shell pipeline.
        
             | starttoaster wrote:
             | That's great for key=value data, but more complex data
             | structures don't work so well in that format, JSON does.
             | "Why would you need to represent data as a complex data
             | structure?" Sometimes attributes are owned by a specific
             | entity, and that entity might own multiple attributes. It
             | might even own other sub-entities. JSON represents that.
             | Key=value does not.
        
               | bsdetector wrote:
               | JSON is literally key=value, just nested. Which you can
               | do with shell variables.
               | 
               | The question was "What's not to like [about JSON output
               | from cmdline tools]?" and the answer is that it's
               | cumbersome to read in a shell and all but requires
               | another pipeline stage.
               | 
               | I didn't even recommend shell variable output and made it
               | clear this isn't today a reasonable solution so I'm not
               | sure where this hostility in the replies comes from, but
               | I assume from recognition that it's a more practical
               | solution to reading data within a shell but not wanting
               | that to be so.
        
               | starttoaster wrote:
               | > JSON is literally key=value, just nested.
               | 
               | The nature of being nested, and also containing
               | structures like lists, maps, etc. All of which makes it
               | more complicated than key=value.
               | 
               | > The question was "What's not to like [about JSON output
               | from cmdline tools]?" and the answer is that it's
               | cumbersome to read in a shell and all but requires
               | another pipeline stage.
               | 
               | It depends on the intended use for your shell program. If
               | you intend the CLI tool to be used in CI pipelines (eg.
               | your CLI tool's output is being read by an automated
               | process on a computer) and the data it outputs is more
               | complicated than a simple key=value, JSON is great for
               | that. Your CI program can pipe to jq. You as a human can
               | pipe to jq, though I agree it's somewhat less desirable.
               | Though just piping to jq without any arguments pretty
               | prints it for you which also makes it fairly readable for
               | humans.
               | 
               | > so I'm not sure where this hostility in the replies
               | comes from
               | 
               | You're reading into hostility where there isn't any.
        
               | bsdetector wrote:
               | > The nature of being nested, and also containing
               | structures like lists, maps, etc. All of which makes it
               | more complicated than key=value.
               | 
               | These are javascript objects, which are key-value. A list
               | array is just keyed by a number instead of a string.
               | They're functionally exactly the same as name=value
               | except JSON is parsed depth-first whereas shell variables
               | are breadth-first parsing (which is way better from
               | shells).
               | 
               | Do you have an example of a CLI tool - intended for human
               | use - that has output so complicated it can't be easily
               | mapped to name=value? I don't think there is one, and
               | it's certainly not common.
               | 
               | > You're reading into hostility where there isn't any.
               | 
               | I think "it seems you're determined not to use jq" is
               | pretty hostile since I made no intimation of that at all.
        
               | starttoaster wrote:
               | > I think "it seems you're determined not to use jq" is
               | pretty hostile since I made no intimation of that at all.
               | 
               | Well, I didn't say that, so I don't know what that other
               | person's feelings or intentions are, to be fair. I
               | personally have no feeling of hostility towards you just
               | because we (apparently) disagree on the usefulness of
               | JSON to represent complex data types, or at least
               | disagree on how often human-usable CLI tools output
               | complex data. But to answer:
               | 
               | > Do you have an example of a CLI tool - intended for
               | human use - that has output so complicated it can't be
               | easily mapped to name=value? I don't think there is one,
               | and it's certainly not common.
               | 
               | kubectl. Which to be fair defaults to output to a table-
               | like format. Though it gets all that data in the table
               | from JSON for you. smartctl is another one, which also
               | defaults to table format. To be honest, I could go on and
               | on if the only qualifier is a CLI tool that emits complex
               | data, not suited for just key=value.
               | 
               | > These are javascript objects, which are key-value. A
               | list array is just keyed by a number instead of a string.
               | They're functionally exactly the same as name=value
               | except JSON is parsed depth-first whereas shell variables
               | are breadth-first parsing (which is way better from
               | shells).
               | 
               | As mentioned before, just because you can compare JSON to
               | key=value, does not mean it's as simple as key=value.
               | It's a data serialization language that builds well on
               | top of simple key=value formats. You're welcome to enjoy
               | other data serialization languages, like yaml, HCL, or
               | PKL. But none of those are simple key=value formats
               | either. They built the ability to represent more complex
               | structures on top of that.
               | 
               | A data serialization language allows the end-user to
               | specify how they would like to use that data, while
               | allowing them to use standard parsing tools like jq.
               | Cramming complex data into a value string in a key=value
               | format gives end users the same allowance to use that
               | data however they want, while also giving them a chore to
               | handle parsing it in custom ways tailored to just your
               | CLI application, likely in ways that would seem far more
               | brittle than parsing a defined language with well defined
               | constraints. That doesn't sound like great UX to me. But
               | to be fair to you, you're not saying that you wish to use
               | key=value to represent complex data. Rather, you're
               | saying there's a general lack of complex data to be
               | found, to which I also disagree with.
        
               | bsdetector wrote:
               | > But none of those are simple key=value formats either.
               | 
               | What is the difference between:                   {
               | object: { name: value }}         { object: "{ name: value
               | }"}         object="name=value"
               | 
               | There's zero difference between any of them except how
               | you parse and process the data.
               | 
               | > kubectl. Which to be fair defaults to output to a
               | table-like format.
               | 
               | With line-based shell-variable output you have a line of
               | variables and you have blocks of lines separated by an
               | empty line (like an HTTP 1 header).
               | 
               | This can easily map to any table, two dimensions, or two
               | levels of data structure without even quoting
               | subvariables like in the example above. So, no, kubectl
               | is not an example at least not how you've described it.
        
               | starttoaster wrote:
               | > What is the difference between .. There's zero
               | difference between any of them except how you parse and
               | process the data.
               | 
               | Answered in the previous message... "A data serialization
               | language allows the end-user to specify how they would
               | like to use that data, while allowing them to use
               | standard parsing tools like jq. Cramming complex data
               | into a value string in a key=value format gives end users
               | the same allowance to use that data however they want,
               | while also giving them a chore to handle parsing it in
               | custom ways tailored to just your CLI application, likely
               | in ways that would seem far more brittle than parsing a
               | defined language with well defined constraints."
               | 
               | > With line-based shell-variable output you have a line
               | of variables and you have blocks of lines separated by an
               | empty line (like an HTTP 1 header)...
               | 
               | I would not choose to write application logic that
               | foregoes defined data serialization languages for parsing
               | barely structured strings the way you seem to prefer. But
               | you go about it the way you prefer, I guess. This whole
               | discussion leaves a lot of room for personal opinions. I
               | think we both agree that the other person's opinion here
               | is subjectively the more annoying route to deal with. But
               | that's the way life is sometimes.
        
         | eternityforest wrote:
         | If you're using UNIX tools to parse it, it sucks, but generally
         | if I'm reading the output of a command, and the command is more
         | than one word, I'm doing the whole thing in Python.
         | 
         | That's real programming, and for that I want type checkers and
         | debuggers and modern syntax and all that. And the performance
         | is often faster because you're not spinning up subprocesses for
         | each command.
        
         | pluto_modadic wrote:
         | ./some_command | jq '.memory_use'
         | 
         | vs:
         | 
         | okay, run it once with head (assuming we have a --dry
         | option...) ah, that's the column. okay cut -d"," -f0, ah whoops
         | it's starting at 1. ah, damn, there's a weird comma in the
         | name/quote. oh weird, that one's null, ah heck.
         | 
         | schemas are cool. JSON extends and builds on the UNIX idea of
         | having things pipe and plumb well together.
        
           | alerighi wrote:
           | Plain text is more easy to reason about, because we are used
           | to process text. A good textual output, that is records
           | delimited by spaces, tabs or a delimiter, to me is all it's
           | needed, for most applications.
           | 
           | An object structure it's much more complex to use. For
           | example an output that is a set of records can be easily
           | imported in an Excel sheet, in an SQL database, processed
           | line by line, without issues. Processing JSON is not straight
           | forward, not all programs support JSON.
           | 
           | Finally JSON can't be processed as a stream, meaning that
           | tools like head, tail, etc. doesn't work on JSON, you have to
           | read it all in memory, or use JSON lines, that is not a
           | standard format, that not all parsers support natively, etc.
           | 
           | JSON is good if for integrating the program inside other
           | programs (as a subprocess), so having an option to
           | input/output JSON in a program is useful, but to me it's not
           | as useful for interactive shell usage. I prefer to use UNIX
           | tools such as grep, cut, head, tail, etc.
        
           | mlhpdx wrote:
           | But as the article points out, the advice to use unbuffered
           | JSON lines for commands that are line oriented is well given.
           | Not doing that can really make life sad.
        
           | dylan604 wrote:
           | you still need to know the schema of the JSON which could
           | also use a dry run as well. not really sure how just because
           | it's JSON solves that in your mind
        
             | Jcowell wrote:
             | Orient the idea be that because it's json there's a schema
             | somewhere the end user can refer to?
        
               | dylan604 wrote:
               | as if the output of the other command also isn't
               | available?
        
         | placatedmayhem wrote:
         | Line-oriented formats, like most traditional Unix-style tools,
         | are for human consumption. JSON is bad at that, thus gron.
         | 
         | On the other hand, structured output formats, like JSON, make
         | it easier to consume with other programs. Standard formats have
         | readily-available and commonly used libraries, whereas line
         | parsing tends to be one-off for every program. Whether JSON is
         | the best format for this is certainly debatable, but it is
         | quite ubiquitous, which is a huge advantage. I doubt many folks
         | would propose XML as a general recommendation.
         | 
         | Tools should have both options on their path to maturity --
         | both human-consumable and computer-consumable output format
         | options.
        
         | janderland wrote:
         | I'd rather use JMES Path (or jq even tho it's messier) to
         | restructure my data rather than some mixture of awk, sed, cut,
         | etc.
         | 
         | Line output often needs restructuring in my experience, though
         | JSON will always need it.
        
         | Spivak wrote:
         | Unix tools on newline delineated list like objects work great,
         | but try them on dict like objects and it becomes clunky really
         | fast. Parsing `ip` output is a good example where the data
         | naturally lends itself to iface:attrs. Plucking the value you
         | want by `.[.name | startswith('eth')].addr` is way easier than
         | pulling this out with grep/awk.
        
           | BenjiWiebe wrote:
           | You probably already know this but just in case: You can
           | specify an interface when using ip, e.g. ip addr show dev
           | enp2s0
        
         | __MatrixMan__ wrote:
         | The caller can just specify the output format that they want,
         | and ever since I switched to nushell I pretty much always want
         | json.
        
         | keybored wrote:
         | What orthogonality? "Line-based data" that varies randomly from
         | tool to tool? Using json between programs is perfectly
         | orthogonal.
        
         | IshKebab wrote:
         | Nonsense. Using a structured format means all tools are using
         | _the same_ format, instead of every tool making up their own
         | (usually broken) system.
         | 
         | Look at how many flags `ls` has to feed filenames into other
         | programs without breaking.
         | 
         | Also, using a proper format like JSON makes it actually robust.
         | Most ad hoc pipelines break if you so much as put a space in a
         | filename.
        
       | pluto_modadic wrote:
       | JSON is so much more usable sometimes. Especially when things
       | don't neatly parse into cut -d splits!
        
       | JohnMakin wrote:
       | Yes - and this is mostly because the jq parser tool is so useful.
       | yq is similarly great.
        
         | rwmj wrote:
         | This is really the only reason, because JSON as a format is
         | otherwise terrible. No good way to represent 64 bit numbers,
         | very limited types and no comments.
        
       | cynicalsecurity wrote:
       | This basically turns command line tolls into APIs. Crazy, in a
       | good way.
        
         | inetknght wrote:
         | Not really that crazy at all. And yes it's definitely good.
        
       | moregrist wrote:
       | I'm a big fan of tools producing json output, but I couldn't help
       | but chuckle a bit at:
       | 
       | > This allows easier parsing in scripts by using JSON parsing
       | tools like jq, jello, jp, etc. without arcane awk, sed, cut, tr,
       | reverse, etc. incantations.
       | 
       | While I love jq, I find its syntax to be utterly baroque. Far
       | more than awk and sed (let alone simple tools like cut and tr).
       | To the point where I end up looking at the jq man page to remind
       | myself how to do even seemingly simple things at times. And often
       | jq is just the beginning of a pipeline to get hierarchical data
       | into columnar format so I can hit it with awk and sed.
       | 
       | But maybe this is all about familiarity; I have years of
       | experience cobbling together shell tools and reach for jq
       | comparatively less. I suspect the author may be the opposite.
        
         | Mordisquitos wrote:
         | I would argue it's not just about familiarity. I had the same
         | reaction to that quote and feel the same as you do, and yet I
         | have been using jq _before_ I first learned and started using
         | awk.
         | 
         | For a long time I had been familiar with jq, but every time I
         | had to use it for any non-trivial task after not touching it
         | for a while I had to go to the docs and really scratch my head.
         | Meanwhile, I had no idea about awk, and for ages I would just
         | see those awk one-liners on Stack Overflow as weird arcane
         | alternatives to "simply" piping sed, grep and tr over bash.
         | 
         | However, one day I finally had the time and the reason to have
         | a look at AWK Programming for a specific use that I needed, and
         | it immediately clicked. Awk makes sense and now I can easily
         | fall back to it every time, even if I haven't used it for ages.
         | At most, all I need to look up is the order of parameters in
         | the gawk regex functions.
         | 
         | And yet, time and time again I am baffled by jq, however much I
         | have used it in the past. Whenever I need to parse a JSON I'm
         | happy enough if I can use jq to get it to a good enough halfway
         | point so that I can pipe it to awk to do the actual work.
        
         | DangitBobby wrote:
         | ChatGPT flawlessly constructs jq commands for whatever banal
         | task you have. Game changer.
        
           | simonw wrote:
           | Completely agree. I use jq multiple times a day now, purely
           | because GPT-4 or Claude Opus mean I never have to think about
           | how to express anything in it.
        
         | candiddevmike wrote:
         | I output JSON and embed a jq-like arg in in all of my projects
         | (https://rotx.dev/docs/references/cli#jq). Makes it really easy
         | to slice and dice output.
        
         | arp242 wrote:
         | It's not you; parsing line-based text is fundamentally simpler
         | than parsing structured data like JSON, especially nested JSON.
         | It's just a lot easier to reason about, and a lot easier to
         | split up in smaller chunks, too. Who hasn't ended up with
         | something like:                 cmd | sed | grep | sed | grep
         | 
         | Each bit gets you closer what you want in small easy-to-
         | understand steps, making it fairly easy to develop. And for
         | things you only do once or twice, who cares as long as it
         | works, right? I find that's a lot harder to do with jq. I feel
         | this is a problem waiting for a better solution (and gron isn't
         | it).
        
           | fuzztester wrote:
           | >Who hasn't ended up with something like:
           | 
           | > cmd | sed | grep | sed | grep
           | 
           | And the decorate-sort-undecorate pattern:
           | 
           | https://en.m.wikipedia.org/wiki/Decorate-sort-undecorate
           | 
           | and the related Schwartzian transform:
           | 
           | https://en.m.wikipedia.org/wiki/Schwartzian_transform
           | 
           | (The former redirects to the latter.)
        
       | ramses0 wrote:
       | I'll my $0.02: when possible, assume/emit "arrays of records".
       | 
       | eg: `ls ...` should be an an array of homogenous entities
       | (objects) with usually simple key=>value relationships.
       | 
       | Slight variations such as "field contains an array of strings"
       | (eg: tags, groups, file path hierarchy, temperature per each CPU)
       | is also totally fine.
       | 
       | Seeing the wildly differing record types and deeply nested
       | "stuff" in some of the examples weirds me out.
       | {         "type_cpu": [ ... ],          "type_ram": [ ... ],
       | "type_network": [ ... ],         ...         }
       | 
       | ...obviously, in some cases a graph/hierarchy is very useful to
       | be able to traverse, but emitting records and "coalescing" the
       | path into a record-field makes simple operations simple.
       | 
       | Example: `hwstats | jq '.type_cpu | .[] | select( .temp > 40 ) |
       | .hierarchy_field`
       | 
       | Don't make me go like `.cpu.bank[0].chip[3].temp` or whatever...
       | just give me "all the CPU's" and then give me a good URI(!!) to
       | describe it or search for it later.
       | 
       | Rationale: coalescing and traversing weird object paths is tough
       | to do dynamically via current shell DSLs, and the URI concept as
       | an ID excellent application.
        
         | jiehong wrote:
         | An array of records, or just 1 record per line, also works
         | nicely with jq.
         | 
         | For that matter, it can even be imported in like duckdb and
         | handled like a table :)
        
       | dec0dedab0de wrote:
       | I prefer jsonl from command line tools, just cause it's nicer to
       | log, and parse using standard tools. Plus it's like a two second
       | function if you're doing it in a real programming language. I
       | think it's the best of both worlds.
       | 
       | They say to use JSON lines for streaming, but even if you're not
       | it's still very nice to be able to do _my_command | grep xyz_ and
       | still have it be parseable
        
       | isatty wrote:
       | Absolutely not a fan of JSON. It's not typed, and I definitely
       | don't want to know about each tools schema (and it'll be worse
       | without the definition).
       | 
       | There is nothing wrong with Linux tools, they do one thing and do
       | it well.
        
         | throw10920 wrote:
         | This is self-contradictory. JSON is _more_ typed and _more_
         | predictable than what comes out of UNIX tools.
        
       | xomodo wrote:
       | I don't mind json outputs, but pls keep cli options as low
       | letters, eg.: 'kubectl get pods -o json'.
        
         | echoangle wrote:
         | Why is that important?
        
           | al_borland wrote:
           | Mixing case in a case sensitive environment is unnecessarily
           | confusing and annoying.
        
           | Karellen wrote:
           | > The file systems of Unix machines all have the same general
           | structure. [...] Note the obsessive use of abbreviations and
           | avoidance of capital letters; this is a system invented by
           | people to whom repetitive stress disorder is what black lung
           | is to miners. Long names get worn down to three-letter
           | nubbins, like stones smoothed by a river.
           | 
           | -- Neal Stephenson, _In the Beginning Was the Command Line_
           | (Ch. 14), 1999
           | 
           | https://steve-
           | parker.org/articles/others/stephenson/oral.sht...
        
       | matheusmoreira wrote:
       | Parsing is hard and this definitely beats having to write ad-hoc
       | parsers for stuff. I worry we're essentially reinventing a less
       | verbose form of SQL though.
        
         | kunley wrote:
         | No, parsing is easy...
         | 
         | I recommend the "Crafting Interpreters" book.
        
       | jiehong wrote:
       | On linux systems, I was thinking that it would be great if there
       | was a dbus wrapper that outputs all data as json by default.
       | 
       | As many things are exposed on dbus (like systemd units, timers,
       | services, etc), and busctl [0] has a --json output. We could have
       | a dbus-services, or dbus-timers, etc to retrieve that on dbus and
       | emit json. If most things were on dbus, one could have access to
       | almost everything on the system as json directly, while talking
       | to the underlying implementation of most cli by default.
       | 
       | I suppose something similar could be done on MacOS with launchctl
       | and others (I don't know much about MacOS, though).
       | 
       | [0]: https://www.man7.org/linux/man-pages/man1/busctl.1.html
        
       ___________________________________________________________________
       (page generated 2024-04-20 23:00 UTC)