[HN Gopher] Tips on adding JSON output to your command line util...
___________________________________________________________________
Tips on adding JSON output to your command line utility. (2021)
Author : fanf2
Score : 77 points
Date : 2024-04-20 16:42 UTC (6 hours ago)
(HTM) web link (blog.kellybrazil.com)
(TXT) w3m dump (blog.kellybrazil.com)
| enriquto wrote:
| But why? Just for people to gron it back into usable form?
|
| Unix tools work best with line-based data. Using some
| "structured" monstrosity like xml or json forces all other tools
| to deal with this particular format, thus breaking the
| orthogonality between programs.
| simonw wrote:
| Newline-delimited JSON is so much more useful to me than weird
| Unix line-formatted output that I have to parse with pattern
| matching or regular expressions.
|
| It's basically a line-based format that can represent all of
| the JSON types and includes support for nested data structures
| where necessary. What's not to like?
| qwertox wrote:
| It's also supported by simdjson [0] (which has a lot of
| language bindings [1]):
|
| > Multithreaded processing of gigantic Newline-Delimited JSON
| (ndjson) and related formats at 3.5 GB/s
|
| [0] https://simdjson.org/
|
| [1] https://github.com/simdjson/simdjson?tab=readme-ov-
| file#bind...
| bsdetector wrote:
| JSON is not immediately usable by and is cumbersome to parse
| correctly in a shell.
|
| A simple line-based shell variable name=value format works
| unreasonably well. For example: # ls
| --shell-var ./thefile dir="/home/user" file="thefile"
| size=1234 ... # eval $(ls --shell-var ./thefile);
| echo $size 1234
|
| If this had been in shells and cmdline tools since the
| beginning it would have saved so much work, and the security
| problems could have been dealt with by an eval that only set
| variables, adding a prefix/scope to variables, and so on.
|
| Unfortunately it's too late for this and today you'll be
| using a pipeline to make the json output shell friendly or
| use some substring hacks that probably work most of the time.
| simonw wrote:
| I use jq - which ChatGPT knows inside out, so I can
| generally get exactly what I want from it with a single
| prompt.
| kitd wrote:
| That's not your original request though, to use line-based
| data. It seems you're determined not to use jq but if
| anything, json output | jq is more the unix way than piping
| everything through shell vars.
| mlhpdx wrote:
| And, the author isn't suggesting only having JSON output,
| but adding it as an option for those of use that would
| make use of it. The plain text should remain as well (and
| has to or many, many things would break).
|
| On a separate point, I find the JSON much easier to
| reason about. The wall of text output doesn't work for my
| brain - I just can't see it all. Structuring/nesting with
| clear delineations makes it far easier for me to grok.
| bsdetector wrote:
| > That's not your original request though, to use line-
| based data.
|
| It wasn't my request and OP (not me) said "line-based
| data" is best. The comment I replied to said "Newline-
| delimited JSON ... a line-based format".
|
| If the only objection you have is "but that's line-
| based!" then you're in a completely different
| conversation.
|
| > if anything, json output | jq is more the unix way than
| piping everything through shell vars.
|
| The unix way is line-based. The comment I replied to is
| talking about line-based output. Line-based output is the
| only structure for data universal to unix cmdline tools -
| even tab/space isn't universal; sending structured non-
| line-delimited data to a program to unpack it is the
| least unix-like way to do it.
|
| Also there's no pipe in the shell-variable output scheme
| I described, whereas "json | jq" is a shell pipeline.
| starttoaster wrote:
| That's great for key=value data, but more complex data
| structures don't work so well in that format, JSON does.
| "Why would you need to represent data as a complex data
| structure?" Sometimes attributes are owned by a specific
| entity, and that entity might own multiple attributes. It
| might even own other sub-entities. JSON represents that.
| Key=value does not.
| bsdetector wrote:
| JSON is literally key=value, just nested. Which you can
| do with shell variables.
|
| The question was "What's not to like [about JSON output
| from cmdline tools]?" and the answer is that it's
| cumbersome to read in a shell and all but requires
| another pipeline stage.
|
| I didn't even recommend shell variable output and made it
| clear this isn't today a reasonable solution so I'm not
| sure where this hostility in the replies comes from, but
| I assume from recognition that it's a more practical
| solution to reading data within a shell but not wanting
| that to be so.
| starttoaster wrote:
| > JSON is literally key=value, just nested.
|
| The nature of being nested, and also containing
| structures like lists, maps, etc. All of which makes it
| more complicated than key=value.
|
| > The question was "What's not to like [about JSON output
| from cmdline tools]?" and the answer is that it's
| cumbersome to read in a shell and all but requires
| another pipeline stage.
|
| It depends on the intended use for your shell program. If
| you intend the CLI tool to be used in CI pipelines (eg.
| your CLI tool's output is being read by an automated
| process on a computer) and the data it outputs is more
| complicated than a simple key=value, JSON is great for
| that. Your CI program can pipe to jq. You as a human can
| pipe to jq, though I agree it's somewhat less desirable.
| Though just piping to jq without any arguments pretty
| prints it for you which also makes it fairly readable for
| humans.
|
| > so I'm not sure where this hostility in the replies
| comes from
|
| You're reading into hostility where there isn't any.
| bsdetector wrote:
| > The nature of being nested, and also containing
| structures like lists, maps, etc. All of which makes it
| more complicated than key=value.
|
| These are javascript objects, which are key-value. A list
| array is just keyed by a number instead of a string.
| They're functionally exactly the same as name=value
| except JSON is parsed depth-first whereas shell variables
| are breadth-first parsing (which is way better from
| shells).
|
| Do you have an example of a CLI tool - intended for human
| use - that has output so complicated it can't be easily
| mapped to name=value? I don't think there is one, and
| it's certainly not common.
|
| > You're reading into hostility where there isn't any.
|
| I think "it seems you're determined not to use jq" is
| pretty hostile since I made no intimation of that at all.
| starttoaster wrote:
| > I think "it seems you're determined not to use jq" is
| pretty hostile since I made no intimation of that at all.
|
| Well, I didn't say that, so I don't know what that other
| person's feelings or intentions are, to be fair. I
| personally have no feeling of hostility towards you just
| because we (apparently) disagree on the usefulness of
| JSON to represent complex data types, or at least
| disagree on how often human-usable CLI tools output
| complex data. But to answer:
|
| > Do you have an example of a CLI tool - intended for
| human use - that has output so complicated it can't be
| easily mapped to name=value? I don't think there is one,
| and it's certainly not common.
|
| kubectl. Which to be fair defaults to output to a table-
| like format. Though it gets all that data in the table
| from JSON for you. smartctl is another one, which also
| defaults to table format. To be honest, I could go on and
| on if the only qualifier is a CLI tool that emits complex
| data, not suited for just key=value.
|
| > These are javascript objects, which are key-value. A
| list array is just keyed by a number instead of a string.
| They're functionally exactly the same as name=value
| except JSON is parsed depth-first whereas shell variables
| are breadth-first parsing (which is way better from
| shells).
|
| As mentioned before, just because you can compare JSON to
| key=value, does not mean it's as simple as key=value.
| It's a data serialization language that builds well on
| top of simple key=value formats. You're welcome to enjoy
| other data serialization languages, like yaml, HCL, or
| PKL. But none of those are simple key=value formats
| either. They built the ability to represent more complex
| structures on top of that.
|
| A data serialization language allows the end-user to
| specify how they would like to use that data, while
| allowing them to use standard parsing tools like jq.
| Cramming complex data into a value string in a key=value
| format gives end users the same allowance to use that
| data however they want, while also giving them a chore to
| handle parsing it in custom ways tailored to just your
| CLI application, likely in ways that would seem far more
| brittle than parsing a defined language with well defined
| constraints. That doesn't sound like great UX to me. But
| to be fair to you, you're not saying that you wish to use
| key=value to represent complex data. Rather, you're
| saying there's a general lack of complex data to be
| found, to which I also disagree with.
| bsdetector wrote:
| > But none of those are simple key=value formats either.
|
| What is the difference between: {
| object: { name: value }} { object: "{ name: value
| }"} object="name=value"
|
| There's zero difference between any of them except how
| you parse and process the data.
|
| > kubectl. Which to be fair defaults to output to a
| table-like format.
|
| With line-based shell-variable output you have a line of
| variables and you have blocks of lines separated by an
| empty line (like an HTTP 1 header).
|
| This can easily map to any table, two dimensions, or two
| levels of data structure without even quoting
| subvariables like in the example above. So, no, kubectl
| is not an example at least not how you've described it.
| starttoaster wrote:
| > What is the difference between .. There's zero
| difference between any of them except how you parse and
| process the data.
|
| Answered in the previous message... "A data serialization
| language allows the end-user to specify how they would
| like to use that data, while allowing them to use
| standard parsing tools like jq. Cramming complex data
| into a value string in a key=value format gives end users
| the same allowance to use that data however they want,
| while also giving them a chore to handle parsing it in
| custom ways tailored to just your CLI application, likely
| in ways that would seem far more brittle than parsing a
| defined language with well defined constraints."
|
| > With line-based shell-variable output you have a line
| of variables and you have blocks of lines separated by an
| empty line (like an HTTP 1 header)...
|
| I would not choose to write application logic that
| foregoes defined data serialization languages for parsing
| barely structured strings the way you seem to prefer. But
| you go about it the way you prefer, I guess. This whole
| discussion leaves a lot of room for personal opinions. I
| think we both agree that the other person's opinion here
| is subjectively the more annoying route to deal with. But
| that's the way life is sometimes.
| eternityforest wrote:
| If you're using UNIX tools to parse it, it sucks, but generally
| if I'm reading the output of a command, and the command is more
| than one word, I'm doing the whole thing in Python.
|
| That's real programming, and for that I want type checkers and
| debuggers and modern syntax and all that. And the performance
| is often faster because you're not spinning up subprocesses for
| each command.
| pluto_modadic wrote:
| ./some_command | jq '.memory_use'
|
| vs:
|
| okay, run it once with head (assuming we have a --dry
| option...) ah, that's the column. okay cut -d"," -f0, ah whoops
| it's starting at 1. ah, damn, there's a weird comma in the
| name/quote. oh weird, that one's null, ah heck.
|
| schemas are cool. JSON extends and builds on the UNIX idea of
| having things pipe and plumb well together.
| alerighi wrote:
| Plain text is more easy to reason about, because we are used
| to process text. A good textual output, that is records
| delimited by spaces, tabs or a delimiter, to me is all it's
| needed, for most applications.
|
| An object structure it's much more complex to use. For
| example an output that is a set of records can be easily
| imported in an Excel sheet, in an SQL database, processed
| line by line, without issues. Processing JSON is not straight
| forward, not all programs support JSON.
|
| Finally JSON can't be processed as a stream, meaning that
| tools like head, tail, etc. doesn't work on JSON, you have to
| read it all in memory, or use JSON lines, that is not a
| standard format, that not all parsers support natively, etc.
|
| JSON is good if for integrating the program inside other
| programs (as a subprocess), so having an option to
| input/output JSON in a program is useful, but to me it's not
| as useful for interactive shell usage. I prefer to use UNIX
| tools such as grep, cut, head, tail, etc.
| mlhpdx wrote:
| But as the article points out, the advice to use unbuffered
| JSON lines for commands that are line oriented is well given.
| Not doing that can really make life sad.
| dylan604 wrote:
| you still need to know the schema of the JSON which could
| also use a dry run as well. not really sure how just because
| it's JSON solves that in your mind
| Jcowell wrote:
| Orient the idea be that because it's json there's a schema
| somewhere the end user can refer to?
| dylan604 wrote:
| as if the output of the other command also isn't
| available?
| placatedmayhem wrote:
| Line-oriented formats, like most traditional Unix-style tools,
| are for human consumption. JSON is bad at that, thus gron.
|
| On the other hand, structured output formats, like JSON, make
| it easier to consume with other programs. Standard formats have
| readily-available and commonly used libraries, whereas line
| parsing tends to be one-off for every program. Whether JSON is
| the best format for this is certainly debatable, but it is
| quite ubiquitous, which is a huge advantage. I doubt many folks
| would propose XML as a general recommendation.
|
| Tools should have both options on their path to maturity --
| both human-consumable and computer-consumable output format
| options.
| janderland wrote:
| I'd rather use JMES Path (or jq even tho it's messier) to
| restructure my data rather than some mixture of awk, sed, cut,
| etc.
|
| Line output often needs restructuring in my experience, though
| JSON will always need it.
| Spivak wrote:
| Unix tools on newline delineated list like objects work great,
| but try them on dict like objects and it becomes clunky really
| fast. Parsing `ip` output is a good example where the data
| naturally lends itself to iface:attrs. Plucking the value you
| want by `.[.name | startswith('eth')].addr` is way easier than
| pulling this out with grep/awk.
| BenjiWiebe wrote:
| You probably already know this but just in case: You can
| specify an interface when using ip, e.g. ip addr show dev
| enp2s0
| __MatrixMan__ wrote:
| The caller can just specify the output format that they want,
| and ever since I switched to nushell I pretty much always want
| json.
| keybored wrote:
| What orthogonality? "Line-based data" that varies randomly from
| tool to tool? Using json between programs is perfectly
| orthogonal.
| IshKebab wrote:
| Nonsense. Using a structured format means all tools are using
| _the same_ format, instead of every tool making up their own
| (usually broken) system.
|
| Look at how many flags `ls` has to feed filenames into other
| programs without breaking.
|
| Also, using a proper format like JSON makes it actually robust.
| Most ad hoc pipelines break if you so much as put a space in a
| filename.
| pluto_modadic wrote:
| JSON is so much more usable sometimes. Especially when things
| don't neatly parse into cut -d splits!
| JohnMakin wrote:
| Yes - and this is mostly because the jq parser tool is so useful.
| yq is similarly great.
| rwmj wrote:
| This is really the only reason, because JSON as a format is
| otherwise terrible. No good way to represent 64 bit numbers,
| very limited types and no comments.
| cynicalsecurity wrote:
| This basically turns command line tolls into APIs. Crazy, in a
| good way.
| inetknght wrote:
| Not really that crazy at all. And yes it's definitely good.
| moregrist wrote:
| I'm a big fan of tools producing json output, but I couldn't help
| but chuckle a bit at:
|
| > This allows easier parsing in scripts by using JSON parsing
| tools like jq, jello, jp, etc. without arcane awk, sed, cut, tr,
| reverse, etc. incantations.
|
| While I love jq, I find its syntax to be utterly baroque. Far
| more than awk and sed (let alone simple tools like cut and tr).
| To the point where I end up looking at the jq man page to remind
| myself how to do even seemingly simple things at times. And often
| jq is just the beginning of a pipeline to get hierarchical data
| into columnar format so I can hit it with awk and sed.
|
| But maybe this is all about familiarity; I have years of
| experience cobbling together shell tools and reach for jq
| comparatively less. I suspect the author may be the opposite.
| Mordisquitos wrote:
| I would argue it's not just about familiarity. I had the same
| reaction to that quote and feel the same as you do, and yet I
| have been using jq _before_ I first learned and started using
| awk.
|
| For a long time I had been familiar with jq, but every time I
| had to use it for any non-trivial task after not touching it
| for a while I had to go to the docs and really scratch my head.
| Meanwhile, I had no idea about awk, and for ages I would just
| see those awk one-liners on Stack Overflow as weird arcane
| alternatives to "simply" piping sed, grep and tr over bash.
|
| However, one day I finally had the time and the reason to have
| a look at AWK Programming for a specific use that I needed, and
| it immediately clicked. Awk makes sense and now I can easily
| fall back to it every time, even if I haven't used it for ages.
| At most, all I need to look up is the order of parameters in
| the gawk regex functions.
|
| And yet, time and time again I am baffled by jq, however much I
| have used it in the past. Whenever I need to parse a JSON I'm
| happy enough if I can use jq to get it to a good enough halfway
| point so that I can pipe it to awk to do the actual work.
| DangitBobby wrote:
| ChatGPT flawlessly constructs jq commands for whatever banal
| task you have. Game changer.
| simonw wrote:
| Completely agree. I use jq multiple times a day now, purely
| because GPT-4 or Claude Opus mean I never have to think about
| how to express anything in it.
| candiddevmike wrote:
| I output JSON and embed a jq-like arg in in all of my projects
| (https://rotx.dev/docs/references/cli#jq). Makes it really easy
| to slice and dice output.
| arp242 wrote:
| It's not you; parsing line-based text is fundamentally simpler
| than parsing structured data like JSON, especially nested JSON.
| It's just a lot easier to reason about, and a lot easier to
| split up in smaller chunks, too. Who hasn't ended up with
| something like: cmd | sed | grep | sed | grep
|
| Each bit gets you closer what you want in small easy-to-
| understand steps, making it fairly easy to develop. And for
| things you only do once or twice, who cares as long as it
| works, right? I find that's a lot harder to do with jq. I feel
| this is a problem waiting for a better solution (and gron isn't
| it).
| fuzztester wrote:
| >Who hasn't ended up with something like:
|
| > cmd | sed | grep | sed | grep
|
| And the decorate-sort-undecorate pattern:
|
| https://en.m.wikipedia.org/wiki/Decorate-sort-undecorate
|
| and the related Schwartzian transform:
|
| https://en.m.wikipedia.org/wiki/Schwartzian_transform
|
| (The former redirects to the latter.)
| ramses0 wrote:
| I'll my $0.02: when possible, assume/emit "arrays of records".
|
| eg: `ls ...` should be an an array of homogenous entities
| (objects) with usually simple key=>value relationships.
|
| Slight variations such as "field contains an array of strings"
| (eg: tags, groups, file path hierarchy, temperature per each CPU)
| is also totally fine.
|
| Seeing the wildly differing record types and deeply nested
| "stuff" in some of the examples weirds me out.
| { "type_cpu": [ ... ], "type_ram": [ ... ],
| "type_network": [ ... ], ... }
|
| ...obviously, in some cases a graph/hierarchy is very useful to
| be able to traverse, but emitting records and "coalescing" the
| path into a record-field makes simple operations simple.
|
| Example: `hwstats | jq '.type_cpu | .[] | select( .temp > 40 ) |
| .hierarchy_field`
|
| Don't make me go like `.cpu.bank[0].chip[3].temp` or whatever...
| just give me "all the CPU's" and then give me a good URI(!!) to
| describe it or search for it later.
|
| Rationale: coalescing and traversing weird object paths is tough
| to do dynamically via current shell DSLs, and the URI concept as
| an ID excellent application.
| jiehong wrote:
| An array of records, or just 1 record per line, also works
| nicely with jq.
|
| For that matter, it can even be imported in like duckdb and
| handled like a table :)
| dec0dedab0de wrote:
| I prefer jsonl from command line tools, just cause it's nicer to
| log, and parse using standard tools. Plus it's like a two second
| function if you're doing it in a real programming language. I
| think it's the best of both worlds.
|
| They say to use JSON lines for streaming, but even if you're not
| it's still very nice to be able to do _my_command | grep xyz_ and
| still have it be parseable
| isatty wrote:
| Absolutely not a fan of JSON. It's not typed, and I definitely
| don't want to know about each tools schema (and it'll be worse
| without the definition).
|
| There is nothing wrong with Linux tools, they do one thing and do
| it well.
| throw10920 wrote:
| This is self-contradictory. JSON is _more_ typed and _more_
| predictable than what comes out of UNIX tools.
| xomodo wrote:
| I don't mind json outputs, but pls keep cli options as low
| letters, eg.: 'kubectl get pods -o json'.
| echoangle wrote:
| Why is that important?
| al_borland wrote:
| Mixing case in a case sensitive environment is unnecessarily
| confusing and annoying.
| Karellen wrote:
| > The file systems of Unix machines all have the same general
| structure. [...] Note the obsessive use of abbreviations and
| avoidance of capital letters; this is a system invented by
| people to whom repetitive stress disorder is what black lung
| is to miners. Long names get worn down to three-letter
| nubbins, like stones smoothed by a river.
|
| -- Neal Stephenson, _In the Beginning Was the Command Line_
| (Ch. 14), 1999
|
| https://steve-
| parker.org/articles/others/stephenson/oral.sht...
| matheusmoreira wrote:
| Parsing is hard and this definitely beats having to write ad-hoc
| parsers for stuff. I worry we're essentially reinventing a less
| verbose form of SQL though.
| kunley wrote:
| No, parsing is easy...
|
| I recommend the "Crafting Interpreters" book.
| jiehong wrote:
| On linux systems, I was thinking that it would be great if there
| was a dbus wrapper that outputs all data as json by default.
|
| As many things are exposed on dbus (like systemd units, timers,
| services, etc), and busctl [0] has a --json output. We could have
| a dbus-services, or dbus-timers, etc to retrieve that on dbus and
| emit json. If most things were on dbus, one could have access to
| almost everything on the system as json directly, while talking
| to the underlying implementation of most cli by default.
|
| I suppose something similar could be done on MacOS with launchctl
| and others (I don't know much about MacOS, though).
|
| [0]: https://www.man7.org/linux/man-pages/man1/busctl.1.html
___________________________________________________________________
(page generated 2024-04-20 23:00 UTC)