[HN Gopher] Learn to Process Text in Linux Using Grep, Sed, and Awk
___________________________________________________________________
Learn to Process Text in Linux Using Grep, Sed, and Awk
Author : qlacus
Score : 51 points
Date : 2023-01-06 20:08 UTC (2 hours ago)
(HTM) web link (www.linode.com)
(TXT) w3m dump (www.linode.com)
| d_watt wrote:
| I feel like I have a weird relationship with these command line
| tools. They're obviously powerful, and grep in particular can be
| a huge time saver to be able to quickly process a file.
|
| Whenever I think Sed or Awk would be correct for a problem, I
| find myself with a clever 1 liner an hour later I look at and
| think "I should have just used python," as it's untestable /
| unmaintainable, and I'll come back the next day and not be able
| to read it.
|
| Are there people out there for whom these tools are daily drivers
| that are actually part of their core toolset, or are these mostly
| just hobbiest things where it feels neat to know. Like HAM radio
| operation.
| marginalia_nu wrote:
| I use sed and awk fairly frequently (and I'm familiar with the
| tools to the point where I got half a dozen problems into AoC
| with just sed once).
|
| Most of what my sed and awk programs I write aren't supposed to
| be maintainable, but rather typically one-off stuff for-the-
| moment sort of thing. I'll grant it's append-only code, but if
| I need to do the same thing later (or something adjacent) I'll
| typically just write another oneliner since there's no real
| effort involved in it.
|
| It's pretty much always more effort to write an actual script
| in some language like python, since it needs to be in a file
| and have a name and take arguments and so on.
| dima55 wrote:
| I sorta, kinda agree. Tools written in AWK (and friends) are
| indeed somewhat unmaintainable, but they're really close to
| being just right for a LOT of applications. The vnlog toolkit
| (https://github.com/dkogan/vnlog) adds just a little bit of
| syntactic sugar to the usual commandline tools to make
| processing scripts robust and easy to read and write. This was
| not my intent initially, but I now do most of my data
| processing with the shell and vnl-wrapped awk (and sort and
| join, ...) It's really nice. If you write stuff in awk, you
| should check it out. (Disclaimer: I'm the author)
| cerved wrote:
| I use sed a lot in shell scripts, it's very powerful at
| massaging output from a stream and input it into something
| else.
| throwaway_au_1 wrote:
| I also don't enjoy coming back to large sed expressions I wrote
| previously but I'm starting to write them in a way that makes
| that much easier. I'm on my phone so this will probably be
| wrong but hopefully the gist is evident. Basically, load the
| expression from a commented 'file': sed
| ./target --file=<(cat <<- SED_EXPR | grep -v -E
| '^\s*#' # do a string replacement
| s#target#replacement#g # then do
| something else .. SED_EXPR )
| ghshephard wrote:
| I've probably used awk in excess of 10-20 times/day
| consistently (and somedays upwards of 100 times a day) for the
| last 15 years, despite having a full complement (and reasonable
| knowledge) of Python, Pandas, etc...
|
| If I'm going to use the tool multiple times - it lands in
| Python, gets checked in, might even get unit tests if I'm
| feeling energetic. But there are tons of times a day when I
| want to ask a question of a 5 GB text file that I can pound out
| in < 60 seconds with a bit of piped awk/sed/tr and some bash
| looping on the results).
| owenmarshall wrote:
| My approach with these tools is to use then for speed, and to
| resist the (very strong) urge to golf with them.
|
| So I always start with `cat filename | some tool`[1] and refine
| until the very moment that I think "damn, I can't remember how
| to do that with sed/awk/jq/xsv/..." and end my pipeline feeding
| into Python. There's no law against it! And on slow days I go
| back and figure out how I can do what I want with the other
| tools.
|
| [1]: Yes, yes, UUOC[2], but I don't care!
|
| [2]:
| https://groups.google.com/g/comp.unix.shell/c/532AcI3-zs4/m/...
| laumars wrote:
| Not just using it daily, but I have scripts I've written
| literally a decade ago, and which haven't been modified at all,
| which still work just as well as the day I wrote them. Whereas
| there's a whole plethora of Python projects I've had to abandon
| in that time because they were never ported to Python 3.
| jamal-kumar wrote:
| I have a ton of custom log monitoring written in awk/sed
| because it's way faster to process it that way in shell scripts
| than it would be in python. The key thing is that it doesn't
| need to allocate any memory and that it's by default processing
| text streams, which means by default you're just zipping the
| log files through the processor cache and outputting the result
| to a file to get emailed off.
|
| It's important to realize that when you have gigabytes of data
| you need to analyze that you can go really fast this way.
| Python sure does have stuff like StringIO, but it just doesn't
| feel like the right kind of tool for that job. A ton of stuff
| is just columnar data like CSV or log files out there, and
| people have been using shell tools to process them for decades.
|
| If you're having a hard time reading your complicated awk one
| liners, I suggest pretty printing it. [1] Plenty of people are
| using this in production by writing scripts in files.
|
| [1] https://stackoverflow.com/questions/55745956/is-it-
| possible-...
| xebus wrote:
| Maybe someone will come up with a wrapper command line
| interface where we can express the desired outcome and GPT-3 /
| ChatGPT can produce (then execute) the "correct" commands. I
| agree with the sentiment of the power of grep/awk/sed (and use
| them in that order) but I also would prefer Python for anything
| I need longer than an hour. Having a natural language command
| -> grep/awk/sed pipeline -> saved as a CLI shortcut could be
| cool.
___________________________________________________________________
(page generated 2023-01-06 23:02 UTC)