[HN Gopher] Is the "modern data stack" still a useful idea?
       ___________________________________________________________________
        
       Is the "modern data stack" still a useful idea?
        
       Author : tim_sw
       Score  : 118 points
       Date   : 2024-02-11 20:58 UTC (1 days ago)
        
 (HTM) web link (roundup.getdbt.com)
 (TXT) w3m dump (roundup.getdbt.com)
        
       | tomrod wrote:
       | Intriguing that the CEO of DBT declares MDS ("modern data stack")
       | to be a meaningless term (i.e. no longer useful). Maybe a good
       | demarcation that we are entering the post-modern phase of cloud
       | tech generally? Cloud still there, but not the only option.
       | 
       | I look forward to this happening with GenAI -- the tools we come
       | up with will be pretty cool in the long run. I hope we can find
       | ways around platform enshittification because that really sucks.
       | Similar for blockchain tech.
       | 
       | I see a common thread in these techs too -- the same type of
       | LinkedIn influencers shop these during their hype phases.
       | Ultimately, useful technology and best practices come out of it!
        
         | actionfromafar wrote:
         | Well put! Nothing's modern anymore, the tech cycles are closing
         | in on one another so much the old one hasn't gone out of phase
         | until it's overcome by not one but several new ones. In a nice
         | touch of self-refentiality, your message itself is very post-
         | modern.
        
           | tomrod wrote:
           | I hadn't considered that. Thanks for pointing that out to me
           | :)
        
       | karakanb wrote:
       | Disclaimer: I am the co-founder of a competitor, Bruin
       | (https://getbruin.com). We are exactly the kind of integrated
       | platform Tristan is talking about.
       | 
       | The article resonates with me a lot, and it is because we called
       | this out months ago. I find the idea of Tristan walking back on
       | the premise of MDS and claiming it to be not useful anymore funny
       | because they were one of the main drivers of the term and the
       | whole hype around it.
       | 
       | He even acknowledges that they played ball with other companies
       | in the space:
       | 
       | > There was a lot of valuable co-marketing, partnership deals,
       | co-sponsored events, and co-selling. This had real value for
       | everyone involved--customers and vendors alike. Companies
       | voluntarily integrated their products together, cross-promoted
       | each other publicly, and built partnerships that made owning and
       | operating these technologies far easier for customers.
       | 
       | Sorry, but no, this didn't have value for the customers, only for
       | the vendors. The customers were left alone by themselves to deal
       | with all of this complexity, and the vendors made a lot of money
       | off of that. They convinced companies that they needed a bunch of
       | different tools to build a simple pipeline, and rode on the wave
       | of huge valuations based on these same ideas that they are
       | walking back on.
       | 
       | The companies prefer integrated solutions now because they woke
       | up. Instead of paying 150k/y each to Fivetran, dbt and whatnot,
       | they realize that they are better of just hiring an engineer or
       | two in the worst case. It is 2024, and none of these tools still
       | properly talk to each other. Do you want to get an end-to-end
       | lineage of your data? Good luck with that. How about quality? How
       | about governance? Companies are left alone with this hype cycle.
       | 
       | I'd claim that a significant part of the blame lies on the
       | executives and leaders in the companies, who just jumped on the
       | ship for the sake of building their CVs and skipped the critical
       | thinking step. None of them seriously asked themselves the
       | question of whether or not it makes sense.
       | 
       | To be honest, I feel sorry for all the money spent on building
       | solutions around all of these hype-driven products.
        
         | neeleshs wrote:
         | Disclaimer : I'm the cofounder of Syncari. I agree 100%. The
         | whole notion of modern data stack is a myopic , neither here
         | nor there philosophy to address data in my opinion.
         | 
         | Just take a look at the sheer number of tools one has to run -
         | ETL(or more commonly, data dumpers), Wearhouse, transform
         | tool,"reverse ETL", data quality tools, BI. And if you want
         | some ML, that's whole another game.
         | 
         | We also believe strongly in an integrated approach to this
         | space, with a data model at the center of it all.
        
           | gigatexal wrote:
           | Deleted
        
         | tomnipotent wrote:
         | He's not "walking back on the premise of MDS", but acknowledges
         | that the term "MDS" had been co-opted and lost its descriptive
         | fidelity.
         | 
         | > Sorry, but no, this didn't have value for the customers, only
         | for the vendors
         | 
         | Absolutely untrue. I've discovered multiple vendors over the
         | last two decades because of co-marketing and partnership deals,
         | and I know many others that have. My success rate with these
         | vendors is no different from vendors I found myself or were
         | referred to me by others.
         | 
         | > The customers were left alone by themselves
         | 
         | This is true with every vendor integration with every company
         | on the globe. You either 1) have internal expertise to push
         | through the pain, 2) hire support from the vendor itself (often
         | X hours come baked into the deal), or 3) hire a consultant.
         | This is as true for the MDS vendors as it is for Salesforce,
         | AWS, or an ERP. It's true for your company, too.
         | 
         | > They convinced companies that they needed a bunch of
         | different tools
         | 
         | Except you do? You need a database, you need something to
         | extract data and move it locally, you need something to
         | transform that data into something useful, and you need the
         | ability to deliver that data to end users. This was true twenty
         | years ago, true ten years ago, and still true today. The all-
         | in-one vendors prior to MDS were horrible, and I'd rather cut
         | off my arm then use something like Pentaho again. I'm also not
         | convinced the current all-in-one vendors are any better.
         | 
         | One of the advantages of the pick-your-stack MDS was that you
         | could tailor the tools to you organizations specific needs,
         | rather than have to fight against whatever your mega-vendor
         | offers. I loved the fact that I could use open source for 90%
         | of my stack, and then use PowerBI for the last 10%. Or I could
         | be 100% FOSS and deliver data with Metabase. Or I could use
         | Tableau for 50%, and FOSS for the other 50%.
         | 
         | > they realize that they are better of just hiring an engineer
         | or two in the worst case
         | 
         | This simply isn't true for every size or class of company. Many
         | businesses are better off with vendor-supported platforms than
         | trying to build something bespoke, which can quickly become a
         | liability with staff turnover or lack of strong internal
         | technical leadership.
         | 
         | > It is 2024, and none of these tools still properly talk to
         | each other
         | 
         | What does this even mean? How does a data extract tool "talk"
         | to a transformation tool? Or "talk" to the database? I can see
         | from the screenshot on your website that it offers a GUI with
         | drag-and-drop - is that your definition of "talking"? Or are
         | you referring to a tight integration that doesn't require
         | additional work to glue them together?
         | 
         | > How about quality? How about governance?
         | 
         | These are management and organization issues, not tooling. I've
         | seen amazing quality and governance with teams working in
         | MSSQL/SSIS, and I've seen horrible quality and governance from
         | teams using Hadoop/Spark. A tool can make a competent
         | practitioners job easier, but it's not going to make an
         | incompetent competent.
        
         | cgio wrote:
         | As someone who has worked in this space, I would intuitively go
         | with the decoupled tools every time. The toughest thing is
         | matching a monolith of opinions with your requirements.
         | Individual tools leave design space in their integration and
         | thus allow proper engineering. That they allow it, does not
         | mean it happens. If you go with a shopping bag architecture
         | then you might be better off results wise with an integrated
         | solution, but still, when the real engineers come along they'll
         | have an uphill battle to uproot an integrated solution vs
         | redesigning the combination of individually good tools.
        
           | CRConrad wrote:
           | > I would intuitively go with the decoupled tools every time.
           | 
           | Seems to me that the problem is, each of these "decoupled
           | tools" you mention isn't some specialised little awl or
           | hacksaw ("the Unix philosophy..."), nor even a do-anything
           | Swiss Army knife -- but a whole stack of cloud services,
           | "integrations", and administration tools in its own right. A
           | whole machine shop of machine shops. Building your own custom
           | setup out of that is hard. You'll end up with something at
           | least as unwieldy as a _really_ big do-anything single-vendor
           | mega-machine shop stack of cloud services,  "integrations",
           | and administration tools.
        
       | extr wrote:
       | What would HN commenters define as a truly modern data stack
       | right now? Mostly asking what this chart
       | https://a16z.com/emerging-architectures-for-modern-data-infr...
       | would look like in 2024.
        
         | rch wrote:
         | - Postgres and Kafka       - Airflow and Flink       -
         | Kudu+Impala or Clickhouse       - Iceberg/Parquet or Delta
         | - Ozone(HDFS) or Ceph       - Spark Rapids and/or Ray+Metaflow
         | - Open Metadata or Atlas+Ranger
        
           | fifilura wrote:
           | TBH that sounds like a heap of buzzwords not useful for
           | anyone (but maybe an architect promoting his career).
        
             | rch wrote:
             | Pretty much, but updated for 2024.
        
         | mr_toad wrote:
         | > What would HN commenters define as a truly modern data stack
         | right now?
         | 
         | Just don't spend too much time debating what it is, or the
         | analysts who need numbers yesterday will develop their own
         | solutions before you've even published your A3 architecture
         | diagram.
         | 
         | A bird in the hand, as they say.
        
         | te_chris wrote:
         | IT DEPENDS. But for standard ecom/ops stuff (and this is
         | heavily biased by the fact this is what I know and have
         | deployed with low overhead): Big query. This is fed by an ETL
         | pipeline of the stuff you need. dbt to ELT into useful,
         | reportable tables. Something to visualise - I still like Mode,
         | but there's lots of new BI toys around. KISS.
         | 
         | This is extremely low maintenance at the actual requirement
         | levels of most businesses with < 50 employees.
        
         | neighbour wrote:
         | - Fivetran
         | 
         | - Snowflake/BigQuery/Redshift
         | 
         | - dbt
         | 
         | - Looker/Metabase/Tableau
        
       | dm03514 wrote:
       | Sorry nothing positive to say here.
       | 
       | I've been using dbt and MDS for nearly 3.5 years and I believe
       | the entire approach is profoundly broken. There's really nothing
       | "modern" about it, especially compared to software engineering.
       | 
       | https://on-systems.tech/blog/135-draining-the-data-swamp/
       | 
       | Building on Extracted operational data is hard at best and a
       | business altering security liability at worse.
       | 
       | I believe The MDS trails at least a decade behind modern software
       | engineering practices, lacking industry guidance and generic
       | tooling to support: CI/CD, versioned deployment artifacts, zero
       | downtime deployments, unit testing, observability, monitoring and
       | alerting.
       | 
       | MDS Data engineering is a meme in the industry, the expectation
       | of "100%" combined with the lack of modern tooling makes success
       | really hard to achieve.
        
         | civilized wrote:
         | I was a fan of dbt for a while, but the shine wore off when I
         | saw one of my smartest coworkers try to use it. His development
         | speed was about an order of magnitude slower than what I expect
         | from data scientists using dataframe packages like R's dplyr,
         | Python's polars, or Spark DataFrames. In my own experience, dbt
         | is significantly better than writing raw SQL, but still nowhere
         | near a normal software development experience. It needs an IDE
         | badly, but the current IDE is cloud-only.
         | 
         | I am currently building my analytics pipelines on top of R's
         | dbplyr package, which allows me to run dplyr operations on
         | database tables just as if they were local data frames. Since
         | it's R, it's not without warts, but at least I get to use a
         | real programming language and the free and fantastic RStudio
         | IDE.
        
           | tomnipotent wrote:
           | > allows me to run dplyr operations on database tables
           | 
           | Which require that data take a round trip between the R
           | process and the database. It's not uncommon for these jobs to
           | spend more time in the read/write step than doing meaningful
           | work, and why I prefer dbt when possible to keep
           | transformations as close to the data as possible.
        
             | nograpes wrote:
             | The second point the grandparent made was that they were
             | using _dbplyr_ , which allows you to avoid having the data
             | take a round trip between the R process and the database.
        
           | golergka wrote:
           | We built the perfect IDE for dbt at Deep Chanel, with
           | autocomplete automatic real time type checking for the whole
           | project.
           | 
           | Sadly, the company closed this June, and even the website is
           | already down.
        
             | neighbour wrote:
             | You need to find a way to release this product. I would pay
             | to use it.
        
               | golergka wrote:
               | It's already been released and publicly available for
               | free for almost a year.
        
               | CRConrad wrote:
               | So where is it, if "even the website is already down"?
               | 
               | ETA: Seems it isn't. https://www.deepchannel.com/ That
               | what you meant?
        
               | golergka wrote:
               | Hmm, may be I've had network issues last time I checked.
               | Anyway, that's the only place you can get it -- it's a
               | completely free desktop app, but it's not open source.
        
             | civilized wrote:
             | This is extremely interesting. You might be onto something.
        
           | jochem9 wrote:
           | Imo dbt sits in the analysts space. SQL is the language they
           | know and dbt is the best solution to scale that.
           | 
           | In that sense it should only be used for last mile data
           | transformations. Not for wrangling raw data into neat tables.
        
             | fifilura wrote:
             | I am curious, what do you perceive as the difference
             | between data transformations and "wrangling raw data into
             | tables"?
             | 
             | Edit: I am also curious why you consider SQL to be inferior
             | for the latter?
        
           | camgunz wrote:
           | The core divide in this space is SQL vs. $OTHER_LANGUAGE. If
           | you're fluent in SQL you'll be great at dbt. If you're only
           | fluent in not-SQL you won't.
           | 
           | A lot of these discussions are basically "can we please just
           | use the language/tooling we know", and sometimes the answer's
           | yes! A lot of people breathe a big sigh of relief when they
           | discover you can use like, Spark through Python and what-not.
           | Regardless of whether or not it's a good idea or good use of
           | resources, this space is advancing because the market demands
           | it. But like, I don't think you're ever gonna use dbt with
           | not-SQL; it's all in on SQL. Feels like maybe it was the
           | wrong fit for your team.
           | 
           | I should say I'm not sure if your coworker was an SQL person
           | so dunno if this directly applies. Just saying what my
           | experience has been.
        
         | gigatexal wrote:
         | Y'all don't have CI/CD? Maybe it's hyped but we do simple
         | stuff. Snowflake schema DWH. Jinja2 templated SQL or dbt.
         | Airflow with a monolith tool that does transforms and such. Is
         | it perfect? No. Is it understandable? Very much so.
         | 
         | Testing is such an interesting concept in data engineering. One
         | needs consistent test data. We aim to implement that with
         | snapshots eventually but now we have sanity checks at each
         | layer.
         | 
         | > lacking industry guidance and generic tooling to support:
         | CI/CD, versioned deployment artifacts, zero downtime
         | deployments, unit testing, observability, monitoring and
         | alerting.
         | 
         | Ci cd is doable and easy: changes are deployed via GitHub
         | actions, the CI part is a bit missing I guess without tests.
         | Versioned artifacts also, add a tag to the airflow job to know
         | what tagged version of the code is running. Zero downtime
         | deployments we do that all day -- I mean using views we can a/b
         | deploy changes to the underlying tables and then do a simple
         | schema change to the view and nobody knows the difference. Unit
         | testing still yet to be done. Observability and alerting we use
         | internal dashboards and sentry.
         | 
         | What am I missing?
        
           | Exoristos wrote:
           | Relevance and actual value, I'd assume.
        
           | antupis wrote:
           | Testing and monitoring in every aspect are the only things
           | that are greatly lacking, you always need to glue something
           | together to get nice tests or some kind of data contract
           | monitoring.
        
         | xyzzy_plugh wrote:
         | The mistake everyone makes is treating the entire space as if
         | it's somehow different than the rest of software engineering,
         | when it is precisely exactly the same.
        
           | datavirtue wrote:
           | This. Nearly everyone is confused.
        
           | esafak wrote:
           | Machine learning people have cottoned on to this, so MLOps
           | was invented.
        
             | tomrod wrote:
             | The training part of MLOps is important to ensure
             | replicability and other desirable properties or the ML
             | artifact, but the rest is clearly good CI/CD and
             | observability.
        
           | wodenokoto wrote:
           | We had an application engine come in and run an ML project.
           | 
           | The first thing he did was remove real data from training, as
           | he was shocked to see that development wasn't done against
           | dummy data.
           | 
           | And once the dev environment is running on real data, the
           | changes in how you develop and operationalize just seems to
           | cascade in my experience.
        
             | fifilura wrote:
             | I am slightly confused by your reply. Can you elaborate a
             | bit, was it good or bad?
        
               | wodenokoto wrote:
               | You can develop a database schema for your application on
               | dummy data, and test your business logic on dummy data.
               | But you can't develop a machine learning model on dummy
               | data.
               | 
               | It is best practice to keep live data out of your
               | development environment for normal software development,
               | but it is impossible for machine learning projects.
        
               | fifilura wrote:
               | Yes, this is what I thought and exactly what I am
               | experiencing right now.
               | 
               | It is somewhat frustrating to see the
               | engineering/platform team spending time on this kind of
               | abstraction that is essentially useless for doing the
               | work we need to do.
        
               | tomrod wrote:
               | Yep. I advise my clients to have a controlled, sandbox
               | partition of prod. Call it the whiteboard or what have
               | you, things can be erased or you can cache model version
               | runs. Your data transforms can happen in dev but training
               | should happen against real data. Synthetic data can be a
               | compromise but that can sharply bias your model and,
               | since it a harder lift, often multiplies data prep time
        
           | appplication wrote:
           | Well... no I would disagree. I am a data engineer who
           | relentlessly pushes for more classic SWE best practices in
           | our engineering teams, but there are some fundamental
           | differences you cannot ignore the data engineering that are
           | not present in standard software engineering.
           | 
           | The most significant of these is that version control is not
           | guaranteed to be the primary gate to changes to your system's
           | state. The size, shape, volume, skew of your data can change
           | independently of any code. You can have the best
           | unit/integration tests and airtight CI/CD but it doesn't
           | matter if your date source suddenly starts populating a
           | column you were not expecting, or the kurtosis of your data
           | changes and now your once performant jobs are timing out.
           | 
           | Compounding this, there is no one answer for how to handle
           | data changes. Sure you can monitor your schema and grind
           | everything to a halt the second something is different in
           | your source, but then you're likely running a high false
           | positive rate, and causing backup or data loss issues for
           | downstream consumers.
           | 
           | The truth is change management and scaling is more or less a
           | predictable and solved problem in traditional SWE. But the
           | same principles can't be applied with the same effectiveness
           | to DE. I have seen a number of SWE-turned-DE struggle
           | precisely because they do not believe there to be any
           | difference between the two. It would be nice if it were so
           | simple.
        
             | xyzzy_plugh wrote:
             | You are massively oversimplifying "classic SWE" here.
             | Ultimately it makes no difference.
             | 
             | How you handle contract violations between dependencies and
             | third parties is not an inherently different problem. What
             | _should_ you do if a data source changes schema? How do you
             | detect it, alert, take action? How are users or customers
             | or consumer jobs impacted? This is all just normal
             | software.
             | 
             | In many ways it feels like you're making my case for me.
        
               | fifilura wrote:
               | Whenever I get confronted with SWE best practices it is
               | all about version control, unit tests, integration tests
               | and CI/CD.
               | 
               | What are some good resources for the other things you
               | describe?
        
               | jimberlage wrote:
               | I don't have great links handy, but read up on
               | observability.
               | 
               | You'll want to read up on what people do on error,
               | especially in the case that contracts change (search "api
               | versioning".)
        
             | coldtea wrote:
             | > _The most significant of these is that version control is
             | not guaranteed to be the primary gate to changes to your
             | system's state. The size, shape, volume, skew of your data
             | can change independently of any code._
             | 
             | That has been true in SWE since the days of ENIAC.
        
         | kgdiem wrote:
         | > CI/CD, versioned deployment artifacts, zero downtime
         | deployments, unit testing, observability, monitoring and
         | alerting.
         | 
         | This definitely stood out to me when I was working with a lot
         | of Snowflake's new products, especially their "native apps".
         | 
         | I started on some bespoke tooling given the lack of everything
         | that you mentioned and at the time was thinking about how to
         | productize it / turn it into some open source tools but I've
         | not gotten back around to it since leaving that job.
         | 
         | I had a really great experience working with sqlglot from
         | Tobiko Data https://tobikodata.com/ -- haven't had a chance to
         | check out their sqlmesh project yet but I have some faith in
         | their work.
        
           | camgunz wrote:
           | Yeah seconding sqlglot, has saved me _tons_ of time. I think
           | yeah, they're smart and onto something here.
        
         | benjaminwootton wrote:
         | dbt was a huge leap forward in this regard.
         | 
         | It enables source code control, modularity, testing, CI/CD,
         | environments, documentation, git branching, local development
         | experience.
         | 
         | Not a fanboy, but it's who reason for being is basically to
         | enable "software engineering practices for data".
        
           | antupis wrote:
           | Yup, it is not perfect but I will take dbt every day against
           | the old way as some random DDL files in git.
        
             | CRConrad wrote:
             | So basically its huge advantage is that it saves its code
             | as text, which makes it gittable?
             | 
             | (Dang, and here the tool I've been designing in my head for
             | several years was going to claim that niche... So unique.
             | ;-)
        
               | camgunz wrote:
               | Nah, the main thing is the dependency resolution. So you
               | make pipelines like:
               | 
               | - employees
               | 
               | - tall_employees
               | 
               | - tall_employees_by_salary
               | 
               | These tables (dbt would call them models because they can
               | also be views depending on how they're configured) depend
               | on each other, so you have to build them in a specific
               | order. Without something like this, you're manually
               | running SQL to build your pipelines, and that's very
               | error prone. With something like this you can run tests
               | (they're also SQL in dbt, which is pretty common in data
               | engineering teams), run your pipeline in CI, etc.
               | 
               | There's a bunch of other features, but they all hang off
               | this basic idea.
        
               | riku_iki wrote:
               | > you're manually running SQL to build your pipelines
               | 
               | you can have some sh file which runs SQL in desired
               | order..
        
       | williamcotton wrote:
       | It seems that there's a philosophy with these kinds od cloud data
       | services: measure everything then figure out what to measure
       | after.
       | 
       | It seems that the better approach is to make a hypothesis first,
       | collect only the data you need, and then analyze.
       | 
       | Otherwise it seems the indirect costs of retooling around
       | "measure all the things indiscriminately" and the direct costs of
       | paying for such services don't seem worth it.
       | 
       | Can someone offer a reasonable rebuttal?
        
         | staticautomatic wrote:
         | It doesn't always make sense to "collect only the data you
         | need." For example, someone in our org will ostensibly have a
         | good reason for setting up an event in Google Analytics. As the
         | head of analytics I have no use for their event but I'm going
         | to end up ingesting it into BigQuery anyway because the built-
         | in GA4->BQ transfer sends ALL the data, and I'm fine with that
         | because it's damn near free for me to ingest and store. I could
         | instead "collect only what I need" by running the transfer
         | through an ELT tool and filtering out that event, but why would
         | I bother doing that work and paying the ELT vendor money when
         | Google will do it for free and charge me like $1K/year to store
         | 100M rows of everything collected in GA4?
        
           | williamcotton wrote:
           | That cost analysis seems incomplete. How much did it cost to
           | implement in the first place? That seems like millions of
           | dollars in capitalization before service expenses are even
           | considered!
        
             | staticautomatic wrote:
             | Sorry but I can't tell if you're being sarcastic or not.
             | Google Analytics and BigQuery are both nominally free, the
             | service expense is $1K/year, and the implementation cost
             | literally a few minutes of my time.
        
               | williamcotton wrote:
               | I'm not being sarcastic. There is engineering labor that
               | requires hooking up all of these events. Then there is
               | maintenance costs. And many more costs, direct or
               | otherwise.
        
               | staticautomatic wrote:
               | With respect to the example I gave, no, there actually
               | isn't. In fact I don't even understand what you think the
               | work and costs involved are.
        
               | williamcotton wrote:
               | I've seen organizations spend millions just in
               | engineering labor costs for such endeavors. But sure, for
               | small companies with limited domain events to begin with,
               | the costs are relatively lower, and probably with the
               | lower domain complexity it is less of a share of
               | engineering time overall.
               | 
               | But there are still opportunity costs! It's not free to
               | install dependencies across numerous code bases and hook
               | into and label the events, even for a small business.
        
               | fifilura wrote:
               | OP suggests using "built-in GA4->BQ transfer [to send]
               | ALL the data".
               | 
               | This seems like an extremely small endeavor.
        
               | williamcotton wrote:
               | So yes, for a small business that just drops a GA snippet
               | into the bottom of their app it is inexpensive. I'm not
               | sure that this qualifies as the "modern data stack".
               | 
               | That's also not what happens when a larger company
               | integrates Segment into their applications! That takes
               | many engineering months to set up and then maintain for
               | the lifetime of the app.
               | 
               | If the full costs are analyzed, is the money spent well?
               | 
               | There's sone irony here in that much effort is put
               | towards analysis of user behavior but very little effort
               | put towards accounting for where the money is spent
               | internally.
               | 
               | That is why there is a slowdown in customers using these
               | MDS products, not what the author of the article claims,
               | rather that when feet are put to the fire, expensive
               | services that haven't paid off their debt are given the
               | axe.
        
               | fifilura wrote:
               | I think i can spot the difference in interpretation here.
               | 
               | GG*GGP "someone in our org will ostensibly have a good
               | reason for setting up an event in Google Analytics"
               | 
               | In this case they already have Google Analytics. The
               | question is whether it was worth exporting that data to
               | BigQuery.
               | 
               | Then whether it was worth setting up GA in the first
               | place is not discussed.
        
               | williamcotton wrote:
               | Isn't the setup cost a key part of the total cost of
               | analytics, especially when it comes to these MDS systems?
               | 
               | I encourage anyone in a position to make purchases or who
               | has a say in what engineer labor focuses on is to get a
               | copy of Horngren's Cost Accounting. It should make my
               | position very clear!
               | 
               | FWIW, there is indeed capitalization occurring when an
               | organization incorporates analytics, even the "measure
               | all the things" approach, but the question is, once the
               | total costs are discovered, does the return outweigh the
               | expenditure?
               | 
               | These things cannot be considered in isolation,
               | especially when using advanced approaches like Activity-
               | Based Costing, something that is basically a requirement
               | for a product or service that is fundamentally software
               | in nature. There is too much complexity to treat
               | managerial accounting in a software setting like a steel
               | mill.
               | 
               | The reason why this doesn't happen is YOLO business
               | management fueled by the casino that is VC backed
               | companies. But we are in a rapidly maturing industry
               | where margins are slimming and costs need to be reigned
               | in, especially in publicly traded companies where
               | investors will punish a company that makes wildly off
               | target financial projections discovered when cash flows
               | are reported in the following quarters.
               | 
               | I encourage everyone in all management positions to start
               | thinking about and tracking actual costs. Not story
               | points, not customer use patterns, but cold hard cash!
        
         | neighbour wrote:
         | The guy you're arguing with in the reply thread has a point but
         | I think you're talking past each other because you're coming at
         | it from two different sides. I understand his point about
         | syncing everything if you're a small enough company or the data
         | is small enough because the cost is immaterial.
         | 
         | I agree with most of what you're saying and have seen the
         | downfalls of "just sync all our data" first hand. At a late
         | stage start-up I worked at, we would often have teams (finance,
         | CX, etc.) implement a shiny new tool and immediately request to
         | have that data accessible for analytics/reporting. If we did
         | blindly agree, the data would often never be used and it would
         | run up our Fivetran costs to sync the data and our Github
         | actions bill to build the dbt models and include it in our
         | modeling.
         | 
         | We eventually landed on an approach that considered the
         | following:
         | 
         | - Does the tool already have reporting baked in? If so, just
         | use that until you need more.
         | 
         | - How often are you really going to be using it and what makes
         | this data so special that you need to join it to our other
         | data?
         | 
         | - How much work will be involved to implement a sync for this
         | data? Does Fivetran or Segment have a connector readily
         | available?
         | 
         | - If the tool doesn't have reporting built-in and the team did
         | have a genuine business case for the data being synced, we
         | would figure out which specific objects from the source that
         | needed to be synced and sync them at a fairly infrequent
         | cadence (once a day).
         | 
         | So while I offer no reasonable rebuttal, I wanted to explain a
         | happy medium that worked for us.
        
       | Animats wrote:
       | Oh, it's something for the ad industry.
       | 
       | Ads are watched only by people who don't have good ad blockers.
        
       | hn72774 wrote:
       | DBT in simplest terms is just a way to orchestrate
       | transformations in dependency order. Table "A" needs to be
       | updated before table "B," because "B" selects from "A." It works
       | well for that.
       | 
       | I've seen it used to get wild west SQL logic into version
       | control. To replace scheduled SQL workbooks running entire data
       | pipelines.
       | 
       | I've also seen over-engineered integrations with other
       | orchestration tools in the "stack."
       | 
       | Once a DBT project grows to a certain size, roughly 500-1000 .SQL
       | files, it gets hard to manage. Not impossible, though it takes
       | intentionality about how to group things together and scale them
       | out operationally.
       | 
       | "Slim CI" is a recent buzzword that has some nice ideas about
       | build automations.
       | 
       | Cosmos is supposed to solve a lot of the automation gripes with
       | dbt. I haven't tried it yet but would like to.
       | https://www.astronomer.io/cosmos/
        
       | gms wrote:
       | I've been in data and analytics for over a decade and co-founded
       | a consolidated-ETL company (https://www.polytomic.com).
       | 
       | Nice to see Tristan realising this (people should be commended
       | for changing their minds). There were three problems with this
       | term:
       | 
       | 1. It mostly resonated with VCs and industry observers whose jobs
       | are to peddle in buzzwords, rather than users who simply want
       | their problems solved.
       | 
       | 2. It was (and is) ill-defined. Ask a group of people in the
       | industry to define it and you're guaranteed to get different
       | answers.
       | 
       | 3. It committed the cardinal sin of using the adjective 'modern'
       | in a noun. At some point, everything today stops being modern.
       | Couple this with (2) and your term is now even more meaningless.
       | 
       | At a mundane level there's no drama here: just another example in
       | the long list of noise contributors to the free-money party that
       | was going on during the Covid years (as the essay acknowledges).
        
       | Lyngbakr wrote:
       | I've worked in data for a while now as an Data Engineer, Data
       | Scientist, and Director and working with experienced software
       | engineers has highlighted to me that most of the data stack is
       | fluff. All the layers/tools that are heaped upon one another just
       | lead to complexity and dependence on paid services. I'm not
       | advocating for reinventing the wheel, but rather that a bespoke
       | solutions seem to be worth the cost. In short, I'd prefer to
       | spend the budget on experienced developers who can build
       | maintainable systems than on a plethora of MDS tools. YMMV,
       | though.
        
         | Waterluvian wrote:
         | If you wouldn't mind: in your experience how often does the
         | problem seem to stem from "we might need that later..." or "we
         | don't have a concrete set of questions we want to answer?"
         | 
         | I have considerably less experience but so far I have
         | experienced a complicated data science stack being used as a
         | way to handle a ton of data without clear questions and desires
         | on how to act on possible answers. As if they're searching a
         | haystack for... something?
        
           | Lyngbakr wrote:
           | IME, it has usually stemmed from "that's how everyone else is
           | doing it, so we should too" without considering specific
           | use/business cases. Senior management were simply not willing
           | to deviate from the norm.
        
             | tharkun__ wrote:
             | And also: I need a job. My job/workplace has these tools.
             | They are usable for my job and if I were to question the
             | tools that would come on top of my workload, so I use them
             | even if I know that cheaper tools were used at my last job.
             | 
             | Do I care if they pay my salary and the tools are usable?
        
               | tadfisher wrote:
               | That's perfectly reasonable, but it sounds like you're
               | saying the tools work for your use case. So long as the
               | executive/management level is evaluating the cost-benefit
               | of other tools, and is receptive to voices from ICs, then
               | there's no problem to solve.
               | 
               | It's just not what the GP is describing. In some orgs you
               | end up with these expensive, standard-ish "suites" that
               | are way too generalized, don't solve business needs as
               | efficiently as other tools, are painful for ICs to use to
               | meet requirements, and it seems like no one can really
               | explain why these tools are being used over something
               | else other than rumors (the previous CTO came from this
               | vendor, we got a sweetheart deal 10 billing cycles ago,
               | etc.).
               | 
               | So really this an issue of ownership; someone needs to be
               | accountable for infrastructure decisions and outcomes,
               | and it's way healthier for an org if ICs are aligned with
               | management on the same. It sounds like you're not in that
               | sort of org, so good for you.
        
               | CRConrad wrote:
               | But do you care if they pay your salary and the tools are
               | _barely_ usable?
        
         | nerdponx wrote:
         | There is no "the" data stack.
         | 
         | Some organizations have complicated orchestration requirements,
         | some don't. Some organizations have large amounts of
         | heterogeneous data, some don't. Some organizations need to run
         | large distributed machine learning tasks, some don't.
         | 
         | If you adopt tools that don't actually solve real problems that
         | your team/org faces, then of course you'll end up with a bunch
         | of extraneous fluff.
         | 
         | Your comment has the same flavor as the frequent HN meme of "I
         | just use Vi and Tmux and a BSD server for my business and
         | anything else is a sign of organizational failure and
         | individual weakness".
         | 
         | All of these tools exist for a reason. Where you get into
         | trouble is losing focus on solving your real practical
         | problems.
        
           | Lyngbakr wrote:
           | > Where you get into trouble is losing focus on solving your
           | real practical problems.
           | 
           | Absolutely agree. As I say YMMV, but _my_ experience is that
           | often people have adopted tools because  "that's what data
           | teams use" rather than because it's the best tool for the
           | job. Often, we could -- and should -- build an in-house
           | solution that fits perfectly rather than buying (every month)
           | an imperfect solution.
        
           | CRConrad wrote:
           | > Some organizations need to run large distributed machine
           | learning tasks, some don't.
           | 
           | I think almost no organizations need to run "large
           | distributed machine learning tasks". Many just think they do,
           | because it's the buzzword of the moment.
        
             | nerdponx wrote:
             | Also true, but I think maybe that's a great example of what
             | I and OP are saying here. Do you _really_ need it? Or do
             | you just think you might need it one day because people are
             | talking about it on Medium, or because you had a nice chat
             | with some sales rep? Or do you just want it on your resume?
        
         | sagarm wrote:
         | > All the layers/tools that are heaped upon one another just
         | lead to complexity and dependence on paid services. I'm not
         | advocating for reinventing the wheel, but rather that a bespoke
         | solutions seem to be worth the cost.
         | 
         | Which third party products do you use in your stack, then?
         | 
         | Do you only use basic cloud building blocks (VMs, EBS, S3), or
         | F/OSS components like Spark and Trino, or are higher level
         | things like Tableau/Looker also part of the solution?
        
       | datadrivenangel wrote:
       | The Modern Data Stack is Dead! Long Live the Analytics Stack!\
       | 
       | Interesting to see one of the strongest popularizers of the term
       | acknowledging this shift. The free money is over, so we need to
       | get back to work and dbt is the least bad way to organize lots of
       | SQL for data management.
        
       | staticautomatic wrote:
       | I'm a newly minted head of analytics who transitioned from a
       | different domain, so I never had to muck my way through the MDS
       | but attentively watched others from the sidelines over the last
       | few years. Best I can tell, "the modern data stack" is just a
       | marketing phrase invented by a cadre of vampire vendors. The
       | lessons I learned watching others translated into a few simple
       | requirements for our nascent "stack" that most importantly
       | include transparent pricing I can reason about and divvy up, as
       | many integrations as possible so I can minimize rolling my own,
       | and a straightforward framework for ETL code. These three
       | requirements plainly disqualify most of the MDS universe.
       | 
       | With the benefit of starting basically from scratch and not
       | having to mess around with real-time analytics, it's pretty easy
       | to ignore the MDS vendors. So far I've landed on BigQuery,
       | AirByte, GitHub, BI Engine, Looker Studio, and Pandas 2.x or
       | DuckDB for local stuff. I send as many things as possible
       | straight to BQ, lock junior analysts out of gigantic tables,
       | archive periodically to partitioned parquet files in cold
       | storage, use mostly turnkey integrations, and ruthlessly
       | prioritize custom ETL jobs. Putting GitHub in the mix isn't super
       | ergonomic and we may be in the market for new tools once we cross
       | the "big data" frontier, but that'll be a while from now. I'll
       | probably never know or care what the MDS vendors think I'm
       | missing.
        
         | iamacyborg wrote:
         | AirByte was definitely one of the MDS vendors.
         | 
         | They literally got a huge investment because their deck bragged
         | about having a few thousand Github stars and a few thousand
         | folks in a Slack channel.
        
           | cookie_monsta wrote:
           | I can't tell if this is envy or scorn
        
             | benjaminwootton wrote:
             | I don't think it's either.
             | 
             | Modern Data Stack is a set of characteristics - cloud
             | based, SaaS, consumption based pricing, ELT bias, open
             | source bias, component based etc.
             | 
             | AirByte is all of them. It's just a fact rather than a
             | criticism.
        
               | staticautomatic wrote:
               | I think this is fair and I don't doubt that AirByte
               | jumped on the MDS marketing bandwagon, but the fact that
               | their pricing isn't inscrutable sets them apart in my
               | view.
        
             | iamacyborg wrote:
             | Neither, just an interesting factoid that highlights just
             | how zany the VC landscape was in the MDS space just a few
             | years ago.
             | 
             | https://airbyte.com/blog/the-deck-we-used-to-raise-
             | our-150m-...
        
             | tomrod wrote:
             | Github stars used to be a signal of quality, and velocity
             | of those stars indicated a winning product. But like all
             | rating systems, they are for sale to the nearest Mechanical
             | Turk or GenAI.
        
         | volderette wrote:
         | Airbyte is definitely one of the MDS vendors. Plus they have a
         | ton of bugs because their only focus is having the most
         | connectors on the market. A lot of them are broken or badly
         | implemented.
        
           | staticautomatic wrote:
           | I don't doubt the bugs. In fact, I expect them because so
           | many of the connectors are community contributed and so I try
           | to think of it as more like GitHub than a curated set of
           | expertly-built connectors. That said, for the time being it's
           | more important that I not break my other rules. The promise
           | of a perfectly working connector (if such a thing exists) in
           | exchange for unknowable pricing or a SDK I'm gonna have a
           | hard time training people on is a tradeoff I feel I can't
           | reasonably make right now, but I'm very open to the idea that
           | my calculus may change.
        
         | kwillets wrote:
         | You're on the right track with pricing. What I find ubiquitous
         | about MDS people is a lack of old-school ideas like
         | benchmarking and price-performance. Cloud salespeople talk
         | endlessly about your company becoming data-driven until you
         | bring up cost data.
        
       | jonmoore wrote:
       | The Modern Data Stack / MLOps product space was succinctly
       | described by one actually-technical CEO as "vending into
       | ignorance"; the author corroborates this with a commendably
       | candid take:
       | 
       | > _Imagine it's 2021, peak MDS, and you meet the CDO of a large
       | bank. "Oh cool," she says, "you're the CEO of a tech company.
       | What does your product do?" What do you say?_
       | 
       | > _"We build a tool that leverages the power of the cloud to
       | apply standard SQL and software engineering best practices to the
       | historically mundane (but critical!) job of data
       | transformation."_
       | 
       | > _"We're the standard for data transformation in the modern data
       | stack."_
       | 
       | > _I will tell you that, empirically, option #2 is more
       | effective._
       | 
       | This tallies with what I've seen from a lot of enterprise CxOs
       | and their teams as technology hype moved from big data and block
       | chain and onto data science/machine learning.
       | 
       | There is so much to write about this, but I'll just recommend
       | "Life Cycle of a Silver Bullet"
       | http://freyr.websages.com/Life_Cycle_of_a_Silver_Bullet.pdf,
       | which deserves more attention than it's had on HN.
        
         | lulznews wrote:
         | Is MLOps more or less of a thing than prompt engineer?
        
           | SheddingPattern wrote:
           | More of a thing, but it's mostly DevOps.
        
           | jochem9 wrote:
           | MLOps is deploying, monitoring and (re)training ML models.
           | Sits in the DevOps and data engineering space.
           | 
           | Prompt engineering is making generative AI do what you want
           | by crafting the right context. I would put it somewhere in
           | the software and data engineering space, given they will most
           | likely integrate applications with it. MLOps comes into play
           | if you have your own trained or tuned model.
        
       | m0llusk wrote:
       | compiled a quick list of terms to help understand this piece:
       | 
       | MDS - modern data stack
       | 
       | BI - business intelligence
       | 
       | Redshift - amazon redshift data warehouse service product handles
       | large scale data sets and database migrations can handle analytic
       | workoads on big data sets with column oriented approach built on
       | top of massive parallel processing (MPP) from data warehouse
       | company ParAccel (later acquired by Actian)
       | 
       | ETL - extract, transform, load
       | 
       | ELT - extract, load, transform -- an alternative to ETL that
       | stores raw data
       | 
       | Looker - does BI, started with looker data sciences, acquired by
       | google
       | 
       | Tableau - data visualization focused on business intelligence
       | 
       | clickstream data - user website navigation records
       | 
       | snowflake - an MDS data platform solution
       | 
       | mongo - nosql database
       | 
       | datadog - monitoring, analytics for devops
       | 
       | confluent - real time data streams
       | 
       | databricks - web based cluster management and data lakes for
       | machine learning
       | 
       | meme-ification - reduction of an idea to cartoons
       | 
       | peak - highest point preceeding drop off
       | 
       | fivetran - data ingestion
       | 
       | dbt - dbt labs, data transformation
       | 
       | CDO - chief data officer
       | 
       | co-marketing - integrated marketing such as linked brands
       | 
       | co-sponsored - cooperative sponsorship
       | 
       | co-selling - shared sales assets including data records
       | 
       | ARR - annual recurring revenue
       | 
       | ZIRP - zero interest rate policy
       | 
       | private multiples - private company valuation calculations,
       | metrics
       | 
       | forward revenue - revenue expected in the future
       | 
       | PowerBI - microsoft business analytics
        
       ___________________________________________________________________
       (page generated 2024-02-12 23:02 UTC)