[HN Gopher] Is the "modern data stack" still a useful idea?
___________________________________________________________________
Is the "modern data stack" still a useful idea?
Author : tim_sw
Score : 118 points
Date : 2024-02-11 20:58 UTC (1 days ago)
(HTM) web link (roundup.getdbt.com)
(TXT) w3m dump (roundup.getdbt.com)
| tomrod wrote:
| Intriguing that the CEO of DBT declares MDS ("modern data stack")
| to be a meaningless term (i.e. no longer useful). Maybe a good
| demarcation that we are entering the post-modern phase of cloud
| tech generally? Cloud still there, but not the only option.
|
| I look forward to this happening with GenAI -- the tools we come
| up with will be pretty cool in the long run. I hope we can find
| ways around platform enshittification because that really sucks.
| Similar for blockchain tech.
|
| I see a common thread in these techs too -- the same type of
| LinkedIn influencers shop these during their hype phases.
| Ultimately, useful technology and best practices come out of it!
| actionfromafar wrote:
| Well put! Nothing's modern anymore, the tech cycles are closing
| in on one another so much the old one hasn't gone out of phase
| until it's overcome by not one but several new ones. In a nice
| touch of self-refentiality, your message itself is very post-
| modern.
| tomrod wrote:
| I hadn't considered that. Thanks for pointing that out to me
| :)
| karakanb wrote:
| Disclaimer: I am the co-founder of a competitor, Bruin
| (https://getbruin.com). We are exactly the kind of integrated
| platform Tristan is talking about.
|
| The article resonates with me a lot, and it is because we called
| this out months ago. I find the idea of Tristan walking back on
| the premise of MDS and claiming it to be not useful anymore funny
| because they were one of the main drivers of the term and the
| whole hype around it.
|
| He even acknowledges that they played ball with other companies
| in the space:
|
| > There was a lot of valuable co-marketing, partnership deals,
| co-sponsored events, and co-selling. This had real value for
| everyone involved--customers and vendors alike. Companies
| voluntarily integrated their products together, cross-promoted
| each other publicly, and built partnerships that made owning and
| operating these technologies far easier for customers.
|
| Sorry, but no, this didn't have value for the customers, only for
| the vendors. The customers were left alone by themselves to deal
| with all of this complexity, and the vendors made a lot of money
| off of that. They convinced companies that they needed a bunch of
| different tools to build a simple pipeline, and rode on the wave
| of huge valuations based on these same ideas that they are
| walking back on.
|
| The companies prefer integrated solutions now because they woke
| up. Instead of paying 150k/y each to Fivetran, dbt and whatnot,
| they realize that they are better of just hiring an engineer or
| two in the worst case. It is 2024, and none of these tools still
| properly talk to each other. Do you want to get an end-to-end
| lineage of your data? Good luck with that. How about quality? How
| about governance? Companies are left alone with this hype cycle.
|
| I'd claim that a significant part of the blame lies on the
| executives and leaders in the companies, who just jumped on the
| ship for the sake of building their CVs and skipped the critical
| thinking step. None of them seriously asked themselves the
| question of whether or not it makes sense.
|
| To be honest, I feel sorry for all the money spent on building
| solutions around all of these hype-driven products.
| neeleshs wrote:
| Disclaimer : I'm the cofounder of Syncari. I agree 100%. The
| whole notion of modern data stack is a myopic , neither here
| nor there philosophy to address data in my opinion.
|
| Just take a look at the sheer number of tools one has to run -
| ETL(or more commonly, data dumpers), Wearhouse, transform
| tool,"reverse ETL", data quality tools, BI. And if you want
| some ML, that's whole another game.
|
| We also believe strongly in an integrated approach to this
| space, with a data model at the center of it all.
| gigatexal wrote:
| Deleted
| tomnipotent wrote:
| He's not "walking back on the premise of MDS", but acknowledges
| that the term "MDS" had been co-opted and lost its descriptive
| fidelity.
|
| > Sorry, but no, this didn't have value for the customers, only
| for the vendors
|
| Absolutely untrue. I've discovered multiple vendors over the
| last two decades because of co-marketing and partnership deals,
| and I know many others that have. My success rate with these
| vendors is no different from vendors I found myself or were
| referred to me by others.
|
| > The customers were left alone by themselves
|
| This is true with every vendor integration with every company
| on the globe. You either 1) have internal expertise to push
| through the pain, 2) hire support from the vendor itself (often
| X hours come baked into the deal), or 3) hire a consultant.
| This is as true for the MDS vendors as it is for Salesforce,
| AWS, or an ERP. It's true for your company, too.
|
| > They convinced companies that they needed a bunch of
| different tools
|
| Except you do? You need a database, you need something to
| extract data and move it locally, you need something to
| transform that data into something useful, and you need the
| ability to deliver that data to end users. This was true twenty
| years ago, true ten years ago, and still true today. The all-
| in-one vendors prior to MDS were horrible, and I'd rather cut
| off my arm then use something like Pentaho again. I'm also not
| convinced the current all-in-one vendors are any better.
|
| One of the advantages of the pick-your-stack MDS was that you
| could tailor the tools to you organizations specific needs,
| rather than have to fight against whatever your mega-vendor
| offers. I loved the fact that I could use open source for 90%
| of my stack, and then use PowerBI for the last 10%. Or I could
| be 100% FOSS and deliver data with Metabase. Or I could use
| Tableau for 50%, and FOSS for the other 50%.
|
| > they realize that they are better of just hiring an engineer
| or two in the worst case
|
| This simply isn't true for every size or class of company. Many
| businesses are better off with vendor-supported platforms than
| trying to build something bespoke, which can quickly become a
| liability with staff turnover or lack of strong internal
| technical leadership.
|
| > It is 2024, and none of these tools still properly talk to
| each other
|
| What does this even mean? How does a data extract tool "talk"
| to a transformation tool? Or "talk" to the database? I can see
| from the screenshot on your website that it offers a GUI with
| drag-and-drop - is that your definition of "talking"? Or are
| you referring to a tight integration that doesn't require
| additional work to glue them together?
|
| > How about quality? How about governance?
|
| These are management and organization issues, not tooling. I've
| seen amazing quality and governance with teams working in
| MSSQL/SSIS, and I've seen horrible quality and governance from
| teams using Hadoop/Spark. A tool can make a competent
| practitioners job easier, but it's not going to make an
| incompetent competent.
| cgio wrote:
| As someone who has worked in this space, I would intuitively go
| with the decoupled tools every time. The toughest thing is
| matching a monolith of opinions with your requirements.
| Individual tools leave design space in their integration and
| thus allow proper engineering. That they allow it, does not
| mean it happens. If you go with a shopping bag architecture
| then you might be better off results wise with an integrated
| solution, but still, when the real engineers come along they'll
| have an uphill battle to uproot an integrated solution vs
| redesigning the combination of individually good tools.
| CRConrad wrote:
| > I would intuitively go with the decoupled tools every time.
|
| Seems to me that the problem is, each of these "decoupled
| tools" you mention isn't some specialised little awl or
| hacksaw ("the Unix philosophy..."), nor even a do-anything
| Swiss Army knife -- but a whole stack of cloud services,
| "integrations", and administration tools in its own right. A
| whole machine shop of machine shops. Building your own custom
| setup out of that is hard. You'll end up with something at
| least as unwieldy as a _really_ big do-anything single-vendor
| mega-machine shop stack of cloud services, "integrations",
| and administration tools.
| extr wrote:
| What would HN commenters define as a truly modern data stack
| right now? Mostly asking what this chart
| https://a16z.com/emerging-architectures-for-modern-data-infr...
| would look like in 2024.
| rch wrote:
| - Postgres and Kafka - Airflow and Flink -
| Kudu+Impala or Clickhouse - Iceberg/Parquet or Delta
| - Ozone(HDFS) or Ceph - Spark Rapids and/or Ray+Metaflow
| - Open Metadata or Atlas+Ranger
| fifilura wrote:
| TBH that sounds like a heap of buzzwords not useful for
| anyone (but maybe an architect promoting his career).
| rch wrote:
| Pretty much, but updated for 2024.
| mr_toad wrote:
| > What would HN commenters define as a truly modern data stack
| right now?
|
| Just don't spend too much time debating what it is, or the
| analysts who need numbers yesterday will develop their own
| solutions before you've even published your A3 architecture
| diagram.
|
| A bird in the hand, as they say.
| te_chris wrote:
| IT DEPENDS. But for standard ecom/ops stuff (and this is
| heavily biased by the fact this is what I know and have
| deployed with low overhead): Big query. This is fed by an ETL
| pipeline of the stuff you need. dbt to ELT into useful,
| reportable tables. Something to visualise - I still like Mode,
| but there's lots of new BI toys around. KISS.
|
| This is extremely low maintenance at the actual requirement
| levels of most businesses with < 50 employees.
| neighbour wrote:
| - Fivetran
|
| - Snowflake/BigQuery/Redshift
|
| - dbt
|
| - Looker/Metabase/Tableau
| dm03514 wrote:
| Sorry nothing positive to say here.
|
| I've been using dbt and MDS for nearly 3.5 years and I believe
| the entire approach is profoundly broken. There's really nothing
| "modern" about it, especially compared to software engineering.
|
| https://on-systems.tech/blog/135-draining-the-data-swamp/
|
| Building on Extracted operational data is hard at best and a
| business altering security liability at worse.
|
| I believe The MDS trails at least a decade behind modern software
| engineering practices, lacking industry guidance and generic
| tooling to support: CI/CD, versioned deployment artifacts, zero
| downtime deployments, unit testing, observability, monitoring and
| alerting.
|
| MDS Data engineering is a meme in the industry, the expectation
| of "100%" combined with the lack of modern tooling makes success
| really hard to achieve.
| civilized wrote:
| I was a fan of dbt for a while, but the shine wore off when I
| saw one of my smartest coworkers try to use it. His development
| speed was about an order of magnitude slower than what I expect
| from data scientists using dataframe packages like R's dplyr,
| Python's polars, or Spark DataFrames. In my own experience, dbt
| is significantly better than writing raw SQL, but still nowhere
| near a normal software development experience. It needs an IDE
| badly, but the current IDE is cloud-only.
|
| I am currently building my analytics pipelines on top of R's
| dbplyr package, which allows me to run dplyr operations on
| database tables just as if they were local data frames. Since
| it's R, it's not without warts, but at least I get to use a
| real programming language and the free and fantastic RStudio
| IDE.
| tomnipotent wrote:
| > allows me to run dplyr operations on database tables
|
| Which require that data take a round trip between the R
| process and the database. It's not uncommon for these jobs to
| spend more time in the read/write step than doing meaningful
| work, and why I prefer dbt when possible to keep
| transformations as close to the data as possible.
| nograpes wrote:
| The second point the grandparent made was that they were
| using _dbplyr_ , which allows you to avoid having the data
| take a round trip between the R process and the database.
| golergka wrote:
| We built the perfect IDE for dbt at Deep Chanel, with
| autocomplete automatic real time type checking for the whole
| project.
|
| Sadly, the company closed this June, and even the website is
| already down.
| neighbour wrote:
| You need to find a way to release this product. I would pay
| to use it.
| golergka wrote:
| It's already been released and publicly available for
| free for almost a year.
| CRConrad wrote:
| So where is it, if "even the website is already down"?
|
| ETA: Seems it isn't. https://www.deepchannel.com/ That
| what you meant?
| golergka wrote:
| Hmm, may be I've had network issues last time I checked.
| Anyway, that's the only place you can get it -- it's a
| completely free desktop app, but it's not open source.
| civilized wrote:
| This is extremely interesting. You might be onto something.
| jochem9 wrote:
| Imo dbt sits in the analysts space. SQL is the language they
| know and dbt is the best solution to scale that.
|
| In that sense it should only be used for last mile data
| transformations. Not for wrangling raw data into neat tables.
| fifilura wrote:
| I am curious, what do you perceive as the difference
| between data transformations and "wrangling raw data into
| tables"?
|
| Edit: I am also curious why you consider SQL to be inferior
| for the latter?
| camgunz wrote:
| The core divide in this space is SQL vs. $OTHER_LANGUAGE. If
| you're fluent in SQL you'll be great at dbt. If you're only
| fluent in not-SQL you won't.
|
| A lot of these discussions are basically "can we please just
| use the language/tooling we know", and sometimes the answer's
| yes! A lot of people breathe a big sigh of relief when they
| discover you can use like, Spark through Python and what-not.
| Regardless of whether or not it's a good idea or good use of
| resources, this space is advancing because the market demands
| it. But like, I don't think you're ever gonna use dbt with
| not-SQL; it's all in on SQL. Feels like maybe it was the
| wrong fit for your team.
|
| I should say I'm not sure if your coworker was an SQL person
| so dunno if this directly applies. Just saying what my
| experience has been.
| gigatexal wrote:
| Y'all don't have CI/CD? Maybe it's hyped but we do simple
| stuff. Snowflake schema DWH. Jinja2 templated SQL or dbt.
| Airflow with a monolith tool that does transforms and such. Is
| it perfect? No. Is it understandable? Very much so.
|
| Testing is such an interesting concept in data engineering. One
| needs consistent test data. We aim to implement that with
| snapshots eventually but now we have sanity checks at each
| layer.
|
| > lacking industry guidance and generic tooling to support:
| CI/CD, versioned deployment artifacts, zero downtime
| deployments, unit testing, observability, monitoring and
| alerting.
|
| Ci cd is doable and easy: changes are deployed via GitHub
| actions, the CI part is a bit missing I guess without tests.
| Versioned artifacts also, add a tag to the airflow job to know
| what tagged version of the code is running. Zero downtime
| deployments we do that all day -- I mean using views we can a/b
| deploy changes to the underlying tables and then do a simple
| schema change to the view and nobody knows the difference. Unit
| testing still yet to be done. Observability and alerting we use
| internal dashboards and sentry.
|
| What am I missing?
| Exoristos wrote:
| Relevance and actual value, I'd assume.
| antupis wrote:
| Testing and monitoring in every aspect are the only things
| that are greatly lacking, you always need to glue something
| together to get nice tests or some kind of data contract
| monitoring.
| xyzzy_plugh wrote:
| The mistake everyone makes is treating the entire space as if
| it's somehow different than the rest of software engineering,
| when it is precisely exactly the same.
| datavirtue wrote:
| This. Nearly everyone is confused.
| esafak wrote:
| Machine learning people have cottoned on to this, so MLOps
| was invented.
| tomrod wrote:
| The training part of MLOps is important to ensure
| replicability and other desirable properties or the ML
| artifact, but the rest is clearly good CI/CD and
| observability.
| wodenokoto wrote:
| We had an application engine come in and run an ML project.
|
| The first thing he did was remove real data from training, as
| he was shocked to see that development wasn't done against
| dummy data.
|
| And once the dev environment is running on real data, the
| changes in how you develop and operationalize just seems to
| cascade in my experience.
| fifilura wrote:
| I am slightly confused by your reply. Can you elaborate a
| bit, was it good or bad?
| wodenokoto wrote:
| You can develop a database schema for your application on
| dummy data, and test your business logic on dummy data.
| But you can't develop a machine learning model on dummy
| data.
|
| It is best practice to keep live data out of your
| development environment for normal software development,
| but it is impossible for machine learning projects.
| fifilura wrote:
| Yes, this is what I thought and exactly what I am
| experiencing right now.
|
| It is somewhat frustrating to see the
| engineering/platform team spending time on this kind of
| abstraction that is essentially useless for doing the
| work we need to do.
| tomrod wrote:
| Yep. I advise my clients to have a controlled, sandbox
| partition of prod. Call it the whiteboard or what have
| you, things can be erased or you can cache model version
| runs. Your data transforms can happen in dev but training
| should happen against real data. Synthetic data can be a
| compromise but that can sharply bias your model and,
| since it a harder lift, often multiplies data prep time
| appplication wrote:
| Well... no I would disagree. I am a data engineer who
| relentlessly pushes for more classic SWE best practices in
| our engineering teams, but there are some fundamental
| differences you cannot ignore the data engineering that are
| not present in standard software engineering.
|
| The most significant of these is that version control is not
| guaranteed to be the primary gate to changes to your system's
| state. The size, shape, volume, skew of your data can change
| independently of any code. You can have the best
| unit/integration tests and airtight CI/CD but it doesn't
| matter if your date source suddenly starts populating a
| column you were not expecting, or the kurtosis of your data
| changes and now your once performant jobs are timing out.
|
| Compounding this, there is no one answer for how to handle
| data changes. Sure you can monitor your schema and grind
| everything to a halt the second something is different in
| your source, but then you're likely running a high false
| positive rate, and causing backup or data loss issues for
| downstream consumers.
|
| The truth is change management and scaling is more or less a
| predictable and solved problem in traditional SWE. But the
| same principles can't be applied with the same effectiveness
| to DE. I have seen a number of SWE-turned-DE struggle
| precisely because they do not believe there to be any
| difference between the two. It would be nice if it were so
| simple.
| xyzzy_plugh wrote:
| You are massively oversimplifying "classic SWE" here.
| Ultimately it makes no difference.
|
| How you handle contract violations between dependencies and
| third parties is not an inherently different problem. What
| _should_ you do if a data source changes schema? How do you
| detect it, alert, take action? How are users or customers
| or consumer jobs impacted? This is all just normal
| software.
|
| In many ways it feels like you're making my case for me.
| fifilura wrote:
| Whenever I get confronted with SWE best practices it is
| all about version control, unit tests, integration tests
| and CI/CD.
|
| What are some good resources for the other things you
| describe?
| jimberlage wrote:
| I don't have great links handy, but read up on
| observability.
|
| You'll want to read up on what people do on error,
| especially in the case that contracts change (search "api
| versioning".)
| coldtea wrote:
| > _The most significant of these is that version control is
| not guaranteed to be the primary gate to changes to your
| system's state. The size, shape, volume, skew of your data
| can change independently of any code._
|
| That has been true in SWE since the days of ENIAC.
| kgdiem wrote:
| > CI/CD, versioned deployment artifacts, zero downtime
| deployments, unit testing, observability, monitoring and
| alerting.
|
| This definitely stood out to me when I was working with a lot
| of Snowflake's new products, especially their "native apps".
|
| I started on some bespoke tooling given the lack of everything
| that you mentioned and at the time was thinking about how to
| productize it / turn it into some open source tools but I've
| not gotten back around to it since leaving that job.
|
| I had a really great experience working with sqlglot from
| Tobiko Data https://tobikodata.com/ -- haven't had a chance to
| check out their sqlmesh project yet but I have some faith in
| their work.
| camgunz wrote:
| Yeah seconding sqlglot, has saved me _tons_ of time. I think
| yeah, they're smart and onto something here.
| benjaminwootton wrote:
| dbt was a huge leap forward in this regard.
|
| It enables source code control, modularity, testing, CI/CD,
| environments, documentation, git branching, local development
| experience.
|
| Not a fanboy, but it's who reason for being is basically to
| enable "software engineering practices for data".
| antupis wrote:
| Yup, it is not perfect but I will take dbt every day against
| the old way as some random DDL files in git.
| CRConrad wrote:
| So basically its huge advantage is that it saves its code
| as text, which makes it gittable?
|
| (Dang, and here the tool I've been designing in my head for
| several years was going to claim that niche... So unique.
| ;-)
| camgunz wrote:
| Nah, the main thing is the dependency resolution. So you
| make pipelines like:
|
| - employees
|
| - tall_employees
|
| - tall_employees_by_salary
|
| These tables (dbt would call them models because they can
| also be views depending on how they're configured) depend
| on each other, so you have to build them in a specific
| order. Without something like this, you're manually
| running SQL to build your pipelines, and that's very
| error prone. With something like this you can run tests
| (they're also SQL in dbt, which is pretty common in data
| engineering teams), run your pipeline in CI, etc.
|
| There's a bunch of other features, but they all hang off
| this basic idea.
| riku_iki wrote:
| > you're manually running SQL to build your pipelines
|
| you can have some sh file which runs SQL in desired
| order..
| williamcotton wrote:
| It seems that there's a philosophy with these kinds od cloud data
| services: measure everything then figure out what to measure
| after.
|
| It seems that the better approach is to make a hypothesis first,
| collect only the data you need, and then analyze.
|
| Otherwise it seems the indirect costs of retooling around
| "measure all the things indiscriminately" and the direct costs of
| paying for such services don't seem worth it.
|
| Can someone offer a reasonable rebuttal?
| staticautomatic wrote:
| It doesn't always make sense to "collect only the data you
| need." For example, someone in our org will ostensibly have a
| good reason for setting up an event in Google Analytics. As the
| head of analytics I have no use for their event but I'm going
| to end up ingesting it into BigQuery anyway because the built-
| in GA4->BQ transfer sends ALL the data, and I'm fine with that
| because it's damn near free for me to ingest and store. I could
| instead "collect only what I need" by running the transfer
| through an ELT tool and filtering out that event, but why would
| I bother doing that work and paying the ELT vendor money when
| Google will do it for free and charge me like $1K/year to store
| 100M rows of everything collected in GA4?
| williamcotton wrote:
| That cost analysis seems incomplete. How much did it cost to
| implement in the first place? That seems like millions of
| dollars in capitalization before service expenses are even
| considered!
| staticautomatic wrote:
| Sorry but I can't tell if you're being sarcastic or not.
| Google Analytics and BigQuery are both nominally free, the
| service expense is $1K/year, and the implementation cost
| literally a few minutes of my time.
| williamcotton wrote:
| I'm not being sarcastic. There is engineering labor that
| requires hooking up all of these events. Then there is
| maintenance costs. And many more costs, direct or
| otherwise.
| staticautomatic wrote:
| With respect to the example I gave, no, there actually
| isn't. In fact I don't even understand what you think the
| work and costs involved are.
| williamcotton wrote:
| I've seen organizations spend millions just in
| engineering labor costs for such endeavors. But sure, for
| small companies with limited domain events to begin with,
| the costs are relatively lower, and probably with the
| lower domain complexity it is less of a share of
| engineering time overall.
|
| But there are still opportunity costs! It's not free to
| install dependencies across numerous code bases and hook
| into and label the events, even for a small business.
| fifilura wrote:
| OP suggests using "built-in GA4->BQ transfer [to send]
| ALL the data".
|
| This seems like an extremely small endeavor.
| williamcotton wrote:
| So yes, for a small business that just drops a GA snippet
| into the bottom of their app it is inexpensive. I'm not
| sure that this qualifies as the "modern data stack".
|
| That's also not what happens when a larger company
| integrates Segment into their applications! That takes
| many engineering months to set up and then maintain for
| the lifetime of the app.
|
| If the full costs are analyzed, is the money spent well?
|
| There's sone irony here in that much effort is put
| towards analysis of user behavior but very little effort
| put towards accounting for where the money is spent
| internally.
|
| That is why there is a slowdown in customers using these
| MDS products, not what the author of the article claims,
| rather that when feet are put to the fire, expensive
| services that haven't paid off their debt are given the
| axe.
| fifilura wrote:
| I think i can spot the difference in interpretation here.
|
| GG*GGP "someone in our org will ostensibly have a good
| reason for setting up an event in Google Analytics"
|
| In this case they already have Google Analytics. The
| question is whether it was worth exporting that data to
| BigQuery.
|
| Then whether it was worth setting up GA in the first
| place is not discussed.
| williamcotton wrote:
| Isn't the setup cost a key part of the total cost of
| analytics, especially when it comes to these MDS systems?
|
| I encourage anyone in a position to make purchases or who
| has a say in what engineer labor focuses on is to get a
| copy of Horngren's Cost Accounting. It should make my
| position very clear!
|
| FWIW, there is indeed capitalization occurring when an
| organization incorporates analytics, even the "measure
| all the things" approach, but the question is, once the
| total costs are discovered, does the return outweigh the
| expenditure?
|
| These things cannot be considered in isolation,
| especially when using advanced approaches like Activity-
| Based Costing, something that is basically a requirement
| for a product or service that is fundamentally software
| in nature. There is too much complexity to treat
| managerial accounting in a software setting like a steel
| mill.
|
| The reason why this doesn't happen is YOLO business
| management fueled by the casino that is VC backed
| companies. But we are in a rapidly maturing industry
| where margins are slimming and costs need to be reigned
| in, especially in publicly traded companies where
| investors will punish a company that makes wildly off
| target financial projections discovered when cash flows
| are reported in the following quarters.
|
| I encourage everyone in all management positions to start
| thinking about and tracking actual costs. Not story
| points, not customer use patterns, but cold hard cash!
| neighbour wrote:
| The guy you're arguing with in the reply thread has a point but
| I think you're talking past each other because you're coming at
| it from two different sides. I understand his point about
| syncing everything if you're a small enough company or the data
| is small enough because the cost is immaterial.
|
| I agree with most of what you're saying and have seen the
| downfalls of "just sync all our data" first hand. At a late
| stage start-up I worked at, we would often have teams (finance,
| CX, etc.) implement a shiny new tool and immediately request to
| have that data accessible for analytics/reporting. If we did
| blindly agree, the data would often never be used and it would
| run up our Fivetran costs to sync the data and our Github
| actions bill to build the dbt models and include it in our
| modeling.
|
| We eventually landed on an approach that considered the
| following:
|
| - Does the tool already have reporting baked in? If so, just
| use that until you need more.
|
| - How often are you really going to be using it and what makes
| this data so special that you need to join it to our other
| data?
|
| - How much work will be involved to implement a sync for this
| data? Does Fivetran or Segment have a connector readily
| available?
|
| - If the tool doesn't have reporting built-in and the team did
| have a genuine business case for the data being synced, we
| would figure out which specific objects from the source that
| needed to be synced and sync them at a fairly infrequent
| cadence (once a day).
|
| So while I offer no reasonable rebuttal, I wanted to explain a
| happy medium that worked for us.
| Animats wrote:
| Oh, it's something for the ad industry.
|
| Ads are watched only by people who don't have good ad blockers.
| hn72774 wrote:
| DBT in simplest terms is just a way to orchestrate
| transformations in dependency order. Table "A" needs to be
| updated before table "B," because "B" selects from "A." It works
| well for that.
|
| I've seen it used to get wild west SQL logic into version
| control. To replace scheduled SQL workbooks running entire data
| pipelines.
|
| I've also seen over-engineered integrations with other
| orchestration tools in the "stack."
|
| Once a DBT project grows to a certain size, roughly 500-1000 .SQL
| files, it gets hard to manage. Not impossible, though it takes
| intentionality about how to group things together and scale them
| out operationally.
|
| "Slim CI" is a recent buzzword that has some nice ideas about
| build automations.
|
| Cosmos is supposed to solve a lot of the automation gripes with
| dbt. I haven't tried it yet but would like to.
| https://www.astronomer.io/cosmos/
| gms wrote:
| I've been in data and analytics for over a decade and co-founded
| a consolidated-ETL company (https://www.polytomic.com).
|
| Nice to see Tristan realising this (people should be commended
| for changing their minds). There were three problems with this
| term:
|
| 1. It mostly resonated with VCs and industry observers whose jobs
| are to peddle in buzzwords, rather than users who simply want
| their problems solved.
|
| 2. It was (and is) ill-defined. Ask a group of people in the
| industry to define it and you're guaranteed to get different
| answers.
|
| 3. It committed the cardinal sin of using the adjective 'modern'
| in a noun. At some point, everything today stops being modern.
| Couple this with (2) and your term is now even more meaningless.
|
| At a mundane level there's no drama here: just another example in
| the long list of noise contributors to the free-money party that
| was going on during the Covid years (as the essay acknowledges).
| Lyngbakr wrote:
| I've worked in data for a while now as an Data Engineer, Data
| Scientist, and Director and working with experienced software
| engineers has highlighted to me that most of the data stack is
| fluff. All the layers/tools that are heaped upon one another just
| lead to complexity and dependence on paid services. I'm not
| advocating for reinventing the wheel, but rather that a bespoke
| solutions seem to be worth the cost. In short, I'd prefer to
| spend the budget on experienced developers who can build
| maintainable systems than on a plethora of MDS tools. YMMV,
| though.
| Waterluvian wrote:
| If you wouldn't mind: in your experience how often does the
| problem seem to stem from "we might need that later..." or "we
| don't have a concrete set of questions we want to answer?"
|
| I have considerably less experience but so far I have
| experienced a complicated data science stack being used as a
| way to handle a ton of data without clear questions and desires
| on how to act on possible answers. As if they're searching a
| haystack for... something?
| Lyngbakr wrote:
| IME, it has usually stemmed from "that's how everyone else is
| doing it, so we should too" without considering specific
| use/business cases. Senior management were simply not willing
| to deviate from the norm.
| tharkun__ wrote:
| And also: I need a job. My job/workplace has these tools.
| They are usable for my job and if I were to question the
| tools that would come on top of my workload, so I use them
| even if I know that cheaper tools were used at my last job.
|
| Do I care if they pay my salary and the tools are usable?
| tadfisher wrote:
| That's perfectly reasonable, but it sounds like you're
| saying the tools work for your use case. So long as the
| executive/management level is evaluating the cost-benefit
| of other tools, and is receptive to voices from ICs, then
| there's no problem to solve.
|
| It's just not what the GP is describing. In some orgs you
| end up with these expensive, standard-ish "suites" that
| are way too generalized, don't solve business needs as
| efficiently as other tools, are painful for ICs to use to
| meet requirements, and it seems like no one can really
| explain why these tools are being used over something
| else other than rumors (the previous CTO came from this
| vendor, we got a sweetheart deal 10 billing cycles ago,
| etc.).
|
| So really this an issue of ownership; someone needs to be
| accountable for infrastructure decisions and outcomes,
| and it's way healthier for an org if ICs are aligned with
| management on the same. It sounds like you're not in that
| sort of org, so good for you.
| CRConrad wrote:
| But do you care if they pay your salary and the tools are
| _barely_ usable?
| nerdponx wrote:
| There is no "the" data stack.
|
| Some organizations have complicated orchestration requirements,
| some don't. Some organizations have large amounts of
| heterogeneous data, some don't. Some organizations need to run
| large distributed machine learning tasks, some don't.
|
| If you adopt tools that don't actually solve real problems that
| your team/org faces, then of course you'll end up with a bunch
| of extraneous fluff.
|
| Your comment has the same flavor as the frequent HN meme of "I
| just use Vi and Tmux and a BSD server for my business and
| anything else is a sign of organizational failure and
| individual weakness".
|
| All of these tools exist for a reason. Where you get into
| trouble is losing focus on solving your real practical
| problems.
| Lyngbakr wrote:
| > Where you get into trouble is losing focus on solving your
| real practical problems.
|
| Absolutely agree. As I say YMMV, but _my_ experience is that
| often people have adopted tools because "that's what data
| teams use" rather than because it's the best tool for the
| job. Often, we could -- and should -- build an in-house
| solution that fits perfectly rather than buying (every month)
| an imperfect solution.
| CRConrad wrote:
| > Some organizations need to run large distributed machine
| learning tasks, some don't.
|
| I think almost no organizations need to run "large
| distributed machine learning tasks". Many just think they do,
| because it's the buzzword of the moment.
| nerdponx wrote:
| Also true, but I think maybe that's a great example of what
| I and OP are saying here. Do you _really_ need it? Or do
| you just think you might need it one day because people are
| talking about it on Medium, or because you had a nice chat
| with some sales rep? Or do you just want it on your resume?
| sagarm wrote:
| > All the layers/tools that are heaped upon one another just
| lead to complexity and dependence on paid services. I'm not
| advocating for reinventing the wheel, but rather that a bespoke
| solutions seem to be worth the cost.
|
| Which third party products do you use in your stack, then?
|
| Do you only use basic cloud building blocks (VMs, EBS, S3), or
| F/OSS components like Spark and Trino, or are higher level
| things like Tableau/Looker also part of the solution?
| datadrivenangel wrote:
| The Modern Data Stack is Dead! Long Live the Analytics Stack!\
|
| Interesting to see one of the strongest popularizers of the term
| acknowledging this shift. The free money is over, so we need to
| get back to work and dbt is the least bad way to organize lots of
| SQL for data management.
| staticautomatic wrote:
| I'm a newly minted head of analytics who transitioned from a
| different domain, so I never had to muck my way through the MDS
| but attentively watched others from the sidelines over the last
| few years. Best I can tell, "the modern data stack" is just a
| marketing phrase invented by a cadre of vampire vendors. The
| lessons I learned watching others translated into a few simple
| requirements for our nascent "stack" that most importantly
| include transparent pricing I can reason about and divvy up, as
| many integrations as possible so I can minimize rolling my own,
| and a straightforward framework for ETL code. These three
| requirements plainly disqualify most of the MDS universe.
|
| With the benefit of starting basically from scratch and not
| having to mess around with real-time analytics, it's pretty easy
| to ignore the MDS vendors. So far I've landed on BigQuery,
| AirByte, GitHub, BI Engine, Looker Studio, and Pandas 2.x or
| DuckDB for local stuff. I send as many things as possible
| straight to BQ, lock junior analysts out of gigantic tables,
| archive periodically to partitioned parquet files in cold
| storage, use mostly turnkey integrations, and ruthlessly
| prioritize custom ETL jobs. Putting GitHub in the mix isn't super
| ergonomic and we may be in the market for new tools once we cross
| the "big data" frontier, but that'll be a while from now. I'll
| probably never know or care what the MDS vendors think I'm
| missing.
| iamacyborg wrote:
| AirByte was definitely one of the MDS vendors.
|
| They literally got a huge investment because their deck bragged
| about having a few thousand Github stars and a few thousand
| folks in a Slack channel.
| cookie_monsta wrote:
| I can't tell if this is envy or scorn
| benjaminwootton wrote:
| I don't think it's either.
|
| Modern Data Stack is a set of characteristics - cloud
| based, SaaS, consumption based pricing, ELT bias, open
| source bias, component based etc.
|
| AirByte is all of them. It's just a fact rather than a
| criticism.
| staticautomatic wrote:
| I think this is fair and I don't doubt that AirByte
| jumped on the MDS marketing bandwagon, but the fact that
| their pricing isn't inscrutable sets them apart in my
| view.
| iamacyborg wrote:
| Neither, just an interesting factoid that highlights just
| how zany the VC landscape was in the MDS space just a few
| years ago.
|
| https://airbyte.com/blog/the-deck-we-used-to-raise-
| our-150m-...
| tomrod wrote:
| Github stars used to be a signal of quality, and velocity
| of those stars indicated a winning product. But like all
| rating systems, they are for sale to the nearest Mechanical
| Turk or GenAI.
| volderette wrote:
| Airbyte is definitely one of the MDS vendors. Plus they have a
| ton of bugs because their only focus is having the most
| connectors on the market. A lot of them are broken or badly
| implemented.
| staticautomatic wrote:
| I don't doubt the bugs. In fact, I expect them because so
| many of the connectors are community contributed and so I try
| to think of it as more like GitHub than a curated set of
| expertly-built connectors. That said, for the time being it's
| more important that I not break my other rules. The promise
| of a perfectly working connector (if such a thing exists) in
| exchange for unknowable pricing or a SDK I'm gonna have a
| hard time training people on is a tradeoff I feel I can't
| reasonably make right now, but I'm very open to the idea that
| my calculus may change.
| kwillets wrote:
| You're on the right track with pricing. What I find ubiquitous
| about MDS people is a lack of old-school ideas like
| benchmarking and price-performance. Cloud salespeople talk
| endlessly about your company becoming data-driven until you
| bring up cost data.
| jonmoore wrote:
| The Modern Data Stack / MLOps product space was succinctly
| described by one actually-technical CEO as "vending into
| ignorance"; the author corroborates this with a commendably
| candid take:
|
| > _Imagine it's 2021, peak MDS, and you meet the CDO of a large
| bank. "Oh cool," she says, "you're the CEO of a tech company.
| What does your product do?" What do you say?_
|
| > _"We build a tool that leverages the power of the cloud to
| apply standard SQL and software engineering best practices to the
| historically mundane (but critical!) job of data
| transformation."_
|
| > _"We're the standard for data transformation in the modern data
| stack."_
|
| > _I will tell you that, empirically, option #2 is more
| effective._
|
| This tallies with what I've seen from a lot of enterprise CxOs
| and their teams as technology hype moved from big data and block
| chain and onto data science/machine learning.
|
| There is so much to write about this, but I'll just recommend
| "Life Cycle of a Silver Bullet"
| http://freyr.websages.com/Life_Cycle_of_a_Silver_Bullet.pdf,
| which deserves more attention than it's had on HN.
| lulznews wrote:
| Is MLOps more or less of a thing than prompt engineer?
| SheddingPattern wrote:
| More of a thing, but it's mostly DevOps.
| jochem9 wrote:
| MLOps is deploying, monitoring and (re)training ML models.
| Sits in the DevOps and data engineering space.
|
| Prompt engineering is making generative AI do what you want
| by crafting the right context. I would put it somewhere in
| the software and data engineering space, given they will most
| likely integrate applications with it. MLOps comes into play
| if you have your own trained or tuned model.
| m0llusk wrote:
| compiled a quick list of terms to help understand this piece:
|
| MDS - modern data stack
|
| BI - business intelligence
|
| Redshift - amazon redshift data warehouse service product handles
| large scale data sets and database migrations can handle analytic
| workoads on big data sets with column oriented approach built on
| top of massive parallel processing (MPP) from data warehouse
| company ParAccel (later acquired by Actian)
|
| ETL - extract, transform, load
|
| ELT - extract, load, transform -- an alternative to ETL that
| stores raw data
|
| Looker - does BI, started with looker data sciences, acquired by
| google
|
| Tableau - data visualization focused on business intelligence
|
| clickstream data - user website navigation records
|
| snowflake - an MDS data platform solution
|
| mongo - nosql database
|
| datadog - monitoring, analytics for devops
|
| confluent - real time data streams
|
| databricks - web based cluster management and data lakes for
| machine learning
|
| meme-ification - reduction of an idea to cartoons
|
| peak - highest point preceeding drop off
|
| fivetran - data ingestion
|
| dbt - dbt labs, data transformation
|
| CDO - chief data officer
|
| co-marketing - integrated marketing such as linked brands
|
| co-sponsored - cooperative sponsorship
|
| co-selling - shared sales assets including data records
|
| ARR - annual recurring revenue
|
| ZIRP - zero interest rate policy
|
| private multiples - private company valuation calculations,
| metrics
|
| forward revenue - revenue expected in the future
|
| PowerBI - microsoft business analytics
___________________________________________________________________
(page generated 2024-02-12 23:02 UTC)