[HN Gopher] Predicting hit songs with 97% accuracy
___________________________________________________________________
Predicting hit songs with 97% accuracy
Author : geox
Score : 78 points
Date : 2023-06-20 12:58 UTC (10 hours ago)
(HTM) web link (www.frontiersin.org)
(TXT) w3m dump (www.frontiersin.org)
| sproketboy wrote:
| [dead]
| barbegal wrote:
| In summary they took 24 songs (13 hits and 11 flops) and
| evaluated these against some neurological parameters from 33
| volunteers, measuring average immersion, peak immersion and
| "Retreat" (low immersion). They then synthesised 10000
| observations which were labelled either hit or flop and had a
| similar distribution of the three parameters as the original 24
| songs.
|
| A machine learning algorithm was trained on 5000 of these
| observations and tested on the other 5000 observations and the 24
| songs.
|
| It got 97% of the synthetic observations correctly labelled and
| 23 out of the 24 songs.
|
| Unfortunately, it can clearly be seen that the generation of the
| synthetic data based on all 24 songs means overfitting to the
| data can easily take place (despite what the authors think their
| 10-fold cross validation proves)
|
| Without proper separation of training and test data this is a
| garbage study and tells us virtually nothing.
| ed-209 wrote:
| Seems this method is only successful at discriminating between
| extremes (of popular/unpopular) ...
|
| >Staff from an online streaming service choose 24 songs for this
| study without input from the researchers. The streaming service
| also provided the definition of hits or flops. This resulted in a
| "clean" experiment as song choice could not be cherry-picked for
| the study and the criterion for a hit was established in advance.
| mrguyorama wrote:
| For reference, a friend and I took a machine learning class in
| college and basically had no idea what we were doing. We were
| able to build a random forest model that predicted hit songs with
| over 90% accuracy. The data set included things like tempo and
| synthetic parameters like "jazziness". This was the million song
| dataset from the people behind Echonest.
|
| On further examination it turns out that nearly all the
| "prediction" was based off a single parameter: The song's artist.
| senthilnayagam wrote:
| I have been generating several thousand music samples with
| musicgen,
|
| I don't care much about hits at this moment, but if it can
| eliminate bad songs even by 50% then it can help build a pipeline
| to bring humans into evaluating the generated songs
| mjburgess wrote:
| Regardless of the severe methodological flaws, you can a-priori
| rule out a study like this: cultural taste is anti-inductive.
|
| ie., whatever the trends are over period X, the subsequent period
| Y will depart from X (due to, eg., bordem, counter-culture,
| etc.).
|
| So you cannot predict taste based on an associative statistical
| model.
| huijzer wrote:
| Maybe the easiest and strongest counterargument against this
| paper is that a 97% accuracy is extremely unlikely because 97% is
| a near perfect score even though hit songs are distributed
| according to a power law [1]. If we could predict this power law
| with 97% accuracy, then we could also predict the next successful
| CEOs, companies, and soccer players with 97% accuracy. But we
| can't.
|
| Anyway, it doesn't matter. Science is measured by popularity
| (citations), and not by truth. With that in mind, it's a nice
| paper. It has a nice story. It has some complicated graphs. It is
| difficult for a layman to figure out the problem. So, yes, it's a
| successful academic paper. It's a great read and very
| interesting, according to academia.
|
| [1]: https://michaeltauberg.medium.com/power-law-in-popular-
| media...
| ryanwaggoner wrote:
| Ah, the final nail in the coffin of art.
| fallingfrog wrote:
| What determines what's a hit? Is it how many times Spotify plays
| it? Because that's probably dominated by restaurants playing the
| song as part of a loop of 10 or 15 songs which are chosen by some
| committee somewhere in the corporate office. So what you're
| really predicting is the taste of these people who are basically
| trying to choose elevator music.
| elektor wrote:
| This is the study population: Thirty-three participants (47%
| female) were recruited from the Claremont Colleges and
| surrounding community. Participants ranged in age from 18 to 57
| (M = 24.25, SD = 10.47).
| wongarsu wrote:
| So the correct conclusion should be "this is a promising
| research direction worthy of a bigger study", not necessarily
| "this is a good way to identify hit songs".
| CaptainNegative wrote:
| I'm not sure even the first statement is entirely fair. As
| far as I can tell this result is indistinguishable from "a
| sufficiently powerful model can distinguish data generated by
| two different popsynth models".
|
| I would have had more faith had they scrambled the labels and
| run the experiment of distinguishing one random subset of
| songs from another random subset. Anything greater than 50%
| would indicate the model's ability to overfit; I'd like to
| know how accurately their model would perform.
| pingou wrote:
| "Every day, 24,000 new songs" "less than 4% of new songs will
| become hits". I doubt that there are almost 1000 new hit songs
| everyday.
| riskable wrote:
| This is something I learned about a long time ago: The
| Billboard 100 is not actually a list of which songs are played
| the most. It's a list of which songs _produced by a certain
| subset of record labels_ that happened to be played a lot _on
| certain radio stations_ and certain streaming services. Nelson
| surveys also play a role but that element of it is basically
| meaningless at this point.
|
| There's another element at work that gets _way_ too much weight
| in the charts: Sales figures. As in, "how many people actually
| bought this single or an album with this single?" This is the
| main reason why only major labels end up in the charts: They're
| the only ones allowed to report figures (for several reasons; a
| few legitimate one's too).
|
| Let's say you make a great song and upload it to Spotify.
| Within a month it has been streamed 100 million times, making
| it one of the most streamed songs that month. Your song won't
| be in the Billboard top 100 _because it 's not on a major
| record label_ and Billboard won't be able to associate the
| pirate uploads of that song to YouTube (where it also tracks
| song data) with you, who never registered it anywhere other
| than Spotify.
|
| There's also entire categories of music that basically have no
| chance of ever making the Billboard Top 100 because they're
| "unofficial" (think: Unauthorized but legal remixes).
|
| I am running out of typing time here but just know that the
| main differentiator is the record label. Unless the song is on
| a label that actually has the necessary resources, getting into
| the Billboard Top 100 isn't likely to happen. There will be
| occasional exceptions of course but the main takeaway is that
| the Billboard Top 100 is primarily a marketing tool for the big
| record labels in the music industry. If it ever stops
| prioritizing their stuff it wouldn't be serving its purpose.
| CPLX wrote:
| As someone with direct experience in all this I can just say
| the best reply to this comment is to point out that it's just
| completely disconnected from reality. There's no rules about
| the charts being reserved for major labels or any of the
| other stuff being said here, it's just nonsense.
| amanaplanacanal wrote:
| It might help if you weigh in with more info: exactly what
| does billboard measure when creating the charts? Radio
| airplay? Streams? Sales? Some combo?
| CPLX wrote:
| I mean Billboard publishes the formulas and methodology.
| It's not even slightly secret or anything. Sources
| include Billboard's own website and dozens of articles.
|
| https://www.google.com/search?rls=en&q=how+does+billboard
| +ca...
|
| The tell that this comment is nonsense is the assertion
| that only major labels count. That's never been true _at
| all_ and anyone saying that is telegraphing that they
| aren 't really connected to this issue.
|
| Registering with SoundScan and getting UPC/ISRC is
| trivial and many services (like HN user founded
| Distrokid) automate portions of it.
|
| None of this is new. Those of us that were on tour
| playing music in the 90's got used to filling out
| Soundscan forms on paper at the merch table and finding a
| way to fax them in. Indie artists have been reporting
| sales to the charts for decades.
|
| Also there's no such thing as an "unauthorized but legal
| remix" that's just someone else's song. The parent
| comment is "not even wrong" levels of wrong.
| boomboomsubban wrote:
| Worldwide? Doesn't seem ridiculous.
| InCityDreams wrote:
| I bounce between three rock-orientated radio stations when in
| the car.
|
| I've only heard Rush once - when Neil Peart died.
|
| Similar for other (dead) artists.
|
| Whatever label Rush are on, they ain't doing the payola to get
| airplay in my locality.
| amanaplanacanal wrote:
| Interesting. In my market there is one "classic rock" station
| and rush has been a staple since the 70's. This is one of the
| stations that didn't play 90's stuff at all until they were
| purchased by one of the big chains a few years ago.
| r3trohack3r wrote:
| I've always enjoyed listening to albums. When I was younger, I'd
| listen to albums of pop music because that's what I'd hear on the
| radio.
|
| The albums often contained at least one track that was "catchier"
| than what was played on the radio. I always wondered why these
| never made the rounds.
|
| Fast forward 20 years and I've seen a handful of those songs come
| back and make the rounds on pop channels. I've always wondered
| why they didn't do numbers out the gate. Maybe they were missed
| at the time or maybe they were ahead of their time and needed
| tastes to catch up.
|
| I could see taking an AI model that predicts a positive response
| from "the masses" being unbelievably valuable to anybody sitting
| on a massive catalogue of historical content. Most of that
| content is only valuable, at this point, in bundled deals.
|
| But the prospect of there being a handful of diamonds sitting in
| there that could be panned for with AI? I imagine that's billions
| of dollars waiting to be tapped at a near zero cost (plus,
| possibly, the cost of remastering).
|
| (P.S. for fellow album lovers out there, I built
| https://audile.blankenship.io for personal use, it's free have at
| it)
| CrazyStat wrote:
| I'm a statistician, not a machine learning person, but
|
| >Small data sets are not appropriate for machine learning as they
| lead to high bias in their results (Vabalas et al., 2019). To
| address this, we created a synthetic set data with 10,000
| observations using the synthpop package in R (Nowok et al.,
| 2016). This standard automated procedure creates observations by
| repeatedly randomly sampling the joint distribution of the data.
| This technique is used when obtaining large datasets is
| infeasible, including analyses of computer vision (Mayer et al.,
| 2018), sensitive information like hospital records (Tucker et
| al., 2020), and with unbalanced data (He et al., 2008; Luo et
| al., 2018).
|
| Isn't this just making up data? How is the analysis of the
| synthetic 10,000 observations (based on observations from 33 real
| people) at all meaningful?
| scotty79 wrote:
| Curious, if it manages to classify real data correctly after
| such learning.
| beerpls wrote:
| Yeah, this is an issue with way more science than anyone (esp
| scientists) is comfortable acknowledging
|
| In grad school I encountered experiment after experiment, paper
| after paper where the authors went around emphasizing the parts
| which were statistically "sexy" while ignoring glaring issues
| like this
|
| It's almost like the methodology for creating science is more
| rigorous than the science itself
|
| Sometimes the sample is too small. Sometimes they assume a
| correlation that doesn't exist. It's troubling and was a large
| part of why I left academia
| CrazyStat wrote:
| I faced similar pressures in graduate school to leave out
| somewhat unfavorable results (which I thought were
| interesting and worth presenting even if they didn't flatter
| our approach). It wasn't the only thing that drove me away
| from academia, but it certainly contributed.
|
| In the statistics department, of all places! Where the
| importance of being forthcoming with your data should be best
| understood!
| mjburgess wrote:
| It's not.
|
| Indeed, "predictive accuracy" on simulated data is largely a
| meaningless metric in any case. The simulation's set up
| _determines_ the predictive accuracy of any model of it.
|
| The only value of this metric, post-simulation, is to determine
| whether or not an experiment is worth the ROI -- ie., it's
| essentially a way of determining whether some set of premises
| (simulated) are _possibly_ consistent with a given predictive
| project.
| j16sdiz wrote:
| In machine learning, it is very common to _train_ with
| synthetic data. This is ok as long as them can evaluate with
| real, observed data.
|
| They included a discussion in appendix, but I cannot open that
| file.
| CrazyStat wrote:
| I looked at the appendix but it's mostly just tables and
| density plots comparing the real and synthetic datasets.
| There's no additional justification of why it makes sense to
| use the synthetic dataset in the first place.
|
| As a statistician it still feels very weird to train on
| synthetic data, because my evaluation of model uncertainty
| (parameter uncertainty and predictive uncertainty) is
| (usually) based on the data the model is trained on. But I
| understand that ML typically evaluates models based on out-
| of-sample data, so training on synthetic data makes sense in
| that case. In this paper it appears they're also _evaluating_
| on the synthetic data, which seems problematic:
|
| > One-half of the synthetic data was used to train the bagged
| ML model and tune the hyperparameters. The other half of the
| synthetic data was used to test it.
| moss2 wrote:
| Evaluating on synthetic data is... I want to say
| "forbidden" but that's not the right word. Basically all of
| their results are scientifically useless.
|
| I can't figure out which journal or conference this article
| was published to, but if I sat on the reviewer board I
| would have declined their article on this fact alone.
|
| Good find.
| mensetmanusman wrote:
| Bootstrapping
|
| https://en.m.wikipedia.org/wiki/Bootstrapping_(statistics)
|
| https://en.m.wikipedia.org/wiki/Bootstrap_aggregating
| CrazyStat wrote:
| I'm familiar with bootstrapping--both the classical
| bootstrap, which fits the model repeatedly on different
| simulated datasets to evaluate model uncertainty, and
| bagging, which fits the model repeatedly on different
| simulated datasets as inputs for ensemble learning.
|
| Neither of those is what was done here, which was fitting the
| model _once_ on a greatly inflated synthetic sample.
| screye wrote:
| Why is this study so underpowered ? It should be trivial to
| assemble a massive labelled dataset of hits/not-hits.
| Additionally, every metric should be compared to the human
| baseline, not an ML baseline. If an average nobody can identify a
| prospective hit-song with higher accuracy than the best ML model,
| then then the ML baseline is completely meaningless.
|
| No offense, but I have seen better ML 101 capstone projects.
| Applejinx wrote:
| I've done this, as just an individual, studying 'evergreen'
| albums and which ones sustained sales over time. All you have
| to do is weight the biggest sellers in the platinum-album class
| and adjust for time on the market (partially compensating in
| this way for the growing industry artificially pushing albums
| to be multiplatinum)
|
| You get predictable results like Zeppelin, The Eagles, etc. as
| the 'most hit-laden' records.
|
| As far as actually measuring the SOUND with machine learning,
| you could just take a short-cut and study Mutt Lange's mixes,
| or you could run the computers on it, but this is not actually
| an interesting question because everyone's been trying to
| converge on the 'hit song sound' for decades and decades.
|
| You'll get pretty generic and unobjectionable results, and
| you'll still get blindsided by outliers such as Bohemian
| Rhapsody (too impossibly long and complicated) and Somebody I
| Used To Know (very odd arrangement).
|
| Targeting the most mass market possible (which is do-able with
| AI) is such a shotgun approach that it'll tend to lead to
| failure because too many talented humans have already beat that
| approach into the ground.
|
| What you'd want is for your hit-song AI to 'hallucinate' in an
| interesting way. You want things a little off-model, a 'hook'.
| In a conceptual sense, not just a 'mathematically optimal
| catchy phrase' sense.
| art187 wrote:
| Because you have to have their device to get the data.
| Zensynthium wrote:
| If every song afterwards is a hit because of this and future
| study's findings being implemented, I wonder if the criteria for
| a song being a hit will change. Is it a predetermined pattern of
| the song, or is it the novelty I wonder.
| HendryM wrote:
| [flagged]
| AndrewKemendo wrote:
| I am guessing that, within my lifetime, recommendation systems
| (music, movies, fashion suggestions) will give perfect
| recommendations from the perspective of the user. It feels like
| that's already true for Spotify, so I'm not surprised at these
| findings.
|
| I'm sure there won't be any challenging existential side effects
| of pervasive corporate surveillance, eliminating what remains of
| the illusion of choice.
| sneak wrote:
| > _It feels like that 's already true for Spotify_
|
| I vehemently disagree. This may be true for the general case of
| music listener but there is a long tail (of both listeners and
| musical works) that Spotify is very very bad at.
| belly_joe wrote:
| In my experience, Spotify generally picks 20-40 top songs
| from each genre and recommends them to any listener that has
| showed interest in that genre.
|
| To a listener new to the genre, this yields a pretty good hit
| rate and the perception that Spotify is great at
| recommendation, but after spending any reasonable amount of
| time listening to the genre, the same 20-40 recommendations
| get stale and Spotify is completely unable to surface
| relevant songs from deeper in the genre catalog.
| AndrewKemendo wrote:
| So it's your perspective that the current limits will
| continue to be limits in some sense? Why wouldn't that tail
| eventually get consumed?
| softsound wrote:
| Yep, I'm not even a big music person, but Spotify is terrible
| at predicting music for me, it's so bad I barely try new
| music because most of it is trash but then again Spotify
| thinks I like 350+ genres so there is a lot of room for
| failure.
|
| There are also two kinds of people, people who like albums
| and people who like singles. I'm not an album person because
| it's often lackluster and repetitive, few bands can really
| create good albums that surprise me but singles on the other
| hand can shine and the band knows they can. Just because I
| like one song doesn't mean I'll like 5 other songs from a
| band. Most of my favorite music is pretty unknown. I'd
| imagine hits are just like that hit or miss and good
| marketing is what really helps a lot of things sell.
| esjeon wrote:
| I doubt it, because the choice can always be impacted by
| outside actors. I believe AIs in the future will just
| manipulate people and keep pumping new stuffs for the sake of
| keeping them interested, instead of struggling to understand
| each individuals' taste. It's much cheaper, easier, practical
| and effective.
| AndrewKemendo wrote:
| (Not joking) Eventually no actors or actions will be outside
| of the data collection environment that feeds into the global
| knowledge graphs.
|
| And yes I agree and expect that non human systems will adjust
| human preferences over the long run. That's the point right?
| A caretaker assistant that is smarter, more capable and can
| better predict your needs than anything ever
| teruakohatu wrote:
| Number of actual songs classified: 24, 13 hits and 11 flops. An
| unnamed streaming service chose the songs and so the sampling
| method is unknown.
| CPLX wrote:
| As someone who worked for many years in the music business and
| has seen all sorts of stuff I am _extremely_ skeptical. But I
| guess anything is possible eventually.
|
| With that said I wonder if they are suffering from a base rate
| problem.
|
| The problem is that hit songs are VERY rare. If you feed people a
| ratio of 50% hits and 50% random stuff you might get solid
| numbers just by asking a guy named Bob what he thinks too. Hits
| are fairly distinctive when compared to random DIY music. So Bob
| will just pick the professional sounding stuff and Bob will be
| mostly right.
|
| To be useful though you'd have to feed the system 20,000 songs,
| of which 100 are highly professionally produced songs from people
| who are competent and then identifying which 5 of those 100 are
| going to be hits, as well as the one fluke hit that's in the DIY
| pool.
|
| I bet it can't do that.
| zeroonetwothree wrote:
| Indeed, they even discuss in the article that hits are only
| about 4% of songs. Still, with 97% accuracy you theoretically
| turn 960 flops+40 hits into 29 flops+39 hits which is massively
| better.
| Applejinx wrote:
| That is wildly, wildly out of line. They're out of their
| minds if they think 4% of songs actively attempting to be hit
| songs, get there. Even if we rule out obvious failures (and
| this class of song produces surprise hits, in part through
| novelty!) the number of songs trying to do this and failing
| is WAY higher than 97%. I'm not sure people quite understand
| how many songs are attempted...
| CPLX wrote:
| Not really.
|
| The real task for those who actually do this is being able to
| sell songs that basically seem like they might be hits from
| ones that actually are.
|
| The reason is that it costs money and time and effort to try
| to get attention for the songs to actually make them hit
| songs.
|
| Most songs are just obviously not going to be hit songs they
| aren't even in the same ballpark.
|
| The classic mistake for a budding indie record label owner or
| producer is to audition 31 people and then say wow this
| person is better than all of them so that's the one that will
| be a success.
|
| Except none of them are ever going anywhere.
|
| Conversely, the problem for a budding major label or high
| profile executive is to audition 31 people all of whom are
| actually quite good but only one or two are going to be hits.
|
| They're all great that's how they got in the door, but which
| one is the star?
|
| These are _hard_ problems. Harder than people realize.
|
| So anyways for a machine system have to use _songs that
| obviously have no real strong appeal to anyone at all_ as the
| _starting point_ for analysis.
|
| Otherwise it's just useless.
| KMnO4 wrote:
| The results are actually really really bad.
|
| From the article:
|
| > _Nevertheless, less than 4% of new songs will become hits_
|
| So 96% of songs are not-hits. If the model predicts that no song
| will be a hit, then its accuracy is already 96%.
| screamingninja wrote:
| Wouldn't the accuracy be zero in that case, since it would
| identify none of the hit songs correctly? I feel that identify
| hit songs is harder than identifying non-hit songs.
| KMnO4 wrote:
| TP = correct hit songs = 0
|
| TN = correct non-hits = 96
|
| FP = incorrect hits = 4
|
| FN = incorrect non-hits = 0
|
| Accuracy = (0+96)/(0+96+4+0) = 96/100
|
| This is just a problem with using the metric of accuracy. If
| your classes are unbalanced, then a few "errors" aren't going
| to affect your accuracy.
| seanw265 wrote:
| As other commenters have pointed out, accuracy is not a great
| metric to use here. A confusion matrix [1] is needed to fully
| capture the results.
|
| I'm surprised that one was not prominently displayed in this
| article. They are standard practice for evaluating binary
| classifiers.
|
| [1] https://www.dataschool.io/simple-guide-to-confusion-
| matrix-t...
| Spivak wrote:
| Which is why you're not interested in the hit/not hit accuracy.
| You're interested in "if model says it's a hit how often is it
| right?"
|
| Imagine a Monty Hall situation with 100 doors, you can predict
| with 99% accuracy which doors have goats but you only have 1%
| accuracy on cars.
| Aachen wrote:
| Relevant
| https://web.archive.org/web/20210112213001/https://dilbert.c...
| beardyw wrote:
| My eyes glazed over before I got very far through that. Can
| anyone tell me, if I gave it songs from the 1940s, or whenever,
| would it work the same for those?
| art187 wrote:
| You would need to hook up the Immersion Neuroscience devices up
| to measure emotional resonance and other neurological measures,
| but yes that is the dubious claim.
| beardyw wrote:
| But that would be taking 2023 responses, which I would have
| thought would not have matched 1940 responses.
|
| Feels like a way to lock us into a stagnant pool.
| zeroonetwothree wrote:
| Regardless the 97% figure, even their basic linear model gets 67%
| correct based on neurophysiology, whereas based on surveys it
| would be 50% (since self reported liking has no relationship with
| hit status). So it's intriguing that your brain can perceive
| things about the song but not tell your conscious self that they
| make it good? Or maybe there are just markers hit songs have that
| aren't related to liking them.
___________________________________________________________________
(page generated 2023-06-20 23:03 UTC)