[HN Gopher] Show HN: Analyzing top HN posts with language models
       ___________________________________________________________________
        
       Show HN: Analyzing top HN posts with language models
        
       Hi HN,  I spent a few weeks looking at the top HN posts of all
       time. This included exploration, clustering, creating
       visualizations, and zooming in on what (to me personally) seems
       like some of the best discussions on here.  Three things in this
       post:  1- The interesting groups of HN posts  2- The interactive
       visualizations that you can explore in your browser  3- The data
       from this exploration -- this includes CSV of the titles as well as
       the text embeddings of 3,000 Ask HN articles.  Blog post about this
       whole process here: [1]  ============  1- The interesting groups of
       HN posts  From the exploration, Ask HN proved the most interesting.
       These are the top four groups of topics I found insightful. Each
       group contains about 400 posts.  - Life experiences and advice
       threads [2]  - Technical and personal development [3]  - Software
       career insights, advice, and discussions [4]  - General content
       recommendations (blogs/podcasts) [5]  ============  2- The
       interactive visualizations that you can explore in your browser  -
       Top 10,000 Hacker News articles of all time [6]  - Top 3,000 posts
       in Ask HN [7]  ============  3- The data from this exploration  CSV
       file of top 3K Ask HN posts: [8]  The sentence embeddings of the
       titles of those posts: [9]  This is a colab notebook containing the
       code examples (including loading these two data files): [10]
       ============  If you've ever wanted to get into language models,
       this is a good place to start. Happy to answer any questions
        
       Author : jayalammar
       Score  : 94 points
       Date   : 2022-06-10 12:47 UTC (10 hours ago)
        
       | victorianoi wrote:
       | Nice idea and analysis! I reproduced it as well with
       | https://graphext.com and got similar clusters
       | https://drive.google.com/file/d/1-kXsKezu2_S07rQn-0bjbHuUXHE...
        
         | victorianoi wrote:
         | BTW there is an implicit recency bias in the dataset, since
         | 2017 the number of top 3K post became more frequently and the
         | avg score is larger year after year as the community in HN
         | grows:
         | 
         | - Number of top 3K per month of publishing -
         | https://drive.google.com/file/d/1beAPP9ijruMUs5DN5wOVsBArvxP...
         | 
         | - Avg score of top 3K per month of publishing -
         | https://drive.google.com/file/d/10nSIgH1a6DN6XrDU2DyMJTCgsIg...
        
           | victorianoi wrote:
           | and it also looks like most topics are constant over the
           | years - https://drive.google.com/file/d/1ilYn9cnEZwiH1FioUtU9
           | ummhmvn...
        
       | jayalammar wrote:
       | [1] https://txt.cohere.ai/combing-for-insight-
       | in-10-000-hacker-n...
       | 
       | [2] https://assets.cohere.ai/blog/text-
       | clustering/askhn_cluster_...
       | 
       | [3] https://assets.cohere.ai/blog/text-
       | clustering/askhn_cluster_...
       | 
       | [4] https://assets.cohere.ai/blog/text-
       | clustering/askhn_cluster_...
       | 
       | [5] https://assets.cohere.ai/blog/text-
       | clustering/askhn_cluster_...
       | 
       | [6] https://assets.cohere.ai/blog/text-
       | clustering/hn10k_clustere...
       | 
       | [7] https://assets.cohere.ai/blog/text-clustering/askhn-3k.html
       | 
       | [8] https://storage.googleapis.com/cohere-assets/blog/text-
       | clust...
       | 
       | [9] https://storage.googleapis.com/cohere-assets/blog/text-
       | clust...
       | 
       | [10] https://colab.research.google.com/github/cohere-
       | ai/notebooks...
        
         | jayalammar wrote:
         | Disclosure: These were made by Cohere's embeddings, a company
         | where I work. The process should work on text embeddings from
         | other sources.
        
       | wpietri wrote:
       | The conflict of interest here concerns me. I don't object to
       | content marketing, but I'd rather a) you were clear from the
       | start that you work for this company and are promoting its
       | product, and b) that this "revolves around [...] using Cohere's
       | Embed endpoint", so that people can judge how much they want to
       | "get into language models" with pay-per-character pricing, as
       | opposed to something more open.
        
         | jayalammar wrote:
         | Thanks. I just added a disclosure to the comment (can't edit
         | the parent anymore). The full embeddings are freely provided
         | here without the need to use the service.
        
           | wpietri wrote:
           | Thanks! I appreciate it.
        
         | rmbyrro wrote:
         | Do you advocate for the disclaimer only because 1) the sample
         | uses their product or 2) just because they sell a product
         | correlated to the topic?
         | 
         | I see a lot of articles that fall into #2 being published here
         | without a disclaimer. And I think a disclaimer isn't necessary
         | for #2. Even for #1 I wouldn't bother, but I understand the
         | expectation.
         | 
         | Many advocate a lot against ads, targeting, etc. If we also
         | advocate against promotional content, what would companies do
         | to get attention and traffic?
        
           | wpietri wrote:
           | I don't understand why you think asking that conflicts of
           | interests being clearly disclosed is me wanting to "advocate
           | against promotional content". Do you believe content
           | marketing only works if it's quietly manipulative? That seems
           | like a pretty grim take.
        
           | notatoad wrote:
           | most content marketing goes to a post on the company's own
           | website, which is a sort of inherent disclosure.
           | 
           | i don't think asking people to disclose that they work for
           | the company whose product they're promoting is "advocating
           | against promotional content".
        
         | mbesto wrote:
         | "Show HN"s regularly have commercial models behind them. Not
         | sure what your stink is...this is normal.
         | 
         | https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
        
           | wpietri wrote:
           | "Normal" doesn't mean good, so that line of argument doesn't
           | work for me.
           | 
           | But the specific "stink" I'm objecting to here is, as I said,
           | a conflict of interest. In specific, saying, "If you've ever
           | wanted to get into language models, this is a good place to
           | start" purports to be neutral and helpful. When instead, this
           | person is promoting a product. Maybe using a (to my eyes
           | expensive) commercial service is the truly the best place to
           | start learning that. Maybe this is truly the best service to
           | learn it with. But we can't expect a fair answer to those
           | questions from a person whose works at the company and whose
           | apparent job is promoting the product that pays their salary.
        
       | toppy wrote:
       | I don't know how HN score metrics work but after some short
       | review of the datafile [1] I've noticed a lot of the posts has
       | the form of a simple questions and as such seems to be naturally
       | biased when comes to user engagement. Have you considered to add
       | additional metrics to remove that bias and re-analyze?
       | 
       | [1] https://storage.googleapis.com/cohere-assets/blog/text-
       | clust...
        
         | jayalammar wrote:
         | What do you mean by naturally biased? That people seem to favor
         | them?
        
           | toppy wrote:
           | People seem to be more engaged in discussions arising from
           | questions rather than statements, no?
        
             | jayalammar wrote:
             | I think that's part of the expectations out of "ask HN". I
             | don't know that the same effect happens outside of Ask HN.
        
       | hoerzu wrote:
       | I grouped posts by topic: https://n3ws.ploomber.io
       | 
       | Explanation: https://ploomber.io/blog/hn_classifier/
        
       | uniqueuid wrote:
       | Interesting, but it doesn't seem like the dimensionality
       | reduction produces a good separation of topics. The UMAP
       | projection looks pretty dense. Did you consider pruning or using
       | something other than embeddings?
        
         | jayalammar wrote:
         | So it really depends on what you use for clustering. In this
         | case, I'm clustering by the original embeddings so the UMAP
         | results are different. I've also seen:
         | 
         | 1- Clustering by UMAP. Here the plot would show clean
         | separation of topics. But the clustering algorithm would be
         | working on highly compressed data (from the 1024 dimensions of
         | the embedding down to the 2 of UMAP).
         | 
         | 2- BERTopic's approach of doing UMAP down to 5 dimensions,
         | using this dimensionality for clustering, then UMAP again from
         | 5 to 2. Which is an interesting approach.
         | 
         | I've heard people having good results with all three. It's
         | kinda hard to objectively compare, but my leaning was to give
         | the clustering algorithm the representation containing the most
         | information about the text.
        
           | uniqueuid wrote:
           | Right, bertopic's double clustering is interesting. I've also
           | seen people combine that with louvain instead of k-means.
           | 
           | My intuition was: UMAP itself tries to optimize for 2d
           | separation in the projection. So we should expect at least
           | some correspondence between the kmeans results and the layout
           | in the UMAP plot (except in some pathological edge cases
           | perhaps).
           | 
           | Nevertheless, nice example and blog post!
        
             | jayalammar wrote:
             | TIL louvain clustering! I see it used for graphs. Can also
             | be used for vectors/points?
             | 
             | Thank you!
        
               | uniqueuid wrote:
               | You're welcome!
               | 
               | You can actually create a graph by using k-means
               | similarities as edge weights. Then you do graph
               | clustering on it. (using any algorithm, but louvain is
               | one of the saner ones ... clique percolation, girvan-
               | newman etc all have known problems).
        
           | PaulHoule wrote:
           | Try t-SNE. I used to scoff at cluster plots until I saw those
           | but with t-SNE... wow, those clusters are actually separated!
        
             | uniqueuid wrote:
             | Are you sure t-SNE and UMAP actually perform very
             | differently? Last I looked, they were somewhat comparable.
             | 
             | [edit]: Seems they are similar for some purposes:
             | https://blog.bioturing.com/2022/01/14/umap-vs-t-sne-
             | single-c...
             | 
             | Also interesting: Rapidsai has a cuda accelerated version
             | of umap that is very fast (hdbscan as well BTW).
        
       | bryanrasmussen wrote:
       | as people upvote other things than your list of relevant links it
       | becomes difficult to find the relevant links. although I guess
       | people can find it by your name.
       | 
       | on edit: so it seems some are upvoting the links to keep them on
       | the top in opposition to those upvoting discussion points.
        
       | natch wrote:
       | There's too much fixation with "top" in our industry. Top voted
       | tends to mostly be a function of early posting. Later posts don't
       | get votes because they simply were not seen. There seems to be a
       | misreading on a mass scale of what "top" really indicates though;
       | people think it means "quality" when it does not. Study after
       | study, website after website, policy after policy, our online
       | world is built on this fundamental misunderstanding of what is
       | really going on. How do you avoid piling on to this
       | misunderstanding?
        
         | tra3 wrote:
         | Is there a name for this phenomenon so I can google further?
         | Intuitively makes sense, because I've seen this before.
        
       ___________________________________________________________________
       (page generated 2022-06-10 23:01 UTC)