https://blog.reyem.dev/post/extracting_hn_book_recommendations_with_chatgpt_api/ reyem.dev blog Menu * Home * About * Posts Extracting Hacker News Book Recommendations with the ChatGPT API Created: 2023-10-04 I love books and I enjoy reading through the Hacker News(HN) book recommendation threads. On HN, there's almost 200 stories so far this year that have the separate word "book" in the title, and aren't linked to another page. I wondered what the most commonly recommended or mention books are. Mainly wondering if SICP or PCL would be the top recommendation. After reading of the man who categorised his favourite podcast into dewey decimal using GPT, I was aware that the GPT API could be used to categorise data and output the information in json format. So using the HN data fetched from the hackernews API, I used the subset of stories that seem to be book recommendation threads and extracted book titles, authors and urls from the text using calls to the Chat Completions API. Here's the top 50 book recommendations: # Title Author Count First Mention Structure and 1 Interpretation of Abelson and Sussman 376 5675 Computer Programs 2 Godel, Escher, Bach Douglas Hofstadter 293 56795 3 How to Win Friends and Dale Carnegie 292 5584 Influence People 4 The C Programming Brian Kernighan, Dennis 284 135262 Language Ritchie 5 Dune Frank Herbert 263 57231 6 Thinking, Fast and Slow Daniel Kahneman 244 3277457 7 Meditations Marcus Aurelius 233 134993 8 Atlas Shrugged Ayn Rand 222 86114 9 The Art of Computer Donald E. Knuth 213 135245 Programming 10 Sapiens: A Brief Yuval Harari 205 10028239 History of Humankind 11 Zen and the Art of Robert M Pirsig 203 56941 Motorcycle Maintenance 12 The Pragmatic Andrew Hunt 203 5704 Programmer Introduction to Charles E. Leiserson, 13 Algorithms Clifford Stein, Ronald 171 55391 Rivest, Thomas H. Cormen 14 The Selfish Gene Richard Dawkins 168 85867 Code: The Hidden 15 Language of Computer Charles Petzold 160 135906 Hardware and Software 16 The Mythical Man-Month Fred Brooks 159 5725 17 The Black Swan Nassim Nicholas Taleb 158 56763 Designing 18 Data-Intensive Martin Kleppman 153 8671875 Applications 19 1984 George Orwell 152 85938 20 Code Complete Steve McConnell 149 56709 21 Snow Crash Neal Stephenson 146 85862 22 The Three-Body Problem Cixin Liu 143 8867599 23 Ender's Game Orson Scott Card 143 56704 24 The Design of Everyday Don Norman 136 85860 Things 25 Bible Unknown 134 85859 26 Founders at Work Jessica Livingston 133 5613 27 Antifragile Nassim Nicholas Taleb 130 4966437 28 Man's Search for Victor E. Frankl 129 1634144 Meaning 29 The Hitchhiker's Guide Douglas Adams 128 56709 to the Galaxy 30 Cryptonomicon Neal Stephenson 127 85940 31 The Fountainhead Ayn Rand 127 135463 32 Surely You're Joking, Richard Feynman 125 85858 Mr. Feynman! 33 Fooled by Randomness Nassim Nicholas Taleb 125 57595 34 Siddhartha Herman Hesse 124 86337 35 Foundation Isaac Asimov 123 140379 36 The Lord of the Rings J. R. R. Tolkien 121 56629 37 Zero to One Peter Thiel 115 7968392 38 Calculus Charles B. Morrey Jr., 114 193554 Murray H. Protter 39 Neuromancer William Gibson 112 56663 40 The Phoenix Project Gene Kim 110 5569687 41 The Lean Startup Eric Ries 110 1570888 42 Never Split the Chris Voss 108 12245967 Difference 43 Design Patterns Addy Osmani 107 80916 44 Guns, Germs, and Steel Jared Diamond 107 56777 45 JavaScript: The Good Douglas Crockford 106 259986 Parts 46 Clean Code Robert C. Martin 106 1945860 47 Deep Work Cal Newport 105 11702897 48 The Elements of Noam Nisan, Shimon Schocken 104 1295307 Computing Systems 49 The Little Schemer Daniel P. Friedman, 102 56629 Matthias Felleisen Influence: The 50 Psychology of Robert B. Cialdini 101 193848 Persuasion Edit: Some corrections, Dune was by Frank not Brian Herbet. And the Meditations listed is by Marcus Aurelius, not Descartes. Thanks to the Hacker News commenters who pointed out these errors. My sql query should have returned the most common author for each title, instead of min(author). Some things I discovered while doing this project: * When the API doesn't return valid JSON, usually this is when chatGPT is saying things like "I apologize for the confusion..." or "You're welcome! If you have any more questions, feel free to ask.", in response to a HN comment that just says "thanks" or asks a question. * Designed the prompt so that I can discard responses with empty titles. This is because I was unable to get chatgpt to stop including mentions of Authors without a title of a particular book. * Processing 57k comments cost about $40 using gpt 3.5 turbo API. * Even with a temperature of 0, GPT's results vary from call to call. Others have noticed this effect (HN discussion), it's not just GPT-4 that is non-deterministic - GPT 3.5 turbo exhibits greater variability compared to earlier GPT-3 models. * It can identify links from the text, but I had to remove the html tag and just leave the url otherwise GPT would pick up the truncated link text instead of the URL. Here's an example of the json output by chatgpt, for this comment, it got everything wrong except the link, but it shows the format of the data: [ { "match": "Hitchhiker's Guide Vms Unsupported Undocumented Can Go Away At Any Time Feature", "title": "The Hitchhiker's Guide to the Galaxy", "author": "Douglas Adams", "link": "http://www.amazon.com/Hitchhikers-Guide-Vms-Unsupported-Undocumented-Can-Go-Away-At-Any-Time-Feature/dp/1878956000" } ] Edit: someone has asked for the prompt, here it is: prompt = [ {"role": "system", "content": "Assistant that identifies book titles and authors in the following document and shows the words you match to a book title from. Some titles may be abbreviated, please expand the abbreviated title. If the document talks about an author but doesn't mention a book, leave \"title\" blank. If you know who the author is, provide the author. Don't include the book's subtitle. If the text is asking for a recommendation, without mentioning a book, then return an empty array. Provide your answer in a json array."} {"role": "user", "content": 'Wren\'s Explosion https://www.amazon.com/gp/395, and any Plath."'}, {"role": "assistant", "content": '''[{"match":"'Wren's Explosion","title":"Explosion","author":"P.C. Wren","link":"https://www.amazon.com/gp/395"}, {"match":"any Plath","title":"", "author":"Sylvia Plath"}]'''}, {"role":"user", "content":"3-days free trial isn't freemium."}, {"role":"assistant", "content":"[]"}, {"role":"user", "content": "Miranda Hamilton"}, {"role":"assistant", "content": '''[{"match":"Miranda Hamilton","title":"Hamilton","author":"Lin-Manuel Miranda, Jeremy McCarter","link":""}]'''}, ] The Data Because I enjoy working with data and think you might find it interesting to analyse the results, here's the raw data produced by GPT , sorted by title. Note there's a match column in there which includes an excerpt of the comment where the book was identified. I also normalised the book titles, lowercasing and removing 'the' if present at the start, and removed any subtitles. This enabled me to query the top books without missing too many items due to inconsistence in the names that gpt came up with. Here is the input data in zipped csv format, it expands out to a 24 MB file. Note I have added an amazon affiliate links to amazon urls in the tables above, mainly as a learning exercise. * hn * ai * python * books Please enable JavaScript to view the comments powered by Disqus. comments powered by Disqus (c) 2023 reyem.dev blog. Generated with Hugo and Mainroad theme.