[HN Gopher] Thefastest.ai
       ___________________________________________________________________
        
       Thefastest.ai
        
       Author : zkoch
       Score  : 46 points
       Date   : 2024-04-23 21:21 UTC (1 hours ago)
        
 (HTM) web link (thefastest.ai)
 (TXT) w3m dump (thefastest.ai)
        
       | CharlesW wrote:
       | Groq really has an unfortunate name. (I assume they had theirs
       | before Grok.)
        
         | jsheard wrote:
         | Yeah they're not happy about that
         | 
         | https://wow.groq.com/hey-elon-its-time-to-cease-de-grok/
        
           | legohead wrote:
           | I'll act as official mediator. Conclusion: both should change
           | their name
        
             | basil-rash wrote:
             | Pretty silly to name your company a common word directly
             | related to the product, then get upset at others using that
             | same word for their product. It's like if Grindr made angle
             | grinders then got mad at a different company releasing an
             | angle grinder they called "Grinder".
        
         | kevindamm wrote:
         | The spelling with a 'k' is more canon (referring to the term
         | from Heinlein) and that was the spelling in the tech culture
         | that borrowed it... what is the reason for choosing a 'q' in
         | theirs, do you know?
        
           | CharlieDigital wrote:
           | I like the "q" as in "query" or "question"; seems a fitting
           | homophone.
        
       | cedws wrote:
       | Any idea which one Copilot uses? I'm interested in exploring ways
       | to get down autocomplete suggestion latency.
        
         | sva_ wrote:
         | I think Github Copilot itself is GPT-3.5, but Copilot Chat is
         | GPT-4.
        
       | akozak wrote:
       | It'd be nice to have a similar site but cost per token.
        
         | zkoch wrote:
         | This is a great idea. We'll add it.
        
           | akozak wrote:
           | Probably highly volume dependent, but still useful!
        
       | jxy wrote:
       | No prompt length? For practical purposes, the prompt processing
       | time would far more important.
        
         | juberti wrote:
         | We're going to add a selector to choose prompt size (and
         | multimedia content in the prompt)
        
       | saltsaman wrote:
       | Couple of things:
       | 
       | 1. Filtering by model should be enabled by default.
       | Mixtral-8x7b-instruct on Perplexity is almost as fast as the 7B
       | Llama 2 on fireworks, but are quite different in sizes.
       | 
       | 2. Pricing is a very important factor that is not included.
       | 
       | 3. Overall service reliability should also be an important
       | signal.
        
         | juberti wrote:
         | Can you describe what you'd like to see for #1? We currently
         | show everything, but let people filter via the UI or URL param,
         | e.g., https://thefastest.ai/?mf=3-70
        
       | passion__desire wrote:
       | I don't understanding why would we need to having similar
       | expectations from systems that we have from humans and building a
       | whole theory on it. I can adjust my behaviour around systems. I
       | am not restricted to operate within default values. e.g Whenever
       | a price is listed as $99, I automatically know it is $100.
       | Marketing gimmicks don't work once you know about them or in
       | other words, expectations can be set in a new environment.
        
         | fragmede wrote:
         | Marketing gimmicks absolutely still work even if you know about
         | them because they take advantage of basic human psychology so
         | when you're tired/hungry/sleepy or otherwise not operating at
         | peak performance, your lizard brain/autopilot takes over and
         | you choose what's been chosen for you.
        
       | ankerbachryhl wrote:
       | Been looking a lot for a simple overview like this, I've spent
       | too much time benchmarking models/regions myself. Thank you for
       | creating!
        
       | pants2 wrote:
       | Another good resource: https://artificialanalysis.ai/
        
       | anonzzzies wrote:
       | Groq with llama3 70b is so fast and good enough for what we do
       | (source code stuff) that it's really quite painful to work with
       | most others now. We replaced most our internal integrations with
       | this and everything is great so far. I guess they will be bought
       | soon?
        
         | m3kw9 wrote:
         | What do you guys do?
        
       | pants2 wrote:
       | There are dozens of AI chip startups out there with wild claims
       | about speed. Groq seems like the first to actually prove it by
       | launching a product. I hope they spur a speed war with other
       | chipmakers to make the fastest inference engine.
        
       | pants2 wrote:
       | I'd be interested to hear how Llama 8B with long chain-of-thought
       | prompts compares to GPT-4 one-shot prompts for real-world tasks.
       | 
       | In classification for example, you could ask Llama 8B to reason
       | through each possibility, rank them, rate them, make
       | counterarguments, etc. - all in the same time that GPT-4 would
       | take to output one classification without reasoning. Which does
       | better?
        
         | rgbrgb wrote:
         | Good idea, that could make for a pretty interesting eval. It's
         | similar to a timed test... we don't really care how long it
         | takes or how much scratch paper you needed as long as you
         | deliver the correct answer within the time limit.
        
       ___________________________________________________________________
       (page generated 2024-04-23 23:01 UTC)