[HN Gopher] Thefastest.ai
___________________________________________________________________
Thefastest.ai
Author : zkoch
Score : 46 points
Date : 2024-04-23 21:21 UTC (1 hours ago)
(HTM) web link (thefastest.ai)
(TXT) w3m dump (thefastest.ai)
| CharlesW wrote:
| Groq really has an unfortunate name. (I assume they had theirs
| before Grok.)
| jsheard wrote:
| Yeah they're not happy about that
|
| https://wow.groq.com/hey-elon-its-time-to-cease-de-grok/
| legohead wrote:
| I'll act as official mediator. Conclusion: both should change
| their name
| basil-rash wrote:
| Pretty silly to name your company a common word directly
| related to the product, then get upset at others using that
| same word for their product. It's like if Grindr made angle
| grinders then got mad at a different company releasing an
| angle grinder they called "Grinder".
| kevindamm wrote:
| The spelling with a 'k' is more canon (referring to the term
| from Heinlein) and that was the spelling in the tech culture
| that borrowed it... what is the reason for choosing a 'q' in
| theirs, do you know?
| CharlieDigital wrote:
| I like the "q" as in "query" or "question"; seems a fitting
| homophone.
| cedws wrote:
| Any idea which one Copilot uses? I'm interested in exploring ways
| to get down autocomplete suggestion latency.
| sva_ wrote:
| I think Github Copilot itself is GPT-3.5, but Copilot Chat is
| GPT-4.
| akozak wrote:
| It'd be nice to have a similar site but cost per token.
| zkoch wrote:
| This is a great idea. We'll add it.
| akozak wrote:
| Probably highly volume dependent, but still useful!
| jxy wrote:
| No prompt length? For practical purposes, the prompt processing
| time would far more important.
| juberti wrote:
| We're going to add a selector to choose prompt size (and
| multimedia content in the prompt)
| saltsaman wrote:
| Couple of things:
|
| 1. Filtering by model should be enabled by default.
| Mixtral-8x7b-instruct on Perplexity is almost as fast as the 7B
| Llama 2 on fireworks, but are quite different in sizes.
|
| 2. Pricing is a very important factor that is not included.
|
| 3. Overall service reliability should also be an important
| signal.
| juberti wrote:
| Can you describe what you'd like to see for #1? We currently
| show everything, but let people filter via the UI or URL param,
| e.g., https://thefastest.ai/?mf=3-70
| passion__desire wrote:
| I don't understanding why would we need to having similar
| expectations from systems that we have from humans and building a
| whole theory on it. I can adjust my behaviour around systems. I
| am not restricted to operate within default values. e.g Whenever
| a price is listed as $99, I automatically know it is $100.
| Marketing gimmicks don't work once you know about them or in
| other words, expectations can be set in a new environment.
| fragmede wrote:
| Marketing gimmicks absolutely still work even if you know about
| them because they take advantage of basic human psychology so
| when you're tired/hungry/sleepy or otherwise not operating at
| peak performance, your lizard brain/autopilot takes over and
| you choose what's been chosen for you.
| ankerbachryhl wrote:
| Been looking a lot for a simple overview like this, I've spent
| too much time benchmarking models/regions myself. Thank you for
| creating!
| pants2 wrote:
| Another good resource: https://artificialanalysis.ai/
| anonzzzies wrote:
| Groq with llama3 70b is so fast and good enough for what we do
| (source code stuff) that it's really quite painful to work with
| most others now. We replaced most our internal integrations with
| this and everything is great so far. I guess they will be bought
| soon?
| m3kw9 wrote:
| What do you guys do?
| pants2 wrote:
| There are dozens of AI chip startups out there with wild claims
| about speed. Groq seems like the first to actually prove it by
| launching a product. I hope they spur a speed war with other
| chipmakers to make the fastest inference engine.
| pants2 wrote:
| I'd be interested to hear how Llama 8B with long chain-of-thought
| prompts compares to GPT-4 one-shot prompts for real-world tasks.
|
| In classification for example, you could ask Llama 8B to reason
| through each possibility, rank them, rate them, make
| counterarguments, etc. - all in the same time that GPT-4 would
| take to output one classification without reasoning. Which does
| better?
| rgbrgb wrote:
| Good idea, that could make for a pretty interesting eval. It's
| similar to a timed test... we don't really care how long it
| takes or how much scratch paper you needed as long as you
| deliver the correct answer within the time limit.
___________________________________________________________________
(page generated 2024-04-23 23:01 UTC)