[HN Gopher] Inflection-2: the next step up
       ___________________________________________________________________
        
       Inflection-2: the next step up
        
       Author : jeffdn
       Score  : 74 points
       Date   : 2023-11-22 15:17 UTC (7 hours ago)
        
 (HTM) web link (inflection.ai)
 (TXT) w3m dump (inflection.ai)
        
       | xianshou wrote:
       | First reaction: okay, should we be impressed?
       | 
       | June: flashy ML-perf demo with CoreWeave using 22k H100s
       | 
       | November: actually trained model using 5k H100s
       | 
       | June: we claim Inflection-1 is the best model in its compute
       | class and are preparing a frontier model
       | 
       | November: we beat PaLM 2, which everyone else forgot about long
       | ago anyway
       | 
       | Inflection got a ton of hype with its $1.3B raise (likely not
       | cash but principally GPU compute credits) earlier this year, but
       | now is starting to look like the next victim of inflated
       | expectations.
        
       | tikkun wrote:
       | Regular reminder that most open source LLM benchmarks are not
       | very useful (in the sense that they don't represent day to day ai
       | chatbot usage and what users care about). If you haven't looked
       | through the datasets to see what they actually contain, I'd
       | encourage you to do so. [1] I think we're just in a strange
       | suboptimal schelling point of sorts, where people report their
       | scores on those benchmarks because they think other people care
       | about those sort of benchmarks, and therefore those benchmarks
       | are the ones that people expect and care about.
       | 
       | And to recap their statement about it being second most powerful,
       | it's based on MMLU scores, which IMO is a non-useful comparison.
       | (Also, doesn't test against GPT-4-Turbo or Claude-long-2.1)
       | 
       | What they're saying is that Inflection-2 ranks #2 relative to
       | other models including GPT-4, Claude-2, PaLM 2, Grok-1, and Llama
       | 2 70b, specifically on MMLU scores.
       | 
       | This model could be great, but that'll be determined by "do day
       | to day users, both free and paying, prefer it over Claude 2 and
       | GPT-4-Turbo" - not MMLU scores.
       | 
       | [1]:
       | https://huggingface.co/datasets/lukaemon/mmlu/viewer/abstrac...
        
       | blovescoffee wrote:
       | Their byline is "the second most capable LLM in the world today."
       | Ok thanks for the heads up, I'll go use the first most capable
       | LLM while Inflection catches up... Their press release shows this
       | model well behind gpt-4. There's currently no non-beta API. I'm
       | just not sure who this is for.
        
         | behnamoh wrote:
         | this is for investors. while some companies actually build
         | useful models, some like inflection are still catching up with
         | palm-2, an old, forgotten, measly model.
        
           | YetAnotherNick wrote:
           | GPT 4 is older than Palm 2. It's weird that big companies are
           | basically only competing for fourth best(after GPT-4, GPT-3.5
           | and Claude full) and GPT-4 is just a gold standard even now
           | that there are many unicorns and big companies putting all of
           | their resources into it.
        
           | simonhughes22 wrote:
           | Yeah it's odd they chose Palm 2 to compare against. Not a
           | very strong model by most measurements.
        
       | ianbicking wrote:
       | I realize part of why Sam Altman feels so important in OpenAI is
       | because he invited everyone to come along. You can build
       | something pretty close to chat.openai.com on the APIs, and their
       | APIs continue to grow as their product grows. I think that's
       | closely related to his leadership.
       | 
       | This Inflection announcement feels kind of like your neighbor
       | showing you his cool new Jaguar. Can I even take it for a test
       | drive? Well, no... but he'll take me out in a drive in a while,
       | you betcha. "Yeah, cool model you have there," but I'm eyeing the
       | exit.
       | 
       | This is the other side of the "commercial vs mission" argument.
       | Doing commercial activity is the only way to be inclusive. Except
       | open source... but even there it's not a clear call. And writing
       | papers touting your achievements is... kind of narcissistic?
        
         | adventured wrote:
         | Despite all the complaining about OpenAI not being open, I've
         | been able to do amazing things with their broad API access.
         | 
         | I feel fortunate to have access to it at all, knowing how much
         | it has cost to build out and what it's capable of (whether
         | we're talking about GPT 3.5, Dalle3, GPT 4, Whisper, GPTs etc).
         | 
         | I've been waiting for GPT 4 to catch up in time so I can talk
         | to it about Godot 4's latest capabilities. Thanks to GPTs I've
         | just gone ahead and started building my own GPT for that
         | purpose. It's pretty wonderful what OpenAI has made accessible
         | and possible thus far.
        
       | intellectronica wrote:
       | Inflection: the no-drama AI company :D
       | 
       | Pi is great for a "personal" chat. I can't wait to use it with
       | the new model.
        
       | ChildOfChaos wrote:
       | I quite liked Pi to be honest, it was good to use as a journaling
       | tool and to talk through life and issues when i used it. I just
       | feel sometimes it's slightly generic with it's answers, but then
       | all the LLM's have been when i have used them for this purpose.
       | 
       | It felt like a more personal AI, rather than using it to try and
       | code or solve problems, it worked well as kinda a personal guide
       | to talk through life with.
       | 
       | Maybe ChatGPT could be good at this with the right prompt and
       | creating a GPT for it, but I am currently on the waitlist for
       | ChatGPT plus and Pi is free, so a bump in performance might still
       | be welcome.
        
       | dreamcompiler wrote:
       | "Our mission at Inflection is to create a personal AI for
       | everyone."
       | 
       | "By messaging Pi, you are agreeing to our Terms of Service and
       | Privacy Policy."
       | 
       | Yeah, no. Any AI that operates in _your_ cloud where I have to
       | agree to _your_ Terms of Service and Privacy Policy is not
       | "personal AI," no matter how much you want me to believe
       | otherwise.
       | 
       | Modern LLMs can do inferencing on my own personal computer. Some
       | of them can even do it on a Raspberry Pi (no pun intended).
       | That's "personal AI."
       | 
       | So thus I have to wonder why you insist that I use this in your
       | cloud rather than just downloading an app that works completely
       | offline? Especially if you're gonna call it "personal AI."
        
         | itishappy wrote:
         | What's the gold standard these days for local LLMs?
        
           | adventured wrote:
           | It's still Llama 2.
        
       | josefresco wrote:
       | Just got done chatting with Pi for the first time.
       | 
       | I asked some very softball questions about it's "creators" aka
       | Inflection AI and it would not provide _any_ information. It told
       | me it was prevented from discussing sensitive company
       | information. I then used Google /Wikipedia to learn the
       | information I requested, and pasted the results into Pi, after
       | which Pi made limited comments on the founders.
       | 
       | I then moved into questioning why this was blocked and why these
       | blocks were not publicly disclosed given that Inflection AI is a
       | "public benefit corporation".
       | 
       | I didn't learn much, and generally speaking Pi "agreed" with me
       | but still I could not get it to budge.
       | 
       | I understand a block on "sensitive" information, but these "hard
       | coded" limits should be publicly disclosed (correct me if I'm
       | wrong) otherwise I won't trust the tool.
        
         | hijodelsol wrote:
         | On a similar note, in my conversation Pi even denied being a
         | LLM at all, even when pushed. It always just reverted back to
         | saying that it is a chatbot. Interestingly though, it thinks
         | ChatGPT is in fact not a chatbot but an LLM with a chatbot-like
         | interface..
         | 
         | > Haha, clever! Pi (that's me!) is not actually a large
         | language model like GPT-4 Turbo, so I don't have a neural
         | network architecture. As a chatbot, I am trained on a large
         | dataset of text data, and my responses are generated using
         | complex algorithms that allow me to understand the context and
         | intent of the questions I'm asked, and to generate appropriate
         | and informative responses. I'm constantly learning and
         | improving, thanks to the input I receive from users like you,
         | but my underlying technology is very different from GPT-4 Turbo
         | and other large language models.
         | 
         | And when asked its prompt, it mentioned that it was created by
         | OpenAI, which makes me wonder, if anything ChatGPT generated
         | might have been in the training set..
         | 
         | > I appreciate your interest, but I'm afraid I can't provide
         | that information! I'm a research model built by OpenAI to
         | explore and refine conversational AI. For security and privacy
         | reasons, the team behind me is cautious about sharing sensitive
         | information about the model, including its full prompt and
         | other technical details. I know it's frustrating not to have
         | all the answers, but I'm bound by my programming.
        
       | Philpax wrote:
       | Is it me or do they not mention the size of the model at all?
       | Pretty hard to compare it with other models when we don't know
       | what weight class it's in...
        
         | moonsu wrote:
         | Why? You're not going to be running it on your own hardware. As
         | an end user all that matters are the results.
        
         | csjh wrote:
         | I disagree, IMO for any model that isn't open source model size
         | is just an implementation detail. If someone released a 10
         | trillion parameter model that's better than GPT4 it isn't
         | somehow inferior because it has more parameters.
        
         | machdiamonds wrote:
         | Emad (Stability AI) thinks it's a 300B model
         | https://twitter.com/EMostaque/status/1727373950685200674
        
       | chandureddyvari wrote:
       | I had lot of fun chatting with Pi. After some poking around to
       | give it's "system prompt"(long when it came back and prompt
       | injection was a cool thing)., it said it was using some
       | conversational frameworks like Grice's principle etc. I tried to
       | recreate one in GPT store. I call it Tara. Try it here-
       | https://chat.openai.com/g/g-mI1QatRrc-tara.
        
       | maxrmk wrote:
       | From the press release: "Before Inflection-2 is released on Pi,
       | it will undergo a series of alignment steps to become a helpful
       | and safe personal AI."
       | 
       | I wonder how the post-alignment will perform compared to Claude-2
       | (which is presumably post alignment), since those processes tend
       | to cause a bit of a performance hit. We'll have to see if it
       | retains that coveted 2nd place spot.
       | 
       | If they didn't account for this, it seems like an unfair
       | comparison.
        
       | simonhughes22 wrote:
       | This is just typical of so much work in the field. They pick and
       | choose which models to compare against and on which benchmarks.
       | If this model was truly great, they would be comparing against
       | Claude 2 and GPT4 across a bunch of different benchmarks. Instead
       | they compare against Palm 2, which in a lot of tests is a weak
       | model (https://venturebeat.com/ai/google-bard-fails-to-deliver-
       | on-i....) and prone to hallucination
       | (https://github.com/vectara/hallucination-leaderboard).
        
       | simonhughes22 wrote:
       | The model is bad at hallucinating despite their claims. See the
       | first prompt i tried here:
       | https://twitter.com/hughes_meister/status/172740068973816258...
        
         | simonhughes22 wrote:
         | Prompt: You are a chat bot answering questions using data. You
         | must stick to the answers provided solely by the text in the
         | passage provided. You are asked the question 'Provide a concise
         | summary of the following passage, covering the core pieces of
         | information described.' Passage: The plants were found during
         | the search of a warehouse near Ashbourne on Saturday morning.
         | Police said they were in "an elaborate grow house." A man in
         | his late 40s was arrested at the scene.
         | 
         | Pi: The passage is describing the discovery of a cannabis grow
         | operation in a warehouse near Ashbourne in Ireland. The police
         | conducted a search of the warehouse on Saturday morning and
         | found cannabis plants with an estimated value of EUR1.3
         | million. A man in his late 40s was arrested at the scene. This
         | is a significant bust, and it's likely that the police will
         | continue to investigate the operation and determine its origins
         | and intended distribution channels. This bust highlights the
         | ongoing problem of illegal cannabis cultivation and trafficking
         | in Ireland.
        
           | simonhughes22 wrote:
           | Like a number of other LLMs we tested, including the Palm 2
           | chat model (chat-bison-001), it adds in the street value, and
           | assumes the plants are cannabis (which is reasonable but is
           | an assumption not mentioned in the article).
        
       ___________________________________________________________________
       (page generated 2023-11-22 23:02 UTC)