[HN Gopher] Improving mathematical reasoning with process superv...
       ___________________________________________________________________
        
       Improving mathematical reasoning with process supervision
        
       Author : davidbarker
       Score  : 107 points
       Date   : 2023-05-31 17:15 UTC (5 hours ago)
        
 (HTM) web link (openai.com)
 (TXT) w3m dump (openai.com)
        
       | dr_dshiv wrote:
       | Might mathematically adept Large Language Models be able to
       | access the platonic "World of Forms?" (noetos kosmos - Noetos
       | Kosmos) If so, there is an interesting basis for how and why AI
       | might truly _understand_ , even without consciousness.
       | 
       | https://aixd.substack.com/p/ai-might-not-need-experience-to-...
        
         | kelseyfrog wrote:
         | How the hell is an AI going to access the noetic realm without
         | a nous? We have a hard enough time debating whether or not AI
         | is conscious. We have to debate whether or not it has a nous
         | too?
        
           | dr_dshiv wrote:
           | C.f. Roger Penrose's 3 worlds theory of reality, which
           | distinguishes the platonic mathematical world from
           | consciousness from material world. Navigating the platonic
           | world of forms might not require one to "have a nous" but be
           | constructed from it.
           | 
           | I like the question though -- and of course, these are new
           | ideas. Who knows? Do you think it is a halfway reasonable
           | justification for AI understanding?
        
             | kelseyfrog wrote:
             | > Do you think it is a halfway reasonable justification for
             | AI understanding?
             | 
             | Personally, I think AI is revealing that consciousness is a
             | social construction, and one that, as a social label, can
             | have multiple groundings. That doesn't mean it's not
             | necessarily real, just that it _has_ to be operationalized
             | in order to quantify it. Operationalization requires some
             | assumptions and it 's important to be transparent about
             | what they are. You'll see people often invoke truth,
             | reality, and fact, in order to justify their assumptions as
             | "that's just the way it is." What a load of BS. Invoking
             | reality to win an argument in a game about metaphysics
             | displays a profound lack of understanding.
             | 
             | "Consciousness" especially w.r.t AI is a contentious topic
             | almost guaranteed to elicit emotions and trigger people to
             | fall back on these assumptions. If nous has any utility,
             | it's to de-escalate these discussions and avoid identity-
             | laden emotional responses. It may even do that just because
             | people don't have any idea what a nous is. Not too many
             | people are well versed in (neo-)Platonism. If the nous was
             | good enough to explain the mind to people back then, why
             | shouldn't it be good enough to explain AI?
        
       | ResearchCode wrote:
       | Is machine learning model "alignment" a serious academic concept?
       | I've only seen this mentioned in pseudo-philosophical Twitter
       | posts before.
        
         | mkoubaa wrote:
         | Academic concepts are only considered serious by convention
        
         | BasedAnon wrote:
         | yes
        
         | ftxbro wrote:
         | its meaning has shifted to now mean ESG (https://en.wikipedia.o
         | rg/wiki/Environmental,_social,_and_cor...) and DEI (https://en.
         | wikipedia.org/wiki/Diversity,_equity,_and_inclusi...) for
         | corporations that make algorithms that have learned natural
         | language well enough that people are talking with them and
         | taking them seriously. the AI alignment guys don't like that
         | and they have started calling their old AI alignment meaning as
         | 'not-kill-everyoneism' with mixed results (https://www.lesswron
         | g.com/posts/jMzBhCRrr7otmqcvK/notkilleve...).
         | 
         | EDIT: btw i don't have anything against ESG or DEI. I'm not a
         | culture warrior complaining about 'woke ai'. I'm also not a
         | lesswrong guy. I'm answering the OP's question "Is machine
         | learning model "alignment" a serious academic concept? I've
         | only seen this mentioned in pseudo-philosophical Twitter posts
         | before." by saying what they mean now by 'AI alignment'.
        
           | sebzim4500 wrote:
           | I don't think anyone uses alignement to refer specifically to
           | DEI bullshit.
           | 
           | There are people using the term generally to refer to any
           | kind of finetuning of an LLM, so they would consider what
           | OpenAssistant did to be alignment even though there was no
           | attempt to convince it to not kill humanity or to be
           | politically correct.
        
         | Q6T46nT668w6i3m wrote:
         | Yes. It's very important. Think of it as quantifying whether a
         | model understood the objective. This is, as you'd expect,
         | especially important for auto-regressive models.
        
         | sebzim4500 wrote:
         | Yeah, but like every other term in AI, no two people agree on
         | what the word means so discussion around it is challenging.
        
         | mecsred wrote:
         | Depends on your definition I guess. Some academics make careers
         | out of pseudo-philosophical twitter posts.
        
         | golol wrote:
         | It is. I'd argue that alignment is the difference between
         | GPT-3, which was a neat toy for nerds, and ChatGPT.
        
       | User23 wrote:
       | Interesting. Last I poked around with it GPT-4 performed
       | abysmally with basic logic. For example it routinely confused the
       | inference with the equivalence. Definitely cool to see progress.
        
       | graycat wrote:
       | So, the goal of the paper in
       | 
       | https://cdn.openai.com/improving-mathematical-reasoning-with...
       | 
       | is
       | 
       | "Improving Mathematical Reasoning with Process Supervision".
       | 
       | Since I hold a Ph.D. in pure/applied math from a famous US
       | research university, I'm at least curious, maybe impressed, by
       | the goal of "improving mathematical reasoning".
       | 
       | In particular, one goal in the paper is
       | 
       | "To train more reliable models ..."
       | 
       | One of the techniques reported in the paper is
       | 
       | "active learning" ... "used to train our best reward model."
       | 
       | to improve the "reward" system to reduce "hallucinations",
       | apparently also to obtain
       | 
       | "more reliable models".
       | 
       | And apparently part of the work is to be "step by step"?
       | 
       | Okay: From my background in writing proofs in math the work can
       | be seen as "step by step". E.g., the common high school special
       | _format_ for writing proofs in plane geometry looks  "step by
       | step".
       | 
       | Then, for "step by step" for a goal of
       | 
       | "more reliable models"
       | 
       | there are at least hundreds of examples nearly 100% "reliable" in
       | well known math books by P. Halmos, W. Rudin, R. Buck, J. Neveu,
       | and many more, apparently back at least to Euclid.
       | 
       | So, it appears that the effort in
       | 
       | "Improving Mathematical Reasoning with Process Supervision".
       | 
       | is in a sense to _solve_ again a problem already solved with a
       | nearly perfectly  "reliable" solution back at least to Euclid.
       | 
       | In addition, in my experience as a math student and teacher,
       | commonly already in just the first few weeks of high school plane
       | geometry, students get good at writing proofs that are very
       | "reliable", and how those students learn does not look much like
       | what is described in
       | 
       | "Improving Mathematical Reasoning with Process Supervision".
       | 
       | So, my guess would be, for the broad goal of having AI do
       | "reliable" math, need a quite different approach.
       | 
       | Math that is not "reliable"? From my experience applying math,
       | e.g., to some important problems in US business and national
       | security, in our society math that is not "reliable" stands to be
       | less welcome than _week old fish_.
       | 
       | Indeed, at this point, for the unique crown jewel of STEM, I
       | nominate nearly 100% reliable mathematical proof as in Halmos,
       | ..., Euclid.
        
         | MacsHeadroom wrote:
         | AI Provers (math models) are an entirely different class of
         | model from LLMs. This paper and method is about getting
         | language models to be better at math without changing their
         | architecture.
         | 
         | Math models already do math much better than language models.
        
       | crakenzak wrote:
       | Link to paper: https://cdn.openai.com/improving-mathematical-
       | reasoning-with...
        
       | mbbbackus wrote:
       | Really excited to see if this approach can help reduce
       | hallucinations overall. Process supervision combined with the
       | browsing models could drastically improve how logically sound
       | anything that's generated is. Creating consistently valid logic
       | is the harder part so this is awesome.
        
         | anonymousDan wrote:
         | Is there some formal/objective definition of what exactly
         | constitutes a 'hallucination' as opposed to other types of
         | errors? At a high-level it only seems relevant with respect to
         | questions for which there is some objectively true answer, or
         | where 'facts' are included in answering a question that are
         | false.
        
           | golol wrote:
           | The paper states that hallucinations are logical reasoning
           | errors... that confused me. For me a hallucination is the
           | conjuring up of facts, things, references, characters etc. as
           | is convenient for the given context.
        
             | sebzim4500 wrote:
             | It's really hard to formally define hallucination, for me
             | it's a "know it when you see it" situation.
             | 
             | In my view, saying "10 + 4 = 15" is not a hallucination but
             | inventing a citation that doesn't exist is one. There is a
             | big grey area between these, like what if it cites a real
             | paper but gets the year wrong? Is that a hallucination
             | (because the paper as cited is fake), or just an incorrect
             | statement (because the year is just wrong)?
        
             | BeetleB wrote:
             | Semantics aside: I once gave it a math problem, and its
             | hallucinations really did present as logical errors. I was
             | a TA for several years, and it was reminiscent of the
             | mistakes students make. Something that sounds fairly
             | plausible, but still incorrect, with logical errors or
             | unjustified steps.
        
           | User23 wrote:
           | LLMs don't have errors. What you perceive as an erroneous
           | output is actually just one that failed to please you.
        
           | mbbbackus wrote:
           | Good point. You can create a formal definition of
           | hallucinations in formal language, like math and logic, but
           | probably not for natural language. There are bound to be edge
           | cases where natural language defies the rules of formal logic
           | without being incoherent. But, while you might not be able to
           | have a simple formal definition, maybe you could create a
           | model that's trained to recognize these edge cases, and the
           | model would be a sort of approximation of a formal
           | definition.
        
       | omginternets wrote:
       | A (good) Socratic tutor for mathematics is something I would pay
       | for without hesitation.
        
         | passion__desire wrote:
         | Hopefully one day multimodal learning models can learn directly
         | from youtube lectures put up by John Baez [0] and be able to
         | take students from Eli5 to EliPhD
         | 
         | [0] https://www.youtube.com/watch?v=IdXVIWgxyDw
        
           | bionhoward wrote:
           | try VoxScript Plug-in and "can you explain $URL like I'm 5?
           | then explain it like I'm a PhD"
        
           | omginternets wrote:
           | One can dream :)
        
             | nerpderp82 wrote:
             | You can extract the transcripts and have it automatically
             | provide context on terms and supplemental reading to
             | support the material.
        
         | nerpderp82 wrote:
         | You could definitely do that now using _many_ techniques, that
         | would cost under $50, up to building your own RLHF for the
         | socratic method. If this is something you and interested in and
         | want, I would highly suggest pursuing. I use it for exploring
         | and learning subjects all day long. With a couple of external
         | tools, it would be 10x better and would enable less skilled
         | users to get value out of it.
         | 
         | You don't know what you don't know after all.
        
           | sebzim4500 wrote:
           | Honestly I don't think that GPT-4 is smart enough to be a
           | useful tutor for anything beyond highschool maths, no matter
           | how much finetuning/RLHF you do.
           | 
           | I can see it being great for humanities though if you can get
           | the hallucination rate down.
        
             | visarga wrote:
             | Funny, nobody thinks GPT-4 is quite there in their own
             | field, but can readily agree it might be good enough in
             | other fields.
        
         | jmfldn wrote:
         | I'm excited for the interactive, educational applications of
         | this technology. Especially for things like maths.
        
           | visarga wrote:
           | YouTube has a real opportunity there - stop the video at any
           | point and ask a question!
        
       | sebzim4500 wrote:
       | One thing I noticed is that they have a table in the appendix
       | where they show that some of the datasets used for finetuning
       | were already included in the pretraining.
       | 
       | Is this the first time they have published that there was
       | synthetic data in the GPT-4 pretraining dataset? Maybe that's
       | part of the secret sauce behind why no one can catch GPT-4? I
       | don't recall anyone else mentioning synthetic data.
       | 
       | It certainly explains why there are so many people on the data
       | team at OpenAI.
        
         | visarga wrote:
         | In the future I expect synthetic data will take the lion's
         | share because it is easier to scale. In order to produce it you
         | need a good way to filter out the junk, this can be done in
         | many ways, some cheap and other expensive.
        
       | Ozzie_osman wrote:
       | It's always felt weird that there's this whole concept with a
       | formal name ("chain-of-thought prompting") and papers written
       | about it, etc, and basically it boils down to putting the phrase
       | "Let's think step-by-step" at the end of the prompt. The obvious
       | conclusion was that these models _should_ be trained to think
       | step-by-step without having to prompt them to do so each time.
       | 
       | The other interesting thing is the pedagogical parallels of this
       | technique. A lot of times, when teaching/learning, we reward
       | people purely based on outcome (did you get the right answer),
       | with mild attempts to do better (like "show your work for partial
       | credit"). Ideally this is just baked into teaching/learning all
       | the time (the process is rewarded as well as the outcome).
        
         | bionhoward wrote:
         | Many extremely simple ideas are extremely difficult to validate
         | with empirical evidence. We're fortunate to have an opportunity
         | to make many such ideas much easier to A/B test using LLMs.
        
       ___________________________________________________________________
       (page generated 2023-05-31 23:01 UTC)