[HN Gopher] Improving mathematical reasoning with process superv...
___________________________________________________________________
Improving mathematical reasoning with process supervision
Author : davidbarker
Score : 107 points
Date : 2023-05-31 17:15 UTC (5 hours ago)
(HTM) web link (openai.com)
(TXT) w3m dump (openai.com)
| dr_dshiv wrote:
| Might mathematically adept Large Language Models be able to
| access the platonic "World of Forms?" (noetos kosmos - Noetos
| Kosmos) If so, there is an interesting basis for how and why AI
| might truly _understand_ , even without consciousness.
|
| https://aixd.substack.com/p/ai-might-not-need-experience-to-...
| kelseyfrog wrote:
| How the hell is an AI going to access the noetic realm without
| a nous? We have a hard enough time debating whether or not AI
| is conscious. We have to debate whether or not it has a nous
| too?
| dr_dshiv wrote:
| C.f. Roger Penrose's 3 worlds theory of reality, which
| distinguishes the platonic mathematical world from
| consciousness from material world. Navigating the platonic
| world of forms might not require one to "have a nous" but be
| constructed from it.
|
| I like the question though -- and of course, these are new
| ideas. Who knows? Do you think it is a halfway reasonable
| justification for AI understanding?
| kelseyfrog wrote:
| > Do you think it is a halfway reasonable justification for
| AI understanding?
|
| Personally, I think AI is revealing that consciousness is a
| social construction, and one that, as a social label, can
| have multiple groundings. That doesn't mean it's not
| necessarily real, just that it _has_ to be operationalized
| in order to quantify it. Operationalization requires some
| assumptions and it 's important to be transparent about
| what they are. You'll see people often invoke truth,
| reality, and fact, in order to justify their assumptions as
| "that's just the way it is." What a load of BS. Invoking
| reality to win an argument in a game about metaphysics
| displays a profound lack of understanding.
|
| "Consciousness" especially w.r.t AI is a contentious topic
| almost guaranteed to elicit emotions and trigger people to
| fall back on these assumptions. If nous has any utility,
| it's to de-escalate these discussions and avoid identity-
| laden emotional responses. It may even do that just because
| people don't have any idea what a nous is. Not too many
| people are well versed in (neo-)Platonism. If the nous was
| good enough to explain the mind to people back then, why
| shouldn't it be good enough to explain AI?
| ResearchCode wrote:
| Is machine learning model "alignment" a serious academic concept?
| I've only seen this mentioned in pseudo-philosophical Twitter
| posts before.
| mkoubaa wrote:
| Academic concepts are only considered serious by convention
| BasedAnon wrote:
| yes
| ftxbro wrote:
| its meaning has shifted to now mean ESG (https://en.wikipedia.o
| rg/wiki/Environmental,_social,_and_cor...) and DEI (https://en.
| wikipedia.org/wiki/Diversity,_equity,_and_inclusi...) for
| corporations that make algorithms that have learned natural
| language well enough that people are talking with them and
| taking them seriously. the AI alignment guys don't like that
| and they have started calling their old AI alignment meaning as
| 'not-kill-everyoneism' with mixed results (https://www.lesswron
| g.com/posts/jMzBhCRrr7otmqcvK/notkilleve...).
|
| EDIT: btw i don't have anything against ESG or DEI. I'm not a
| culture warrior complaining about 'woke ai'. I'm also not a
| lesswrong guy. I'm answering the OP's question "Is machine
| learning model "alignment" a serious academic concept? I've
| only seen this mentioned in pseudo-philosophical Twitter posts
| before." by saying what they mean now by 'AI alignment'.
| sebzim4500 wrote:
| I don't think anyone uses alignement to refer specifically to
| DEI bullshit.
|
| There are people using the term generally to refer to any
| kind of finetuning of an LLM, so they would consider what
| OpenAssistant did to be alignment even though there was no
| attempt to convince it to not kill humanity or to be
| politically correct.
| Q6T46nT668w6i3m wrote:
| Yes. It's very important. Think of it as quantifying whether a
| model understood the objective. This is, as you'd expect,
| especially important for auto-regressive models.
| sebzim4500 wrote:
| Yeah, but like every other term in AI, no two people agree on
| what the word means so discussion around it is challenging.
| mecsred wrote:
| Depends on your definition I guess. Some academics make careers
| out of pseudo-philosophical twitter posts.
| golol wrote:
| It is. I'd argue that alignment is the difference between
| GPT-3, which was a neat toy for nerds, and ChatGPT.
| User23 wrote:
| Interesting. Last I poked around with it GPT-4 performed
| abysmally with basic logic. For example it routinely confused the
| inference with the equivalence. Definitely cool to see progress.
| graycat wrote:
| So, the goal of the paper in
|
| https://cdn.openai.com/improving-mathematical-reasoning-with...
|
| is
|
| "Improving Mathematical Reasoning with Process Supervision".
|
| Since I hold a Ph.D. in pure/applied math from a famous US
| research university, I'm at least curious, maybe impressed, by
| the goal of "improving mathematical reasoning".
|
| In particular, one goal in the paper is
|
| "To train more reliable models ..."
|
| One of the techniques reported in the paper is
|
| "active learning" ... "used to train our best reward model."
|
| to improve the "reward" system to reduce "hallucinations",
| apparently also to obtain
|
| "more reliable models".
|
| And apparently part of the work is to be "step by step"?
|
| Okay: From my background in writing proofs in math the work can
| be seen as "step by step". E.g., the common high school special
| _format_ for writing proofs in plane geometry looks "step by
| step".
|
| Then, for "step by step" for a goal of
|
| "more reliable models"
|
| there are at least hundreds of examples nearly 100% "reliable" in
| well known math books by P. Halmos, W. Rudin, R. Buck, J. Neveu,
| and many more, apparently back at least to Euclid.
|
| So, it appears that the effort in
|
| "Improving Mathematical Reasoning with Process Supervision".
|
| is in a sense to _solve_ again a problem already solved with a
| nearly perfectly "reliable" solution back at least to Euclid.
|
| In addition, in my experience as a math student and teacher,
| commonly already in just the first few weeks of high school plane
| geometry, students get good at writing proofs that are very
| "reliable", and how those students learn does not look much like
| what is described in
|
| "Improving Mathematical Reasoning with Process Supervision".
|
| So, my guess would be, for the broad goal of having AI do
| "reliable" math, need a quite different approach.
|
| Math that is not "reliable"? From my experience applying math,
| e.g., to some important problems in US business and national
| security, in our society math that is not "reliable" stands to be
| less welcome than _week old fish_.
|
| Indeed, at this point, for the unique crown jewel of STEM, I
| nominate nearly 100% reliable mathematical proof as in Halmos,
| ..., Euclid.
| MacsHeadroom wrote:
| AI Provers (math models) are an entirely different class of
| model from LLMs. This paper and method is about getting
| language models to be better at math without changing their
| architecture.
|
| Math models already do math much better than language models.
| crakenzak wrote:
| Link to paper: https://cdn.openai.com/improving-mathematical-
| reasoning-with...
| mbbbackus wrote:
| Really excited to see if this approach can help reduce
| hallucinations overall. Process supervision combined with the
| browsing models could drastically improve how logically sound
| anything that's generated is. Creating consistently valid logic
| is the harder part so this is awesome.
| anonymousDan wrote:
| Is there some formal/objective definition of what exactly
| constitutes a 'hallucination' as opposed to other types of
| errors? At a high-level it only seems relevant with respect to
| questions for which there is some objectively true answer, or
| where 'facts' are included in answering a question that are
| false.
| golol wrote:
| The paper states that hallucinations are logical reasoning
| errors... that confused me. For me a hallucination is the
| conjuring up of facts, things, references, characters etc. as
| is convenient for the given context.
| sebzim4500 wrote:
| It's really hard to formally define hallucination, for me
| it's a "know it when you see it" situation.
|
| In my view, saying "10 + 4 = 15" is not a hallucination but
| inventing a citation that doesn't exist is one. There is a
| big grey area between these, like what if it cites a real
| paper but gets the year wrong? Is that a hallucination
| (because the paper as cited is fake), or just an incorrect
| statement (because the year is just wrong)?
| BeetleB wrote:
| Semantics aside: I once gave it a math problem, and its
| hallucinations really did present as logical errors. I was
| a TA for several years, and it was reminiscent of the
| mistakes students make. Something that sounds fairly
| plausible, but still incorrect, with logical errors or
| unjustified steps.
| User23 wrote:
| LLMs don't have errors. What you perceive as an erroneous
| output is actually just one that failed to please you.
| mbbbackus wrote:
| Good point. You can create a formal definition of
| hallucinations in formal language, like math and logic, but
| probably not for natural language. There are bound to be edge
| cases where natural language defies the rules of formal logic
| without being incoherent. But, while you might not be able to
| have a simple formal definition, maybe you could create a
| model that's trained to recognize these edge cases, and the
| model would be a sort of approximation of a formal
| definition.
| omginternets wrote:
| A (good) Socratic tutor for mathematics is something I would pay
| for without hesitation.
| passion__desire wrote:
| Hopefully one day multimodal learning models can learn directly
| from youtube lectures put up by John Baez [0] and be able to
| take students from Eli5 to EliPhD
|
| [0] https://www.youtube.com/watch?v=IdXVIWgxyDw
| bionhoward wrote:
| try VoxScript Plug-in and "can you explain $URL like I'm 5?
| then explain it like I'm a PhD"
| omginternets wrote:
| One can dream :)
| nerpderp82 wrote:
| You can extract the transcripts and have it automatically
| provide context on terms and supplemental reading to
| support the material.
| nerpderp82 wrote:
| You could definitely do that now using _many_ techniques, that
| would cost under $50, up to building your own RLHF for the
| socratic method. If this is something you and interested in and
| want, I would highly suggest pursuing. I use it for exploring
| and learning subjects all day long. With a couple of external
| tools, it would be 10x better and would enable less skilled
| users to get value out of it.
|
| You don't know what you don't know after all.
| sebzim4500 wrote:
| Honestly I don't think that GPT-4 is smart enough to be a
| useful tutor for anything beyond highschool maths, no matter
| how much finetuning/RLHF you do.
|
| I can see it being great for humanities though if you can get
| the hallucination rate down.
| visarga wrote:
| Funny, nobody thinks GPT-4 is quite there in their own
| field, but can readily agree it might be good enough in
| other fields.
| jmfldn wrote:
| I'm excited for the interactive, educational applications of
| this technology. Especially for things like maths.
| visarga wrote:
| YouTube has a real opportunity there - stop the video at any
| point and ask a question!
| sebzim4500 wrote:
| One thing I noticed is that they have a table in the appendix
| where they show that some of the datasets used for finetuning
| were already included in the pretraining.
|
| Is this the first time they have published that there was
| synthetic data in the GPT-4 pretraining dataset? Maybe that's
| part of the secret sauce behind why no one can catch GPT-4? I
| don't recall anyone else mentioning synthetic data.
|
| It certainly explains why there are so many people on the data
| team at OpenAI.
| visarga wrote:
| In the future I expect synthetic data will take the lion's
| share because it is easier to scale. In order to produce it you
| need a good way to filter out the junk, this can be done in
| many ways, some cheap and other expensive.
| Ozzie_osman wrote:
| It's always felt weird that there's this whole concept with a
| formal name ("chain-of-thought prompting") and papers written
| about it, etc, and basically it boils down to putting the phrase
| "Let's think step-by-step" at the end of the prompt. The obvious
| conclusion was that these models _should_ be trained to think
| step-by-step without having to prompt them to do so each time.
|
| The other interesting thing is the pedagogical parallels of this
| technique. A lot of times, when teaching/learning, we reward
| people purely based on outcome (did you get the right answer),
| with mild attempts to do better (like "show your work for partial
| credit"). Ideally this is just baked into teaching/learning all
| the time (the process is rewarded as well as the outcome).
| bionhoward wrote:
| Many extremely simple ideas are extremely difficult to validate
| with empirical evidence. We're fortunate to have an opportunity
| to make many such ideas much easier to A/B test using LLMs.
___________________________________________________________________
(page generated 2023-05-31 23:01 UTC)