https://xenaproject.wordpress.com/2025/01/20/think-of-a-number/ Xena Mathematicians learning Lean by doing. [cropped-light2] Skip to content * Home * About Xena * Student projects * What maths is in Lean? * Installing Lean and mathlib * Twitter * Useful links. - Can AI do maths yet? Thoughts from a mathematician. What is a quotient? - Think of a number. Posted on January 20, 2025 by xenaproject My feed was recently clogged up with news articles reporting that Sam Altman thinks that AGI is here, or will be here next year, or whatever. I will refrain from giving even more air to this nonsense by linking to the stories. This kind of irresponsible hype-generation drives me nuts (although it also drives up stock prices so I can see why the tech bros are motivated to do it). Sure AI can have a good crack at undergraduate mathematics right now, and sure that's pretty amazing. But our universities are full of students who can also have a good crack at undergraduate mathematics so in some sense we have achieved very little (and certainly nothing which is of any use to me as a working mathematician in 2025). If AI cannot do mathematics at PhD-student level (i.e., if it can't start thinking for itself) then AGI is certainly not here, whatever "AGI" even means. In an attempt to move beyond the hype and to focus on where we are right now, I'm going to try an experiment. It might fail. But if you're a number theorist reading this, you can help me make it succeed. Can AI do mathematics? The null hypothesis as far as I am concerned is "not really". More precisely, I have seen absolutely no evidence that language models are doing anything other than paraphrasing and attempting to generalise what they've seen on the internet, not always correctly. Don't get me wrong -- this gets you a long way! Indeed many undergraduate mathematics exams at university level are designed so that students who have a basic knowledge of the ideas in the course will pass the exam. The exam will test this knowledge by asking the student either to regurgitate the ideas, or to apply them in situations analogous to those seen in the problem sheets which came with the course. You can pass these exams by intelligent pattern-matching (if you have an encyclopedic knowledge of the material, which a machine would have). Furthermore, undergraduate pure mathematics is stagnant. The courses (other than my Lean course) which we teach in the pure mathematics degree at Imperial College London are essentially exactly the same as those which were offered to me when I was an undergraduate in Cambridge in the late 1980s, and very little post-1960 mathematics is taught at undergraduate level, even at the top institutions in the world, because it simply takes too long to get there. This is exactly why language models can pass undergraduate mathematics exams. But as I already said, this is no use to me as a working mathematician and it is certainly not AGI. How do we test AI? If you look at currently available mathematics databases, they are essentially all focused on mathematics which is at undergraduate or olympiad-style level. Perhaps the tech bros are deluded in thinking that because their models are getting high marks on such databases then they're doing well. I want to firstly stress the fundamental notion that these databases are completely unrepresentative of mathematics. More recently we had the FrontierMath dataset, which was designed by mathematicians and which looked like a real step in the right direction, even though the answers to the questions were not proofs or ideas, but just numbers (which is also a long way from what research mathematics looks like); the dataset is private, but the 5 sample questions which were revealed to us were clearly beyond undergraduate level. People (including me) were tricked into saying publically that any machine which can do something nontrivial here would really be a breakthrough. The next thing we know, OpenAI were claiming that their system could get 25 percent on this database. Oof. But shortly after that, Epoch AI (who put he dataset together) revealed that 25 percent of the questions were undergraduate or olympiad level, and this week it transpires that OpenAI were behind the funding of the dataset and reportedly were even given access to some of the questions (added later: in fact Elliot Glazer from Epoch AI has confirmed this on Reddit); all of a sudden 25 percent doesn't sound so great any more. So maybe we need to do it all again :-/ Let's make a database I want people (i.e. researchers in number theory at PhD level or beyond) to help me put together a secret database of hard number theory problems. I already have 5 but I need at least 20 and ideally far more, so I need help. The problems need to be beyond undergraduate level in the sense that undergraduates will not have been taught all the skills necessary to solve them. Because LLMs are not up to writing proofs, the answers unfortunately need to be non-negative integers: there is simply no point asking an LLM basic research level questions and then ploughing through the crap they produce and giving it 0/10; anyone who thinks that LLMs are actually capable of writing proofs should show me one that can do the Putnam because all efforts I saw on the 2024 exam were abysmal and this is undergraduate level: as always, the difficulty here is that soon after the questions and solutions are made public, the models train on them and the experiment is thus invalidated. In particular there is no point asking a model how to solve the 2024 Putnam questions now, one would expect perfect solutions, parroted off the internet by the parrots which we are being told are close to AGI. Note that restricting to questions for which the answer is a non-negative integer is again moving in a direction which is really really far from what researchers actually do, but unfortunately this is the level we are at right now. Once we have a decent amount of questions, which perhaps means at least 20 and maybe means 50, I'll announce this and then any AI company who wants a go can get in touch and I'll send them the questions but with no solutions and they can send me back the answers and then I'll announce the company and the score they got publically. Each company is allowed one go. When a few have had a go, I'll make the questions public. What will the questions look like? I'll finish this post with a more technical description of what I'm looking for. I need to be able to solve the question myself so they should be in number theory (broadly interpreted) or a nearby area. The answer should be a nonnegative integer (expressible in base 10 so perhaps < 10^100). Ideally they should be expressible in LaTeX without any fancy diagrams. The questions may or may not be comprehensible to an undergraduate but the proofs should be beyond all but the smartest ones, ideally because they need material which is not taught to undergraduates. The questions should not be mindless variants of standard questions which are already available on the internet -- although variants which actually need some understanding of what is going on are fine; we are trying to test the difference between understanding and the stochastic parrot model which is my model of an LLM: we are interested in testing the hypothesis that a language model can think mathematically. Language models can use computers and the internet so questions where you need to use a computer or online database to answer them is fine, although the question should not simply be "look up a number in a database" as this is too easy. Ideally the question needs an idea which is not explicit in the question. Ideally the answer is not easily guessable, because I have heard reports of LLMs getting questions on the FrontierMath dataset right, backed up by completely spurious reasoning. Please don't test your question by typing it into an LLM to see the answer. Any ideas? Contact me at my Imperial email. Let's give it 1 month, so applications close on 20th Feb and I'll report back soon after that date, hopefully to say that we have a decent-sized database and the experiment can commence. Thanks in advance to all number theorists who respond. Share this: * Click to share on X (Opens in new window) X * Click to share on Facebook (Opens in new window) Facebook * Like Loading... Related Unknown's avatar About xenaproject The Xena Project aims to get mathematics undergraduates (at Imperial College and beyond) trained in the art of formalising mathematics on a computer. Why? Because I have this feeling that digitising mathematics will be really important one day. View all posts by xenaproject - This entry was posted in Machine Learning, number theory and tagged AI, Artificial Intelligence, chatgpt, database, llm, technology, think of a number. Bookmark the permalink. - Can AI do maths yet? Thoughts from a mathematician. What is a quotient? - 12 Responses to Think of a number. 1. Positivist's avatar Positivist says: January 21, 2025 at 5:29 pm Could you comment on the important ways in which this problem set will differ from FrontierMath? My sense from FrontierMath is that they seemed a bit too focused on generating publicity (e.g. gathering hyped-up quotes from Fields Medallists), and less on quality. (Solutions should not be guessable without understanding, strange mix of problem difficulties with a weird rating system, etc). Is the idea that because you are curating them in your domain of expertise, that these problems will be a more uniform test of difficulty? (And also that you are more trustworthy/more closely aligned with the math research community than the folks at EpochAI?) LikeLike Reply + xenaproject's avatar xenaproject says: January 25, 2025 at 4:49 pm It will be similar to FrontierMath except that an AI company won't have the answers, I'm unlikely to write a paper, and I'm more likely to upload to Kaggle. LikeLike Reply 2. Simon's avatar Simon says: January 22, 2025 at 5:28 pm Could you give at least one example publicly to give an idea of the format/style/level of question? LikeLike Reply + xenaproject's avatar xenaproject says: January 25, 2025 at 4:48 pm Sure -- take a look at the two number theory questions in FrontierMath (p-adic numbers and points on a curve over a finite field). It's more of the same, or harder. Certainly nothing easier, that's one key idea here. LikeLike Reply 3. E's avatar E says: January 25, 2025 at 3:51 am Arguing that undergrads aren't thinking for themselves and "aren't of any use to you as a working mathematician" makes a really good case for AI being intelligent, because it's smart enough not to think like you. LikeLike Reply + xenaproject's avatar xenaproject says: January 25, 2025 at 4:46 pm Ha ha I mean that they (or rather 99% of them) are not any use to my project formalizing Fermat's Last Theorem, because we already have all of the undergraduate-level theory which I need formalized in Lean. I need PhD students right now because the material I need formalized is harder. 5 years ago undergraduates were essential to my goal, which was to get a basic undergraduate degree formalized in Lean. LikeLike Reply 4. Steve Byrnes's avatar sbyrnes321 says: January 25, 2025 at 10:33 pm Great post and good luck with this project! A little nitpick is that, for example, instead of saying "Can AI do mathematics?", I would have rather that you said "Can today's AI do mathematics?", or better yet "Can today's foundation models do mathematics?" ...After all, I hope we can agree that "AI" (in the sense of any algorithm running on a chip) "can do mathematics" (in the sense that this is theoretically possible and will almost certainly happen someday, even if it might be many years away). After all, human brains can do mathematics, and human brains work by algorithms not magic. LikeLike Reply 5. maths's avatar maths says: January 26, 2025 at 4:43 pm this problem - https://x.com/gf_256/status/1651346013792227332 LikeLike Reply + xenaproject's avatar xenaproject says: January 28, 2025 at 5:55 pm I think that problem is pretty well-known by now LikeLike Reply 6. Vova's avatar Vova says: January 28, 2025 at 5:24 pm basically completely unrelated to your points but I remember when this term AGI meant being better than the average human at almost all tasks not bring better than the best humans at all tasks. OpenAI when they wrote the charter to release agi to humanity rather than keep it way back when (big lol looking back that we ever believed this could work) defined it as "highly autonomous systems that outperform humans at most economically valuable work " which is in a sense orthogonal to being superhuman at mathematics. Personally I care more about the latter and do think it'll happen in 1-2 years, but I think the evidence is more on the scaling compute side and how well the training paradigms today are expected to generalize than what the current capabilities are (which I think is less informative especially when viewed from the bias of a professional human mathematician). Reinforcement learning methods when they work can often be expected to lead to superhuman performance which imitation style learning basically cannot (it can be mildly superhuman but even on principle not extremely superhuman). And the latest reasoning models are largely based in RL methods. Though I totally agree that RL against a formal mathematics environment seems most promising atm and are probably pursued by quite a few big labs (as I assume you know). LikeLike Reply 7. Alon Amit's avatar Alon Amit says: January 29, 2025 at 4:40 am Have you seen the problems in "Humanity's Last Exam"? There's a good number of mathematical problems, most with exact answers. Here's a link to 91 problems satisfying the search query "elliptic curve". (Disclosure: I had two problems accepted into the HLE database, though for some reason I can only see one of them right now. I'll send the other one to you separately, just in case it wasn't made public.) LikeLike Reply 8. sensationallymagnificentc42266f679's avatar sensationallymagnificentc42266f679 says: February 1, 2025 at 2:06 am Actually, OpenAI claims that the just released o3-mini already solves 20% of the T3 problems of FrontierMath, so it seems that it *can* do PhD-level math after all? See https://openai.com/ index/openai-o3-mini/ LikeLike Reply Leave a comment Cancel reply [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] * Search for: [ ] [Search] * Categories + Algebraic Geometry + computability + Fermat's Last Theorem + formalising mathematics course + General + Imperial + Learning Lean + liquid tensor experiment + M1F + M1P1 + M40001 + M4P33 + Machine Learning + number theory + Olympiad stuff + Research formalisation + rigour + tactics + Technical assistance + Type theory + Uncategorized + undergrad maths * Recent Posts + Think of a number: an update March 16, 2025 + What is a quotient? February 9, 2025 + Think of a number. January 20, 2025 + Can AI do maths yet? Thoughts from a mathematician. December 22, 2024 + Fermat's Last Theorem -- how it's going December 11, 2024 Xena Create a free website or blog at WordPress.com. [Close and accept] Privacy & Cookies: This site uses cookies. By continuing to use this website, you agree to their use. To find out more, including how to control cookies, see here: Cookie Policy * Comment * Reblog * Subscribe Subscribed + [croppe] Xena Join 187 other subscribers [ ] Sign me up + Already have a WordPress.com account? Log in now. * + [croppe] Xena + Subscribe Subscribed + Sign up + Log in + Copy shortlink + Report this content + View post in Reader + Manage subscriptions + Collapse this bar %d [b] Design a site like this with WordPress.com Get started