https://terrytao.wordpress.com/2026/03/13/mathematics-distillation-challenge-equational-theories/ [cropped-co] What's new Updates on my research and expository papers, discussion of open problems, and other maths-related topics. By Terence Tao * Home * About * Career advice * On writing * Books * Mastodon+ * Applets * Subscribe to feed Mathematics Distillation Challenge - Equational Theories 13 March, 2026 in advertising, math.GR | Tags: competition, Equational Theory Project, SAIR | by Terence Tao Mathematical research traditionally involves a small number of professional mathematicians working closely on difficult problems. However, I have long believed that there is a complementary way to do mathematics, in which one works with a broad community of mathematically minded people on problems which may not be as deep as the problems one traditionally works on, but still are of mathematical interest; and that modern technologies, including AI, are more suitable for contributing the latter type of workflow. The " Polymath projects" were one example of this broad type of collaboration, where internet platforms such as blogs and wikis were used to facilitate such collaboration. Some years later, collaborative formalization projects (such as the one to formalize the Polynomial Freiman-Ruzsa conjecture of Marton, discussed previously on this blog here) became popular in some circles. And in 2024, I launched the Equational Theories Project (ETP) (discussed on this blog here and here), combining the rigor of Lean formalization with "good old fashioned AI" (in the form of automated theorem provers) to settle (with formal verification) over 22 million true-false problems in universal algebra. Continuing in this spirit, Damek Davis and I are launching a new project, in the form of an experimental competitive challenge hosted by the SAIR Foundation (where I serve as a board member, and which is supplying technical support and compute). The idea of this challenge, motivated in part by this recent paper of Honda, Murakami, and Zhang, is to measure the extent to which the 22 million universal algebra true-false results obtained by the ETP can be "distilled" into a short, human-readable "cheat sheet", similar to how a student in an undergraduate math class might distill the knowledge learned from that class into a single sheet of paper that the student is permitted to bring into an exam. Here is a typical problem in universal algebra that the ETP was able to answer: Problem 1 Suppose that {*: M \times M \rightarrow M} is a binary operation such that {x * (y * z) = (z * w) * w} for all {x,y,z,w} . Is it true that {x * (y * x) = (x * y) * z} for all {x,y,z}? Such a problem can be settled either by algebraically manipulating the initial equation to deduce the target equation, or by finding a counterexample to the target equation that still satisfies the initial equation. There are a variety of techniques to achieve either of these, but this sort of problem is difficult, and even undecidable in some cases; see this paper of the ETP collaborators for more discussion. Nevertheless, many of these problems can be settled with some effort by humans, by automated theorem provers, or by frontier AI systems; here for instance is an AI-generated solution to the above problem. However, these AI models are expensive, and do not reveal much insight as to where their answers come from. If one instead tries a smaller and cheaper model, such as one of the many open-source models available, it turns out that these models basically perform no better than random chance, in that when asked to say whether the answer to a question such as the above is true or false, they only answer correctly about 50% of the time. But, similarly to how a student struggling with the material for a math class can perform better on an exam when provided the right guidance, it turns out that such cheap models can perform at least modestly better on this task (with success rates increasing to about 55%-60%) if given the right prompt or "cheat sheet". "Stage 1" of the distillation challenge, which we launched today, asks for contestants to design a cheat sheet (of at most 10 kilobytes in size) that can increase the performance of these models on the above true-false problems to as high a level as possible. We have provided a "playground" with which to test one's cheat sheet (or a small number of example cheat sheets) some cheap models against a public set of 1200 problems (1000 of which were randomly selected, and rather easy, together with 200 "hard" problems that were selected to resist the more obvious strategies for resolving these questions); a brief video explaining how to use the playground can be found here. Submissions stage will end on April 20, after which we will evaluate the submissions against a private subset of test questions. The top 1000 submissions will advance to a second stage which we are currently in the process of designing, which will involve more advanced models, but also the more difficult task of not just providing a true-false answer, but also a proof or counterexample to the problem. The competition will be coordinated on this Zulip channel, where I hope there will be a lively and informative discussion. My hope is that the winning submissions will capture the most productive techniques for solving these problems, and/or provide general problem-solving techniques that would also be applicable to other types of mathematical problems. We started with the equational theory project data set for this pilot competition due to its availability and spectrum of difficulty levels, but if this type of distillation process leads to interesting results, one could certainly run in on many other types of mathematical problem classes to get some empirical data on how readily they can be solved, particularly after we learn from this pilot competition on how to encourage participation and share of best practices. SAIR will also launch some other mathematical challenges in the coming months that will be of a more cooperative nature than this particular competitive challenge; stay tuned for further announcements. Share this: * Print (Opens in new window) Print * Email a link to a friend (Opens in new window) Email * More * * Share on X (Opens in new window) X * Share on Facebook (Opens in new window) Facebook * Share on Reddit (Opens in new window) Reddit * Share on Pinterest (Opens in new window) Pinterest * Like Loading... Recent Comments Terence Tao's avatar Terence Tao on Mathematics Distillation Chall... Terence Tao's avatar Terence Tao on Mathematics Distillation Chall... J's avatar J on 245A, Notes 5: Differentiation... Unknown's avatar Anonymous on Mathematics Distillation Chall... Unknown's avatar Anonymous on Mathematics Distillation Chall... stacey8szmy's avatar stacey8szmy on Mathematics Distillation Chall... Terence Tao's avatar Terence Tao on Mathematics Distillation Chall... Unknown's avatar Anonymous on 245A, Notes 5: Differentiation... Unknown's avatar Anonymous on Mathematics Distillation Chall... Unknown's avatar Anonymous on Mathematics Distillation Chall... Terence Tao's avatar Terence Tao on 245A, Notes 5: Differentiation... Unknown's avatar Anonymous on 245A, Notes 5: Differentiation... Terence Tao's avatar Terence Tao on Mathematics Distillation Chall... Terence Tao's avatar Terence Tao on Mathematics Distillation Chall... Mikhled's avatar Mikhled on Mathematics Distillation Chall... [ ] [Search] Top Posts * Mathematics Distillation Challenge - Equational Theories * Career advice * The Equational Theories Project: Advancing Collaborative Mathematical Research at Scale * 245A, Notes 5: Differentiation theorems * Does one have to be a genius to do maths? * Books * On writing * Six Math Essentials * About * Cosmic Distance Ladder videos with Grant Sanderson (3blue1brown): commentary and corrections Archives * March 2026 (1) * February 2026 (3) * January 2026 (4) * December 2025 (5) * November 2025 (5) * September 2025 (1) * August 2025 (3) * July 2025 (1) * June 2025 (2) * May 2025 (5) * April 2025 (2) * March 2025 (1) * February 2025 (3) * January 2025 (1) * December 2024 (3) * November 2024 (4) * October 2024 (1) * September 2024 (4) * August 2024 (3) * July 2024 (3) * June 2024 (1) * May 2024 (1) * April 2024 (5) * March 2024 (1) * December 2023 (2) * November 2023 (2) * October 2023 (1) * September 2023 (3) * August 2023 (3) * June 2023 (8) * May 2023 (1) * April 2023 (1) * March 2023 (2) * February 2023 (1) * January 2023 (2) * December 2022 (3) * November 2022 (3) * October 2022 (3) * September 2022 (1) * July 2022 (3) * June 2022 (1) * May 2022 (2) * April 2022 (2) * March 2022 (5) * February 2022 (3) * January 2022 (1) * December 2021 (2) * November 2021 (2) * October 2021 (1) * September 2021 (2) * August 2021 (1) * July 2021 (3) * June 2021 (1) * May 2021 (2) * February 2021 (6) * January 2021 (2) * December 2020 (4) * November 2020 (2) * October 2020 (4) * September 2020 (5) * August 2020 (2) * July 2020 (2) * June 2020 (1) * May 2020 (2) * April 2020 (3) * March 2020 (9) * February 2020 (1) * January 2020 (3) * December 2019 (4) * November 2019 (2) * September 2019 (2) * August 2019 (3) * July 2019 (2) * June 2019 (4) * May 2019 (6) * April 2019 (4) * March 2019 (2) * February 2019 (5) * January 2019 (1) * December 2018 (6) * November 2018 (2) * October 2018 (2) * September 2018 (5) * August 2018 (3) * July 2018 (3) * June 2018 (1) * May 2018 (4) * April 2018 (4) * March 2018 (5) * February 2018 (4) * January 2018 (5) * December 2017 (5) * November 2017 (3) * October 2017 (4) * September 2017 (4) * August 2017 (5) * July 2017 (5) * June 2017 (1) * May 2017 (3) * April 2017 (2) * March 2017 (3) * February 2017 (1) * January 2017 (2) * December 2016 (2) * November 2016 (2) * October 2016 (5) * September 2016 (4) * August 2016 (4) * July 2016 (1) * June 2016 (3) * May 2016 (5) * April 2016 (2) * March 2016 (6) * February 2016 (2) * January 2016 (1) * December 2015 (4) * November 2015 (6) * October 2015 (5) * September 2015 (5) * August 2015 (4) * July 2015 (7) * June 2015 (1) * May 2015 (5) * April 2015 (4) * March 2015 (3) * February 2015 (4) * January 2015 (4) * December 2014 (6) * November 2014 (5) * October 2014 (4) * September 2014 (3) * August 2014 (4) * July 2014 (5) * June 2014 (5) * May 2014 (5) * April 2014 (2) * March 2014 (4) * February 2014 (5) * January 2014 (4) * December 2013 (4) * November 2013 (5) * October 2013 (4) * September 2013 (5) * August 2013 (1) * July 2013 (7) * June 2013 (12) * May 2013 (4) * April 2013 (2) * March 2013 (2) * February 2013 (6) * January 2013 (1) * December 2012 (4) * November 2012 (7) * October 2012 (6) * September 2012 (4) * August 2012 (3) * July 2012 (4) * June 2012 (3) * May 2012 (3) * April 2012 (4) * March 2012 (5) * February 2012 (5) * January 2012 (4) * December 2011 (8) * November 2011 (8) * October 2011 (7) * September 2011 (6) * August 2011 (8) * July 2011 (9) * June 2011 (8) * May 2011 (11) * April 2011 (3) * March 2011 (10) * February 2011 (3) * January 2011 (5) * December 2010 (5) * November 2010 (6) * October 2010 (9) * September 2010 (9) * August 2010 (3) * July 2010 (4) * June 2010 (8) * May 2010 (8) * April 2010 (8) * March 2010 (8) * February 2010 (10) * January 2010 (12) * December 2009 (11) * November 2009 (8) * October 2009 (15) * September 2009 (6) * August 2009 (13) * July 2009 (10) * June 2009 (11) * May 2009 (9) * April 2009 (11) * March 2009 (14) * February 2009 (13) * January 2009 (18) * December 2008 (8) * November 2008 (9) * October 2008 (10) * September 2008 (5) * August 2008 (6) * July 2008 (7) * June 2008 (8) * May 2008 (11) * April 2008 (12) * March 2008 (12) * February 2008 (13) * January 2008 (17) * December 2007 (10) * November 2007 (9) * October 2007 (9) * September 2007 (7) * August 2007 (9) * July 2007 (9) * June 2007 (6) * May 2007 (10) * April 2007 (11) * March 2007 (9) * February 2007 (4) Categories * expository (320) + tricks (13) * guest blog (10) * Mathematics (910) + math.AC (9) + math.AG (42) + math.AP (115) + math.AT (17) + math.CA (195) + math.CO (205) + math.CT (9) + math.CV (38) + math.DG (37) + math.DS (90) + math.FA (24) + math.GM (16) + math.GN (21) + math.GR (89) + math.GT (17) + math.HO (13) + math.IT (13) + math.LO (54) + math.MG (48) + math.MP (31) + math.NA (25) + math.NT (207) + math.OA (22) + math.PR (110) + math.QA (6) + math.RA (49) + math.RT (21) + math.SG (4) + math.SP (48) + math.ST (11) * non-technical (205) + admin (48) + advertising (75) + diversions (7) + media (14) o journals (3) + obituary (15) * opinion (36) * paper (268) + book (23) + Companion (13) + update (26) * question (128) + polymath (87) * talk (69) + DLS (20) * teaching (190) + 245A - Real analysis (11) + 245B - Real analysis (22) + 245C - Real analysis (6) + 246A - complex analysis (11) + 246B - complex analysis (5) + 246C - complex analysis (5) + 247B - Classical Fourier Analysis (5) + 254A - analytic prime number theory (19) + 254A - ergodic theory (18) + 254A - Hilbert's fifth problem (12) + 254A - Incompressible fluid equations (5) + 254A - random matrices (14) + 254B - expansion in groups (8) + 254B - Higher order Fourier analysis (9) + 255B - incompressible Euler equations (2) + 275A - probability theory (6) + 285G - poincare conjecture (20) + Logic reading seminar (8) * The sciences (1) * travel (26) * Uncategorized (1) additive combinatorics approximate groups arithmetic progressions Artificial Intelligence Ben Green Cauchy-Schwarz Cayley graphs central limit theorem Chowla conjecture compressed sensing correlation correspondence principle cosmic distance ladder distributions divisor function eigenvalues Elias Stein Emmanuel Breuillard entropy equidistribution ergodic theory Euler equations exponential sums finite fields Fourier transform Freiman's theorem Gowers uniformity norm Gowers uniformity norms graph theory Gromov's theorem GUE Hilbert's fifth problem ICM incompressible Euler equations inverse conjecture Joni Teravainen Kaisa Matomaki Kakeya conjecture Lie algebras Lie groups Liouville function Littlewood-Offord problem Maksym Radziwill Mobius function Navier-Stokes equations nilpotent groups nilsequences nonstandard analysis Paul Erdos politics polymath1 polymath8 Polymath15 polynomial method polynomials prime gaps prime numbers prime number theorem random matrices randomness Ratner's theorem regularity lemma Ricci flow Riemann zeta function Schrodinger equation Shannon entropy sieve theory structure Szemeredi's theorem Tamar Ziegler ultrafilters universality Van Vu wave maps Yitang Zhang RSS The Polymath Blog * Polymath projects 2021 * A sort of Polymath on a famous MathOverflow problem * Ten Years of Polymath * Updates and Pictures * Polymath proposal: finding simpler unit distance graphs of chromatic number 5 * A new polymath proposal (related to the Riemann Hypothesis) over Tao's blog * Spontaneous Polymath 14 - A success! * Polymath 13 - a success! * Non-transitive Dice over Gowers's Blog * Rota's Basis Conjecture: Polymath 12, post 3 13 comments Comments feed for this article 14 March, 2026 at 1:35 am hopefula0ac40f0f2 hopefula0ac40f0f2's avatar Interesting project. Once enough submissions are there, one might ask for pairs (or triplets or...) of cheat sheets which score best. Reply 14 March, 2026 at 2:00 am Anonymous Unknown's avatar The Zulip channel is invite only I think Reply 14 March, 2026 at 7:12 am Terence Tao Terence Tao's avatar Oops, we forgot to open it up with the competition. Account registration should now be open. Reply 14 March, 2026 at 4:29 am Mikhled Mikhled's avatar prof. Only 10 credits per day at the playground?!! Can u scale up test My cheatsheet solved 10 normal and hard200 already in seconds, though in compliance with all stipulated technical, mathematical, and timing criteria yield 100% And please, regarding the instructions to AI by you in the box under the participants' prompt box, Does what u wrote - u started with the word you are mathematician - affect our cheatsheet prompt efficiency of (the participant's code)? My cheat title is DRCM Mikhled Reply 14 March, 2026 at 7:19 am Terence Tao Terence Tao's avatar Unfortunately with the compute budget we have, spread out over the anticipated number of contestants (which we are not currently capping), this is close to the maximum amount of free compute we can offer for this stage. Stage 2 will have a cap on the number of participants and we will be able to offer more compute per participant. The playground is mostly intended to explain the contest format; I imagine that many contestants will develop their own custom testing frameworks with their own compute resources to be able to fine-tune their cheatsheets. Reply 14 March, 2026 at 3:41 pm Anonymous Unknown's avatar This project reminds me of the complexity class P/Poly. The cheat sheet is the "polynomial-length advice string", which doesn't depend on the specific input but instead is generically helpful across the whole set of (bounded length) inputs, in concert with a relatively efficient algorithm. Reply 14 March, 2026 at 6:03 pm Anonymous Unknown's avatar Are there any monetary prizes planned for Stage 1/2? Reply 14 March, 2026 at 7:12 pm Terence Tao Terence Tao's avatar Not for Stage 1, but we are planning some modest prizes for Stage 2, and possibly some prize ceremony as well. Details to be determined - this is our first competition, and we are still working things out as we go along! Reply 14 March, 2026 at 10:02 pm stacey8szmy stacey8szmy's avatar Hey Prof. Tao Most entries for the ETP Challenge are just static lists of laws. But a list isn't an intelligence; it's a script. If we really want to see which AI-human collaboration has 'solved' Magma logic, could / should we be doing Head-to-Head Adjudication? Pit the top frameworks against each other. Give OS 'A' a complex implication generated by OS 'B' and see if it can maintain logical sovereignty or if it collapses. A framework could/should be a Governance System; it should be able to adjudicate 'illegal' or 'impossible' structures without crashing. Why aren't we testing whose architecture actually holds the realm the longest? I know you said next phases will be developed, I hope this is something we can be anticipating. Thanks for the Mathematics Distillation Challenge. ~Team Zer00logy Reply 14 March, 2026 at 10:27 pm Anonymous Unknown's avatar It is possible that the provided prompts are just an elaborate way of providing a random seed to a stochastic sampling algorithm. Because of that, encouraging submissions with > 50% success rate might introduce a selection bias. I would, for the final assessment, include a large set of "null-prompts", e.g. prompts that sound like being mathematically helpful but in the end aren't, and score success rates against the obtained null distribution. Reply 15 March, 2026 at 7:43 am Terence Tao Terence Tao's avatar We actually are willing to accept prompts that encourage non-deterministic strategies that increase the success rate to something intermediate between 50% and 100%, though contestants will have to be responsible for avoiding overfitting such a prompt on the given test set (which will likely cause deterioration in performance when we test against the final evaluation set of problems). In Stage 2 we will in fact explicitly ask the LLM to produce a probability between 0 and 1 that a given implication is true (to be scored by log-loss), though to get full credit they will need to supply a deterministically verifiable certificate such as a Lean proof or an explicit counterexample. Reply 14 March, 2026 at 11:35 pm Anonymous Unknown's avatar I'm currently at 93.3% for the 1,000 "normal" questions, and 79.9% for the 200 "hard" questions, in terms of accuracy, when using weak model LLMs (not frontier). I'm trying to figure out if the juice is worth the squeeze for further optimization. Do you have ballpark figures for what is good or not? Thanks. -Kendon Reply 15 March, 2026 at 7:36 am Terence Tao Terence Tao's avatar We have various benchmark results for this challenge here that you can compare against. We are currently in the process of implementing a way in which competitors can "publish" their cheatsheets and performance results and hopefully build upon each other's work; we plan to introduce a rule that if a cheatsheet of one contestant is cited as an inspiration for another contestant that qualifies for Stage 2, then the first contestant will also automatically qualify, to encourage such sharing. Hopefully this will also give more clarity as to what kind of performance is achievable. Our final evaluation will be on a different subset of the 22 million equational implications than the 1200 problems provided. The entire set can be downloaded from this tool to use for further testing and to avoid overfitting. If it turns out that our initial set becomes saturated, we will tune our evaluation set to be more challenging than that set. Reply Leave a comment Cancel reply [ ] [ ] [ ] [ ] [ ] [ ] [ ] D[ ] For commenters To enter in LaTeX in comments, use $latex $ (without the < and > signs, of course; in fact, these signs should be avoided as they can cause formatting errors). Also, backslashes \ need to be doubled as \\. See the about page for details and for other commenting policy. << SLMath Deputy Director Search Blog at WordPress.com.Ben Eastaugh and Chris Sternal-Johnson. Subscribe to feed. * Comment * Reblog * Subscribe Subscribed + [bd4bda] What's new Join 12,193 other subscribers [ ] Sign me up + Already have a WordPress.com account? Log in now. * + [bd4bda] What's new + Subscribe Subscribed + Sign up + Log in + Copy shortlink + Report this content + View post in Reader + Manage subscriptions + Collapse this bar %d [b]