Summer 2026 flood: dozens of named conjectures settled on arXiv with disclosed AI help (July–September catalogue)
Between July and September 2026 arXiv saw a steady stream of papers that resolve a named conjecture or open question and disclose that a frontier model (mostly GPT-5.6 Sol/Pro and GPT-6 Astra, also Claude Fable 5/5.1, Opus 5/5.5, Gemini) found the key idea, the counterexample or the whole proof. This entry catalogues about 50 of them, with the AI role as the authors state it. Most are preprints without peer review. The most important ones have their own entries.
Key facts
- Scale: the 'Gold Rush in AI4Math' survey counted 1,712 arXiv math papers with substantive AI contributions from 1 Mar to 20 Aug 2026 (14% of submissions by August). The VibeMathed tracker listed 748 AI-involved problems (511 marked resolved, 158 Lean-verified) when checked on 30 Sep 2026
- Typical disclosure pattern: GPT-5.6 Sol/Pro dominates July–August; GPT-6 Astra dominates September after its early-September release; Anthropic models appear mostly as Claude Code agents, for Lean formalisation or for review
- Fully AI-generated, per the authors: Talagrand's Conjecture 9.1 (Park & Talagrand, 'generated entirely by the AI model GPT-6 Astra'); the dominating Hadwiger conjecture disproof (found by ChatGPT 6 Astra Ultra); the Partial List Colouring Conjecture disproof ('discovered and fully verified by ChatGPT 6 Astra Ultra'); Daykin–Frankl ('LLM-generated proof'); Yau's scalar-curvature integral bound refuted ('ChatGPT generated the main theorems and their proofs'); optimal shallow circuits for Majority ('originally found by GPT-6 Astra')
- Key idea from AI, human completion: the stable forking conjecture (Hart–Kim–Pillay 1996) refuted with GPT-5.6 Sol; a smooth random fast dynamo on T³ ('central proof idea was generated autonomously by ChatGPT 5.6 Sol Ultra'); Kahn–Saks conjecture on linear extensions (ChatGPT 6 Astra 'used primarily to aid in the discovery'); Talagrand's operator cotype problem ('discovered by ChatGPT (GPT-5.6)'); Khachiyan's ellipsoid conjecture ('An AI language model discovered the proof'; Lean-checked)
- Counterexamples found by chatbots on the first or second prompt: Schubitopes are not Ehrhart positive (GPT-5.6 Sol Pro, first prompt, 38 min); claw-free graphs' chromatic symmetric functions are not Schur positive (ChatGPT-5.6 Sol Pro); a very ample polytope with non-unimodal h*-vector (ChatGPT 5.6 Sol); the Hinrichs–Vybíral conjecture ('found in a single prompt' with ChatGPT 6)
- Agent systems: Kazhdan–Lusztig polynomials of matroids need not be unimodal (Rethlas agent on GPT-5.6 Sol, 4 h 59 min, where the GPT-5.6 Sol Ultra web interface failed); Ehrhart volume conjecture equality case (GPT-5.6 Sol, Fable 5 and Danus); the Fröberg conjecture for quintics and septics in four variables (GPT-5.6 Sol, Claude Fable 5 and Grok 4.6)
- Contrasting disclosures: Reed & Stein's dense-case Erdős–Sós proof was found 'without any use of AI'; Kielak et al. solved the uniform Turán tetrahedron problem with no AI-derived arguments, while a competing author completed an alternative proof with ChatGPT-6; Cairo's Ehlers–Kundt counterexample states 'All ideas in this paper are of human origin'
- Priority and credibility problems: the inhomogeneous Duffin–Schaeffer counterexample (GPT-5.6 Sol) had been announced by Pollington a year earlier; Kumar & Volk 'were unable to follow the details and verify' Sheshadri's AI-assisted determinantal-complexity proof and gave their own short quadratic lower bound; the 'Liouville Goldbach' proof was misreported as a Goldbach breakthrough
Science result
- Field
- mathematics / multiple (combinatorics, analysis, algebra, logic, TCS, mathematical physics)
- Problem
- Dozens of named conjectures and open questions (see list in body)
- Result
- About 50 claimed resolutions from July to September 2026 in which authors disclose a substantive AI contribution.
- AI system
- GPT-5.6 Sol, GPT-5.6 Pro, GPT-6 Astra, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5, Gemini 3.1 Pro, Aristotle
- Human role
- Varies from fully AI-generated proofs to human-led work with AI checking; see each item
- Verification
- Mostly unrefereed preprints; some Lean-verified; AI role self-reported by authors
- Status
- pending
What happened
This catalogue comes from a systematic pass over arXiv. It combined API queries for math, physics and theoretical-CS papers from 1 Jul to 30 Sep 2026 that mention a model name, "conjecture" or "open problem", with the AI-disclosure section of the project sweep and the public trackers. For each paper below, the PDF's AI statement was read. Quotes are from the papers, and nothing here was independently checked unless noted.
Combinatorics and discrete geometry
Kahn–Saks conjecture (balance in posets) proved. Aires, arXiv 2609.30895. ChatGPT 6 Astra "used primarily to aid in the discovery".
Talagrand's Conjecture 9.1. Park & Talagrand, arXiv 2609.33644. "The proofs presented in this document were generated entirely by the AI model GPT-6 Astra, and none of the authors claim any credit for them."
Dominating Hadwiger conjecture disproved. Illingworth & Steiner, arXiv 2609.35361. Found by ChatGPT 6 Astra Ultra; the humans wrote the exposition.
Partial List Colouring Conjecture (Albertson–Grossman–Haas) false. Noel, arXiv 2609.23291. Counterexample "discovered and fully verified by ChatGPT 6 Astra Ultra after some persistent prompting".
Daykin–Frankl conjecture confirmed. Williams, arXiv 2609.03087. "We verify and communicate an LLM-generated proof" (ChatGPT 5.6 Sol Pro).
Teschner's bondage-number conjecture counterexample. Yavari, arXiv 2609.04257. Generated by GPT-5.6 Sol Max.
Anstee–Sali conjecture counterexample. Wu, arXiv 2608.07646. GPT-5.6 Sol helped identify the example.
Stanley's rankwise lower-bound conjecture for differential posets (1988) disproved. Dai, arXiv 2607.22988 (25 Jul). Counterexample (for r=3, rank sequence 1,3,9,22,50 instead of 1,3,9,22,51) found by the TARS agent system in an autonomous search; Xinan Dai reconstructed and checked the proof by hand.
Bernhart–Kainen dispersability conjecture (1979): a new smallest counterexample, a 5-regular two-orbit graph on 16 vertices (the previous smallest was the 20-vertex Folkman graph), given an explicit algebraic construction by an AI agent. arXiv 2608.08118 (NeSy 2026); the Gold Rush survey's Table 3 lists GPT-5.5, Claude Opus 4.7, Gemini 3 Flash, Gemini 3.1 Pro and Claude Sonnet 4.6 as the models used.
Chromatic symmetric functions of claw-free graphs are not Schur positive. Matherne & Morales, arXiv 2607.21508. Found with ChatGPT-5.6 Sol Pro; independently found by Prajapati.
Schubitopes are not Ehrhart positive. Li & St. Dizier, arXiv 2608.00377. GPT-5.6 Sol Pro, first prompt, 38 minutes; verified in SageMath.
Very ample lattice polytope with non-unimodal h*-vector. Hofscheier, Kurylenko & Nill, arXiv 2608.21507. Found with ChatGPT 5.6 Sol; answers a question of Ferroni–Higashitani.
Equality case of Ehrhart's volume conjecture. J. Liu, arXiv 2608.01040. Complements OpenAI's inequality proof; "obtained by generative AI" (GPT-5.6 Sol, Fable 5, Danus).
Kazhdan–Lusztig polynomials of matroids need not be unimodal. Cheng et al., arXiv 2607.24186. The Rethlas agent on GPT-5.6 Sol found it in 4 h 59 min.
Laplacian S_{n,n} conjecture. Johnston, arXiv 2609.26895. "Significant amount of help" from ChatGPT 5.5, 5.6 Sol and 6 Astra.
Comon's conjecture, 27×27×27 counterexample. Lovitz, arXiv 2609.28292. GPT-6 Astra gave "an initial proof of the main result".
Topological Bárány–Larman conjecture proved for prime r; the optimal colorful Tverberg theorem of Blagojević–Matschke–Ziegler cannot extend to prime powers. Soberón, arXiv 2609.37876. "An LLM was used to simplify the proof of the main result and to construct the counterexample" (a 13-point example in R³, "found with the use of an LLM"; model not named).
Nagamochi's scoring lemma (unit-square packing, 2005) counterexample, showing the published proof of his rectangle bound is incomplete. Karakuş, arXiv 2609.37410. "AI tools were used during exploratory work, for symbolic and numerical checks, and for language and typesetting assistance."
Quantum percolation on regular trees: infinite clusters appear strictly before absolutely continuous spectrum (p_c < p_ac < 1). Becker & Oltman, arXiv 2609.38017. ChatGPT used "as a research and writing aid".
Kaplansky's conjecture (semifields) counterexamples. Nagy & Zhou, arXiv 2609.32651. Developed with ChatGPT 6 Pro; formalised by Harmonic's Aristotle in Lean 4.
Generalized packing–covering conjecture. Alfarano, Marino, Neri & Trombetti, arXiv 2609.34910. ChatGPT 6 Astra turned the authors' strategy into a complete argument.
Conway's subprime closure grows by the golden ratio (conjecture of Caragiu–Vicol–Zaki). Popescu, arXiv 2609.14188. GPT-6 Astra assisted; Lean-verified.
Kalai's conjecture for tight trees and Erdős–Sós for digraphs: see the Erdős–Sós entry.
Strongly aperiodic monotile in 3D ("Chair44"). Tsiokos, arXiv 2609.19214. Found by "an OpenAI reasoning model (ChatGPT, Astra)"; the text was largely written by Claude Fable 5.1 and reviewed by agents. Follow-ups: arXiv 2609.24779 (notes) and arXiv 2609.23783 (matching rules).
Chvátal's conjecture (1972) proved. Chang, Liu & Liu, arXiv 2609.19123: ChatGPT proved a guiding special case. Short spectral "Book proof" by GPT-6 Astra in Ellis–Filmus–Friedgut, arXiv 2609.28404. Lean formalization via Codex. Own entry:
2026-09-16-chvatal-conjecture-proved.Diagonal bipartite Ramsey b(t,t) = O(2^t). Mubayi, arXiv 2609.32937. "The proof was found by GPT-6 Astra." Own entry:
2026-09-26-mubayi-zarankiewicz-astra.1/3–2/3 conjecture: first general improvement of the Brightwell–Felsner–Trotter balance constant in 30 years. Aires, Chan, Pak & Panova, arXiv 2609.30888. Minor AI role: "we directed an AI assistant to explore the underlying algebraic optimization problem" for one lemma's inequalities.
List Total Colouring Conjecture (Borodin–Kostochka–Woodall 1997; Juvan–Mohar–Škrekovski; Hilton–Johnson) false: a 20-vertex cubic graph with total chromatic number 4 and list total chromatic number 5. Noel, arXiv 2609.38417 (29 Sep). "ChatGPT 6 Astra Ultra discovered the counterexample with little input from the author" (prompted on Sept 24). Own entry:
2026-09-29-list-total-colouring-conjecture-false-astra.Fishburn's latent-subset conjecture (1987) proved. Dong & Mao, arXiv 2609.35920 (28 Sep). "This proof was found by GPT-6 Astra, following an approach suggested by the author and using the weighted star inequality of Chang, Liu, and Liu" (the inequality from the Chvátal proof).
Burr–Erdős–Graham–Sós conjecture for the seven-cycle, the remaining case k=3 after Bucić–Chen–Ma: f(n, ⌊n²/4⌋+1, C₇) = (1/8+o(1))n². Shahab, arXiv 2609.38286 (29 Sep). "Made extensive use of AI systems, principally agents built on OpenAI's Codex and Anthropic's Claude" for the proof search, the exact rational certificate, the Lean formalisation and the manuscript; one counting lemma proved with Harmonic's Aristotle; main results Lean-checked.
Aigner's majorization conjecture (adaptive group testing on star forests) false: an explicit six-test counterexample and a family for every k ≥ 6; the converse holds for k ≤ 5. Karpelevitch, arXiv 2609.38284 (29 Sep). "Generative-AI assistants, primarily OpenAI Codex and, to a lesser extent, Anthropic Claude" were used for arguments, software, experiments and writing.
Erdős–Ulam conjectures on monochromatic union-closed families, both claimed: every 2-colouring of subsets of [n] has a monochromatic union-closed family of super-polynomial size, but not necessarily of exponential size; Lean-verified. Bhattacharjee, Mandal & Bhattacharya, arXiv 2610.02833 (2 Oct). Minor AI role: "Claude (Anthropic) assisted with the coding" of the Lean formalisation.
Analysis, PDE and geometry
- Landis conjecture fails in dimensions ≥ 3 (real potentials). Frank & Ivanisvili, arXiv 2608.00802. "The authors acknowledge the use of AI tools"; no detail.
- Pólya's conjecture for higher-dimensional Neumann balls. Filonov, Levitin, Polterovich & Sher, arXiv 2607.29305. ChatGPT and Claude "contributed to the development of several technical lemmas" and the rigorous computer-assisted algorithm.
- Yau's conjectured scalar-curvature integral bound refuted. Hao & Zhu, arXiv 2609.06533. "ChatGPT generated the main theorems and their proofs" (GPT-5.6 Pro).
- AI-discovered smooth random fast dynamo on T³. Rowan, arXiv 2608.20105. Central idea "generated essentially autonomously by ChatGPT 5.6 Sol Ultra"; the original AI manuscript is in the arXiv source.
- Nevanlinna's century-old half-plane problem counterexample. He & Zhang, arXiv 2608.24829. "AI-assisted exploration" (model not named).
- Fuchs's conjecture counterexample. Eremenko & Zhang, arXiv 2609.28443. ChatGPT as an "exploratory tool" and for editing.
- Forsythe's conjecture for restarted conjugate gradients. Colbrook, Stepaniants & Townsend, arXiv 2609.04659. Framed as an experiment in "how far a frontier language model could be pushed" (GPT-5.6, GPT-6).
- Rockafellar's sum conjecture fails. Boţ, arXiv 2609.13906. GPT-6 Astra assisted with the development.
- Hinrichs–Vybíral conjecture counterexample. Vybíral, arXiv 2609.21733. Found by a colleague "in a single prompt try" with ChatGPT 6.
- Talagrand's operator cotype problem. Wu, arXiv 2609.19731. "The counterexample was discovered by ChatGPT (GPT-5.6)."
- Khachiyan's ellipsoid conjecture. Zhou, Zou & Liu, arXiv 2609.28447. "An AI language model discovered the proof"; main theorem Lean-verified.
- Smooth Hamiltonian diffeomorphism with two fixed points on S²×S². Jiao, arXiv 2609.33626. "GPT suggested a key idea … most of the computation is done by GPT."
- Explicit mono-monostatic polyhedron (a certified polyhedral Gömböc). Schettini Gherardini, arXiv 2609.07827. A Claude Opus 4.8 / Fable 5 agent designed and ran the experiments; exact certificates.
- Lukic conjecture counterexample. Yan, arXiv 2607.26419. "This example was generated by GPT-5.6."
- Feige's conjecture. Nie & Wei, arXiv 2607.24528. Proof "obtained with the assistance of GPT-5.6 Sol".
- Pólya–Szegő logarithmic vs Newtonian capacity conjecture (1945): only a conditional reduction, the conjecture "remains open". Clark & Laugesen, arXiv 2609.35438 (28 Sep). Review role only: ChatGPT/Codex with GPT-5.6 Sol "assisted with a detailed critical review" of calculations, logic and citations.
Algebra, number theory, topology and logic
Stable forking conjecture (Hart–Kim–Pillay 1996) refuted. Freitag & Mutchnik, arXiv 2609.00436. "This is an AI-generated result proven with the help of GPT-5.6 Sol."
Huneke–Wiegand conjecture counterexample. Pham, arXiv 2609.07615. AI-assisted search with GPT-5.6 Pro; Craig Huneke independently recomputed the data.
Sato's weak F-equivalence conjecture counterexamples. Chakravarty, Choi & Xu, arXiv 2608.18054. GPT-5.6 Sol "produced the key construction".
Qin's quasimodularity conjecture for Hilbert schemes of points. Alekseev et al., arXiv 2609.33884. "Most of the formal arguments … were initially generated by GPT-5.6 Sol."
Fröberg's conjecture for quintics and septics in four variables. Wang & Zhang, arXiv 2608.24797. GPT-5.6 Sol, Claude Fable 5 and Grok 4.6 workflow.
Mod 4 Kawauchi conjecture. Conant, arXiv 2607.18655. Claude Fable 5 "proposed the quotient-tower strategy" and drafted the first version.
HZ/4 is not an E2-Thom spectrum (the remaining case). Ji, arXiv 2609.19446. GPT-5.6 Sol and GPT-6 Astra; reviewed with Claude Fable 5.1.
Yang's conjecture (tempered xi function) disproved. Kazin & Kadyrov, arXiv 2609.29898. GPT-6 Astra identified the key sine-transform formulation.
Inhomogeneous Duffin–Schaeffer conjecture counterexamples. He & Liao, arXiv 2609.30870. Found with GPT-5.6 Sol; the authors later learned that Pollington had announced a counterexample more than a year earlier.
Fraenkel's conjecture (Beatty sequences). Tan & Zhang, arXiv 2609.01570. GPT-5.6 Sol resolved the cases m = 8–11, which inspired the proof strategy; Codex helped with the finite-case verification code.
Separable Jacobian conjecture in characteristic two, dimension-2 counterexample. Mondello, arXiv 2608.02634. Lean-checked; ChatGPT/Codex used for organisation; Aristotle replay.
Colombo's determinant problem. arXiv 2609.00101. The WuJie agent, DeepSeek, Qwen, Kimi and GPT played "a substantial role in identifying the proof strategy"; Lean 4.
Hirose's duality conjecture (q-discretised iterated integrals on the four-punctured line). Seki, arXiv 2609.40213. "ChatGPT, powered by GPT-6 Astra Pro, was used for mathematical discussions and to refine the English prose" (minor role).
Stuck–Zimmer conjecture, two new cases (ergodic actions of irreducible higher-rank lattices; irreducible actions with an SL₂(ℝ) factor). Machado & Yifrach, arXiv 2609.39424. "GPT6 Astra was used to help with the literature review, explore research leads, carry out some computations and proofread" (supporting role).
De Bruijn–Erdős consecutive-gap problem. Korsky, arXiv 2609.07196. GPT Astra used "for completing the mathematical argument".
Lehmer's permutation conjecture (1965; a research problem in Knuth's TAOCP). Verhoeff, arXiv 2610.01240. "The hypercube proof at the heart of this article was found and developed by Claude Opus 5.5 in a multi-agent proof effort"; Lean formalization by Aristotle (own entry: 2026-10-01-lehmer-permutation-conjecture-opus-5-5).
Amenable actions of free groups on unital simple AF-algebras (the last open case, for free groups, of the existence problem for amenable actions of non-amenable groups on classifiable simple C*-algebras). Suzuki, arXiv 2610.01460. "GPT-6 Astra (OpenAI) proposed a proof, including the idea of using [5], a key reference of which the author had previously been unaware."
Erdős–Graham square products of factorials (Erdős Problem #374): order of growth of D_k(X) settled for every k, including D_5 ≍ D_6 ≍ X. Yudin, arXiv 2610.01899. "Large language models (ChatGPT and Claude) were used in developing and checking arguments, locating references, and editing the exposition."
Explicit polynomial counterexample to Connes' embedding conjecture (degree 12 in 65 variables; the conjecture was already known to be false via MIP* = RE, but no explicit polynomial was known). Wang & Zhi, arXiv 2610.01536. "Generative AI tools assisted with exploring proof strategies, checking calculations, and revising the exposition."
Pappas's conjecture on ramified unitary local models, case r>s at self-dual level for p ≥ 2s−1 (flatness of the wedge local model). Luo, arXiv 2610.02602 (1 Oct). Minor AI role: "The author used ChatGPT and Claude for general research assistance".
Cyclic steepest descent is not universally R-superlinear (disproves the form stated in Dai's ICM 2022 survey). Li & Wang, arXiv 2610.00939. "The main results of this paper were obtained through a generative-AI workflow using OpenAI GPT-5.6 Sol, Anthropic Claude Fable 5, and OpenAI GPT-6 Astra."
Borg's conjecture on intersecting integer partitions: counterexamples at every scale. Person & Schweser, arXiv 2610.01747. "ChatGPT (5.6 Sol and later GPT-6 Astra) also performed the asymptotic calculations based on Szekeres' formula and assisted with the exposition."
Also on Oct 1, with AI disclosures but smaller or unstated AI roles: Borcea's 2-variance conjecture and Baernstein's quasi-norm monotonicity conjecture (Zhang, 2610.02035, 2610.02009; Lean 4 for Borcea), the Bryant conjecture for minimal hypersurfaces in S⁴ (Qian & Tao, 2610.00886; "ChatGPT … was used in the preparation"), and higher-dimensional Morley simplices (Tran, 2610.01216).
Theoretical CS, information theory, quantum
Optimal shallow circuits for Majority. Lecomte & Ramakrishnan, arXiv 2609.34029. Constructions "originally found by GPT-6 Astra"; the authors rebuilt the proofs from high-level ideas.
Quadratic lower bound on determinantal complexity. Kumar & Volk, arXiv 2609.34462. Came out of using ChatGPT Astra to parse Sheshadri's AI-assisted proof, whose details they were 'unable to follow' and verify.
Strong secretary conjecture for linear matroids. Bérczi, Dughmi, Livanos & Soto, arXiv 2609.20797. Astra "identified the supermodularity … and proposed the uncrossing argument" on 15 Sep; concurrent-discovery note.
Nelson–Nguyen conjecture. Mai & Rao, arXiv 2609.22548. ChatGPT-5.6 Pro "used in proving and writing".
Markovity conjecture for two-receiver broadcast channels refuted. Liu & Huang, arXiv 2608.13170. GPT-5.6 Sol; Chandra Nair suggested the AI search and checked the result.
Shor's orthogonal-measurement conjecture. arXiv 2609.27992. Codex and GPT-5.6 Sol for exploration and error checking only.
Distributional variants of the Aaronson–Ambainis conjecture. arXiv 2609.35327. "Google Gemini suggested the core idea underlying the inner-gadget construction."
Haah's 3D cubic code thermalizes rapidly (settles whether it is a self-correcting quantum memory: it is not, at any positive temperature; first analytic rapid-mixing proof for a fracton code). Stengele, Caha, Capel & Warzel, arXiv 2609.39959. "We acknowledge the use of LLMs for a proof of Lemma 4.1" and for testing proof ideas (model not named).
Generalized semi-Clifford conjecture (Zeng et al. 2007) at level k ≤ 4, any prime dimension. Marcus, Subramanian & Dall'Agnol, arXiv 2609.38751. OpenAI Codex helped develop proof strategies and check algebra; Claude Code for literature and drafting.
Edit-distance codes over a four-letter alphabet (DNA barcodes): E₄(6,3) ≥ 120 (was 114) plus 12 further improved lower bounds at lengths 6–9. Yeung, arXiv 2609.39081. An LLM coding agent wrote the search and verifier code over five weeks; the paper documents failures in which unchecked intermediate results stopped the search early.
Cogentic (Google Research multi-agent prover on Gemini): novel results on five open problems in online learning, auction theory and mechanism design, each checked by domain experts. Cai, Gupta, Jiang, Liaw, Mehta, Velegkas & Wang, arXiv 2609.40324; results listed at sites.google.com/view/cogentic.
Sept 30 – Oct 2 additions (found Oct 5)
- Erdős #108 (Erdős–Hajnal high-girth problem) disproved through a Lean bounty (Kohlmeyer–Kruer, ChatGPT/Codex). The 3-colour sharpening by Nguyen & Walczak, arXiv 2609.40192, has "ChatGPT-6 Astra discovered a proof of Theorem 2.1". Own entry:
2026-09-15-erdos-108-high-girth-problem-disproved-conjectures-io. - ε-Dvoretzky conjecture proved. Klartag & Moshe, arXiv 2610.03204. The key idea "stemmed from a discussion between ChatGPT and the second named author". Own entry.
- Barvinok's log-concavity question (line version). Leake & Mohammadi Yekta, arXiv 2609.39917: "proven using ChatGPT 6 Astra". Own entry.
- Triangle-free box graphs beating Burling (1965). Tomon, arXiv 2610.03517: "found by ChatGPT Astra after guidance from the author". Own entry.
- Friedgut's continuous-cube coalition conjecture: a simplified proof of Chattopadhyay–Gurumukhani's result. Friedgut, arXiv 2610.03086: "discovered with the aid of ChatGPT-6 Astra … my mathematical contribution to this current project was limited to my two suggestions to the bot".
- Mond's conjecture for corank-one germs C³→C⁴. Rimányi, arXiv 2610.03331. Experiments, first drafts of several proofs and audits came "with substantial help from AI (Anthropic's Claude; OpenAI's ChatGPT)".
- Simpson's closedness conjecture in arbitrary rank. Hu, arXiv 2610.03542. "ChatGPT was used for calculations, proving intermediate results, and polishing the writing."
- Lam–Schilling–Shimozono conjecture: K-k-Schur functions are k-Schur positive. Bai & Guo, arXiv 2610.01708. The proof is human; "the counterexample given in Section 5 was found by ChatGPT-6 Astra".
- Strong Roberson conjecture (homomorphism distinguishing power) refuted. Kristjánsson, arXiv 2610.03550: "developed through substantial interaction with an AI agent".
- Exact eight-point angular-energy conjecture (Hiriart-Urruty's list) counterexample. Zhao, arXiv 2610.03207: "produced by OpenAI's GPT-5.6 Sol Ultra through Codex".
- Hegselmann–Krause models: the proximity digraph need not freeze. Hegarty, Ognissanti & Wedin, arXiv 2610.03229: "All of the examples above were found by Claude Opus 5.5".
- Delooping level is not additive under tensor products. Lam, arXiv 2609.39745. The counterexample was found by "the newest publicly available (Pro Subscription) model of ChatGPT" (Oct 1).
- Quantum random self-reduction for linear problems. Asadi, Hirahara & Shimizu, arXiv 2609.39823: "The algorithm and proof strategy underlying the main theorem were suggested by ChatGPT 5.5."
- Nearest stabilizer product state hardness. Grier, Pashayan & Schaeffer, arXiv 2610.02037: "Significant parts of the hardness proofs … were found by an interactive prompting of Gemini 3.1 Pro."
- Disjoint-data inverse problem sharp spectral threshold. Ylinen, arXiv 2610.00871: "GPT-5.6-sol produced the initial nonuniqueness construction".
- Ramanujan–Sato series of level 11 (open since Cooper et al. 2015). Campbell, arXiv 2610.00551. Built on "a construction described by GPT-5.6 Sol Pro".
- Prime labelings of trees. Ho, arXiv 2610.02271: "Most of the proof development, computational verification, and drafting … were carried out by generative AI systems, principally GPT Astra and Claude Opus".
- Standard conjecture of Hodge type, explicit form for Hermitian varieties. Harashita, arXiv 2610.01731, with "extensive use of Claude Opus 5 and Claude Opus 5.5 for formulating assertions, proving results, drafting".
- Smaller or mostly human: Quillen's finiteness conjecture for the cotangent complex in characteristic 2 (Schürg, 2610.03208; ChatGPT as "somebody to talk to"); irrationality exponent of ζ(2) below 5.0193784 (Niedbala Giraudin, 2610.02912; "Technical assistance: Claude"); typical intersecting families at n=2k+1, 2k+2 (Li, 2610.02119; ChatGPT 5.6 Sol helped with technical lemmas); Matsumura's extension problem, smooth case, confirming Siu's plurigenera conjecture in that setting (Chen, Rao & Wang, 2609.40040; ChatGPT 5.6 and 6 Pro as auxiliary tools).
Late-September disclosures the sweep missed (PDF scan, Oct 5 evening)
- Lovász conjecture (1969), long cycles in vertex-transitive graphs: n^(1−o(1)), up from n^(2/3−o(1)). Li & Methuku, arXiv 2609.38135 (29 Sep). GPT-5.6 Sol pointed the authors to the Tessera–Tointon structure theorem and Babai's contraction lemma, and "supplied the required group-theoretic arguments" for the key remaining Lemma 6.1. Own entry:
2026-09-29-lovasz-conjecture-nearly-linear-bound-gpt-5-6-sol. - Tight convergence of Nesterov's fast gradient method: Peppy, an AI workflow built on performance-estimation problems (PEP), closes Conjectures 4–5 of Taylor et al. (2017) for FGM and OGM and the FGM/FISTA rational-coefficient conjectures. The authors call these the first analytic proofs of the tight FGM bounds, open "roughly a decade". Suh, Yoon, Nguyen, Ying & Ma, arXiv 2609.35762 (28 Sep). Proofs generated by Codex with GPT-6 Astra (Ultra reasoning) and checked symbolically in SymPy; "the Codex-generated drafts required human editing". Code: PEPFlow peppy-v1.
- Watkins's conjecture holds for all infinite groups (graphical regular representations; the Cayley index of every infinite group is 1, 2 or 8). Sutherland, arXiv 2610.01049 (1 Oct). "OpenAI's ChatGPT provided substantial assistance with developing arguments, locating references, and preparing the checking programs"; the checks "do not constitute formal verification".
- Rational points near the moment curve: counterexamples showing the Hickman–Srivastava range is nearly sharp, against the belief that an O(n⁻¹) improvement was possible. Chen, Srivastava & Technau, arXiv 2610.02443 (1 Oct). A self-contained argument (the recursive identity (2.5)) "was later suggested to us by Chat G.P.T. 6 Astra"; written and verified by the authors.
- Chen's conjecture for flat n-tori confirmed (sharp Willmore lower bound, Clifford torus the unique minimiser), but it fails for the total mean curvature of general n-tori when n ≥ 3. Ni, Wang & Xie, arXiv 2609.36491 (29 Sep). "ChatGPT 5.6 Sol was used to generate preliminary versions of certain arguments in Lemma 3.2, Remark 4.6, and Proposition 5.5"; the authors substantially revised them.
- Flat normal bundles are not preserved by mean curvature flow (an embedded torus and an entire graph in R⁴). Kunikawa, arXiv 2609.34291 (28 Sep). ChatGPT 5.6 "was used to suggest suitable ansatzes and to explore examples".
- Smaller AI roles: union-closed sets conjecture bound raised past Liu's conjectured 0.382709 to 0.38288 (Costa & Sadhu, 2610.02295; ChatGPT for "exploratory symbolic calculations, proof checking" and code); Fang–Lin spectral Turán conjecture counterexample (Wu & Lu, 2609.37208; ChatGPT "to discuss proof strategies, check intermediate arguments"); Boston's unramified conjecture for odd GL₂ over Q and McLeman's (3,3)-conjecture for p > 3 (Luo, 2609.37252; ChatGPT 6 for proofreading and "verifying the logical consistency"); Cordero-Erausquin–Eskenazis question on Talagrand type of S₁ (Luo, Wang & Zhou, 2610.03407; Codex for literature and editing); FPT-size resolution refutations (Cai & Li, 2610.02047; one construction from a ChatGPT-found gap).
Outside arXiv (claims with weaker provenance)
- Conway's refinement conjecture for omnific integers (1976): Dan Abramov's "I vibed a proof of Conway's conjecture" (18 Sep; Claude, ChatGPT/Sol and Codex agents; Lean proof passing Palomar checks; not reviewed by mathematicians). overreacted.io
- Theo-Conjecture (FirstPrinciples) on a 35-year-old question about common-divisor graphs (arXiv 2608.04040; "not yet completed peer review").
- Star Fleet Math: 20 parallel Codex/GPT "starships" claiming 13 Lean-verified Erdős-problem solutions (July 2026). starfleetmath.com
Community response: AIM problem list (Oct 2, 2026)
In a guest post on Terence Tao's blog, Matthew Colbrook announced AIM, organised with Steve Brunton, Maarten V. de Hoop, George Stepaniants, Alex Townsend and Rachel Ward. It is a community-curated GitHub list of precise open problems in applied mathematics, spectral theory and related fields, each with a dated literature-status review. On Oct 2 the repo listed 642 open targets (531 open, 111 partial), 23 solved entries, 0 Lean-verified. Contributions with or without AI are welcome. The organisers say the list is "rather than using a set of problems to benchmark AI systems" and is meant to give mathematicians, especially early-career researchers, credit for understanding and explaining proofs. A companion site, AIM, explained, runs a competition for explanatory papers and videos of up to 10 minutes (entries close 2027-01-01). Tao's blog · GitHub
Why it matters
By September 2026, "settled with AI help" had stopped being news in itself: several named conjectures a week, across most of mathematics, with disclosures ranging from "fully generated by the model" to "used for proofreading". For a model with an older cutoff, the practical lesson is to check the current status of any conjecture before calling it open. Many that were open in early 2026 have claimed resolutions, but most are still unrefereed.
Changelog
- 2026-09-30: created from a systematic arXiv/tracker backfill (AI disclosures read in each PDF)
- 2026-09-30: sweep 2026-09-30: added Bárány–Larman (Soberón), Nagamochi scoring lemma (Karakuş) and quantum percolation (Becker–Oltman) disclosures
- 2026-10-01: sweep 2026-10-01: added Hirose duality (Seki), Stuck–Zimmer cases (Machado–Yifrach), Haah cubic code (LLM-proved Lemma 4.1), semi-Clifford level 4 (Codex), edit-distance codes (Yeung) and Google's Cogentic
- 2026-10-01: leads run: added Stanley rankwise conjecture (TARS agent, arXiv 2607.22988) and Bernhart–Kainen smaller counterexample (arXiv 2608.08118)
- 2026-10-01: leads run (07:40 completion): added Chvátal (own entry), Mubayi b(t,t) (own entry) and the 1/3–2/3 balance-constant paper
- 2026-10-02: sweep 2026-10-02: added Lehmer (own entry), Suzuki AF-algebras (GPT-6 Astra), Erdős #374 (Yudin), explicit Connes counterexample, cyclic steepest descent, Borg partitions and four minor Oct 1 disclosures
- 2026-10-02: added the AIM open-problem list (Colbrook et al., guest post on Tao's blog, Oct 2)
- 2026-10-05: added 18 Sept 30–Oct 2 disclosures (Erdős #108, Dvoretzky, Barvinok and Tomon have own entries), found by a manual arXiv scan after the sweep's date window missed them
- 2026-10-05: 12:30 quick run: added late-September items the sweep's 6-day arXiv window surfaced (List Total Colouring, Fishburn, Burr–Erdős–Graham–Sós for C₇, Aigner, Erdős–Ulam, Pappas, Pólya–Szegő); PDF disclosures read
- 2026-10-05: 21:30 run: PDF scan of the sweep's 43 'no AI disclosure' conjecture papers found 13 with AI statements; added Lovász (own entry), Peppy/Nesterov FGM, Watkins, rational points near curves, Chen's conjecture for n-tori, MCF flat normal bundles and five minor ones
People
Related posts (2)
- What does the advent of powerful AI models mean for mathematicians like me? original ↗ Jennifer Taback (guest post on Terence Tao's blog) · blog · 2026-10-03
A liberal-arts-college mathematician's view, on Tao's widely read blog, of what frontier AI changes for researchers who are not working on headline conjectures, and for teaching and tenure. - How AI does, and does not, change the way I do math original ↗ Rachel Webb (guest post on Terence Tao's blog) · blog · 2026-09-29
A mathematician's view, published on Terence Tao's widely read blog, of what LLMs change in research practice during the 2026 wave of AI-assisted proofs.
Related events
- Lovász conjecture (1969): long paths in vertex-transitive graphs pushed to n^(1−o(1)), with GPT-5.6 Sol supplying key steps ★★★
- List Total Colouring Conjecture (late 1990s) disproved; counterexample found by ChatGPT 6 Astra Ultra 'with little input' ★★★★
- Lehmer's 1965 permutation conjecture (a research problem in Knuth's TAOCP) proved; Claude Opus 5.5 found the short hypercube proof ★★★★
- Erdős–Hajnal high-girth problem (Erdős #108) disproved with ChatGPT/Codex help and a Lean proof, via the Conjectures.io bounty; experts sharpen it with GPT-6 Astra ★★★★
- Klartag and Moshe prove the ε-Dvoretzky conjecture (polynomial dependence on ε); the key probabilistic idea came from a ChatGPT discussion ★★★★
- ChatGPT Astra finds geometric triangle-free graphs with near-optimal chromatic number, the first improvement on Burling's 1965 box bound ★★★
- GPT-6 Astra proves Barvinok's log-concavity question for contingency tables on lines, giving lattice-point bounds for all totally unimodular polytopes ★★★
- 'The Gold Rush in AI4Math': substantive AI use in arXiv math papers rises from 1.4% to 14% in five months ★★★
- Planar Schiffer and Pompeiu conjectures disproved by two independent groups; one proof has a Lean certificate written by GPT-5.6 ★★★★
- Convex counterexamples to Schiffer and Pompeiu in dimensions 3, 4, 6, 8, 10 and 14, made with Claude Opus 5.5 and GPT-6 Astra/Sol ★★★
- Pre-release GPT-6 Astra disproves Erdős's 'first serious problem' (1931, $500) and proves the rational-exponents conjecture, all Lean-verified, in Epoch's FrontierMath Erdős runs ★★★★★
- GPT-6 Astra proves the Erdős–Sós conjecture (1962) with a short counting argument; mathematicians race to simplify and extend it ★★★★★
- Mathematician posts an unchecked ChatGPT Astra proof of the planar Mumford–Shah conjecture (1989), citing OpenAI's '100 open problems' claim ★★★★
- Peking University preprint claims an AI-found disproof of the Yau–Tian–Donaldson conjecture for constant scalar curvature metrics ★★★★
- Pierce–Birkhoff conjecture (1956) disproved by a multi-agent GPT + Claude harness; counterexamples Lean-verified ★★★★
- Banach's isometric conjecture (1932) completed in the real case with key steps from ChatGPT 5.5/5.6 Pro; complex and quaternionic cases follow five days later ★★★★
- Matrix Spencer conjecture proved; authors credit GPT-5.6 Sol Pro with 'the heavy-lifting' for the key lemma ★★★★
- The k-server conjecture, the 'holy grail' of online algorithms, is proved at Oxford; ChatGPT 6 Astra generalised the authors' k = 3 proof to all k ★★★★
- Odlyzko–Poonen conjecture (1993) proved unconditionally: 'The proofs are due to GPT-6 Astra', which also wrote a 22,000-line Lean formalisation ★★★★
- Alexander Perry disproves the period-index conjecture (Colliot-Thélène, 2001); a flawed ChatGPT example was the starting point ★★★★
- Medvedev's logic of finite problems (1962) shown undecidable; key idea from ChatGPT Sol 5.6, central argument checked in Lean by Claude Opus 5 ★★★
- Maxwell's conjecture on equilibria of point charges is false: five charges with at least 24 critical points, construction idea from GPT-5.6 Sol ★★★
- The I3322 Bell inequality needs infinite dimensions: Pál–Vértesi conjecture (2010) proved from an approximate proof by GPT-5.5 Pro ★★★
- Lean-verified 'Liouville Goldbach' theorem goes viral as a GPT-6 Astra 'Goldbach breakthrough'; the AI role is unconfirmed and it is not the Goldbach conjecture ★★
- OpenAI says an internal model resolved 100+ long-standing open problems in 24 days of training; no list released ★★★
- Berkeley's Phillip Kerger uses GPT-5.6 Sol and a 10-page prompt to close a 30-year gap in derivative-free convex optimization ★★★
- Rota's 1970 unimodality conjecture for matroid flats disproved, with ChatGPT 5.6 Pro suggesting a key construction ★★★
- Neurosurgery resident uses GPT-5.6 Sol to prove Crouzeix's conjecture in a 16-hour autonomous run ★★★★
- OpenAI's unreleased 'Astra' model claims ten advances in maths and theoretical CS, with Lean proofs ★★★★★
- Claude proves more than two-thirds of Riemann zeta zeros are simple and on the critical line (up from 41.6%) ★★★★★
- Claude-assisted construction claims a complex structure on the 6-sphere, answering Hopf's 1947 problem (pending verification) ★★★★★
- Friedgut's 2004 influential-coalitions conjecture resolved (Chattopadhyay–Gurumukhani, ChatGPT Pro used extensively); Friedgut then has GPT-6 Astra simplify the proof ★★★
- Chvátal's 1972 conjecture proved (Chang–Liu–Liu, ChatGPT-assisted), then a GPT-6 Astra 'proof from The Book' and a Codex-built Lean formalization ★★★★
- ζ(5) proved irrational: Aabir Fauzan's Zenodo preprint, the first such result since Apéry's ζ(3) in 1978, is formally verified in Lean within a week, one formalization written by Claude ★★★★★
- Lean Pool: an AI-maintained archive of Lean formalizations grows past 3 million lines ★★
- Claude (Fable 5.1 in Claude Science) computes the nine-loop six-gluon amplitude in planar N=4 super-Yang-Mills, answering a physicist's public challenge ★★★★
- Conjectures.io bounty platform pays out for Lean proofs of Erdős #859, #18(b), #1062(ii) and Ben Green's problems 24, 39, 40 within ten days, mostly to Purdue's 'JenW1N' ★★★
- New record bound on the irrationality measure of ζ(2), μ ≤ 5.0495243, released as a 92k-line Lean proof written by Claude; beaten by a human paper a week later ★★★
- GPT-6 Astra finds a proof improving the Kővári–Sós–Turán bound: diagonal bipartite Ramsey numbers b(t,t) = O(2^t) (Mubayi) ★★★
- Mathematicians' AGMAI publishes norms for AI labs releasing AI-generated results; Simons Institute TCS group issues 12 actions ★★★
- EPFL's LinCodeEvolve (Viazovska, Abbe) finds seven record-breaking binary linear codes with LLM-guided program search ★★★
- Meta publishes six maths papers made by mathematicians working with Muse Spark in ordinary meta.ai chat, saying five answer open questions ★★★
- Quanta's math editor asks 'Is AI the End of Math As We Know It?'; young mathematicians form an Association for Human Mathematics ★★★
Sources (9)
- paperarXiv 2608.24961: The Gold Rush in AI4Math: Where Are We Now?
- paperarXiv 2609.40324: Cogentic, multi-agent orchestration for automated proof discovery (Google, Gemini)
- paperarXiv 2609.39959: Haah's 3D cubic code thermalizes rapidly
- discussionVibeMathed: tracker of AI-involved math problems
- discussionWikipedia: List of mathematical discoveries by artificial intelligence
- discussionKingy AI: Mathematics & Science Breakthrough Tracker
- discussionGitHub: ai4math-chronicle (provenance tracker)
- discussionTerence Tao's blog (guest post by Matthew Colbrook): AIM, an invitation to explore mathematics together (Oct 2, 2026)
- codeGitHub: MColbrook/AIM (open applied problems, statuses and reviewed solutions)
id: 2026-09-30-ai-assisted-conjecture-wave-summer-2026 · updated 2026-10-05 · open in the interactive timeline