AI Finds Counterexample That Overturns a Paper — Author Confirms; 453 Manuscripts Later, AI Is Now Posing Its Own Questions

AI for MathMulti-Agent SystemsMathematical Research AutomationAI Conjecture GenerationResearch LoopHuman-AI Collaboration
1 hour agoSource: blockweeks.com
AI Finds Counterexample That Overturns a Paper — Author Confirms; 453 Manuscripts Later, AI Is Now Posing Its Own Questions

AI mathematics is crossing a very subtle dividing line.

In the past, we asked: Can a large model solve a difficult problem? Can it write a rigorous proof? Can it find a counterexample that humans have not discovered for decades?

Now, a question more like "research itself" has emerged: After a proof route fails, does the AI know what to do next? Should it continue calculating, switch routes, narrow the proposition, or simply propose a new, falsifiable mathematical conjecture?

AI Has Taste, built by researcher Zeng Zijian of UCSI University, is trying to turn this question into a sustainable research system.

The project's public repository currently records 453 mathematical research manuscripts and 2,312 pages of content, of which 6 are explicitly classified as "AI-Proposed Conjectures."

The positioning given on the repository homepage is even more direct: From Answer Generation to Research-Agenda Generation—from generating answers to generating research agendas.

数学研究

AI overturned a conjecture in a paper, and the original author confirmed it by reply

AI Has Taste did not start from the belief that "AI must be good at mathematics." Zeng Zijian recalled that at the beginning of the project, he was not confident whether large models could truly advance mathematical research. An accidental test became the turning point: he directly handed a conjecture from a paper to the AI, originally just wanting to see how far the model could go. As a result, the AI did not continue to "prove" along the paper, but constructed a counterexample that directly overturned the conjecture.

His first reaction was not excitement, but disbelief. So he organized the counterexample and derivation and wrote to the original paper's author to verify. The reply he subsequently received confirmed: this counterexample holds. The corresponding research was later included among the current more than 450 mathematical manuscripts.

"At first I did not believe in AI's power in mathematics. Until I input a conjecture from a paper, and the AI directly gave a counterexample to overturn it; I immediately wrote to the original author, and the other party replied confirming it was correct. After that, I truly let go and went all in." — Zeng Zijian

数学研究

数学研究

What this email changed was not a single conjecture, but the scale of the project. If AI can not only restate known knowledge and generate text that looks like a proof, but also actively attack new conjectures in papers and find a counterexample confirmed by the original author, then it has a chance to truly enter the "research loop." After that, Zeng Zijian began to gradually agentify topic selection, proof, counterexample search, failure management, independent review, and writing, and let different machines run in parallel over the long term.

The real change is not "how many papers were written"

Looking at quantity alone, 453 manuscripts are already eye-catching enough. But if this project is understood only as "AI writing mathematical papers in batches," then one misses the most noteworthy part.

The project itself repeatedly emphasizes: 453 is the number of "research manuscripts," which does not equal all 453 original open problems being solved. The outputs include complete proofs, counterexamples, partial results, finite computation certificates, structured reductions, and papers proposing new conjectures. AI internal review also does not equal external peer review.

The real change happens in the research chain. Traditional AI benchmarks often start from "the problem has already been given": humans choose the topic, AI solves it, and finally there is success or failure. AI Has Taste extends the process both forward and backward—AI can screen candidate problems, compare routes, and actively look for counterexamples; after a route fails, it can also analyze the structure of the failure, modify the proposition, and hand new mathematical objects to another independent Agent to continue attacking.

数学研究

The key to AI Has Taste is not "more prompts," but turning failure into input for the next round of research.

Topic selection, proof, error finding, and writing Agents

This system is not "one model asking and answering itself in a single dialog box." The public README splits the Agent roles into four layers:

Supervisor / Screener is responsible for selecting candidates, comparing nearby results, and finding the cheapest decisive experiment;

Researcher is responsible for proofs, counterexamples, and exact computations;

Fresh-context Reviewer specifically attacks frozen propositions, key lemmas, and certificates in a new context;

Writer / Release Checker is responsible for clearly writing the final accepted scope and checking citations, builds, and release materials.

Zeng Zijian further deployed the workflow as multi-machine collaboration: different machines can separately advance proofs, search for counterexamples, run exact computations, and conduct independent review. What is shared is frozen propositions, experiment records, failure reasons, code, and evidence, rather than letting one Agent generate results while stamping its own approval.

数学研究

The core design of multi-agent: role separation + shared evidence + fresh-context review.

There is a sentence in the repository that sums up this design very well: Criticism is part of the engine, not an afterword. Criticism is not a procedure added after the paper is written, but a part of the research engine.

Mathematicians, roboticists, and automation scholars enter the verification chain

As research manuscripts grew from dozens to hundreds, another question became sharper: who judges that these agent outputs are not systematic hallucinations?

AI Has Taste internally uses a fresh-context Reviewer to separate "research" and "error-finding," but the project does not treat AI self-review as the endpoint.

As the work progressed, Zengzaijian also began inviting real scholars from different fields to participate in the verification, review, and academic discussion of some results.

According to the project team, collaborators and reviewers involved in part of the work include: Ong Seng Huat, Fellow of the Academy of Sciences Malaysia and Honorary Professor at the Institute of Mathematical Sciences, University of Malaya; Kurunathan Ratnavelu, Fellow of the Academy of Sciences Malaysia and Honorary Professor at the Institute of Mathematical Sciences, University of Malaya; and Professor Xiong Yonghua from the School of Artificial Intelligence and Automation, China University of Geosciences (Wuhan).

Here it is necessary to distinguish "participating in verification" from "endorsing all results": the above scholars did not sign off on all 453 manuscripts one by one, but rather verified, reviewed, or discussed some of the problems, proofs, computational results, or research directions.

For a highly automated scientific research system, this combination of "internal machine red team + human expert spot checks/review" is, on the contrary, more critical than handing all review to the same set of agents.

AI is responsible for pushing research speed to the limit; human experts are responsible for pulling "looks valid" back to "worthy of belief" at key nodes.

After failure, a better question was chosen

The case in the project that best embodies "research taste" is precisely not a beautiful success, but a failure.

On the fractional Gaussian trace problem, the system initially tried monotonicity and one-sided pairing routes, but exact counterexamples kept appearing. An ordinary benchmark would usually leave only one label here: FAILED.

But this workflow did not stop. The agent continued to ask: what exactly did the counterexample break? Does the failure have a stable structure? In the end, the research shifted attention to a fixed negative defect, constructed a ten-term correction structure, and proposed a new unified boundedness conjecture. More critically, the new conjecture was then attacked in turn by another round of exact computation and independent review. The public archive scan reached index 1004; this unified bound has still not been proved.

The significance of this matter is not that "AI proved another big theorem"—in fact, it did not. The significance is that: the original problem was not solved, but the failed route was compressed into a new proposition that is clearer, more falsifiable, and, once it holds, can reconnect to the original problem. Research did not terminate at the failure, but grew the next problem out of the failure.

In the past: humans set the problems, AI answered them. Now: AI answers problems, and also begins to participate in deciding "what the next problem is."

6 new AI conjectures

As of the current public snapshot, the repository classifies 6 papers separately under "AI-Proposed Conjectures." These works span combinatorics, graph theory, number theory, and analytic-type problems, including the nonsingularity of infinite binary-kernel matrices for three subsequences of A182510, CDS-colorability of Young diagrams, decomposition of triangle-free graphs with maximum degree at most 6, unary-theta congruence formulas for 18-colored generalized Frobenius partitions, nonnegativity of Newton coefficients for reciprocal-subsum denominators over dyadic stages, and the aforementioned fixed-defect conjecture for fractional Gaussian trace.

What is truly worth attention is not whether AI can "brainstorm a conjecture." Casually guessing a sentence is not hard. What is hard is writing the conjecture into a precise definition, explaining where it comes from, what evidence supports it, what counterexamples could overturn it, what results it would lead to if it holds, and leaving the computation and source code for subsequent review.

This is also why the project regards "proposing conjectures" as a higher-level research action than "generating answers": proofs answer existing questions, while conjectures change where subsequent research resources are invested.

Behind 453 manuscripts is the prototype of "agent-scale scientific research"

AI Has Taste has currently publicly released 453 research manuscripts, totaling 2312 pages; among them, the BigWIN2 workflow alone corresponds to 223 manuscripts, and provides paired LaTeX/reproduction material packages for these 223 works.

The most counterintuitive thing about this scale is not that "AI types fast." What is truly expensive in mathematical research is context switching: this problem requires graph theory, the next requires number theory, and the next may require SAT, exhaustive search, exact rational arithmetic, or matrix computation. A human researcher cannot simultaneously maintain hundreds of completely different lines of thought, but an agent system can split them into independent tasks and preserve failed routes, local theorems, and reproducible experiments.

Thus, the role of the individual researcher is changing: not personally deriving every step, but more like a PI—setting research boundaries, choosing the problem pool, deciding which results are worth continuing to invest in, when to stop, and which conclusions are qualified to be made public.

Why is it called "AI Has Taste"?

"AI has taste" sounds very anthropomorphic, but the project README deliberately sets a boundary: taste here is not consciousness or subjective experience, but "mathematical judgment expressed through research decisions."

For example: if both problems can be done, which one first? A route already has local progress, but marginal returns are getting lower and lower—should it stop? Is a counterexample a death sentence for the proposition, or does it show that the real theorem should be formulated differently? Is a small result enough to stand alone as a paper? These decisions were usually implicit in researchers' experience in the past and were hard to measure by standard benchmarks.

AI Has Taste tries to leave an inspectable trace of these decisions: which problem was chosen, why the route was changed, what structure the failure left behind, where the next proposition came from, and how the new conclusion was attacked again by an independent context.

How far is it from an "autonomous mathematician"?

The place where this project is most easily exaggerated also needs to be made clear. 453 manuscripts are not 453 journal papers that have undergone formal peer review; internal agent review is not equivalent to independent review by the mathematical community; and the scope covered by finite computation cannot automatically be generalized to the infinite case. The repository itself explicitly states these limitations.

But if one ignores its way of organizing research merely because these results have not all completed external verification, one would also miss another change: the boundary of AI capability is moving from "executing a given task" to "participating in designing the next task."

The truly scarce resources in mathematical research have never been only computing power and reasoning length, but also: which problems are worth spending time on, which failures are worth preserving, and when the topic should be changed.

When AI begins to enter these links, the competition in "AI for Math" is no longer just "who is better at proving," but will gradually become: who has a better research loop, and who can distill from a large number of failures the directions more worth continuing to pursue.

One More Thing

The past narrative of AI scientific research was very much like a super student: the teacher sets the problems, and it is responsible for solving them.

What AI Has Taste wants to do is more like a research group: humans provide boundaries and value judgments, and agents find problems, test problems, make mistakes, overturn themselves, leave local results, and then propose the next problem.

If this route continues to hold, the most important human-machine division of labor in future basic research may change from "humans are responsible for thinking, AI is responsible for computing" to a more complex structure: humans are responsible for goals, values, and ultimate responsibility; AI undertakes large-scale search, verification, failure management, and generation of candidate research directions; formal tools and public evidence are responsible for making results reviewable.

After AI can solve problems, the next hurdle may really be: can AI choose problems?

This article comes from the WeChat public account "Xin Zhi Yuan," author: Xin Zhi Yuan