AI in academia this week: Astra, fake references, strikes, and AI‑checked proofs
AI is reshaping academia at every level right now—from how we do research, to how we teach, to how we prove the hardest problems in mathematics. This week brought big model news, a worrying case of fake references, staff strikes over AI, and some eye-opening developments in computer-checked proofs.
Astra (GPT‑6) takes the lead on science benchmarks
OpenAI’s new Astra model, often referred to as GPT‑6 Astra, has jumped straight to the top of a key scientific benchmark called TerminalBench Science. This benchmark compares how well different AI models handle complex scientific and research-focused tasks.
Until recently, one of the strongest models on this benchmark was a system known as Fable 5.1. Astra has now overtaken it, claiming the number one spot. For researchers and academics, this matters because it suggests Astra is especially strong at tasks like understanding scientific text, reasoning through research problems, and working with technical content.
In practical terms, this kind of performance could make Astra a powerful assistant for literature reviews, brainstorming experiments, or explaining difficult concepts. It also raises bigger questions about how close these systems might be getting to more general scientific reasoning, a topic we explored in more detail in our breakdown of OpenAI’s Astra model and AGI.
Journal editors caught using hallucinated references
On the other side of the spectrum, we saw a worrying example of how not to use AI in academia. Editors of a higher education journal published an editorial that included fake or incorrect references—many of which appear to have been generated by an AI tool and never checked.
Out of 31 references, at least nine were either completely non-existent or had wrong journal titles, incorrect author names, or other major errors. On the journal’s own website, some of the reference links even return a message like “does not exist” when clicked.
This is a textbook case of AI hallucination: large language models are “plausibility machines.” If you ask them for references without giving them real citations or connecting them to a verified database, they will often invent sources that look real but aren’t.
The troubling part is not that AI hallucinated—this is a known limitation—but that journal editors, who are supposed to set the standard for research integrity, failed to catch it. Simple automated checks (like verifying that each reference actually exists) could have prevented this. The incident is now under investigation, and it raises a hard question: if even editors misuse AI, how can the rest of academia be expected to do better?
University staff strikes over AI and job security
While AI is boosting research capabilities, it’s also creating real anxiety among university staff. In Sydney, thousands of university employees recently walked off the job in a strike focused on AI and job security.
Universities are under pressure to cut costs, and many administrators see AI as a way to automate or replace parts of teaching. Some institutions have already experimented with AI chatbots to support or even deliver content in online-only courses. At the same time, new buildings are being designed with fewer or no lecture halls, shifting more learning online.
For many academics, this feels like a step too far. In-person lectures and human-led teaching are still deeply valued—not just for content delivery, but for mentorship, discussion, and community. Early evidence suggests that when AI is used as a primary teacher rather than a support tool, student outcomes often suffer.
The strike in Sydney is essentially a demand for a human-centered approach to AI in education. Staff want universities to use AI to support teaching, not to quietly replace teachers or erode working conditions.
MIT’s balanced approach to AI in education
As universities scramble to respond to AI, MIT has released a set of guiding principles and recommendations for AI in education that offer a more measured path forward.
MIT’s approach starts from a simple idea: AI isn’t going away, and students need AI literacy across disciplines—not just in computer science. Instead of trying to “AI-proof” classes by banning tools outright, MIT encourages instructors to rethink what they’re really assessing and how AI changes the learning process.
From banning AI to designing AI-aware courses
One of MIT’s key recommendations is to move beyond knee-jerk bans. Simply reverting to in-person, no-tech exams may be part of the solution in some cases, but it shouldn’t be the whole strategy.
Instead, MIT urges instructors to:
• Consider how AI changes social aspects of learning—such as office hours, study groups, and interactions with teaching assistants.
• Revisit course goals before redesigning assessments. If the only goal is to produce a polished essay or report, AI can already do that fairly well. But if the goal is to develop understanding, critical thinking, synthesis, and original insight, then assessments need to directly test those skills rather than just the final document.
This shift—from grading products to assessing thinking and process—is likely to become a central theme in AI-era education. For a broader look at how tools fit into this new landscape, you can also check out our guide to the best AI tools for academia in 2026.
AI and the new era of math proofs
Beyond teaching and policy, AI is also starting to play a striking role in pure mathematics, especially in tackling long-standing open problems and formalizing complex proofs.
OpenAI and the Navier–Stokes problem
OpenAI has claimed progress on one of the famous Millennium Prize Problems: the Navier–Stokes existence and smoothness problem. This problem is central to understanding fluid dynamics and has resisted proof for decades.
According to a Nature article, OpenAI researchers have produced a proof that shows there exist fluid flows that start out normal but later behave in a way that challenges the standard assumptions of the equations. If this work stands up to scrutiny, it would be a major milestone—both for mathematics and for AI-assisted research.
Claude and the first fully computer-checked proof
Anthropic’s Claude model has also made waves by helping to formalize a significant theorem into a fully computer-checked proof. What makes this notable is not just the theorem itself, but the fact that the entire proof has been encoded in a way that a computer can rigorously verify.
There was even a human mathematician who had been funded—reportedly with a £1 million, multi-year grant—to work on formalizing this same area. Claude’s work effectively “beat him to it” in about 11 days.
Instead of reacting defensively, the mathematician wrote a thoughtful blog post titled along the lines of “Anthropic has beaten me to it.” He pointed out that, mathematically, the AI-assisted formalization doesn’t introduce new ideas; it largely follows the existing literature. But he’s excited about what this means for the field:
• Auto-formalization could make the review process for math papers far less painful.
• It can help keep the community honest by clarifying exactly which results are being assumed and which are being proved.
In other words, AI might not replace mathematicians’ creativity, but it could transform how rigor, verification, and peer review are done.
What this all means for the future of academia
This week’s stories capture the tension running through academia right now. On one hand, models like Astra and tools like Claude are unlocking new capabilities in research and proof verification. On the other, careless use of AI is already damaging trust in scholarly publishing, and staff are striking to protect the human core of teaching.
The path forward seems to require three things: stronger guardrails and verification (especially around references and proofs), thoughtful policies like MIT’s that treat AI as a tool rather than a cheat code or a replacement, and a renewed focus on what only humans can do—mentoring, critical thinking, original insight, and building academic communities.
AI won’t stop moving quickly. The question is how fast universities, journals, and researchers can adapt their practices to match.
Comments
No comments yet. Be the first to share your thoughts!