2 Sources
[1]
Our research shows how AI is getting better at exams. Here's how unis can respond
When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat. There were some big claims by AI companies at the time, including that tools showed "human-level performance on various
[2]
Research shows how AI is getting better at exams: How universities can respond
When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat. There were some big claims by AI companies at the time, including that tools showed "human-level performance on various
Share
Copy Link
University of Wollongong research shows AI models achieved 76.3% in criminal law exams, dramatically outperforming 82.5% of students—a massive jump from 52.5% in 2023. Seven of 18 AI papers ranked at or above the 90th percentile. The findings expose urgent academic integrity concerns as generative AI models demonstrate capabilities that could enable widespread cheating in unsupervised assessments.

Generative AI models have achieved a remarkable breakthrough in academic performance, with new University of Wollongong research revealing that AI in exams now poses serious challenges to academic integrity. When researchers tested nine models from five AI providers across criminal law and torts subjects, the results were striking: AI papers averaged 76.3% in criminal law, outperforming 82.5% of students, while in torts they averaged 66%, surpassing 61% of students
1
2
.This represents a dramatic shift from 2023, when AI papers averaged just 52.5% and performed around the 22nd percentile. Seven of the 18 AI papers now ranked at or above the 90th percentile of students, demonstrating capabilities that call into question the integrity of results from unsupervised assessments
1
.The University of Wollongong research methodology involved testing generative AI models on actual final exam questions in two compulsory law subjects. The models received exam questions and prompts but no subject-specific lecture notes, textbooks or curated legal materials. However, they did have internet access, and some featured an "enhanced reasoning" capability that gives AI more time to plan, consider different approaches and check its reasoning before answering
2
.Six of the 18 papers were graded by subject coordinators, while the remaining 12 were mixed among genuine student papers and blind-graded by tutors who had no knowledge of AI involvement. This rigorous approach ensured the AI performance in law subjects could be accurately compared against human students
1
.The most significant development wasn't just the improved grades—it was the models' performance on critical analysis of hypothetical legal scenarios, which had been a clear weakness in 2023. In criminal law, AI performed substantially better with an average of 76% compared to 59.6% for students on these complex analytical tasks. In torts, AI models and students performed at roughly the same level, while AI models performed dramatically better on essay questions
2
.This shift demonstrates that generative AI models have overcome their previous limitations in tasks requiring sophisticated legal reasoning, raising immediate concerns about their potential use for cheating in unsupervised assessments.
Related Stories
Despite impressive overall performance, the research revealed that AI models exhibit jagged capabilities—excelling in one subject while performing weakly in another. Strong analysis could sit alongside poor citations, weak source selection or fabricated authorities. Researchers identified lower rates of hallucinations in some models, though these also cited fewer sources
1
.Performance varied substantially between models and subjects, confirming that while AI has become capable enough to threaten academic integrity, it hasn't yet become a reliable legal expert. The study's scope—limited to two law subjects at one Australian university using 18 AI-generated papers—means results may differ in other disciplines or with different prompting strategies
2
.Facing this new reality, researchers propose that universities must adopt complementary approaches balancing AI literacy with foundational knowledge. Today's job market requires graduates to work with AI responsibly and effectively, but students need enough foundational knowledge and skills to detect errors, sharpen weak points and improve AI output. If students cannot add value beyond what AI produces alone, their work may become replaceable
1
.The recommended university response includes implementing AI-free supervised assessments such as in-person exams, oral assessments and supervised problem-solving. These ensure students develop independent knowledge and reasoning ability to recognize when AI is wrong. For first-year law students, researchers suggest these tasks should comprise the bulk of assessment
2
.This approach acknowledges that no single assessment format can achieve both aims of teaching AI literacy while ensuring students master core competencies independently. Universities must act decisively to maintain academic standards while preparing graduates for an AI-integrated workplace.
Summarized by
Navi
[1]
03 Dec 2025•Entertainment and Society

18 Jul 2024

20 Jul 2026•Entertainment and Society
