University of Wollongong research shows AI models achieved 76.3% in criminal law exams, dramatically outperforming 82.5% of students—a massive jump from 52.5% in 2023. Seven of 18 AI papers ranked at or above the 90th percentile. The findings expose urgent academic integrity concerns as generative AI models demonstrate capabilities that could enable widespread cheating in unsupervised assessments.

News article

AI Models Demonstrate Dramatic Performance Leap in University Exams

Generative AI models have achieved a remarkable breakthrough in academic performance, with new University of Wollongong research revealing that AI in exams now poses serious challenges to academic integrity. When researchers tested nine models from five AI providers across criminal law and torts subjects, the results were striking: AI papers averaged 76.3% in criminal law, outperforming 82.5% of students, while in torts they averaged 66%, surpassing 61% of students

1

2

.

This represents a dramatic shift from 2023, when AI papers averaged just 52.5% and performed around the 22nd percentile. Seven of the 18 AI papers now ranked at or above the 90th percentile of students, demonstrating capabilities that call into question the integrity of results from unsupervised assessments

1

.

AI Getting Better at Exams Through Enhanced Reasoning

The University of Wollongong research methodology involved testing generative AI models on actual final exam questions in two compulsory law subjects. The models received exam questions and prompts but no subject-specific lecture notes, textbooks or curated legal materials. However, they did have internet access, and some featured an "enhanced reasoning" capability that gives AI more time to plan, consider different approaches and check its reasoning before answering

2

.

Six of the 18 papers were graded by subject coordinators, while the remaining 12 were mixed among genuine student papers and blind-graded by tutors who had no knowledge of AI involvement. This rigorous approach ensured the AI performance in law subjects could be accurately compared against human students

1

.

Critical Analysis Breakthrough Changes the Game

The most significant development wasn't just the improved grades—it was the models' performance on critical analysis of hypothetical legal scenarios, which had been a clear weakness in 2023. In criminal law, AI performed substantially better with an average of 76% compared to 59.6% for students on these complex analytical tasks. In torts, AI models and students performed at roughly the same level, while AI models performed dramatically better on essay questions

2

.

This shift demonstrates that generative AI models have overcome their previous limitations in tasks requiring sophisticated legal reasoning, raising immediate concerns about their potential use for cheating in unsupervised assessments.

Jagged Capabilities and Persistent Limitations

Despite impressive overall performance, the research revealed that AI models exhibit jagged capabilities—excelling in one subject while performing weakly in another. Strong analysis could sit alongside poor citations, weak source selection or fabricated authorities. Researchers identified lower rates of hallucinations in some models, though these also cited fewer sources

1

.

Performance varied substantially between models and subjects, confirming that while AI has become capable enough to threaten academic integrity, it hasn't yet become a reliable legal expert. The study's scope—limited to two law subjects at one Australian university using 18 AI-generated papers—means results may differ in other disciplines or with different prompting strategies

2

.

University Response: Three-Pronged Assessment Strategy

Facing this new reality, researchers propose that universities must adopt complementary approaches balancing AI literacy with foundational knowledge. Today's job market requires graduates to work with AI responsibly and effectively, but students need enough foundational knowledge and skills to detect errors, sharpen weak points and improve AI output. If students cannot add value beyond what AI produces alone, their work may become replaceable

1

.

The recommended university response includes implementing AI-free supervised assessments such as in-person exams, oral assessments and supervised problem-solving. These ensure students develop independent knowledge and reasoning ability to recognize when AI is wrong. For first-year law students, researchers suggest these tasks should comprise the bulk of assessment

2

.

This approach acknowledges that no single assessment format can achieve both aims of teaching AI literacy while ensuring students master core competencies independently. Universities must act decisively to maintain academic standards while preparing graduates for an AI-integrated workplace.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved