Skip to main content

The future of AI essay grading

22 Nov 2024 | A new generation of AI essay-grading tools is coming. These systems will be able to compete with human grading at university level for the first time. We need to think about what we will do when that happens.

In 1966, the educational psychologist Ellis Page set a new goal for computer science: automated essay scoring. 

So far, Page’s vision has been only partially realised, especially at university level. Some admissions tests, such as the Graduate Records Examination, combine AES with human grading. But these essays are short and highly constrained. Existing systems are not up to regular university-level essay assessments.

The 'deep learning revolution' in AI could change that. Since the early 2010s, systems known as 'deep neural networks' have made multiple breakthroughs in AI. Transformer architecture, a particular kind of neural network invented in 2017 has made possible language processing AIs, such as OpenAI’s GPT4, that have astonished experts with their capabilities. 

In the past, automated essay scorers graded texts on surface-level features like word count, sentence length, and basic grammar. The next generation of automated essay scorers will respond to the deep, semantic content of essays. 

New language processing AIs have already produced promising results in automated essay scoring. Before long, an organisation such as TurnItIn, with its massive dataset of graded essays and expertise in AI, will apply this new technology to the challenge of automated essay scoring at university level. 

The likely result will be a system that returns average grades approximately consistent with those given by human experts, at least for some kinds of university-level essay assessments, and that provides helpful feedback at the same time.

Is AI essay grading in universities a good idea? 

Some will feel uneasy about AI essay grading in higher education – perhaps even a sense that this would betray the very purpose of universities, reducing an essentially humanistic activity to the machine-facilitated pursuit of numerical targets. 

There are obvious arguments in favour of AI grading in universities, however.

Grading is time consuming and labour-intensive. If we could automate a substantial portion of it, our lives would be easier, and we would have more time for tasks such as teaching, research, and public engagement. 

Furthermore, AI essay grading could turn out better than human grading in some respects. Human grading can suffer from limited subject knowledge, time pressures, and individual bias. An automated system trained on a suitably selective dataset might end up approximating the performance of minimally biased and highly experienced subject experts better than the average human examiner does.

It is because of these advantages, in relieving human examiners and improving accuracy, that assessments that can already be graded by machines, such as multiple-choice tests, arithmetic problems, and coding tasks, usually are. Why should essays be any different?

A number of answers suggest themselves. 

Some involve technical problems familiar from other applications of AI, from self-driving vehicles to computer-assisted legal sentencing. 

The 'long-tail' problem is that machine-learning algorithms tend to outperform humans in commonplace scenarios but perform badly in those that are poorly covered by their training data. The 'black box' problem refers to the fact that the ways deep-learning systems arrive at their outputs are often opaque, even to their developers. The 'feedback loop' problem is the concern that as generative AI proliferates, it becomes harder to get non-AI-generated datasets on which to train new algorithms. 

Each of these applies to AI essay grading. 

A system that performs well on average might overlook an ingenious assignment that resembles nothing in the training data, but whose merit a human expert would recognise. Students could be unhappy with an automated grading system that cannot justify its outputs. And overuse of AI grading could mean that when we want a new system to keep up with developments in our disciplines, there will be no suitable dataset on which to train it. 

Developers are working to resolve these problems. But in the meantime, a significant amount of human grading will still be needed, and an appeals system could be necessary so that unsatisfied students can demand a human opinion. 

What about the worry that there is something intrinsically objectionable about AI grading, even if the outputs are indistinguishable from those of human examiners? 

Well, for one thing, teaching is facilitated by social relationships that AIs are not in a position to replicate. Attending jointly to a topic with a human being who is aware of you, and of that joint attention, is very different from interacting with an AI that only simulates intelligence. Students will find writing for human examiners incomparably more engaging, meaningful, and motivating than writing for AI.

And perhaps some tasks should just be kept sacrosanct from automation. Art lovers tend to be uneasy about the idea of AI-generated art. Religious communities have similar concerns about AI ministers. Essay grading might be an area where even flawless automation is undesirable. I don’t know if this is true of first-year undergraduate assessments, where many students are still acquiring the elementary skills of essay composition, but it may be true of much university-level essay grading.

Conclusion

Essays are only one form of university assessment. But for subjects like mine, they remain extremely important. And while there is widespread concern about generative AI being used to write essay assessments at university level, there has been far less discussion about the potential of new AI to grade these assessments.

Meanwhile, a new generation of AI essay grading tools is on its way. The new systems will be incomparably more capable than those that exist today. It is time to start discussing what we should do about these systems while we are still ahead of the technological curve.

Ralph Weir is a Senior Lecturer in Philosophy at the University of Lincoln and a Fellow of AdvanceHE. He works on philosophical issues to do with the mind, metaphysics, AI, religion, art, and culture. He is currently co-editing an issue of the journal Philosophical Education on AI Ethics. 

Want to learn more about enhancing your assessment practice? Join us on 4 December for a webinar on Using the Framework for Enhancing Assessment.

Do you have innovative practice using AI to share? Submit your paper for our AI Symposium by 25 November.

Let's talk about student success podcast

Join our monthly conversations with higher education's thought leaders. We're creating a space where insights flow as freely as coffee, and where experience meets innovation in supporting student success. 

Listen wherever you get your podcasts

You can listen to our podcast on the major podcast platforms. Search for "Let's Talk About Student Success."

Spotify
Zencastr
Amazon Music
Apple Podcasts

We feel it is important for voices to be heard to stimulate debate and share good practice. Blogs on our website are the views of the author and don’t necessarily represent those of Advance HE.

Keep up to date - Sign up to Advance HE communications

Our monthly newsletter contains the latest news from Advance HE, updates from around the sector, links to articles sharing knowledge and best practice and information on our services and upcoming events. Don't miss out, sign up to our newsletter now.

Sign up to our enewsletter