Laura Kelly reflects on the broader implications of Gen AI-resilient assessment strategies with reference to current scholarship and her own large group teaching in the social sciences.
The ethical use of Generative AI (Gen AI) in higher education environments is a matter of considerable debate. The best approaches for preserving academic integrity, avoiding discrimination and retaining a commitment to sustainability, are being discussed across the higher education sector. In this blogpost, I argue that we must also consider workforce implications, including the ways in which rapid pedagogical change might impact relatively powerless academic actors.
The shift from detection to redesign
Research increasingly challenges the idea that student use of Gen AI in higher education assessment can be successfully monitored by identifying hallmarks, such as fabricated references to scholarly publications or ‘hallucinations’. There are growing calls to move from the ‘detection’ of AI use in existing assignments to the redesign of assessment strategies. There is limited evidence to support strong claims for Gen AI-resilient assessment. However, as Grove (2024) has recently summarised, assessment strategies that may be more ‘resilient’ to Gen AI-related plagiarism include:
- ‘authentic’ assessment that mirrors ‘real-world’ tasks
- ‘synoptic’ assessments that require students to synthesise learning from across their programme of study
- multi-step ‘patchwork’ assessments that reduce the assessment anxiety associated with single tasks (and thus the temptation to seek support from Gen AI tools) and offer opportunities to use feedback for improvement across a single module or unit.
Given these approaches may also bring broader pedagogical benefits, the promotion of such assessment types is increasingly embraced by the sector. While this moment offers possibilities, the unintended consequences of far-reaching changes to assessment strategies must also be explored.
Hidden costs: PGTAs and the implementation challenge
In the context of the ‘massification’ of higher education, many institutions in the UK have followed the North American example and looked to employ doctoral students as part-time educators, particularly for small group teaching on large undergraduate modules. However, the realities of teaching and assessing in sometimes adjacent areas of expertise can be difficult for new educators. Postgraduate Teaching Assistants (PGTAs) often have limited opportunity to influence syllabus or assessment design and report variable support from their Departments.
Studies that directly consider PGTA experiences of Gen AI-resilient assessment identify concerns already raised by explorations of authentic assessments. In short: scaling from small to large groups presents implementation challenges. Any increase in marking hours or assessment support demands produced by new assessment types risks being felt most keenly by staff who often have extensive student contact, least assessment experience, heavy marking loads and potentially increased pressure to secure positive student appraisals. Thus, while it has been argued that authentic assessment can support socially valuable and transformative education by prompting social connection, critical reflection and critical hope, there is potential to embed exploitation.
Recommendations for equitable implementation
I offer six suggestions to educators redesigning large group assessment strategies for the future:
1. Revisions to assessment strategies intended to build Gen AI-resilience should not be implemented without a workload impact assessment that considers the potentially different challenges of new markers (eg taking longer due to inexperience with marking criteria or with subject matter). While expected time spent should be clearly communicated, module leads should also seek feedback from PGTA markers about experiences of assessment and use available (eg online) tools to gauge the time spent on marking and feedback.
2. Building Gen AI use into some assessment tasks, for example through the generation and critical analysis of Gen AI-produced text, will also produce new training and resource needs. PGTAs will need support and training in the ethical use of Gen AI. They will also need stronger subject knowledge for in-class assessment of student efforts than less Gen AI-resilient asynchronous marking. This will increase the preparation time required by PGTAs and/or the preparatory work of subject experts such as module leads. Both have workload implications.
3. Mark standardisation or ‘benchmarking’ should be routine practice for any revised assessment. This will support consistency of marking and feedback, while allowing review of the marking guidance provided.
4. Gen AI offers an opportunity to consider teaching, learning and assessment in the round. Introducing multiple concurrent changes on single modules may introduce too much complexity and be less appropriate for those that involve PGTA markers or large marking teams.
5. Education leaders must support programme level review and be receptive to requests to amend broader teaching and learning strategies and/or workload allocation models if the proportion of time spent on assessment changes or the boundaries between education and assessment become more blurred. This is consistent with calls for communities of practice for assessment and the need to involve stakeholders in discussions of standards and the development of assessment criteria.
6. There is relatively little research that seeks to understand how PGTAs experiences of marking or Gen AI-resilient assessment strategies. We do know that PGTA involvement in pedagogical discussions or education development beyond centralised teaching courses is inconsistent. If the ‘generational incentive’ to rethink assessment is to proceed in a genuinely inclusive and equitable way, the PGTA and otherwise casualised academic workforce must be part of the conversation.
Dr Laura Kelly is Associate Professor in Criminal Justice at the University of Birmingham. She leads the Pedagogical Research Group in the School of Social Policy and Society. A longer version of this piece is available in the University of Birmingham Journal ‘Education in Practice’, 6(1): 15-26.