Assessing students has always been a key aspect of education, but the arrival of large language models has caused concern amongst educators that existing approaches will not work. Tools such as ChatGPT can perform well on many standardised tests and essay-based assessment tasks, although results vary by model, subject and assessment design. This raises the question of how we should adapt our approach to assessments.
The first thing to recognize is that ChatGPT is not a replacement for traditional assessments but rather a tool that can raise the baseline standard of attainment. Like other study tools, it can help students work with information more efficiently, but its answers can be wrong. It does not replace the need for critical thinking, analysis, verification and creativity. As with the arrival of prior new tools, we need to change how we assess. This is a challenge the education sector has faced many times: some tasks that once warranted substantial credit, such as particular kinds of recall or routine calculation, may deserve less weight when reliable tools are available. The key difference this time is the pace of change.
There are several good reasons for making students sit a test, but key amongst them are:
Testing can be a helpful tool for directing learning. By checking what a student knows and whether they can apply the information we can focus additional attention on areas of weakness.
If students fake the answers on tests like these then their use as a knowledge check is destroyed. In contrast, if it is clear that these tests are solely to direct learning, with no value placed on the results, then students' incentives should be aligned to the teachers. Both can view these checkins as stepping stones towards a final goal, be that a higher final grade or a better standard of learning - or hopefully both.
It is crucial that students do not feel judged for submitting the wrong answers. We need to be explicit that we are testing for the purpose of directing learning in order to remove the incentive for students to use tools like ChatGPT to obscure their true level of knowledge. Where teachers are not able to build that expectation then tests and environment can be designed to minimise the risk. This could be done through supervised work, oral follow-up questions or asking students to explain how they reached an answer. Applying knowledge or formulating problems should not, by themselves, be treated as AI-proof tasks. Other changes, like making these tests no notice would require students to retain necessary knowledge and practise retrieving it, rather than cramming for tests, or searching information online when needed.
Sometimes, in contrast, we test in order to rank students’ abilities and to sort them into streams, between universities or within the job market. Sometimes we award prizes. It is this type of assessment that, on the surface, feels most threatened by Generative AI. However, we will see that they need not be.
Again we must be explicit about why we are testing. In this case, we assess how good a student is at a task, say writing a critical essay, in order to establish how good they will be at doing similar challenges in the future. It follows that we should make the test as similar to the future tasks as possible, while recognising that foundational knowledge may also need to be tested separately.
Some people may view exams more as measures of general aptitude, or intelligence. If this is really what we want to take from exam results then surely it would be better to test for this directly, rather than via a proxy? If ChatGPT is now able to outperform students on essay based assessments they probably weren’t much use as measures of general intelligence or aptitude in the first place.
Where the purpose is to assess performance in a setting that permits AI tools, denying access to those tools can give an unrepresentative picture. Separate assessment of unaided knowledge may still be necessary. It would be like assessing engineers, but forcing them to do calculations with pen and paper. Not predictive of future results and therefore not useful. Instead, we should simulate real world situations and assess students' ability to perform, whilst using all available tools, on skills such as forming and defending opinions, analysis of novel sources, being concise and identifying key points in an argument. We could even go further and assess their capacity to be interesting or show good taste and judgement. All these would be useful predictors of success in employment or at advanced levels of education, where crucially they will have access to current tools like ChatGPT and those which emerge in the coming years.
Assessing these skills would require assessors to go beyond a mark scheme and use their judgement to evaluate students' work. In the past, producing a 2000-word essay with clear structure and a coherent argument was considered creditworthy. However, with the availability of ChatGPT and other tools, producing such a draft is now much easier for students with suitable tools and support. As such, we need to raise our baseline of performance and credit only work above this level.
Being much more demanding of students and examiners may feel daunting. However, viewed differently these tools could raise the quality of written work if they are used to support learning and held to demanding standards. Adapting will take time, but will eventually allow students - and the rest of us - to focus on higher order skills, like synthesis and adapting an argument for a specific audience.
Adapting to change caused by technology is a crucial skill for the next generation. As educators and assessors, we need to model this adaptability for current students. We cannot resist the change brought about by technology but rather need to find ways to incorporate it into our approach to education and assessment, and quickly. Failure to do so would be a disservice to our students, as they need to develop the skills necessary to navigate a rapidly changing world.
The impact of ChatGPT on assessment should not be viewed as a threat but rather as an opportunity to improve our approach to education. We should adjust our expectations of what constitutes creditworthy work and test skills that require critical thinking, creativity, and problem-solving. By doing so, we can ensure that our evaluations are not only fair and accurate but also prepare students for the challenges of the future.
