FutureLearn AI in Education Course Week 2

As part of our training in Artificial Intelligence, a number of staff in the Faculty of Creative Arts and Humanities at Liverpool Hope University were given the opportunity to undertake the King’s College, London, FutureLearn course on AI in Education. It’s a subject about which I have very mixed feelings, so here are my reflections on the second week of the course.

The first activity in Week 2 was to think about the pedagogical frameworks for using AI in teaching and learning. We read the LSE’s Manifesto for the Essay in the Age of AI, and thought about which of its points resonated most strongly with us. For me, it was the importance of the essay as a valuable tool for teaching and assessing high level cognitive skills such as research and argument; the need for essay driven assignments that students can personalise (something that I trialled on my Tudor course at Lancaster, where students had to take one of the given historiographical claims and test it on a case study of their own choice); the need for institutions to invest in resources and training (I would be much happier, for example, if we had an in-house/private model that we could train on our own materials as well as the large corpora of commercial GenAI); and finally the issue of equitable access to appropriate AI, which would be partly mitigated if the previous point were addressed.

This is because I use long form written assessments to test not just subject knowledge, but also the ability to find, summarise and apply appropriate theories and historiography; whether students can constuct an appropriate argument; and how well they have understood the issues at stake. They are also, of course, the main ways by which a professional historian is judged, and as it stands, this forms one barrier to changing assessments. There is also student resistance, because although many don’t really know how to write a well-crafted essay. I have opined at length to colleagues over the years about the 3 paragraph essay, which is a product of A-level teaching mainly because that is all you can produce under exam conditions in the time allowed for the higher-mark questions. Some students come to university thinking that every essay has an introduction, three paragraphs and a conclusion. When your word count is 2500 words, those paragraphs get very, very, very long!

We were also introduced to the PAIR framework for getting students to work with AI, and think about how we might use it in our teaching. I could see this being potentially useful in working on literature for essays – you could get the AI to suggest a reading list, then ask students to think about whether what it produces is the best material for the job, and how they would find out. If the university library system is going to suggest material students might read for any given topic using AI, then I guess we might as well bend into it. I’ve already got a session on finding secondary literature planned for my first years, so I think I will revisit it and introduce this framework. We’ll see how it goes sometime next year!

The next activity was to think about chatbots in teaching. It explained how some GenAI models are being created with ‘guardrails’ to make them suitable for children, and that children are being encouraged to use them as teaching assistants that are available 24/7, but it also noted how difficult it is to recreate the pedagogical skills of a good teacher (who wouldn’t always answer a question directly but would encourage students to find out for themselves or think the problem through), and the problems of ensuring the relevance of answers.

Interestingly, this latter aspect is something I was discussing with a colleague at work yesterday – it’s not just the fear that students will see GenAI as a shortcut that gives them answers without them having to think about it, but that without them having to do that thinking, they are not developing the skills to allow them to evaluate whether what they are being told is accurate, appropriate or the best answer. Take for example a student who has ideas for an essay, and asks GenAI to create a structure for it – something that, as it stands, students are being encouraged to do. Let’s park for a moment the fact that our grading criteria assume students are doing this for themselves, and therefore demonstrating a skill that is essential for a historian (and I’m not just saying that – the QAA benchmarks for history graduates also list this). Does the student know whether that structure makes sense? Can they see where it becomes circular (in our experience GenAI structures often do, presumably because they are working only with the information they are given)? Are they developing the skills that would allow them to adjust the structure to make it more logical, more nuanced or just downright clearer? And if the GenAI structure looks plausible, would they realise if the problems that they are having with structure might mean that there are wider difficulties, such as an argument that isn’t entirely logical? None of this is clear to me yet, and I find it problematic. On the one hand, we have to assume that students need to understand what AI can do, and where it might help (going back to that ‘what problem does AI solve?’ question from week 1), but I also naturally resist something that appears to prevent rather than facilitate students learning the higher level cognitive skills required of graduates. And it isn’t as simple as saying that they don’t need them anymore because AI can do it for them – they do, for all the reasons listed above.

And let’s not forget – the companies are not in this for the greater good. They are in it to make money. The course points out that ‘Anthropic (Claude) and OpenAI (ChatGPT) recently launched education-focused services, with the clear intention of capturing the student market through university partnerships.’ Perhaps without realising it, they have hit on one very important point – they see students as a market and universities as business partners. If they can persuade us that they are essential, they make more money.

Activity 5 was centred around AI in assessment, and I actually found this quite interesting. On the first task, we were encouraged to ask AI to help us think about changing our assessments, and this is something that I am really interested in anyway – although I think that the ability to write an essay is actually a really important one for students (after all, it is basically what actual historians do for a living, if you accept that it is the first step on the road to writing long form history of the type that historians are assessed on – the PhD, the monograph, the book chapter and the article), it is only one of several skills that a history graduate needs. Being honest, I was much more enthusiastic about this aspect of the course. This is much more the sort of task I think GenAI is quite good at – it is the teaching equivalent of the question I asked it about the range of my electric car. It is something that doesn’t require specific, detailed knowledge of the period I teach and it is essentially giving me ideas for an admin task I would probably undertake at some point anyway. It came up with some interesting options, and there are a few there that I would like to explore in future. However, we can’t ALL do away with our essays. They might not be the most exciting assessments, but without them the students aren’t gaining the skills they need in long-form writing (and yes, they really do need them, especially if they want to be a historian). So this raises another issue – in order to ensure that our students are assessed in a range of ways, and cover all the skills we need them to demonstrate, we actually need to sit down as a team to think about whose modules suit which types of assessment best, and who is prepared to change what. (There are actually good reasons for doing this even without the rise of AI, but it’s time that is the issue). Another potential issue is the expectations on which the AI is based – can I be sure that the assessments it suggests won’t privilege one group of students over another?

Another task was to think about where AI might assist (very much rather than replace) human assessments. This was absolutely not about using AI to mark student work, but about getting AI to write a bank of comments you could use that fit with the grading criteria for your assessment. It sounds awful, but the reality is that you often end up regurgitating the same comments for the same problems, and that doesn’t mean that you don’t or wouldn’t write specific things where specifics were appropriate. They weren’t bad, and I might use them as a starting point, but they do tend to read like generic comments. CoPilot (which I used because the information remains local rather than feeding the LLM – maybe ChatGPT does too but I wasn’t sure) also turned out to be quite good at generating essay titles from lecture notes…

So in many ways this was a more interesting and more practical week, that did make me think a lot about how I might be able to use AI to help me, even if it did little to answer questions or address fears I might have about how it might be integrated into teaching itself in the future. I still have very mixed feelings, and I’m still struggling with the ethics…

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.