What on earth is "supported evaluation"?
In 1964, US Supreme Court Justice Potter Stewart was asked to decide whether a French film counted as hard-core pornography. He declined to define the term and wrote simply, "I know it when I see it."
It's a great line, and as long as he was the only one deciding, it more or less worked. The trouble starts when more than one person is doing the recognising. If the "it" in my head differs from the "it" in yours, we can both be confident, both be internally consistent and still disagree about the same thing. Neither of us is wrong by our own standard, and nobody on the outside can work out what the standard is.
I think "supported evaluation" is exactly this kind of "it". AQA's 25-mark grid mentions it several times, but never actually defines it. When I went through the mark schemes, examiner reports and exemplar commentaries, the word "supported" turned out to be doing at least six different jobs. This, at least in theory, means examiners could be looking for different things.
The official wording
Let's take a step back for a moment and look at just the "evaluation" side of it. The assessment objectives are set by Ofqual, so they're the same whichever board you teach. AO4 reads:
"Evaluate economic arguments and use qualitative and quantitative evidence to support informed judgements relating to economic issues."
Two things in that sentence are easy to skim past.
The first is what gets evaluated. AO3 says "analyse issues within economics", while AO4 says "evaluate economic arguments". I don't think that's an accident. Nothing in AO4 asks for a second issue, an opposing case or a list of disadvantages. It asks students to assess arguments that are already on the table: how strong they are, when they hold and when they don't. If we were evaluating issues, then that would, I think, be more like looking at different points of view, or different sides of a coin.
The second is that the support requirement sits in the same sentence as the objective. Evidence is used to support judgements. A judgement on its own doesn't meet the objective, and neither does evidence on its own. It's the combination that counts.
The AQA 25-mark question grid
Here are the top three levels of the marking grid for 25-mark questions. I put the rows in to make it easier to separate the skills, and I've formatted to highlight the distinctions between each level, but otherwise it's verbatim.
Level 3 (11–15) | Level 4 (16–20) | Level 5 (21–25) | |
Heading | "Some reasonable analysis but generally unsupported evaluation that:" | "Sound, focused analysis and some supported evaluation that:" | "Sound, focused analysis and well-supported evaluation that:" |
Knowledge | "focuses on issues that are relevant to the question, showing satisfactory knowledge and understanding of economic terminology, concepts and principles but some weaknesses may be present" | "is well organised, showing sound knowledge and understanding of economic terminology, concepts and principles with few, if any, errors" | "is well organised, showing sound knowledge and understanding of economic terminology, concepts and principles with few, if any, errors" |
Application | "includes reasonable application of relevant economic principles to the given context and, where appropriate, some use of data to support the response" | "includes some good application of relevant economic principles to the given context and, where appropriate, some good use of data to support the response" | "includes good application of relevant economic principles to the given context and, where appropriate, good use of data to support the response" |
Analysis | "includes some reasonable analysis but which might not be adequately developed or becomes confused in places" | "includes some well-focused analysis with clear, logical chains of reasoning" | "includes well-focused analysis with clear, logical chains of reasoning" |
Evaluation | "includes fairly superficial evaluation; there is likely to be some attempt to make relevant judgements but these aren't well-supported by arguments and/or data" | "includes some reasonable, supported evaluation" | "includes supported evaluation throughout the response and in a final conclusion" |
Two things jump out when you read it as a ladder rather than one level at a time.
The word that changes is "supported". Level 3 students are already making judgements, and the grid says so. What they aren't doing is supporting them.
Level 5 adds coverage. Supported evaluation "throughout the response and in a final conclusion" is what separates the top two bands, so interim judgements are a requirement rather than a matter of style.
AQA's Paper 3 exemplar scripts from June 2022 illustrate the difference well. Two Paper 3 answers are both described as presenting arguments for and against, and both reach a recommendation, yet they're ten marks apart. The difference is whether the evaluation is "backed up by well-developed analysis and use of the data" or "superficial".
"Supported" evaluation therefore clearly matters a lot.
The six readings of "supported"
The grid never defines the word, so the only way to find out what AQA means is to look at how it's used in the reports and commentaries. Each of the six readings below is inferred from places where examiners' reports and commentaries have either praised the presence or criticised its absence.
1. Support as evidence
A judgement is supported if it carries data, a figure or a real-world example. This is the reading closest to AO4's own wording.
"Different final judgements, provided they were supported by the data presented, were also rewarded." (Paper 3, June 2022, Q31)
It's worth knowing that the Paper 1 report's Level 5 formula (June 2022, repeated almost word for word in 2023, 2024 and 2026) asks for all three sources, joined by "and": "supported by theoretical analysis and by the use of data from the extracts (if applicable) and the candidates' own examples and contexts." Only the extract data is qualified by "if applicable". That "contexts" bit is important, because I don't think this necessarily means finding facts and figures (especially in Papers 1 and 2). I think it can mean applying the features of the industry or economy. For example, "Yeah, taxes can be great at reducing consumption, but they don't do much in the case of cigarettes because nicotine is addictive..."
2. Support as mechanism
A judgement is supported if it comes with a "because", a chain of reasoning. Under this reading no citation is needed, because the theory does the supporting.
"They make some evaluative comments, but these tend to be based on assertion rather than supported by analysis." (Paper 1 exemplar, Q14, Response B, 14 marks)
Here "support" means "tell me why" - we're essentially looking for another chain of analysis.
3. Support as condition
A judgement is supported if it says when it holds and when it doesn't. The indicative content asks for this one most often ("the extent to which" turns up 35 times across the mark schemes).
"The strongest answers considered whether the solution to the deficit may depend upon whether the deficit was cyclical or structural and how this may shape policy." (Paper 2, June 2024)
This is the classic "it depends on...". When is this true?
4. Support as calibration
A judgement is supported if it claims no more than its backing can carry. Interestingly, AQA's own advice for moving one script from Level 4 to Level 5 relies entirely on this reading.
"...with a view to suggesting one of the ways it might move into Level 5, the student might have softened their language. It was quite assertive in places." (Paper 1 exemplar, Q12, Level 4 response, 18 marks)
In my first year of teaching Economics A-Level, I went to an Edexcel training day where a lovely lady in the loos suggested "How big? How likely? How important?" as a catchphrase to get students thinking evaluatively. This sums up this reading of "support" nicely. Often it means softening a view or adding a bit of hedging. It also addresses the "spurious", "contrived" and "unconvincing" criticisms that crop up in examiner reports, especially on Paper 3.
5. Support as comparison
A judgement is supported if it has been weighed against the alternative and found bigger or smaller, not just stated on its own.
"The best responses finished with a final recommendation that was supported by bringing together and assessing the relative merits of the various arguments for and against a substantial increase in the NLW." (Paper 3, November 2020)
This is super common in conclusions, for example, in policy essays, where students write, "Policy A has these advantages and disadvantages. Policy B has these advantages and disadvantages. Overall, I think policy A is best because it has this advantage," and they're not really weighing up those two against each other. Of all of the readings of "supported evaluation", I think this is one of the trickiest to do successfully within the confines of a 40-minute essay. This is often because we are weighing up in two different currencies: trying to compare apples and oranges. Often, changing them both to "fruit" is possible, but that takes some real skill and a significant amount of time.
6. Support as derivation
A judgement is supported if it follows from what the essay has already established. It shouldn't appear out of nowhere in the final paragraph, and it shouldn't just repeat the previous one. This reading hardly appears in the mark schemes but is all over the exemplar commentaries.
"...the evaluation is fully supported by the conclusion itself and the earlier analysis and data." (Paper 1 exemplar, Q4, Response A, 23 marks)
"It is not fully consistent with the earlier discussion..." (Paper 3 exemplar, Q31, Response B, 6 marks)
The idea here is that a judgement needs to be supported by what has come before it. We often tell students that it doesn't matter what they conclude, as long as they've justified it, but I wonder if students interpret that differently from how we mean it. We are trying to communicate that an examiner doesn't have a "correct conclusion" and an "incorrect conclusion" on a mark scheme somewhere. However, once they've actually written their essay, there is a possible wrong conclusion: if they have spent their essay talking about the vast problems and issues with scrapping inheritance tax, the conclusion that it should be abolished is no longer valid without some serious gymnastics. In that case, either the body of the essay needs to be changed, or the conclusion does. No amount of adding to the conclusion is going to fix that problem.
Why this matters
Evidence and mechanism are about what's attached to a claim. Condition, calibration, comparison and derivation are about the shape of the claim: whether it's conditional, proportionate, comparative and derived. Either way, it is plausible to think that one examiner might be looking for evidence and a mechanism, while another is looking for comparison and calibration. A student has no way of knowing which examiner they will get. Both are reading the grid, and on the same script they could end up in different levels entirely, making a big difference to marks. That's the Potter Stewart problem in a nutshell.
At this point, I should calibrate my own argument: I think it's unlikely that most examiners will only recognise one of these six readings of "supported", so that chances of a 'false negative' across a whole essay are pretty slim. However, I think it's highly unlikely that most examiners are recognising and rewarding all six readings of "supported evaluation" equally and interchangeably. I wonder if this is resolved, at least a little, depending on where we are looking. Readings 1-4 (evidence, mechanism, condition and calibration) are very doable in the body of the essay, but 5 and 6 (comparison and derivation) are potentially better suited to the conclusion, so an examiner might be valuing one reading in the body and another in the final judgement. I hope to write another post soon on how we might scaffold students to meet readings 1-4 in the evaluative paragraphs and 5 and 6 in the conclusion.