Mean board score
Board results for the cohort that sat two terms under the new process, against the previous year’s cohort at the same school. Not adjusted for intake, which the school considered broadly unchanged.
A school that stopped comparing divisions on totals and started comparing them on chapters — then reordered six weeks of teaching around what it found.
12 divisions · Standards 4–12 · about 2,400 students. Measured across two full terms, against the same school’s prior year.
Twelve teachers set twelve papers. Nobody could say whether 9-A and 9-C had been tested on the same thing, so a gap between them was always arguable.
One paper, one rubric, one chapter map per division — and a January staff meeting that was about Chapter 7 rather than about who marks harder.
“We had been comparing divisions on totals for years. It turns out we were comparing twelve different papers and calling it data.”
Head of academics
Not a plan — what they did, including the parts that were not the plan. Pick a moment.
They did not roll it out. One head of science ran it on 9-B for a single unit test, without telling anyone, to see whether the chapter map said anything she did not already know.
The map showed 9-B losing eleven marks on Chapter 7 — a chapter she was confident they had covered well. That disagreement is what got the rest of the department interested.
The first paper generated once and sat by four divisions. This is where the comparison stopped being arguable, because for the first time the paper and the rubric were genuinely the same.
Standards 4 to 12, twelve divisions, on one licence. The import took an afternoon; agreeing the marks schemes between heads of department took three weeks.
The January scheme of work was rewritten around the three chapters the reports kept surfacing, rather than around what had been planned the previous June.
Measured across two full terms, against the same school’s prior year. A figure without a method is a figure you cannot check, so each one says what it counted and what it did not.
Board results for the cohort that sat two terms under the new process, against the previous year’s cohort at the same school. Not adjusted for intake, which the school considered broadly unchanged.
The gap between a weak chapter first appearing in a report and the same chapter appearing in the old process — which was the pre-board mock. Measured on three chapters across two terms.
How many chapters the January scheme of work actually changed for. The school set the threshold themselves: any chapter costing a division more than eight marks.
Every division in Standards 4–12, including the two that only sit internal assessments.
Every rollout has one, and a case study without it is marketing. These are theirs, in their own account.
Not the setting — the marks schemes. Three heads of department had to write down what earns the second mark on a three-mark question, and none of them had ever done it in writing. That took three weeks of meetings and was, by their own account, overdue regardless of the software.
It is also the reason the drift number moved at all. Two teachers marking to a written rubric land in the same place; two teachers marking to a remembered one do not.
One taught a subject where the marking is the teaching — extended writing — and correctly judged that a rubric-driven first pass added nothing. The other did not want AI near her marking at all and was not asked to change her mind. Both still use the generator to set papers and mark by hand.
The +14% is one cohort against one prior cohort at one school. It is the figure the school leads with and the one we would defend least strongly, because two terms is a short series and intake varies. The detection and reordering figures are process measurements and are much harder to argue with.
If any of that would be a problem where you are, it is worth raising on the first call rather than in the second term.
All six are measured the same way. See all six
The ones that come up almost every time.
Partly. A single paper gives a usable chapter breakdown of that paper; trends need three or four before they mean anything, and chapter flags stay marked provisional until a second result agrees. The division-level view is useful immediately, the per-student trend from about the second month.
It cost them three weeks of departmental meetings before a single lesson moved, which the headline figure does not show. The reordering itself was two chapters swapped forward, not a rewritten scheme of work.
No. The gain was concentrated in the divisions whose teachers acted on the chapter report. Where it was read and filed, the second-term numbers were flat, which is the honest limit of any analytics.
Book a free 30-minute demo. We’ll build the same breakdown from a paper you already ran, so the first number you see is your own.