Marking time
Total department hours logged against marking for one term, against the same term the previous year with the same number of papers. Self-reported by the six teachers on their own timesheets.
Six teachers, forty scripts each, every fortnight. The marking was not the job — it was the thing standing between them and the job.
6 teachers · Standards 6–10 · fortnightly papers. Measured across one full term, against the same department’s prior year.
A fortnightly paper meant a lost Sunday, and marking drifted depending on which teacher picked up which pile.
Eight minutes to set, twenty-eight to review forty scripts, and a written rubric that made two teachers mark the same script the same way.
“I did not mind the marking. I minded that the marking was the reason I never got to the part where I work out what to teach next.”
Head of science
Not a plan — what they did, including the parts that were not the plan. Pick a moment.
Before anything else. Six teachers, one afternoon, deciding in writing what earns the second mark on a three-mark question. Nobody in the department had ever written it down, and the first attempt disagreed with itself.
Forty scripts marked by hand as usual, and marked again by the platform, without either seeing the other. The comparison is what convinced the two sceptics, in both directions.
The change in the job. Scripts arrive marked against the rubric with the phrase that earned each point, sorted so the least confident ones are first. Twenty-eight minutes for forty.
Eight minutes from scoping chapters to a paper the teacher was willing to publish. The department stopped reusing last year’s papers, which had been the real cost of setting being slow.
Marking time down 83% across the term. The department’s own framing: the hours did not go into more marking, they went into the reteaching conversation that marking had always crowded out.
Measured across one full term, against the same department’s prior year. A figure without a method is a figure you cannot check, so each one says what it counted and what it did not.
Total department hours logged against marking for one term, against the same term the previous year with the same number of papers. Self-reported by the six teachers on their own timesheets.
Median time for one teacher to review a marked set of forty scripts, including every override. The slowest was 51 minutes; the fastest 19.
The proportion of AI-proposed scores the reviewing teacher left unchanged, across roughly 4,800 marked questions. Objective sections are excluded, since agreement there is trivially exact.
Median from opening the generator to publishing, including edits. Excludes the term where the rubric was being written, which is a one-off.
Every rollout has one, and a case study without it is marketing. These are theirs, in their own account.
An afternoon of six teachers arguing about what earns a second mark, then a second draft because the first one contradicted itself. That is genuine work and it happened before a single script was marked. It is also the thing that made the marking consistent between teachers, which the department had been quietly failing at for years.
They asked for forty scripts to be marked both ways, blind, before agreeing to anything. Roughly one script in twenty-five came back materially different, and reading those individually is what changed their minds — in both directions. Two of the disputed marks were the platform’s, and three were the teacher’s.
One teacher in the department marks extended argument, where the marking is the teaching. After a term she stopped using it for that and kept it for everything else. We would rather report that than average it away — the 96% figure excludes her subject entirely.
If any of that would be a problem where you are, it is worth raising on the first call rather than in the second term.
All six are measured the same way. See all six
The ones that come up almost every time.
No. It marked first and the teachers reviewed, which is the whole design. The saving is in not reading forty scripts from scratch — the reviewing, and the final say, stayed with the department.
Written English composition. The rubric that worked for Science and Social Science could not capture what that department was actually rewarding, and they went back to marking it by hand. The 96% figure excludes it.
Their own moderation sample said no, but they ran that sample for two terms before trusting it — which is what we would suggest to anyone.
Book a free 30-minute demo. We’ll build the same breakdown from a paper you already ran, so the first number you see is your own.