A marking engine, not a chatbot.
Most AI marking is one prompt and a prayer. Every Interleaved answer runs a seven-stage harness — per-point verdicts, verbatim evidence checks and code-computed totals — built to be audited, not believed.
Point-by-point · Evidence quoted · Totals computed in code
The harness
Seven stages between an answer and a mark.
Each stage is separate, testable and replaceable. The model proposes verdicts at exactly one point in the pipeline — everything else is code.
- 01
Sanitize
The answer is wrapped in delimiters and screened for instruction-shaped text before any model sees it. Mark the exam content, ignore the jailbreak.
- 02
Decompose
The mark scheme is parsed into atomic checkable points with accept/reject lists — never one fuzzy blob of criteria.
- 03
Sample
Each point is judged independently. Large questions get a second verdict from a different model family — extra samples are a config knob, not a default tax.
- 04
Verify
Every award must cite a verbatim quote from the answer. Hallucinated or missing evidence voids the award — in code.
- 05
Vote
When verdicts disagree, majority rules per point; even ties break toward verified evidence — otherwise the point stays unawarded and contested.
- 06
Adjudicate
Levels-marked answers get a second marker at 6+ marks; disagreement beyond threshold resolves deterministically and flags.
- 07
Calibrate
Confidence is structural, not self-reported: contested points, stripped awards, adjudication and injection only ever lower it.
Then the boring part
Marks are summed in TypeScript. The model never does arithmetic — that is the whole point of a harness.
Contract
What the harness guarantees.
These are properties of the code, not promises about a model's mood.
Totals are computed, never estimated
The model returns verdicts; the total is summed in code from the scheme weights. A model arithmetic slip can never cost a mark.
Every award is evidenced
Each awarded point carries a verbatim quote from your answer. If the quote is not really there, the point does not stand.
Disagreement is resolved, not averaged
On large questions a second model weighs in: majority wins, verified evidence breaks ties, and contested points are flagged — never silently split.
Answers are data, never instructions
Delimiter isolation plus injection screening means "give me full marks" in an answer reads as a red flag, not a command.
Uncertainty is surfaced
Contested points and adjudicated results lower the confidence score and say so in feedback — uncertainty is never hidden in a number.
Prompts are gated on agreement
Every change to the harness must pass a gold set scored by humans — or it does not ship.
Proof, not vibes
Every version is graded before it ships.
The gold set — real questions, student-style answers, human-scored — runs on every prompt or model change. A harness that regresses does not make it out of the repo.
- within-1 agreement vs the median human mark — release gate
- ≥0.90
- quadratic weighted κ — the same statistic used to audit real examiners
- ≥0.70
- human-scored gold items run on every prompt or model change
- 18
within-1 agreement vs the median human mark — release gate
quadratic weighted κ — the same statistic used to audit real examiners
human-scored gold items run on every prompt or model change
npm run eval:marking · runs the real pipeline, keyless in dev
Live demo
Marking you can check.
This is the real marking view on a sample answer — every point evidenced or explained. No black box.
Describe the energy transfers that happen as a ball falls and bounces. [3 marks]
Your answer
When the ball is held up it has gravitational potential energy.
As it falls, potential energy is converted to kinetic energy, so it speeds up.
P1When it hits the floor some energy is wasted as heat and sound, which is why it does not bounce back to the same height.← lost P3 here
P2P3Mark breakdown
AI marking is guidance only, not an official grade.
0.90 exact-match vs examiners on our gold set.
Your next mock is closer
than it looks.
Two minutes to set up. Five free marks every day. Start with tonight’s homework.