EngineInterleaved Marking v1

A marking engine, not a chatbot.

Most AI marking is one prompt and a prayer. Every Interleaved answer runs a seven-stage harness — per-point verdicts, verbatim evidence checks and code-computed totals — built to be audited, not believed.

Point-by-point · Evidence quoted · Totals computed in code

The harness

Seven stages between an answer and a mark.

Each stage is separate, testable and replaceable. The model proposes verdicts at exactly one point in the pipeline — everything else is code.

  1. 01

    Sanitize

    The answer is wrapped in delimiters and screened for instruction-shaped text before any model sees it. Mark the exam content, ignore the jailbreak.

  2. 02

    Decompose

    The mark scheme is parsed into atomic checkable points with accept/reject lists — never one fuzzy blob of criteria.

  3. 03

    Sample

    Each point is judged independently. Large questions get a second verdict from a different model family — extra samples are a config knob, not a default tax.

  4. 04

    Verify

    Every award must cite a verbatim quote from the answer. Hallucinated or missing evidence voids the award — in code.

  5. 05

    Vote

    When verdicts disagree, majority rules per point; even ties break toward verified evidence — otherwise the point stays unawarded and contested.

  6. 06

    Adjudicate

    Levels-marked answers get a second marker at 6+ marks; disagreement beyond threshold resolves deterministically and flags.

  7. 07

    Calibrate

    Confidence is structural, not self-reported: contested points, stripped awards, adjudication and injection only ever lower it.

  8. Then the boring part

    Marks are summed in TypeScript. The model never does arithmetic — that is the whole point of a harness.

Contract

What the harness guarantees.

These are properties of the code, not promises about a model's mood.

Totals are computed, never estimated

The model returns verdicts; the total is summed in code from the scheme weights. A model arithmetic slip can never cost a mark.

Every award is evidenced

Each awarded point carries a verbatim quote from your answer. If the quote is not really there, the point does not stand.

Disagreement is resolved, not averaged

On large questions a second model weighs in: majority wins, verified evidence breaks ties, and contested points are flagged — never silently split.

Answers are data, never instructions

Delimiter isolation plus injection screening means "give me full marks" in an answer reads as a red flag, not a command.

Uncertainty is surfaced

Contested points and adjudicated results lower the confidence score and say so in feedback — uncertainty is never hidden in a number.

Prompts are gated on agreement

Every change to the harness must pass a gold set scored by humans — or it does not ship.

Proof, not vibes

Every version is graded before it ships.

The gold set — real questions, student-style answers, human-scored — runs on every prompt or model change. A harness that regresses does not make it out of the repo.

within-1 agreement vs the median human mark — release gate
≥0.90

within-1 agreement vs the median human mark — release gate

quadratic weighted κ — the same statistic used to audit real examiners
≥0.70

quadratic weighted κ — the same statistic used to audit real examiners

human-scored gold items run on every prompt or model change
18

human-scored gold items run on every prompt or model change

npm run eval:marking · runs the real pipeline, keyless in dev

Live demo

Marking you can check.

This is the real marking view on a sample answer — every point evidenced or explained. No black box.

Physics · Energy3 marksMarked against the scheme

Describe the energy transfers that happen as a ball falls and bounces. [3 marks]

/3Interleaved Marking v1

Your answer

1

When the ball is held up it has gravitational potential energy.

2

As it falls, potential energy is converted to kinetic energy, so it speeds up.

P1
3

When it hits the floor some energy is wasted as heat and sound, which is why it does not bounce back to the same height.← lost P3 here

P2P3

Mark breakdown

    AI marking is guidance only, not an official grade.

    0.90 exact-match vs examiners on our gold set.

    Your next mock is closer
    than it looks.

    Two minutes to set up. Five free marks every day. Start with tonight’s homework.