Continuation Drafter
The benchmark

What was measured, and what it does not tell you

Twenty-nine models drafted continuation claims from the same ten real parent and continuation pairs, scored by the same four-vendor panel.

Ten-specification means only

Panel scores are comparative within a batch, and a single-specification score moves far too much to cite: adding one draft to a twelve-draft batch shifted individual scores by up to 12.3 points.

Coverage sits beside the score

A model is averaged over the specifications it completed, and the one it failed is usually the hardest. That is why the tier with the higher score is not always the one to pick.

A comparative index, nothing more

Not an absolute quality measure, and it says nothing about patentability, novelty or non-obviousness, which would need a prior-art search this tool does not perform.

Local models by memory tier

MachineModelIndexCoverageResident
A server with 512 GB of memoryglm-5.288.3 10 of 10about 309 GB at 4-bit
A workstation with 64 GBgemma4:31b79.8 9 of 10about 30 GB, flat
A desktop or laptop with 32 GBgemma4:26b65.0 10 of 10about 20 GB, flat
A laptop with 16 GBgemma4:12b62.9 9 of 10about 12 GB, flat

A firm running the top row sits about 3.3 points behind the strongest closed-weight model measured, with nothing leaving the building. The gap between memory tiers is larger than the gap between open and closed weights.

The honest limit on all of it

These numbers are machines grading machines. They are useful for choosing between models and they are not a measure of whether a draft is worth filing. The one human read we have, a licensed practitioner reviewing three drafts blind, ranked them differently from the panel. We report the spread rather than a ranking for that reason.