Continuation Drafter
The benchmark

What was measured, and what it does not tell you

51 models drafted continuation claims from the same ten real parent and continuation pairs, scored by the same four-vendor panel, across 5 rounds of testing between July and September 2026.

For the short answer, which AI model is best for patent drafting reads this table for you. Every model below has its own link: each row's anchor is its name.

Ten-specification means only

Panel scores are comparative within a batch, and a single-specification score moves far too much to cite: adding one draft to a twelve-draft batch shifted individual scores by up to 12.3 points.

Coverage sits beside the score

A model is averaged over the specifications it completed, and the one it failed is usually the hardest. That is why the tier with the higher score is not always the one to pick.

A comparative index, nothing more

Not an absolute quality measure, and it says nothing about patentability, novelty or non-obviousness, which would need a prior-art search this tool does not perform.

Which model to run on your own computer

One recommendation per machine size, so nothing here leaves your office.

Sizes are unified memory, measured on a Mac. On a PC with a separate graphics card, the card's own memory decides which model runs at a usable speed; a model larger than the card spills onto the processor and a draft that takes minutes on a GPU can time out. With an 8 GB card, use the 16 GB tier's models whatever the system memory.

More memory does not buy a better draft until the server tier, and the reason is worth a sentence before you spend anything. The best model that fits 32 GB scored 88; the best at 64 GB scored 81, at 96 GB scored 71, and at 192 GB scored 77. That is not a gap in what we tested, it is what sits at those sizes: nearly every model in the middle is a mixture-of-experts design, which keeps a large number of specialists in memory and uses only a few of them for any given word. One of them holds 235 billion parameters and thinks with 22 billion at a time. The 27-billion model that beats it uses all 27 billion, every word. What you are paying for with memory is capacity you mostly do not use, until the server tier, where the models are large enough that the part in use is bigger too.

Your machineModelScoreFinishedMemory used
A server with 512 GB of memoryglm-5.288.4 10 of 10roughly 450 GB at 4-bit, estimated from its size
A workstation with 64 GBgemma4:31b81.3 9 of 10about 30 GB, flat
A desktop or laptop with 32 GBqwen3.8:27b85.5 8 of 9 triedabout 20 GB, flat
A laptop with 16 GBgemma4:12b67.3 9 of 10about 12 GB, flat

A firm running the top row sits about 5.0 points behind the strongest closed-weight model measured, with nothing leaving the building. The gap between memory tiers is larger than the gap between open and closed weights.

Once a row looks right for your machine, set up a local model. If none of them fits, or you would rather not install one at all, the tool can run through your own provider key instead.

If a model returns nothing after a long wait

It is probably a thinking model, and one setting fixes it.

Some models deliberate before they answer, and that deliberation is charged against the same budget as the answer. Given a budget they can exhaust, they will: they reason until it is gone and return an empty result. More budget makes it slower, not more likely to work.

Measured on glm-5.3 against the same 49 KB specification, at the same 16,000-token budget:

SettingResult
provider default13,604 tokens spent thinking, no draft at all
--reasoning-effort higha complete 16-claim draft, 13 seconds

It rescues, it does not improve. Across ten specifications in one scoring round, gpt-5.5 scored 88.8 at high against 88.6 at its own default, and glm-5.2 85.5 against 84.4. Both differences are too small for this benchmark to tell apart. So the setting is the difference between a draft and nothing on the models that need it, and no measurable difference on the models that do not. Models measured to need it: glm-5.3, glm-5.3-flash, kimi-k3, kimi-k2.6, nemotron-3-super, and the qwen3.8 family. One model refuses the setting outright: mistral-medium-3-5 answers it with an error and was measured at its own default, so if a Mistral model returns a bad-request error, remove the flag.

Every model measured

The four rows above are the recommendation, one per machine size. This is every model tested: 51 of them, one row each, best first.

Ten real patent specifications went to each model, from about 17 pages to about 240. Score is how a panel of four independent AI reviewers rated the claim drafting, out of 100. Finished is how many of the ten it completed: a model that skipped some was usually beaten by the longest ones, so read it alongside the score rather than after it.

How much memory can your computer give a model?

On a Mac, that is the memory the machine has (Apple menu → About This Mac): the processor and graphics chip share one pool, so all of it is available to a model. On a PC with a separate graphics card, it is the card's own memory, not the system memory (Task Manager → Performance → GPU, "Dedicated GPU memory"): a model bigger than the card runs partly on the processor, many times slower, and a draft that takes minutes on a graphics chip can take hours. Most laptops have 16 or 32 GB of system memory; most graphics cards have 8, 12, 16 or 24. Cloud only shows the reverse: models that run on someone else's computer, which means the specification leaves your office.

No model we have measured drafts a long specification inside 8 GB. The smallest one that finished every specification, qwen3.5:9b, needs about 7 GB to load and grows to 13.7 GB on the longest one in the corpus, because the memory a model needs grows with the document you give it. On an 8 GB card it fits whole up to roughly 30,000 tokens, about 120,000 characters, and spills onto the processor past that. That is usable for shorter specifications and slow for long ones. For a model that fits with room to spare, the practical floor is a 16 GB card.

ScoreModelNeedsFinishedLicenceNotes
93.4 claude-opus-5 cloud only 9 of 10 closed
91.5 gpt-5.5 cloud only 10 of 10 closed anchor
91.1 claude-opus-4.7 cloud only 10 of 10 closed
90.8 minimax-m3 cloud only 10 of 10 open M3, newer than the m2.5 row above
90.4 claude-fable-5.1 cloud only 10 of 10 closed
90.4 gpt-6-astra cloud only 10 of 10 closed 1.05M context; succeeds gpt-5.5 and did not beat it
90.2 kimi-k3 cloud only 10 of 10 open 2.8T, unhostable
89.3 glm-5.3 512 GB server 10 of 10 open 1.31M context; needs --reasoning-effort
88.6 qwen3.6-max-preview cloud only 10 of 10 open
88.4 glm-5.2 512 GB server 10 of 10 open 744B, anchor
88.3 qwen3.8-max cloud only 10 of 10 open 27-32 min per draft
88.0 nemotron-3-ultra-550b-a55b 512 GB server 10 of 10 open 550B/55B active, ~309 GB at Q4
88.0 qwen3.8-27b 32 GB desktop 10 of 10 open hosted copy; run locally at 4-bit: 85.5, 20.3 GB, failed the 240-page spec
87.2 glm-5.3-flash cloud only 10 of 10 open needs --reasoning-effort
87.0 claude-sonnet-5 cloud only 10 of 10 closed mid tier
86.4 deepseek-v4.1-flash cloud only 10 of 10 open V4.1, newer than the v4-flash row above
86.4 gemini-3.6-flash cloud only 10 of 10 closed
84.5 glm-5.1 512 GB server 9 of 10 open
84.0 minimax-m2.5 cloud only 9 of 10 open
82.8 kimi-k2.5 cloud only 10 of 10 open
82.4 qwen3.6-27b 32 GB desktop 10 of 10 open
81.9 grok-4.6 cloud only 10 of 10 closed
81.4 deepseek-v4-pro cloud only 10 of 10 open 1.6T total / 49B active
81.3 gemma4:31b 64 GB desktop 9 of 10 open 30 GB resident
81.1 kimi-k2.6 cloud only 10 of 10 open needs --reasoning-effort
80.1 deepseek-v4-flash cloud only 10 of 10 open
80.0 glm-5 512 GB server 8 of 10 open
79.9 gemini-3.5-flash-lite cloud only 10 of 10 closed mid tier; 56-85 per-specification range
79.1 qwen3.6-35b-a3b 64 GB desktop 10 of 10 open
77.2 qwen3-235b-a22b-2507 192 GB server 10 of 10 open 235B/22B active
76.7 muse-glimmer-30b 32 GB desktop 8 of 10 open 30B dense, estimate; 128k context refused the two largest specifications
76.1 mistral-medium-3-5 128 GB server 10 of 10 open 128B dense; refuses --reasoning-effort, drafted at default
76.1 qwen3-next-80b-a3b 64 GB desktop 10 of 10 open 80B/3B active
70.7 llama-4-maverick 96 GB 10 of 10 open Maverick, not the Scout row above
68.5 qwen3-coder:30b 64-128 GB 10 of 10 open 44-122 GB resident
67.9 command-a 96 GB server 10 of 10 open 111B, citation-grounded
67.8 gemma4:26b 32 GB desktop 10 of 10 open re-measured locally 2026-09-21
67.4 gemma3:27b 32 GB desktop 8 of 10 open 20-23 GB resident
67.3 gemma4:12b 16 GB 9 of 10 open re-measured locally 2026-09-21; 12 GB resident, flat
66.1 nemotron-3-super-120b-a12b 96 GB server 10 of 10 open hybrid Mamba; 10/10 after re-run
65.0 qwen3.5:9b 16 GB 10 of 10 open 6.6 GB to 13.7 GB resident; cache grows with the specification
62.3 nemotron-3-nano-30b-a3b 32 GB desktop 10 of 10 open hybrid Mamba, 30B/3B active
57.3 gemma-3-27b-it (cloud) 32 GB desktop 10 of 10 open cloud copy of gemma3:27b
56.9 gpt-oss:120b 96 GB 9 of 10 open 65 GB resident
51.5 llama4:latest (Scout) 96 GB 10 of 10 open 108B, 74 GB resident
6 more models were tested but are not ranked

Each finished fewer than eight of the ten specifications, and the ones they skipped were the longest, except one that answered quickly in a form the tool could not read. Any score they earned would be an average over the easier half, so showing one would flatter them. They are listed anyway, because a model that cannot finish a specification your size is worth knowing about.

ModelFinishedNeedsNotes
glm-4.71 of 10512 GB server 1/10 coverage, NOT rankable
mistral-large-25125 of 10128 GB server 5/10 coverage
qwen3.5:4b2 of 1016 GB 2/10 coverage, NOT rankable; answers but not in the form the first pass needs
deepseek-r1:70b6 of 10128-256 GB 111-198 GB resident
devstral-small-2:24b7 of 1032-256 GB 20-187 GB, steep KV growth
granite-4.1-8b5 of 1016 GB hybrid Mamba, 5/10 coverage

Two things to read it with, both measured rather than cautionary. A difference under about four points is not a ranking. The index is a comparative score, and the same model measured in different rounds of testing varied by up to three points without changing at all. Models tested more than once are averaged, and every score is placed on one common scale so the column can be read straight down; the arithmetic and the raw per-round numbers are in the benchmark write-up. Coverage matters as much as the score. A model is averaged over the specifications it drafted, and the one it failed is usually the hardest, so anything below eight of ten is excluded from the ranking rather than shown with a flattering number. Local figures carry an unmeasured quantization penalty: local models ran at four-bit, cloud copies at unspecified precision, and an attempt to measure that difference gave a backwards result, so no figure is claimed for it.

The honest limit on all of it

These numbers are machines grading machines. They are useful for choosing between models and they are not a measure of whether a draft is worth filing. The one human read we have, a licensed practitioner reviewing three drafts blind, ranked them differently from the panel. We report the spread rather than a ranking for that reason.