What was measured, and what it does not tell you
51 models drafted continuation claims from the same ten real parent and continuation pairs, scored by the same four-vendor panel, across 5 rounds of testing between July and September 2026.
For the short answer, which AI model is best for patent drafting reads this table for you. Every model below has its own link: each row's anchor is its name.
Ten-specification means only
Panel scores are comparative within a batch, and a single-specification score moves far too much to cite: adding one draft to a twelve-draft batch shifted individual scores by up to 12.3 points.
Coverage sits beside the score
A model is averaged over the specifications it completed, and the one it failed is usually the hardest. That is why the tier with the higher score is not always the one to pick.
A comparative index, nothing more
Not an absolute quality measure, and it says nothing about patentability, novelty or non-obviousness, which would need a prior-art search this tool does not perform.
Which model to run on your own computer
One recommendation per machine size, so nothing here leaves your office.
Sizes are unified memory, measured on a Mac. On a PC with a separate graphics card, the card's own memory decides which model runs at a usable speed; a model larger than the card spills onto the processor and a draft that takes minutes on a GPU can time out. With an 8 GB card, use the 16 GB tier's models whatever the system memory.
More memory does not buy a better draft until the server tier, and the reason is worth a sentence before you spend anything. The best model that fits 32 GB scored 88; the best at 64 GB scored 81, at 96 GB scored 71, and at 192 GB scored 77. That is not a gap in what we tested, it is what sits at those sizes: nearly every model in the middle is a mixture-of-experts design, which keeps a large number of specialists in memory and uses only a few of them for any given word. One of them holds 235 billion parameters and thinks with 22 billion at a time. The 27-billion model that beats it uses all 27 billion, every word. What you are paying for with memory is capacity you mostly do not use, until the server tier, where the models are large enough that the part in use is bigger too.
| Your machine | Model | Score | Finished | Memory used |
|---|---|---|---|---|
| A server with 512 GB of memory | glm-5.2 | 88.4 | 10 of 10 | roughly 450 GB at 4-bit, estimated from its size |
| A workstation with 64 GB | gemma4:31b | 81.3 | 9 of 10 | about 30 GB, flat |
| A desktop or laptop with 32 GB | qwen3.8:27b | 85.5 | 8 of 9 tried | about 20 GB, flat |
| A laptop with 16 GB | gemma4:12b | 67.3 | 9 of 10 | about 12 GB, flat |
A firm running the top row sits about 5.0 points behind the strongest closed-weight model measured, with nothing leaving the building. The gap between memory tiers is larger than the gap between open and closed weights.
Once a row looks right for your machine, set up a local model. If none of them fits, or you would rather not install one at all, the tool can run through your own provider key instead.
If a model returns nothing after a long wait
It is probably a thinking model, and one setting fixes it.
Some models deliberate before they answer, and that deliberation is charged against the same budget as the answer. Given a budget they can exhaust, they will: they reason until it is gone and return an empty result. More budget makes it slower, not more likely to work.
Measured on glm-5.3 against the same 49 KB specification, at the same
16,000-token budget:
| Setting | Result |
|---|---|
| provider default | 13,604 tokens spent thinking, no draft at all |
--reasoning-effort high | a complete 16-claim draft, 13 seconds |
It rescues, it does not improve. Across ten
specifications in one scoring round, gpt-5.5 scored 88.8 at high
against 88.6 at its own default, and glm-5.2 85.5 against 84.4. Both differences
are too small for this benchmark to tell apart. So the setting is the difference between a draft and
nothing on the models that need it, and no measurable difference on the models that do not.
Models measured to need it: glm-5.3, glm-5.3-flash,
kimi-k3, kimi-k2.6, nemotron-3-super, and the
qwen3.8 family. One model refuses the setting outright:
mistral-medium-3-5 answers it with an error and was measured at its own
default, so if a Mistral model returns a bad-request error, remove the flag.
Every model measured
The four rows above are the recommendation, one per machine size. This is every model tested: 51 of them, one row each, best first.
Ten real patent specifications went to each model, from about 17 pages to about 240. Score is how a panel of four independent AI reviewers rated the claim drafting, out of 100. Finished is how many of the ten it completed: a model that skipped some was usually beaten by the longest ones, so read it alongside the score rather than after it.
| Score | Model | Needs | Finished | Licence | Notes |
|---|---|---|---|---|---|
| 93.4 | claude-opus-5 |
cloud only | 9 of 10 | closed | |
| 91.5 | gpt-5.5 |
cloud only | 10 of 10 | closed | anchor |
| 91.1 | claude-opus-4.7 |
cloud only | 10 of 10 | closed | |
| 90.8 | minimax-m3 |
cloud only | 10 of 10 | open | M3, newer than the m2.5 row above |
| 90.4 | claude-fable-5.1 |
cloud only | 10 of 10 | closed | |
| 90.4 | gpt-6-astra |
cloud only | 10 of 10 | closed | 1.05M context; succeeds gpt-5.5 and did not beat it |
| 90.2 | kimi-k3 |
cloud only | 10 of 10 | open | 2.8T, unhostable |
| 89.3 | glm-5.3 |
512 GB server | 10 of 10 | open | 1.31M context; needs --reasoning-effort |
| 88.6 | qwen3.6-max-preview |
cloud only | 10 of 10 | open | |
| 88.4 | glm-5.2 |
512 GB server | 10 of 10 | open | 744B, anchor |
| 88.3 | qwen3.8-max |
cloud only | 10 of 10 | open | 27-32 min per draft |
| 88.0 | nemotron-3-ultra-550b-a55b |
512 GB server | 10 of 10 | open | 550B/55B active, ~309 GB at Q4 |
| 88.0 | qwen3.8-27b |
32 GB desktop | 10 of 10 | open | hosted copy; run locally at 4-bit: 85.5, 20.3 GB, failed the 240-page spec |
| 87.2 | glm-5.3-flash |
cloud only | 10 of 10 | open | needs --reasoning-effort |
| 87.0 | claude-sonnet-5 |
cloud only | 10 of 10 | closed | mid tier |
| 86.4 | deepseek-v4.1-flash |
cloud only | 10 of 10 | open | V4.1, newer than the v4-flash row above |
| 86.4 | gemini-3.6-flash |
cloud only | 10 of 10 | closed | |
| 84.5 | glm-5.1 |
512 GB server | 9 of 10 | open | |
| 84.0 | minimax-m2.5 |
cloud only | 9 of 10 | open | |
| 82.8 | kimi-k2.5 |
cloud only | 10 of 10 | open | |
| 82.4 | qwen3.6-27b |
32 GB desktop | 10 of 10 | open | |
| 81.9 | grok-4.6 |
cloud only | 10 of 10 | closed | |
| 81.4 | deepseek-v4-pro |
cloud only | 10 of 10 | open | 1.6T total / 49B active |
| 81.3 | gemma4:31b |
64 GB desktop | 9 of 10 | open | 30 GB resident |
| 81.1 | kimi-k2.6 |
cloud only | 10 of 10 | open | needs --reasoning-effort |
| 80.1 | deepseek-v4-flash |
cloud only | 10 of 10 | open | |
| 80.0 | glm-5 |
512 GB server | 8 of 10 | open | |
| 79.9 | gemini-3.5-flash-lite |
cloud only | 10 of 10 | closed | mid tier; 56-85 per-specification range |
| 79.1 | qwen3.6-35b-a3b |
64 GB desktop | 10 of 10 | open | |
| 77.2 | qwen3-235b-a22b-2507 |
192 GB server | 10 of 10 | open | 235B/22B active |
| 76.7 | muse-glimmer-30b |
32 GB desktop | 8 of 10 | open | 30B dense, estimate; 128k context refused the two largest specifications |
| 76.1 | mistral-medium-3-5 |
128 GB server | 10 of 10 | open | 128B dense; refuses --reasoning-effort, drafted at default |
| 76.1 | qwen3-next-80b-a3b |
64 GB desktop | 10 of 10 | open | 80B/3B active |
| 70.7 | llama-4-maverick |
96 GB | 10 of 10 | open | Maverick, not the Scout row above |
| 68.5 | qwen3-coder:30b |
64-128 GB | 10 of 10 | open | 44-122 GB resident |
| 67.9 | command-a |
96 GB server | 10 of 10 | open | 111B, citation-grounded |
| 67.8 | gemma4:26b |
32 GB desktop | 10 of 10 | open | re-measured locally 2026-09-21 |
| 67.4 | gemma3:27b |
32 GB desktop | 8 of 10 | open | 20-23 GB resident |
| 67.3 | gemma4:12b |
16 GB | 9 of 10 | open | re-measured locally 2026-09-21; 12 GB resident, flat |
| 66.1 | nemotron-3-super-120b-a12b |
96 GB server | 10 of 10 | open | hybrid Mamba; 10/10 after re-run |
| 65.0 | qwen3.5:9b |
16 GB | 10 of 10 | open | 6.6 GB to 13.7 GB resident; cache grows with the specification |
| 62.3 | nemotron-3-nano-30b-a3b |
32 GB desktop | 10 of 10 | open | hybrid Mamba, 30B/3B active |
| 57.3 | gemma-3-27b-it (cloud) |
32 GB desktop | 10 of 10 | open | cloud copy of gemma3:27b |
| 56.9 | gpt-oss:120b |
96 GB | 9 of 10 | open | 65 GB resident |
| 51.5 | llama4:latest (Scout) |
96 GB | 10 of 10 | open | 108B, 74 GB resident |
6 more models were tested but are not ranked
Each finished fewer than eight of the ten specifications, and the ones they skipped were the longest, except one that answered quickly in a form the tool could not read. Any score they earned would be an average over the easier half, so showing one would flatter them. They are listed anyway, because a model that cannot finish a specification your size is worth knowing about.
| Model | Finished | Needs | Notes |
|---|---|---|---|
glm-4.7 | 1 of 10 | 512 GB server | 1/10 coverage, NOT rankable |
mistral-large-2512 | 5 of 10 | 128 GB server | 5/10 coverage |
qwen3.5:4b | 2 of 10 | 16 GB | 2/10 coverage, NOT rankable; answers but not in the form the first pass needs |
deepseek-r1:70b | 6 of 10 | 128-256 GB | 111-198 GB resident |
devstral-small-2:24b | 7 of 10 | 32-256 GB | 20-187 GB, steep KV growth |
granite-4.1-8b | 5 of 10 | 16 GB | hybrid Mamba, 5/10 coverage |
Two things to read it with, both measured rather than cautionary. A difference under about four points is not a ranking. The index is a comparative score, and the same model measured in different rounds of testing varied by up to three points without changing at all. Models tested more than once are averaged, and every score is placed on one common scale so the column can be read straight down; the arithmetic and the raw per-round numbers are in the benchmark write-up. Coverage matters as much as the score. A model is averaged over the specifications it drafted, and the one it failed is usually the hardest, so anything below eight of ten is excluded from the ranking rather than shown with a flattering number. Local figures carry an unmeasured quantization penalty: local models ran at four-bit, cloud copies at unspecified precision, and an attempt to measure that difference gave a backwards result, so no figure is claimed for it.
The honest limit on all of it
These numbers are machines grading machines. They are useful for choosing between models and they are not a measure of whether a draft is worth filing. The one human read we have, a licensed practitioner reviewing three drafts blind, ranked them differently from the panel. We report the spread rather than a ranking for that reason.