Continuation Drafter
Choosing a model

Which AI model is best for patent drafting?

We measured it, on continuation claims. 7 hosted models are effectively tied at the top, a few run well on your own computer, and several are worth avoiding. The answer also depends on whether the specification may leave your office.

Every model below drafted continuation claims from the same ten real patent specifications, and a panel of four AI reviewers from different vendors scored the drafts. Each name links to its row in the full benchmark. The scores rate claim drafting, not whether anything is patentable.

The short answer

Which AI model is best for patent drafting?

On a benchmark of ten real patent specifications, these hosted models scored within 4 points of the best, closer than the benchmark can tell apart, so treat them as tied: claude-opus-5, gpt-5.5, claude-opus-4.7, minimax-m3, claude-fable-5.1, gpt-6-astra, kimi-k3. Each sends the specification to its provider. The scores measure claim drafting quality only, not patentability.

What is the best local AI model for lawyers?

A model on your own computer keeps a client's specification on that computer. By the memory the computer can give a model, the measured recommendations are: server with 512 GB of memory: glm-5.2; workstation with 64 GB: gemma4:31b; desktop or laptop with 32 GB: qwen3.8:27b; laptop with 16 GB: gemma4:12b.

The best hosted models

If the specification may go to a provider, these scored within 4 points of the best hosted result (93.4 out of 100). That is closer than the benchmark can tell apart, so treat them as tied and choose on price, speed or the provider you already use.

ModelScoreFinishedLicenceWorth knowing
claude-opus-5 93.49 of 10closed
gpt-5.5 91.510 of 10closed
claude-opus-4.7 91.110 of 10closed
minimax-m3 90.810 of 10open
claude-fable-5.1 90.410 of 10closed
gpt-6-astra 90.410 of 10closed
kimi-k3 90.210 of 10openAmong the slowest models measured, about half an hour per draft, and it timed out once. kimi-k2.5 takes about a quarter of the time at a fifth of the cost, and scored lower (82.8).

Each of these sends the specification to the company that runs it. What that company keeps is set out below, and using your own provider key covers the setup.

The best model on your own computer

If the specification must stay in the office, this is the model we recommend for the memory your computer can give it, with the trade-off each one carries. On a Mac that is the machine's memory; on a PC with a separate graphics card it is the card's memory.

Your machineModelScoreFinishedWorth knowing
A server with 512 GB of memory glm-5.2 88.410 of 10The strongest of these tiers. It needs server hardware, and it is the one tier without a one-command install. Its newer sibling glm-5.3 scored 0.9 higher, too close to call, but returns no draft unless a reasoning setting is passed, which your server has to support; if yours does, it is the one to try.
A workstation with 64 GB gemma4:31b 81.39 of 10Finished nine of the ten specifications tested, failing only the longest at about 240 pages. Worth knowing before you buy hardware: the 32 GB model below actually scores a little higher on our tests, and the gap is too small for our benchmark to tell apart, so more memory does not buy you a better local model until you reach the server tier. Pick this one if you want the most thoroughly tested option; pick the 32 GB one if you want the higher score in less memory.
A desktop or laptop with 32 GB qwen3.8:27b 85.58 of 9 triedMuch stronger drafting than the model this tier used to name, at the same memory. Two trade-offs, both measured on a 32 GB machine: it is slow, taking 2 to 12 minutes for one draft depending on length, and it could not finish a 240-page specification at all. It handled everything up to about 145 pages. For one longer than that, use gemma4:26b instead, which finished every specification tested but drafts less well.
A laptop with 16 GB gemma4:12b 67.39 of 10Marginal, and uneven across specifications. Workable for trying the tool; treat its output as a rougher first pass than the tiers above. It finished nine of the ten specifications tested; qwen3.5:9b scored 65.0 beside it, too close to call, and finished all ten, so use that for a specification this one refuses. On a PC, note that this model needs about 9 GB of graphics memory and the 9b about 14 GB on a long specification. Reading a parent's file history is lighter work than drafting, and there this tier is not marginal: qwen3.5:9b found every rejection in the prosecution records we measured it on.

Buying more memory does not buy a better draft until the server tier: most models in the middle keep many specialists in memory and use a few at a time, and the benchmark explains the measurement. With an 8 GB graphics card, nothing measured drafts a long specification at full speed. Set up a local model for the steps.

What to avoid, and why

Models that could not finish the specifications. Each of these completed fewer than eight of the ten, and the ones they skipped were usually the longest, so an average over the rest would flatter them: glm-4.7, mistral-large-2512, qwen3.5:4b, deepseek-r1:70b, devstral-small-2:24b, granite-4.1-8b.

A reasoning model left at its provider's default. Some spend their whole output allowance thinking and return no draft at all; the benchmark shows the setting that fixes it.

Choosing on a small difference in score. The same model measured twice varied by up to three points without changing, so a gap under 4 points is not a ranking. Coverage, speed and where the specification goes decide more than the last point of score.

Keeping a client's specification private

The best local model for a lawyer is the one that keeps an unpublished application on the lawyer's own computer. For every other route, this is what the provider holds.

For a continuation the specification is usually published already, so secrecy is often not the deciding reason. The practical one is that a model on your own computer adds nobody to the matter: no provider’s usage charge to account for on the client’s bill, and no further company handling the file.

How you run itWhat is kept
A model on your own computerNothing leaves the computer, as long as the model server runs on it too. The specification goes only to that server.
Claude through OpenRouterAnthropic keeps no prompts or answers by default, except on its Fable and Mythos models, which it keeps for 30 days for safety monitoring. The tool also asks for a prompt cache, which holds the start of each prompt, mostly the specification, in memory for up to an hour after its last use.
Claude with your own Anthropic keyNo prompt cache: the Anthropic endpoint this tool uses does not support one. Otherwise as above: nothing kept by default, except on Fable and Mythos models, kept for 30 days.
OpenAI, with your own OpenAI keyThe tool asks for the shortest prompt cache OpenAI offers, which depends on the model: minutes on older models, up to 24 hours on gpt-5.5, and 30 minutes after last use on gpt-5.6 and later. OpenAI says cached data may be kept encrypted in GPU-local storage.
OpenAI models through OpenRouterOpenRouter does not pass the cache setting on, so OpenAI's own default for the model applies.
Other models through OpenRouterThe model runs on whichever host OpenRouter chooses, and that host's policy decides what it keeps. One question can reach more than one host: in a test, four calls about one specification went to four different companies.

This is what running the tool causes a provider to keep. Each provider also has its own terms on logging, abuse monitoring and training, which depend on your account and are yours to check. Checked against each provider’s documentation on 29 September 2026.

Zero data retention is an arrangement on your own provider account, and setting up a provider key says how each offers it. A published patent is public, but what you ask about it and the claims you draft from it can still show a client’s strategy. Whether that, or an unpublished application, may go to a provider is a decision for you to make against your own professional obligations.

What these numbers are

Machines grading machines. They are useful for choosing between models and they are not a measure of whether a draft is worth filing, which is a practitioner's judgement. The one human check so far, a licensed practitioner reviewing three drafts blind, ranked them differently from the AI panel. The full benchmark has every model tested, the memory each needs, and the limits of the method.