Continuation Drafter
Step one

Choose a model for your machine

Model choice matters more than it looks, and memory is the first filter. Pick the card that matches your hardware.

These sizes are unified memory, measured on a Mac, where the processor and the graphics chip share one pool. On a PC with a separate graphics card, the card's own memory is the number that decides which model runs at a usable speed: a model larger than the card spills onto the processor, which is many times slower, and a draft that takes minutes on a GPU can time out. With an 8 GB card, use the 16 GB card's models regardless of how much system memory the machine has.

A server with 512 GB of memory glm-5.2
Score 88.4 Drafted 10 of 10 roughly 450 GB at 4-bit, estimated from its size

The strongest of these tiers. It needs server hardware, and it is the one tier without a one-command install. Its newer sibling glm-5.3 scored 0.9 higher, too close to call, but returns no draft unless a reasoning setting is passed, which your server has to support; if yours does, it is the one to try.

A workstation with 64 GB gemma4:31b
Score 81.3 Drafted 9 of 10 about 30 GB, flat

Finished nine of the ten specifications tested, failing only the longest at about 240 pages. Worth knowing before you buy hardware: the 32 GB model below actually scores a little higher on our tests, and the gap is too small for our benchmark to tell apart, so more memory does not buy you a better local model until you reach the server tier. Pick this one if you want the most thoroughly tested option; pick the 32 GB one if you want the higher score in less memory.

A desktop or laptop with 32 GB qwen3.8:27b
Score 85.5 Drafted 8 of 9 tried about 20 GB, flat

Much stronger drafting than the model this tier used to name, at the same memory. Two trade-offs, both measured on a 32 GB machine: it is slow, taking 2 to 12 minutes for one draft depending on length, and it could not finish a 240-page specification at all. It handled everything up to about 145 pages. For one longer than that, use gemma4:26b instead, which finished every specification tested but drafts less well.

A laptop with 16 GB gemma4:12b
Score 67.3 Drafted 9 of 10 about 12 GB, flat

Marginal, and uneven across specifications. Workable for trying the tool; treat its output as a rougher first pass than the tiers above. It finished nine of the ten specifications tested; qwen3.5:9b scored 65.0 beside it, too close to call, and finished all ten, so use that for a specification this one refuses. On a PC, note that this model needs about 9 GB of graphics memory and the 9b about 14 GB on a long specification. Reading a parent's file history is lighter work than drafting, and there this tier is not marginal: qwen3.5:9b found every rejection in the prosecution records we measured it on.

Each card's index and drafted figure comes from the same benchmark run: 51 models over ten real patent families, with what the numbers do and do not support stated there.

Setting up gemma4:12b

Four commands and you are drafting

Install Ollama

Get it from ollama.com/download and start it. The desktop app starts the server for you; otherwise run ollama serve.

Pull the model

A large download, once.

ollama pull gemma4:12b

Get the binary

From the download page, then make it executable. On macOS, clear the quarantine mark a browser download carries, or the first run is stopped with no message: the download page gives that command too.

Draft

Point it at your specification and the filed parent claims.

DRAFT_MODEL=gemma4:12b ./continuation-drafter draft \
  --spec spec.txt --parent parent-claims.txt

One caveat for this model

It did not finish the longest specification tested, about 240 pages. If yours is near that long, use gemma4:26b instead, which finished every one.

If a model instead returns nothing at all after a long wait, it is a thinking model that has spent its whole budget deliberating before writing anything. Add --reasoning-effort high. Raising --max-tokens does not help, and makes it slower.

Not enough memory, or would rather not install a model?

Run it through your own provider account instead. That sends the specification to that provider, so it suits published material, and for anything unpublished it is a judgment you make per matter. How that works.