How do you write patent claims for a machine learning invention?
Claim the mechanism in the training or serving pipeline, keep training and inference claims apart, and decide what stays secret before anything is written up.
Publisher: patentagents.ai
Short answer
Claim the specific computation that makes the model or pipeline work better, written as method steps, and restate those steps as system and storage-medium claims. Back each claimed step with a write-up that explains how the improvement happens in enough detail for another engineer to rebuild it.
What part of an ML system is worth claiming?
Most ML products have three layers: the interface, the model and its weights, and the pipeline that trains, serves, and feeds the model. The pipeline holds more of the engineering choices: data preparation, training updates, and inference scheduling, caching, and batching.
Ask what the system does differently in tensors, memory, or compute. "Faster answers" is an outcome; "reuses stored attention keys for repeated document chunks instead of recomputing them" is a mechanism. Claims are written around mechanisms.
Can you claim a model architecture, or only how it is used?
Section 101 covers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement of one. An architecture change is usually claimed as the operations a computer performs when it runs the model, or as a system configured to perform them. A new attention variant becomes steps on inputs, keys, values, and outputs.
That reaches the model itself, not only its uses. A December 5, 2025 USPTO memo to examiners adds, as an example that may show an improvement in computer functionality, an improved way of training a machine learning model that protected earlier-task knowledge while learning new tasks. The example comes from Ex parte Desjardins, a precedential Appeals Review Panel decision the December memo summarizes.
What makes an ML claim read as a technical improvement?
That December 5 examiner memo restates the guidance in two steps. First, the specification should describe the invention so a skilled reader would recognize the improvement; a bare assertion without that detail does not count. Second, the claim must include the components or steps that provide the improvement, though it need not state the improvement in words.
The training claims did recite a mathematical concept. What carried them was a specification explaining how the model itself operated better (learning new tasks while protecting earlier ones) and a claim containing the parameter-adjustment step that produced that result.
Write down the mechanism and why it produces the result, then check that the claimed steps contain it. "Updating parameters for a new task while limiting changes to parameters that mattered for the old one" names a mechanism. The memo guides examiners and does not predict outcomes.
Which claim do you draft first: method, system, or training data?
Most ML inventions start with the method claim, because the invention is a sequence of operations. System and storage-medium claims restate those steps as processors, memory, and stored instructions.
Keep training and inference apart. They often run on different hardware operated by different parties, so a single claim that requires both may describe a pipeline no one system runs end to end. Separate independent claims can reduce that problem when each claim's required operations match the relevant implementation and actor allocation.
Training data is harder. Section 101 lists processes, machines, manufactures, and compositions of matter; a dataset by itself is information. Ideas about data are usually written as a method that selects, labels, filters, or generates data and then trains on it.
What does a worked example claim set look like?
These claims are invented for this article and are not taken from any real application or patent. They cover a hypothetical transformer serving system using rotary position encoding.
1. A computer-implemented method comprising: receiving a prompt of text segments for a transformer that applies rotary position encoding to attention keys; retrieving cached keys and values computed for a first segment at a stored offset; rotating the cached keys by the shift from that offset to the first segment's position; recomputing keys and values for boundary tokens at the start of the first segment with attention to preceding segments; and generating an output token from the rotated keys, cached values, and recomputed keys and values.
2. The method of claim 1, wherein the cache is keyed by a hash of the first segment's token identifiers and a model-version identifier.
3. The method of claim 1, wherein the number of boundary tokens is a configured value smaller than the length of the first segment.
4. A system comprising one or more processors and memory storing instructions that, when executed, cause the system to perform the method of claim 1.
What is honestly new in that example?
Claim 1 reuses cached attention state when a chunk of text reappears at a new place in a prompt. Caching keys and values for repeated prefixes is a known serving technique; re-rotating keys for a new position may be known too. Claim 1 may not be new as written; no search has been run.
The candidate contribution is the boundary recomputation: correcting a few tokens at the start of a reused chunk, on the theory that those tokens depend most on the text before them. If that is the real advance, useful drafting questions include why those tokens, how many, and what happens to output quality without them. Required detail depends on the claimed invention and on written description, enablement, and best mode. Claim 3 adds only a boundary-token count limit: a configured value smaller than the segment length.
Claim 2 has a practical job: a cache that ignores the model version could serve stale state unless invalidated or otherwise separated after an update. The set leaves out the chat interface, the product name, and outcomes stated in place of steps.
Should weights and the training pipeline stay secret instead?
That depends on what outsiders can see. Federal trade secret law covers technical information when the owner takes reasonable measures to keep it secret and it derives economic value from not being generally known or readily ascertainable through proper means. Trained weights that never leave your servers can fit that description.
Under 35 U.S.C. 122(b), applications are generally published promptly after 18 months from the earliest filing date claimed, so assume the application will become public. A method others could work out from product behavior is a candidate for claims; weights nobody outside can inspect may stay private.
Demos raise a separate question. Section 102(a)(1) counts public use and anything otherwise available to the public before the effective filing date of the claimed invention. Section 102(b)(1) sets aside some of the inventor's own disclosures made one year or less before the effective filing date of the claimed invention. Other countries set their own rules. Record what a demo showed, to whom, and when.
What belongs in the write-up, and what stays out?
Section 112(a) asks for a written description of the invention and of how to make and use it, clear enough that a skilled person could do so. For ML, "a trained model scores the request" says little. Say what goes in and its shape, which component transforms it, what comes out, and whether the change is in training or inference.
Include alternatives you considered, the failure the change fixes, and how you measured the effect. The memo's first step depends on that kind of detail; a bare assertion of improvement does not count.
Customer prompts and credentials usually do not belong in the write-up. Leave out trained weights, raw training data, and other material you hope to keep secret only when they are not needed for written description, enablement, or best mode under section 112(a). Secrecy preferences do not override those requirements. Resolve any conflict between secrecy and disclosure before deciding what to claim and file. Where you omit raw data, describe the kind of data and how it was prepared.
Where does the write-up go next?
Send the write-up, diagrams, and code context to whoever prepares the application, with open questions marked. PatentAgents.ai can read that set, start an editable claim draft and specification draft, and export either as a Word .docx. Filing is for you and that person to decide.
ML claim drafting flow
- 1Find the mechanism in the pipeline
- 2Write it as method steps
- 3Restate it as system claims
- 4Decide what stays secret
Sources
- 35 U.S.C. 101 (inventions patentable)
- 35 U.S.C. 102 (novelty and exceptions)
- 35 U.S.C. 112 (specification)
- 35 U.S.C. 122 (confidentiality and publication of applications)
- 18 U.S.C. 1839 (definitions, including trade secret)
- USPTO Appeals Review Panel machine learning training claims (precedential designation notice, November 2025)
- USPTO updates subject matter eligibility guidance in the MPEP (December 2025)
Tradeoffs and limits
- Writing the mechanism out takes engineering time, and claims depend on it.
- Everything in an application may become public, so a fuller write-up can expose once-secret methods.
- Separate training and inference claims make the set longer, and each must stand on its own.