Outlines stops your AI from ever writing broken JSON

Outlines structured generation blocks any token that would break your schema before the model can pick it. So the JSON comes out valid by construction, with no retry loop to clean up after a bad response. The catch is that the guarantee only covers shape, so the values inside can still be wrong.
Key Takeaways
- Outlines makes the model unable to write output that breaks your schema.
- You pass a Python type instead of begging for good JSON in the prompt.
- Simple choices, plain numbers, and nested objects all use ordinary Python types.
- Hosted APIs get a much weaker version of the guarantee than local models.
- A valid shape can still hold a wrong answer, so check the values.
What outlines structured generation does differently
A language model writes one token at a time. Nothing in that loop knows what a valid JSON object looks like. It has seen plenty of JSON, so it usually gets the braces right, and the usually is where your pipeline breaks.
There are two ways to close that gap. One is to let the model finish, parse the result, and ask again when parsing fails. Instructor is the best-known library in that camp. It checks a finished response against a Pydantic model and retries on error, which works most of the time.
Outlines takes the other route. At every decoding step it works out which tokens could still lead to a valid result, then hides the rest from the sampler. A token that would break the schema is never on the menu, so nothing malformed reaches you later. The failure mode disappears instead of getting rarer, so there is no retry loop to write.
| Validate and retry | Constrain while generating | |
|---|---|---|
| Example library | Instructor | Outlines |
| When the check runs | After the model finishes | At every token |
| Failure handling | Parse, catch, ask again | No failure to handle |
| Needs sampling access | No | Yes, for the full guarantee |
| Extra API calls on failure | One or more | None |
Masking tokens needs the model’s raw probabilities, so the guarantee is strongest where you run the model yourself. Aidan Cooper draws the same boundary in his guide to constrained decoding : the trick only works on models that hand you their full next-token distribution.
The interface mirrors Python’s own type system. You call model(prompt, output_type), and the type you pass becomes the constraint.
The types you will actually reach for
Start with a fixed choice. Literal["Yes", "No"] kills a whole class of bug. The model answers “Yes, because the invoice…” and your equality check fails. With a literal there’s no third option to produce.
Passing int gets you a number back instead of a sentence with a number buried in it, and saves you the regex you would write to dig out the digits.
Real work needs objects. A Pydantic model can hold an enum rating, two string lists, and a summary field. That turns a free-text product review into a row you can store:
from enum import Enum
from pydantic import BaseModel
class Rating(Enum):
poor = 1
fair = 2
good = 3
excellent = 4
class ProductReview(BaseModel):
rating: Rating
pros: list[str]
cons: list[str]
summary: str
review = model(prompt, ProductReview, max_new_tokens=200)Support triage is the classic production case. A ticket model carries priority, category, an escalation flag, and a list of action items. That feeds a queue directly, with no human reading the email first.
The most useful type is the union. Wrap your schema as Union[EventInfo, Literal["I don't know"]] and the model gets a legal way to admit it has nothing. That beats forcing it to invent a date.
Function calling uses the same machinery: pass a typed function, and Outlines reads the argument schema off the signature. Regex patterns and context-free grammars plug into the same call, and the core features docs cover both. A grammar is also how you pin a model to a wire format nobody wrote for it. OpenMMO makes that case plainly. Its agents send the same raw messages the browser client sends, with no tool layer in between.
Field names are prompt surface, so they change the answer. A study of schema key wording held the prompt, the model, and the structure fixed. Renaming the keys alone still moved accuracy. So keep the field count low and the names obvious.
Real schemas stay close to that advice. Across 10,000 collected from the wild, required, items, additionalProperties, and enum dominate, and the exotic keywords barely register.

How to get guaranteed-valid structured output from a model
Install it
Run pip install outlines. The package needs Python 3.10 or newer.
Connect a model
Use outlines.from_transformers(...) for a local model, or the matching constructor for your provider. The model object has the same shape either way.
import outlines
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
AutoTokenizer.from_pretrained(MODEL_NAME),
)Start with the simplest type
Ask for a classification and pass Literal["Positive", "Negative", "Neutral"] as the output type. The model has no way to answer anything else.
sentiment = model("Analyze: 'This changed my life!'", Literal["Positive", "Negative", "Neutral"])Try a number
Pass int when you want a numeric answer. You get 100 back instead of a sentence about boiling water.
Move to a real object
Define a Pydantic model with the fields you need, including enums and lists. Then pass the class itself as the output type.
Parse the result
The call returns a JSON string that matches your schema. Run ProductReview.model_validate_json(result) to get a typed object back.
Raise the token ceiling for larger objects
Set max_new_tokens high enough for the whole structure. An object that gets cut off mid-structure is useless, even though everything before the cut matched the schema.
Swap the model without touching the schema
Change the model constructor and leave the output types alone. If the types are your only contract, nothing else in the code should move.
Which models and providers it works with
The pitch is that the same code runs across local transformers models, llama.cpp, Ollama, vLLM, and hosted APIs. That holds for the code itself, though the guarantee behind it doesn’t travel as well.
Run the model yourself and generation happens inside the library object you handed to Outlines. It attaches a logits processor, sees every step, and masks tokens directly, so every output type is available. Point it at a hosted provider and it forwards your schema to that vendor’s own structured output feature, so you inherit whatever that vendor supports.
The official model integrations matrix spells out the gap, and it is wider than the README suggests:
| Provider | Simple types | JSON schema | Multiple choice | Regex | Grammar |
|---|---|---|---|---|---|
| Transformers (local) | Yes | Yes | Yes | Yes | Yes |
| vLLM | Yes | Yes | Yes | Yes | Yes |
| llama.cpp | Yes | Yes | Yes | Yes | No |
| SGLang | Yes | Yes | Yes | Yes | Partial |
| OpenAI | No | Yes | No | No | No |
| Ollama | No | Yes | No | No | No |
| Gemini | No | Partial | Yes | No | No |
| Anthropic | No | No | No | No | No |
So the sentiment classifier from the quickstart, built on a plain Literal, won’t run against OpenAI or Ollama, and against Anthropic nothing runs at all.
The project lists NVIDIA, Cohere, Hugging Face and vLLM among its users. vLLM’s own structured decoding grew out of the same lineage.
The 1.x line replaced the old generate.json() and generate.choice() helpers with a single Generator object and the model(prompt, output_type) call. Any tutorial written against 0.x will not run as written
.
Where a schema guarantee stops helping
A valid schema and a correct answer are different things. A model can hand you a flawless object full of wrong values, and the polish makes that worse. Errors that would have been obvious as garbled text now arrive as clean, confident data that walks straight into your database.
Enums sharpen the problem. Force a choice, and a model with no good answer picks the least bad one, which then reads exactly like knowledge. Give the schema somewhere to put doubt: an optional field, a confidence score, or an explicit I don't know branch. Some models now build that doubt in. TypeSafe’s Jev
hands back a confidence value beside every answer, and it still cannot promise the answer is right.
There is a measured quality cost too. Blocking tokens at every step can steer the model down a path that is valid but wrong. The team behind draft-conditioned constrained decoding put a 1B model on GSM8K. Plain constrained decoding scored 15.2%, while letting the model draft freely first and constraining afterwards reached 39.0%. Bigger models suffer less, though the effect isn’t zero.
JSONSchemaBench ran six frameworks against 10,000 real-world schemas. Grammar compilation alone cost Outlines 3.5 to 8 seconds before the first token appeared. Each output token then took 30 to 47ms, against 15 to 17ms with no constraint at all.
Compilation runs once per schema, so reusing a Generator spreads that cost. A fresh schema on every request pays it again.
The same paper ran every framework through the official JSON Schema test suite, keyword by keyword. Outlines passes the common ones and scores zero on plenty of the rest, including pattern, minLength, and maxItems. Keep your schemas plain and you never meet that wall.

Outlines is Apache-2.0 licensed, sits above 15,000 GitHub stars, carries roughly 90 open issues, and ships regularly on the 1.3.x line. The company behind it, .txt , sells reliability tooling for LLMs and funds the open-source work.
Reach for it when the output shape is a hard requirement, and a malformed response would break something downstream. Then add your own value checks on top, since nothing in the schema tells you whether the answer is right.
Botmonster Tech