What you need
| Model size | Rough VRAM |
|---|---|
| 3B to 9B | 8 GB |
| 12B to 14B | 12 to 16 GB |
| 27B to 32B | 24 GB, or heavy offloading to system RAM |
CPU-only works, but slowly. Local Studio in the plugin does most of the setup: it can install Ollama, start it, pull a model and point a slot at it with one click (Use in Architect).
What small local models can do
Local models are not as good as cloud models. That is true of every tool that uses them, and it is the trade for running privately on your own hardware. We tested 7 to 9B models on a laptop with a 12 GB GPU in a fresh Unreal Engine 5.7 project. It is a first sheet from a handful of prompts, not a benchmark.
| Task | Small local model (7 to 9B) |
|---|---|
| Answer an Unreal question | Works, in about 3 minutes |
| Read your project ("what actors are in my level?") | Works, in 2 to 5 minutes |
| Place or move things in the level | Unreliable: it acts, got distances wrong and still reported success |
| Build a material | Failed: it repeated the same step in a loop |
| Build Blueprint logic | Failed: the chat stopped itself after repeated failed builds |
- Small models are good for questions and for reading your project.
- Check anything they change: a small model can say "done" when the result is wrong.
- For Blueprint logic and materials use a cloud model, or a larger local model.
- We have not yet tested 27B and larger models on a 24 GB card, so we cannot promise either way.
Models worth trying
| Model | Why |
|---|---|
| ornith:9b | Small and agentic; runs on an 8 GB laptop GPU |
| granite4.1:8b | Compact, built for tool use |
| qwen3:14b | Qwen has the best-proven tool calling for this work |
| mistral-nemo:12b | Reliable, long context, undemanding |
| gemma4:12b | Thinks before acting, understands images |
| qwen3.6:27b | Stronger; wants 24 GB |
Being able to call tools is not the same as building well with them, so treat this as a starting list.
The context window trap
Ollama loads models with a small context window unless told otherwise, and the CoPilot's instructions alone are around 13,500 tokens. A small window silently cuts them off and the model gets confused about everything. The plugin handles it: Use in Architect points the slot at a copy of your model with a 32,768-token window, and a banner warns you whenever the running window is under 16,384.
Setting things up by hand? Raise the model's context window to at least 16,384, and 32,768 if you can. LM Studio works too: load a model and start its server (Developer tab, port 1234).
Frequently asked questions
Does anything leave my PC with a local model?
With a local model, your prompts go to the model on your own machine. Other features you turn on, such as cloud asset generation, are separate.
Which is better, Ollama or LM Studio?
Both work. Ollama has the most direct setup through Local Studio. LM Studio suits you if you already use it: load a model and start its server on port 1234.
Can I mix local and cloud models?
Yes. Each chat slot has its own model, so you can keep a local model for questions and use a cloud model for Blueprint logic.