V2 launch onlyPlans from $25/mo (was $35), 30% off the license, bigger token packs and a token giveaway.30% off the license, lower plans, token giveawayEnds in …See prices → See what's new in V2 · A rebuilt interface, Crew, Blueprint to C++, multiplayer play tests and in-editor updates Read the changelog →
Solutions

Local LLMs in Unreal Engine: what works today

You can run the AI on your own graphics card. No API key, no cost per message, and your prompts never leave your PC. Here is how to set it up, and an honest sheet of what small local models managed in our tests.

What you need

Model sizeRough VRAM
3B to 9B8 GB
12B to 14B12 to 16 GB
27B to 32B24 GB, or heavy offloading to system RAM

CPU-only works, but slowly. Local Studio in the plugin does most of the setup: it can install Ollama, start it, pull a model and point a slot at it with one click (Use in Architect).

What small local models can do

Local models are not as good as cloud models. That is true of every tool that uses them, and it is the trade for running privately on your own hardware. We tested 7 to 9B models on a laptop with a 12 GB GPU in a fresh Unreal Engine 5.7 project. It is a first sheet from a handful of prompts, not a benchmark.

TaskSmall local model (7 to 9B)
Answer an Unreal questionWorks, in about 3 minutes
Read your project ("what actors are in my level?")Works, in 2 to 5 minutes
Place or move things in the levelUnreliable: it acts, got distances wrong and still reported success
Build a materialFailed: it repeated the same step in a loop
Build Blueprint logicFailed: the chat stopped itself after repeated failed builds
  • Small models are good for questions and for reading your project.
  • Check anything they change: a small model can say "done" when the result is wrong.
  • For Blueprint logic and materials use a cloud model, or a larger local model.
  • We have not yet tested 27B and larger models on a 24 GB card, so we cannot promise either way.

Models worth trying

ModelWhy
ornith:9bSmall and agentic; runs on an 8 GB laptop GPU
granite4.1:8bCompact, built for tool use
qwen3:14bQwen has the best-proven tool calling for this work
mistral-nemo:12bReliable, long context, undemanding
gemma4:12bThinks before acting, understands images
qwen3.6:27bStronger; wants 24 GB

Being able to call tools is not the same as building well with them, so treat this as a starting list.

The context window trap

Ollama loads models with a small context window unless told otherwise, and the CoPilot's instructions alone are around 13,500 tokens. A small window silently cuts them off and the model gets confused about everything. The plugin handles it: Use in Architect points the slot at a copy of your model with a 32,768-token window, and a banner warns you whenever the running window is under 16,384.

Setting things up by hand? Raise the model's context window to at least 16,384, and 32,768 if you can. LM Studio works too: load a model and start its server (Developer tab, port 1234).

Frequently asked questions

Does anything leave my PC with a local model?

With a local model, your prompts go to the model on your own machine. Other features you turn on, such as cloud asset generation, are separate.

Which is better, Ollama or LM Studio?

Both work. Ollama has the most direct setup through Local Studio. LM Studio suits you if you already use it: load a model and start its server on port 1234.

Can I mix local and cloud models?

Yes. Each chat slot has its own model, so you can keep a local model for questions and use a cloud model for Blueprint logic.