How to use Laya for local game AI in Unity

Laya is an open source, Apache-2.0 offline decision model (421M) that runs on GPU, Mac, or CPU with under 1.5 GB RAM. We show its hardware needs, speed, accuracy, and how to use it as a Unity RTS AI commander in C#.

By Tim UhlottFounder|Last updated: October 5, 2026|28 minutes read
game developmentailocal
How to use Laya for local game AI in Unity
In our last article we looked at Game Development with Jev, the hosted "decision model" from TypeSafe AI. Jev operates exclusively on TypeSafe's servers and is accessible only via their API. This means it cannot be used for offline games. For offline games, Jev isn’t an option yet. Laya is the open source answer here. It is a free, Apache-2.0 decision model from Convai Innovations that does the same job (state in, typed decisions out), but runs on your player's PC, on a 4 GB graphics card. Let's have a look at what Laya is, what hardware it needs, how fast and accurate it is, and how to wire it into a Unity game. Laya is a Python project, but it runs as a small local service and everything else is C#. Our example throughout is a real-time strategy game we will call Ironfront.

What Laya is

Laya (opens in a new tab) was released on September 18, 2026, three days after Jev, by Convai Innovations (opens in a new tab), a small company from Kerala, India. Within two weeks it passed 20,000 GitHub stars. Laya reads a short description of a situation and answers typed questions about it in one pass through a small network. It never writes text, same as Jev. It picks options (Choice), rates on a scale you define (Score), or returns a yes/no probability (Noul). Those are Jev's three question types on purpose: Laya's server speaks Jev's wire format, POST /v1/systemone, so a Jev client only needs a new URL.
CheckpointBase modelParametersContextGood for
layaModernBERT-large421M512 tokensEnglish. The default. About 800 MB
laya-multilingualmmBERT-base322M1,024 tokens (up to 8,192)100+ languages, about twice as fast. About 650 MB
Under the hood, Laya is a text encoder with a small decision head that scores one marker token per option. Unlike TypeSafe, Convai published the training code and a notebook that runs on free Kaggle GPUs. For game developers, that notebook is the most interesting part of the project.

Laya versus Jev at a glance

Jev 1.13 (TypeSafe)Laya (Convai)
How you get itHosted API, $0.042 per million input tokenspip install laya, free, Apache 2.0
Runs offlineNoYes, under 1.5 GB of VRAM or RAM
Latency, one question70 to 500 ms in the US, 600 to 700 ms from Europe5 to 40 ms on a GPU or Mac, 200 to 600 ms on CPU
Zero-shot accuracyHigher on most independent testsLower, often much lower
Fine-tuningNot offeredYes, with a published notebook
DeterministicNo, about 1% of answers change on rerunYes, on the same hardware and precision
Laya wins "runs offline" by default and Jev wins "zero-shot accuracy" clearly. The rest of this article is about making that second row not matter.

What hardware you need

Decision models like Laya are not so memory hungry as LLMs like ChatGPT. Laya requires about 1.3 GB for one request, so a 4 GB graphics card, a 16 GB MacBook Air, or 3 GB of free RAM with no GPU is enough. But you have to look at the speed when you choose running on GPU or CPU. These are public median latencies from Convai's benchmarks (opens in a new tab) and community ports already have done that for you:
HardwareRuntimeOne questionSeveral in one call
Tesla T4 (2018 datacenter card)PyTorch, English / multilingual39.5 / 32.8 ms10 questions: 158.6 / 72.3 ms
RTX 4070 Ti SUPERPyTorch, stock / laya[fast] kernels17.7 / 4.6 ms30 questions: 43.2 / 35.7 ms
Apple M3 MaxMLX port (opens in a new tab), English / multilingual13.4 / 7.4 ms50 questions: 340 / 127 ms
AMD EPYC, 4 cores, no GPUPyTorch, English / multilingual580 / 193 ms10 questions: 6,244 / 1,842 ms
On any GPU, Laya needs 5 to 40 ms, less than a frame at 30 FPS, where Jev's best case was 70 ms. Batching is nearly free on a GPU and linear on a CPU. The multilingual checkpoint is about twice as fast everywhere, at a small cost in English accuracy. For our example game Ironfront, a commander decision every 10 to 15 seconds plus threat ratings and squad orders add up to one or two calls per second with 3 to 10 questions each. Any GPU from the last six years handles that; CPU only is enough for the commander alone. And since the GPU is also drawing your game (Convai saw 100 ms instead of 40 on a shared GPU), measure on your minimum spec with the renderer running.
Takeaway: 1 GB of memory, any GPU, a few milliseconds per decision. On a CPU it works for slow decisions only.

How fast is it compared to Jev?

Independent head-to-heads agree with Convai's "7.8x faster" headline for small calls. laya-jev-lab (opens in a new tab) measured 7.6 ms on an M4 Max against 588 ms for Jev. Two Tetris matches (atomic.chat (opens in a new tab) and one posted on X) had Laya deciding ten times faster, 21.8 ms per move versus 220.9 ms, and clearing more lines before dying. The browser Tetris demo (opens in a new tab) lists every legal placement and asks Laya to pick one, the same pattern we use for Ironfront. The catch is big batches. On an M3 Pro, 50 questions on one state took Laya 1,002 ms and Jev 170 ms, because Jev's servers batch well and a Mac does not. For 3 to 10 questions per call Laya wins; for 200, Jev wins.

How accurate is it?

Convai's README says it plainly: "Laya is a fast base to specialise, not a zero-shot decision engine." On the Decision Index (opens in a new tab), 38 benchmarks where 0 is random and 100 is perfect, Laya scores 6.04 and Jev 57.91. On JevBench (opens in a new tab) Laya ranks 41st, Jev 2nd. Zero-shot, Laya is close to random on broad tests. Fine-tuning changes that. On Convai's typed-decisions benchmark the base model scored 0.362, the fine-tuned one 0.766, above Jev's zero-shot 0.727. A community browser agent went from 0.10 to 0.66 at picking the right web element out of 45, on one 16 GB GPU. Not a fair fight, since the fine-tuned model saw the training split, but for a game developer that is the point: you control the task and the data. Calibration is the same story. As shipped, the English checkpoint is badly overconfident (calibration error 0.466). After fitting one temperature per question type on a few hundred labeled examples, which the notebook does, it drops to 0.081. Do not use a confidence threshold before that step. Pitfalls from the docs and the game demos:
PitfallWhat to do
Numbers. In the Pong demo the model "never moved when the state was three numbers"Write "our army is much smaller than theirs", not army_value: 1420
Order and count. At 20 options, shuffling flipped 15% of answersShort phrases, short lists, fixed order
yes/no keys pull the answer toward the label textUse game verbs like attack and hold
512 tokens total, about 320 to 460 for your stateTwo paragraphs at most. Never send unit lists
Takeaway: Zero-shot, Laya is not Jev. Fine-tuned on your own game it can beat Jev on your own game, in 20 ms with no bill. Plan the fine-tuning step from day one.

How to use Laya in a Unity game

Laya is a Python project that, when started, brings up a small local service. However, this does not mean your users have to install or run Python. For local development, I recommend using the Python project locally, and when you ship, you can include the built project in your game as an executable.

Step 1: run Laya as a local service

During development, install the official Laya server. It answers on POST /v1/systemone and reports GET /health:
pip install "laya[serve]" LAYA_HOST=127.0.0.1 LAYA_MODELS=english LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve
The first launch downloads the checkpoint; after that, it works offline. On NVIDIA, running pip install "laya[fast]" adds kernels that improve performance on NVIDIA hardware from 17.7 ms to 4.6 ms. For shipping, put the community ggmlc laya executable and a GGUF file in StreamingAssets. The player does not install Python. That binary answers on POST /api/decide. This is for Windows, macOS, and Linux standalone builds. Consoles, mobile, and WebGL cannot start a sidecar process. Here is what the game sends, in Ironfront terms:
{ "state": { "time": "early game, about 6 minutes in", "our_bases": "one main base, expansion under construction", "our_army": "small, mostly infantry, a few tanks. Much smaller than the enemy's", "enemy": "scout saw a second enemy base and a barracks with many infantry", "resources": "ore is low, energy is fine", "last_enemy_action": "two scouts probed our east wall and left" }, "questions": { "plan": { "type": "choice", "instructions": "Pick the best plan for the next minute.", "criteria": { "defend_and_expand": "finish the expansion and keep the army home", "tech_up": "build the factory upgrade before more units", "build_army": "spend everything on more combat units", "attack_now": "send the current army at the nearest enemy base", "harass_economy": "send a small fast group to kill enemy harvesters" } }, "threat": { "type": "score", "instructions": "How much danger is our main base in right now?", "criteria": ["safe", "being watched", "attack likely soon", "under attack"] }, "enemy_rushing": { "type": "noul", "instructions": "The enemy is preparing an early rush.", "criteria": { "true": "the enemy is massing cheap units for an early attack", "false": "the enemy is playing a normal, slower build" } } } }
And what comes back, about 40 ms later on a T4 class GPU:
{ "answers": { "plan": { "type": "choice", "choice": "defend_and_expand", "probabilities": { "defend_and_expand": 0.47, "tech_up": 0.09, "build_army": 0.31, "attack_now": 0.04, "harass_economy": 0.09 }, "confidence": 0.38, "answer_confidence": 0.47 }, "threat": { "type": "score", "score": 1.6, "probabilities": { "0": 0.10, "1": 0.35, "2": 0.40, "3": 0.15 }, "confidence": 0.29 }, "enemy_rushing": { "type": "noul", "noul": 0.62 } }, "routing": { "model": "english" }, "usage": { "input_tokens": 231, "output_tokens": 0 } }
Every state value is words, every option is a phrase, and the Noul has its own criteria text, because without it the English checkpoint tends to answer "no" regardless of state. In the response, read answer_confidence (the probability of the chosen option), not confidence, which Convai says does not transfer between option counts.

Step 2: typed questions in C#

Three helper functions keep questions readable. They build plain dictionaries instead of anonymous types, so IL2CPP stripping won't break Newtonsoft serialization.
using System.Collections.Generic; public static class LayaQ { public static Dictionary<string, object> Choice(string instructions, IReadOnlyDictionary<string, string> options) => new Dictionary<string, object> { ["type"] = "choice", ["instructions"] = instructions, ["criteria"] = options }; public static Dictionary<string, object> Score(string instructions, params string[] levelsLowToHigh) => new Dictionary<string, object> { ["type"] = "score", ["instructions"] = instructions, ["criteria"] = levelsLowToHigh }; public static Dictionary<string, object> Noul(string statement, string ifTrue, string ifFalse) => new Dictionary<string, object> { ["type"] = "noul", ["instructions"] = statement, ["criteria"] = new Dictionary<string, string> { ["true"] = ifTrue, ["false"] = ifFalse } }; }

Step 3: the client

One HttpClient call with a short timeout. It returns null on any failure, and the caller treats null as "run the old scripted AI this tick". HttpClient and Newtonsoft (com.unity.nuget.newtonsoft-json) both work on Unity's .NET Standard 2.1 profile for desktop players.
using System; using System.Net.Http; using System.Text; using System.Threading; using System.Threading.Tasks; using System.Collections.Generic; using Newtonsoft.Json; using Newtonsoft.Json.Linq; public sealed class LayaClient { private static readonly HttpClient Http = new HttpClient(); private readonly string _endpoint; // Same wire format as TypeSafe's Jev. If a server build ever wants a hosted model instead, // this URL is the only thing that changes. public LayaClient(string endpoint = "http://127.0.0.1:8000/v1/systemone") { _endpoint = endpoint; } // Returns the "answers" object, or null if the sidecar did not answer in time. public async Task<JObject> AskAsync(object state, Dictionary<string, object> questions, int timeoutMs = 250) { var body = new JObject { ["state"] = JToken.FromObject(state), ["questions"] = JToken.FromObject(questions) }; using var cts = new CancellationTokenSource(timeoutMs); using var content = new StringContent(body.ToString(Formatting.None), Encoding.UTF8, "application/json"); try { var response = await Http.PostAsync(_endpoint, content, cts.Token); if (!response.IsSuccessStatusCode) return null; var json = JObject.Parse(await response.Content.ReadAsStringAsync()); return json["answers"] as JObject; } catch (Exception) { return null; // timeout, sidecar not up yet, or a bad response: the caller falls back } } } public static class LayaAnswers { public static string Choice(JObject answers, string id) => (string)answers[id]?["choice"]; public static float AnswerConfidence(JObject answers, string id) => (float?)answers[id]?["answer_confidence"] ?? 0f; public static float Score(JObject answers, string id) => (float?)answers[id]?["score"] ?? 0f; public static float Noul(JObject answers, string id) => (float?)answers[id]?["noul"] ?? 0f; }
The timeout is 250 ms, not Jev's 600 ms. The server is local, so a slow answer means something is wrong, and the right move is to fall back, not to wait.

Step 4: starting and stopping the sidecar from Unity

A small MonoBehaviour starts the sidecar when the match loads, waits for GET /health, and kills it on quit. System.Diagnostics.Process works in Windows, macOS, and Linux standalone builds, not on consoles, mobile, or WebGL.
using System.Diagnostics; using System.Net.Http; using System.Threading.Tasks; using UnityEngine; public sealed class LayaSidecar : MonoBehaviour { // During development this is "laya-serve" from your Python environment. // For shipping, point it at the ggmlc "laya" binary in StreamingAssets and pass // "serve laya_english_q8_0.gguf --port 8000" as arguments (and adapt the client route). [SerializeField] private string executable = "laya-serve"; [SerializeField] private string arguments = ""; [SerializeField] private int threadsIfCpu = 4; // physical cores, not logical. See the hardware section public bool Ready { get; private set; } private Process _process; private async void Start() { var psi = new ProcessStartInfo { FileName = executable, Arguments = arguments, UseShellExecute = false, CreateNoWindow = true, }; psi.Environment["LAYA_HOST"] = "127.0.0.1"; psi.Environment["LAYA_MODELS"] = "english"; psi.Environment["LAYA_PRELOAD"] = "1"; psi.Environment["LAYA_THREADS"] = threadsIfCpu.ToString(); _process = Process.Start(psi); // Cold start is a few seconds on a GPU and longer on a CPU. Poll health, do not guess. using var http = new HttpClient(); for (int i = 0; i < 120 && !Ready; i++) { try { var r = await http.GetAsync("http://127.0.0.1:8000/health"); Ready = r.IsSuccessStatusCode; } catch { /* not listening yet */ } if (!Ready) await Task.Delay(500); } } private void OnApplicationQuit() { if (_process != null && !_process.HasExited) _process.Kill(); } }
Until Ready is true, the commander runs your scripted AI. Nothing in the game ever waits on Laya.

Step 5: the commander

The Ironfront commander runs on a 12-second timer: build the state in words, build the legal plan list, ask, validate, fall back. After an await in a MonoBehaviour, Unity's synchronization context brings you back to the main thread, so touching game objects after the call is safe.
using System.Collections; using System.Collections.Generic; using System.Threading.Tasks; using Newtonsoft.Json.Linq; using UnityEngine; public sealed class AICommander : MonoBehaviour { [SerializeField] private LayaSidecar sidecar; [SerializeField] private float intervalSeconds = 12f; // Measure this after fitting temperatures on your own data. The shipped checkpoints are // overconfident, so do not copy a number from a blog post, including this one. [SerializeField, Range(0f, 1f)] private float confidenceThreshold = 0.5f; private LayaClient _laya; private ScriptedCommander _fallback; // the AI you already have private BuildOrders _buildOrders; private BaseDefense _defense; private MatchReader _match; // whatever reads your match state private void Awake() { _laya = new LayaClient(); } private void Start() { StartCoroutine(Loop()); } private IEnumerator Loop() { var wait = new WaitForSeconds(intervalSeconds); while (true) { yield return wait; if (sidecar.Ready) _ = ThinkAsync(); else _fallback.Tick(); } } private async Task ThinkAsync() { // 1. Words, not numbers. The game does the math and writes the sentence. var state = new Dictionary<string, string> { ["time"] = _match.PhaseInWords(), // "early game, about 6 minutes in" ["our_bases"] = _match.BasesInWords(), // "one main base, expansion under construction" ["our_army"] = _match.ArmyComparedToEnemyInWords(), // "much smaller than the enemy's, mostly infantry" ["enemy"] = _match.LatestScoutReportInWords(), // "a second enemy base and a barracks with many infantry" ["resources"] = _match.ResourcesInWords(), // "ore is low, energy is fine" ["last_enemy_action"] = _match.LastEnemyActionInWords() }; // 2. Only offer plans the AI can actually execute right now. var legal = new Dictionary<string, string>(); if (_match.CanAffordExpansion) legal["defend_and_expand"] = "finish the expansion and keep the army home"; if (_match.FactoryUpgradeReady) legal["tech_up"] = "build the factory upgrade before more units"; legal["build_army"] = "spend everything on more combat units"; if (_match.HasArmy) legal["attack_now"] = "send the current army at the nearest enemy base"; if (_match.HasFastUnits) legal["harass_economy"] = "send a small fast group to kill enemy harvesters"; var questions = new Dictionary<string, object> { ["plan"] = LayaQ.Choice("Pick the best plan for the next minute.", legal), ["threat"] = LayaQ.Score("How much danger is our main base in right now?", "safe", "being watched", "attack likely soon", "under attack"), ["enemy_rushing"] = LayaQ.Noul("The enemy is preparing an early rush.", "the enemy is massing cheap units for an early attack", "the enemy is playing a normal, slower build") }; // 3. Ask. The previous plan keeps running while this is in flight. JObject answers = await _laya.AskAsync(state, questions); // 4. Validate, then fall back to the scripted commander if anything looks off. string plan = answers != null ? LayaAnswers.Choice(answers, "plan") : null; if (plan == null || !legal.ContainsKey(plan) || LayaAnswers.AnswerConfidence(answers, "plan") < confidenceThreshold) { _fallback.Tick(); return; } // 5. Hand the plan to the systems that already know how to do it. _buildOrders.Switch(plan); _defense.SetAlertLevel(LayaAnswers.Score(answers, "threat")); if (LayaAnswers.Noul(answers, "enemy_rushing") > 0.7f) _buildOrders.QueueEmergencyDefenses(); } }
Every helper ending in InWords() is where the real work is. Writing good state sentences is the design task; the model is the easy part. Commander personalities are just different instructions on the same question.

Step 6: bases and squads in the same call

A commander, three bases, and four squads are eight questions. On a GPU they share one forward pass: about 70 to 160 ms on a T4, well under 40 ms on a modern desktop card.
private Dictionary<string, object> OperationsQuestions(IEnumerable<PlayerBase> bases, IEnumerable<Squad> squads) { var questions = new Dictionary<string, object>(); foreach (var b in bases) { questions[$"threat_{b.Id}"] = LayaQ.Score( $"How much danger is the {b.NameInWords} in right now?", "safe", "being watched", "attack likely soon", "under attack"); } foreach (var s in squads) { // Legal orders only. If there is nowhere to retreat to, pull_back is not on the list. var legal = new Dictionary<string, string>(); if (s.HasVisibleEnemy) legal["engage"] = "attack the enemy group in front of us"; legal["hold"] = "hold this position"; if (s.HasRetreatPoint) legal["pull_back"] = "fall back to the nearest friendly base"; if (s.HasFastUnits) legal["flank"] = "circle around and hit the enemy from the side"; questions[$"order_{s.Id}"] = LayaQ.Choice($"Pick the best order for the {s.NameInWords}.", legal); } return questions; }
Read them back by id with LayaAnswers.Choice(answers, $"order_{s.Id}") and validate each squad's pick against its own legal list. One catch: all questions share one state of roughly 320 to 460 tokens. One sentence per base and squad fits for a handful of them, not for forty.

Step 7: recording training data from Unity

This is where Laya pulls ahead of Jev, and it is a C# job too. Every playtest is a sequence of states and human decisions. Every 15 seconds, write the state sentences the commander would see, the legal plan list, and the plan the human actually executed next. One JSON line per decision.
using System.Collections.Generic; using System.IO; using Newtonsoft.Json; using Newtonsoft.Json.Linq; public sealed class DecisionRecorder : System.IDisposable { private readonly StreamWriter _writer; public DecisionRecorder(string path) { _writer = new StreamWriter(path, append: true); } public void Record(Dictionary<string, string> state, Dictionary<string, string> legalPlans, string planThePlayerChose) { var row = new JObject { ["state"] = JToken.FromObject(state), ["questions"] = JToken.FromObject(new Dictionary<string, object> { ["plan"] = LayaQ.Choice("Pick the best plan for the next minute.", legalPlans) }), ["labels"] = new JObject { ["plan"] = planThePlayerChose } }; _writer.WriteLine(row.ToString(Formatting.None)); } public void Dispose() => _writer.Dispose(); }
Two thousand replays is tens of thousands of human commander decisions in exactly the sentence format your InWords() helpers produce at runtime. Write those helpers once and use them in both places. Training is the one Python step, on a build machine or a free Kaggle account. Convai's laya_finetune_typed_decisions_2xT4_kaggle.ipynb takes rows of state, question, and label, trains, fits calibration temperatures, and evaluates in 4 to 5 hours on Kaggle's free pair of T4s. The output is a checkpoint folder. Point the sidecar at it; the C# does not change. The same client also covers reading player chat to an allied commander, a pacing director, and picking which human-written bark fits the moment.

Running Laya inside the Unity process

A sidecar is the right first step, but a second process costs you: another executable to sign, antivirus prompts, and no path to consoles. The alternative is ONNX Runtime inside Unity, through onnxruntime-unity (opens in a new tab) and the receptron/laya-onnx (opens in a new tab) export (fp32, 1.3 to 1.7 GB). A Unity developer (@unitycoder_com (opens in a new tab)) has already played tic-tac-toe against it that way. The missing piece is a C# tokenizer matching ModernBERT's tokenizer.json. The official laya-dotnet (opens in a new tab) port targets .NET 10, so it does not drop into Unity, but its sequence builder and calibration code are portable. Do this port once Laya has earned its place: a few hundred lines of C#, a 1.7 GB asset, and a native plugin per platform. WebGL would need the browser port laya-ts (opens in a new tab) behind a .jslib bridge nobody has published.

Multiplayer and replays

Laya is deterministic on the same hardware and precision, which Jev is not, so a decision can be part of a replay. Across machines precision differs (bf16 flipped 3 of 864 answers against fp32), and in a lockstep RTS that is a desync. Run the AI commander on the host and broadcast its decisions as commands.
Takeaway: Laya as a local service, everything else in C#. Batch decisions into one call, validate against the legal list, fall back to your scripted AI, and record the same state sentences during playtests for fine-tuning.

Other local options

The System One Models directory (opens in a new tab) tracks the open decision models. Kev (0.8B to 9B, Apache 2.0) is better out of the box and ships training code. Von (395M) scores each option alone so order does not matter. Clef from Cloudflare (27B) is far more accurate but only fits on a server. Next to a renderer, Laya and Von are the only ones small enough.

Is it worth looking into?

You are buildingVerdictWhy
Offline RTS, 4X, or strategy gameYes, with fine-tuningCommander decisions every 10 seconds are the perfect shape. Replays are free training data
Turn-based, card, roguelike, auto-battlerYesOne decision per turn. Enumerate legal moves and let it pick
Real-time action with a few smart NPCs or a directorYes, on a GPUSame rules as Jev, without the network
Twitch shooter with per-frame controlNoTactics, never steering
Mobile, Switch, consolesNot yetNo sidecar, and the in-process path is unfinished
No GPU, no time to fine-tune, need good answers nowNo, use Jev or KevLaya zero-shot on CPU is the worst of both worlds
Be aware of the limitations: Laya is text-based, still maturing, and requires a 0.5 to 0.9 GB model weight on each client. Decision models like Laya fill the gap between rigid behavior trees and heavyweight LLMs. Where Jev runs remotely, Laya brings smarter decision-making to your player's GPU, training included. Out of the box, it offers plausible guesses, but if you have replay data of real, human-made decisions, you can fine-tune it over a weekend using the Kaggle notebook. The result? A skirmish commander that understands your game like your testers do, answers in 20 ms, requires no Internet, no API key, and no billing surprises. If you're building strategy games in Unity, start by spinning up the sidecar, drop LayaClient into your scene, and wire up a single commander decision. No matter your architecture, keep your behavior trees for fallback and validation. Laya decides, and your C# stays in charge.
00 views
00 shares

Discussion about this post

Comments are reviewed before they appear on the article.

No comments yet. Be the first to start the discussion.

Frequently asked questions

Newsletter

Stay in the Loop.

Subscribe to our newsletter to receive the latest news, updates, and special offers directly in your inbox. Don't miss out!