Laya and Unsloth: Run a Local, Jev-Compatible Decision API
Run Laya locally with Unsloth Desktop's Decision API: GUI setup, model sizes and RAM, a curl example, and migrating existing Jev code with two environment variables.
Unsloth (76,000+ stars on GitHub, best known for local LLM fine-tuning and inference) shipped a Decision API in Unsloth Desktop that serves Laya behind the same wire format as TypeSafe's Jev, POST /v1/systemone. You get a GUI: turn on "Serve requests," pick a model size, and point your existing Jev code at localhost — no separate Python install or server process to manage if you already use Unsloth.
This page covers setup, the model choices, a working request, and what changes if you are migrating existing Jev code. Everything here is sourced from Unsloth's own documentation.
Model options
Unsloth Desktop offers three Laya checkpoints. All run on CPU by default; you can switch to GPU in settings.
| Model | Minimum RAM | Download size | Best for |
|---|---|---|---|
| Multilingual (default) | 4 GB | 678 MB | Most use cases. 100+ languages, 1024-token context. |
| English | 5 GB | 846 MB | English only. Larger encoder, 512-token context. |
| Typed decisions | 5 GB | 846 MB | English only. Invoices, security incidents, customer service, agent traces. |
CPU is the default; switching to GPU keeps the model resident in memory until you unload it or restart Unsloth, and Laya falls back to CPU automatically if the GPU runs out of memory.
Quickstart
1. Install Unsloth Desktop
Download the app for macOS, Windows or Linux from unsloth.ai, or install from the command line:
# macOS, Linux, WSL
curl -fsSL https://unsloth.ai/install.sh | sh
# Windows PowerShell
irm https://unsloth.ai/install.ps1 | iex
2. Turn on the Decision API
Go to Settings → API → Decision API and turn on Serve requests. Confirm the download of the default 678 MB multilingual model.
Your API key lives on the same settings page under Access tokens. To skip the key entirely on localhost, turn on Keyless API access → Chat and inference.
3. Send a decision request
curl http://localhost:8888/v1/systemone \
-H "Authorization: Bearer sk-unsloth-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "laya",
"state": "Hi, I was charged twice for my March invoice (#4411). Please refund the duplicate today.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, errors",
"sales": "pricing, new plans",
"other": "everything else"}},
"refund": {"type": "noul", "instructions": "Does the customer ask for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["not urgent", "soon", "today"]}
}
}'
{
"model": "laya-multilingual",
"answers": {
"team": {"type": "choice", "choice": "billing", "confidence": 1.0,
"probabilities": {"billing": 1.0, "technical": 0.0, "sales": 0.0, "other": 0.0}},
"refund": {"type": "noul", "noul": 0.9794},
"urgency": {"type": "score", "score": 1.9005, "confidence": 0.698,
"legend": {"0": "not urgent", "1": "soon", "2": "today"},
"probabilities": {"0": 0.004, "1": 0.0915, "2": 0.9045}}
},
"usage": {"input_tokens": 172, "output_tokens": 0}
}
The first request loads the model, which takes 10 to 20 seconds. After that, requests take well under a second on most CPUs. model accepts laya, default or jev-latest for whichever model you picked in settings, or a specific name: laya-multilingual, laya-english, laya-typed-decisions.
Question types
| Type | Use it for | You get back |
|---|---|---|
noul | Yes or no | The probability of yes, from 0 to 1 |
choice | Picking one option | The most likely option, plus a probability for each |
score | Rating on a scale | Where the text falls on your scale, counting from 0 |
Each request supports up to 64 questions, up to 255 options per choice, and up to 10 levels per score. choice and score also return confidence: close to 1 means one answer stands out, close to 0 means the probabilities are spread evenly. To act on a choice, threshold on its value in probabilities, not on confidence — Laya and Jev calculate confidence differently, so a threshold tuned on Jev does not transfer.
Unsloth's docs page also embeds a live demo: a travel-packing list that updates its suggested items as you type, powered by local Laya through the Decision API.
Migrating from Jev
Existing TypeSafe SDK code needs two environment variables and nothing else:
export TYPESAFE_BASE_URL=http://localhost:8888
export TYPESAFE_API_KEY=sk-unsloth-YOUR_KEY
The SDK's default model, jev-latest, maps to whichever model you picked in Unsloth's settings. Two things do not carry over automatically: confidence uses a different formula than Jev's, so re-threshold on your own data instead of reusing a Jev-tuned cutoff; and options in a choice share a fixed token budget, so lists longer than about 20 described options get trimmed.
How this compares to other ways to run Laya locally
| Option | Best for | Install | Runs on |
|---|---|---|---|
| Unsloth Desktop | A GUI, especially if you already use Unsloth for local LLMs | Desktop app or one-line installer | CPU, GPU |
| ollaya | Ollama-style CLI, daemon and desktop app | One-line installer | CPU, NVIDIA CUDA |
Python package (pip install laya) | Python code and notebooks; the reference implementation | pip | CPU, CUDA, Apple MPS |
| laya-serve | A Jev-compatible HTTP API from the official package | pip install "laya[serve]" | CPU, CUDA, MPS |
| laya-mlx | Native Apple Silicon inference | pip | Apple MLX |
See Laya and Ollama, Install Laya, Self-host Laya and all runtimes for the others.
Which one should you pick?
- You already run Unsloth Desktop for local LLMs and want decisions in the same app: Unsloth's Decision API.
- You like an Ollama-style CLI and daemon: ollaya.
- You write Python or want the official, most current code:
pip install laya. - You have existing Jev client code and want the smallest possible footprint: laya-serve or ollaya; both speak the same API as Unsloth.
- You are on Apple Silicon and want native speed without a GUI: laya-mlx.
Try Laya without installing anything
The Laya Playground runs the model in your browser, so you can test your own text and questions before choosing a local setup.
FAQ
Is Unsloth's Decision API made by the Laya team?
No. Laya comes from ConvAI Innovations. Unsloth is an independent, widely used tool for running and training local models that added Laya support as a Decision API. See Unsloth's documentation for their side of the integration.
Do I need a GPU to run Laya through Unsloth?
No. CPU is the default, and the smallest model needs only 4 GB of RAM. You can switch to GPU in settings for faster answers; Laya falls back to CPU automatically if the GPU runs out of memory.
Does my data leave my machine?
No. Unsloth's Decision API runs Laya locally; requests go to localhost, and nothing is sent to a remote server.
Can I point my existing Jev SDK code at Unsloth?
Yes. Set TYPESAFE_BASE_URL=http://localhost:8888 and TYPESAFE_API_KEY to your Unsloth key. Re-check your confidence thresholds afterward, since Unsloth's Laya and Jev calculate confidence differently.
How much RAM does the default model need?
4 GB for the default multilingual model (678 MB download). The English and typed-decisions checkpoints need 5 GB.
Related
- Laya and Ollama: another Jev-compatible local option, Ollama-style
- Install Laya: the official Python quickstart
- Self-host Laya: laya-serve, Docker and MCP
- Jev alternatives: open models and servers that accept Jev requests unchanged
- What is Jev?: the API this local server is compatible with
Last verified against Unsloth's documentation: September 29, 2026.