Laya and Unsloth: Run a Local, Jev-Compatible Decision API

Run Laya locally with Unsloth Desktop's Decision API: GUI setup, model sizes and RAM, a curl example, and migrating existing Jev code with two environment variables.

Last updated: Sep 29, 2026

Unsloth (76,000+ stars on GitHub, best known for local LLM fine-tuning and inference) shipped a Decision API in Unsloth Desktop that serves Laya behind the same wire format as TypeSafe's Jev, POST /v1/systemone. You get a GUI: turn on "Serve requests," pick a model size, and point your existing Jev code at localhost — no separate Python install or server process to manage if you already use Unsloth.

This page covers setup, the model choices, a working request, and what changes if you are migrating existing Jev code. Everything here is sourced from Unsloth's own documentation.

Model options

Unsloth Desktop offers three Laya checkpoints. All run on CPU by default; you can switch to GPU in settings.

ModelMinimum RAMDownload sizeBest for
Multilingual (default)4 GB678 MBMost use cases. 100+ languages, 1024-token context.
English5 GB846 MBEnglish only. Larger encoder, 512-token context.
Typed decisions5 GB846 MBEnglish only. Invoices, security incidents, customer service, agent traces.

CPU is the default; switching to GPU keeps the model resident in memory until you unload it or restart Unsloth, and Laya falls back to CPU automatically if the GPU runs out of memory.

Quickstart

1. Install Unsloth Desktop

Download the app for macOS, Windows or Linux from unsloth.ai, or install from the command line:

# macOS, Linux, WSL
curl -fsSL https://unsloth.ai/install.sh | sh
# Windows PowerShell
irm https://unsloth.ai/install.ps1 | iex

2. Turn on the Decision API

Go to Settings → API → Decision API and turn on Serve requests. Confirm the download of the default 678 MB multilingual model.

Your API key lives on the same settings page under Access tokens. To skip the key entirely on localhost, turn on Keyless API access → Chat and inference.

3. Send a decision request

curl http://localhost:8888/v1/systemone \
  -H "Authorization: Bearer sk-unsloth-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "laya",
    "state": "Hi, I was charged twice for my March invoice (#4411). Please refund the duplicate today.",
    "questions": {
      "team": {"type": "choice", "instructions": "Which team should handle this?",
        "criteria": {"billing": "invoices, payments, refunds",
                     "technical": "bugs, outages, errors",
                     "sales": "pricing, new plans",
                     "other": "everything else"}},
      "refund": {"type": "noul", "instructions": "Does the customer ask for a refund?"},
      "urgency": {"type": "score", "instructions": "How urgent is this?",
                  "criteria": ["not urgent", "soon", "today"]}
    }
  }'
{
  "model": "laya-multilingual",
  "answers": {
    "team": {"type": "choice", "choice": "billing", "confidence": 1.0,
              "probabilities": {"billing": 1.0, "technical": 0.0, "sales": 0.0, "other": 0.0}},
    "refund": {"type": "noul", "noul": 0.9794},
    "urgency": {"type": "score", "score": 1.9005, "confidence": 0.698,
                "legend": {"0": "not urgent", "1": "soon", "2": "today"},
                "probabilities": {"0": 0.004, "1": 0.0915, "2": 0.9045}}
  },
  "usage": {"input_tokens": 172, "output_tokens": 0}
}

The first request loads the model, which takes 10 to 20 seconds. After that, requests take well under a second on most CPUs. model accepts laya, default or jev-latest for whichever model you picked in settings, or a specific name: laya-multilingual, laya-english, laya-typed-decisions.

Question types

TypeUse it forYou get back
noulYes or noThe probability of yes, from 0 to 1
choicePicking one optionThe most likely option, plus a probability for each
scoreRating on a scaleWhere the text falls on your scale, counting from 0

Each request supports up to 64 questions, up to 255 options per choice, and up to 10 levels per score. choice and score also return confidence: close to 1 means one answer stands out, close to 0 means the probabilities are spread evenly. To act on a choice, threshold on its value in probabilities, not on confidence — Laya and Jev calculate confidence differently, so a threshold tuned on Jev does not transfer.

Unsloth's docs page also embeds a live demo: a travel-packing list that updates its suggested items as you type, powered by local Laya through the Decision API.

Migrating from Jev

Existing TypeSafe SDK code needs two environment variables and nothing else:

export TYPESAFE_BASE_URL=http://localhost:8888
export TYPESAFE_API_KEY=sk-unsloth-YOUR_KEY

The SDK's default model, jev-latest, maps to whichever model you picked in Unsloth's settings. Two things do not carry over automatically: confidence uses a different formula than Jev's, so re-threshold on your own data instead of reusing a Jev-tuned cutoff; and options in a choice share a fixed token budget, so lists longer than about 20 described options get trimmed.

How this compares to other ways to run Laya locally

OptionBest forInstallRuns on
Unsloth DesktopA GUI, especially if you already use Unsloth for local LLMsDesktop app or one-line installerCPU, GPU
ollayaOllama-style CLI, daemon and desktop appOne-line installerCPU, NVIDIA CUDA
Python package (pip install laya)Python code and notebooks; the reference implementationpipCPU, CUDA, Apple MPS
laya-serveA Jev-compatible HTTP API from the official packagepip install "laya[serve]"CPU, CUDA, MPS
laya-mlxNative Apple Silicon inferencepipApple MLX

See Laya and Ollama, Install Laya, Self-host Laya and all runtimes for the others.

Which one should you pick?

  • You already run Unsloth Desktop for local LLMs and want decisions in the same app: Unsloth's Decision API.
  • You like an Ollama-style CLI and daemon: ollaya.
  • You write Python or want the official, most current code: pip install laya.
  • You have existing Jev client code and want the smallest possible footprint: laya-serve or ollaya; both speak the same API as Unsloth.
  • You are on Apple Silicon and want native speed without a GUI: laya-mlx.

Try Laya without installing anything

The Laya Playground runs the model in your browser, so you can test your own text and questions before choosing a local setup.

FAQ

Is Unsloth's Decision API made by the Laya team?

No. Laya comes from ConvAI Innovations. Unsloth is an independent, widely used tool for running and training local models that added Laya support as a Decision API. See Unsloth's documentation for their side of the integration.

Do I need a GPU to run Laya through Unsloth?

No. CPU is the default, and the smallest model needs only 4 GB of RAM. You can switch to GPU in settings for faster answers; Laya falls back to CPU automatically if the GPU runs out of memory.

Does my data leave my machine?

No. Unsloth's Decision API runs Laya locally; requests go to localhost, and nothing is sent to a remote server.

Can I point my existing Jev SDK code at Unsloth?

Yes. Set TYPESAFE_BASE_URL=http://localhost:8888 and TYPESAFE_API_KEY to your Unsloth key. Re-check your confidence thresholds afterward, since Unsloth's Laya and Jev calculate confidence differently.

How much RAM does the default model need?

4 GB for the default multilingual model (678 MB download). The English and typed-decisions checkpoints need 5 GB.

Last verified against Unsloth's documentation: September 29, 2026.