Laya Requirements: RAM, GPU, Model Size and Supported Systems
What you need to run Laya: Python 3.10+, about 3 GB of RAM per checkpoint on CPU, an optional GPU, 647-808 MB model files and supported hardware. Free.
What do you need to run Laya? Python 3.10 or newer, about 1 GB of disk per checkpoint, and about 3 GB of free RAM for one checkpoint on CPU. A GPU is optional: Laya runs on any modern CPU, and in upstream tests a GPU made it roughly 6 to 40 times faster, depending on the checkpoint and the number of questions. Laya is free and open source under Apache 2.0.
This page collects the requirements in one place. Figures come from the Laya README, the upstream BENCHMARKS.md, the Hugging Face model files and PyPI. Where a number is our own estimate, we say so.
Laya requirements at a glance
| Minimum | Recommended | |
|---|---|---|
| Python | 3.10 | 3.11 or 3.12 |
| Operating system | macOS, Linux or Windows | Any of the three |
| RAM (CPU inference) | About 3 GB free for one checkpoint | 8 GB for the default Router, 16 GB for a server with all three checkpoints loaded |
| GPU | Not required | Any NVIDIA, Apple Silicon or Intel GPU with 4 GB or more |
| Disk | About 2 GB (PyTorch + one checkpoint; more on Linux, see below) | 3 to 5 GB (all three checkpoints, CUDA libraries on Linux) |
| Internet | Only for the first download | Works offline after that |
| Cost | Free (Apache 2.0) | Free; you pay only for your own hardware |
Software: Python, PyTorch and operating systems
- Python 3.10 or newer. The dependencies set this floor:
transformers5.x,huggingface_hub1.x andtorch2.14 all need 3.10. - Dependencies.
pip install layapulls in PyTorch, Transformers, safetensors, huggingface_hub and NumPy. Extras such aslaya[serve](HTTP server) andlaya[mcp](MCP server) add only what those features need. - Operating systems. Upstream documents installs on macOS, Linux and Windows (PowerShell). See Install Laya for the exact commands.
PyTorch is the largest part of the install. For torch 2.14 on Python 3.12, the PyPI wheel is about 127 MB on Apple Silicon Macs, about 124 MB on Windows and about 555 MB on Linux x86_64. On Linux, the default wheel also pulls in NVIDIA CUDA libraries (cuDNN alone is about 519 MB), even on a machine without a GPU. On a CPU-only Linux server, install the CPU build of PyTorch first to skip them:
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install laya
Laya model size and disk space
Laya ships three checkpoints. Each is downloaded from Hugging Face the first time you load it and cached on disk (by default under ~/.cache/huggingface).
| Checkpoint | Encoder | Parameters | Download size | Use it for |
|---|---|---|---|---|
laya | ModernBERT-large | 421M | ~808 MB | English |
laya-multilingual | mmBERT-base | 322M | ~647 MB | 100+ languages, 2 to 3 times faster |
laya-typed-decisions | ModernBERT-large | 421M | ~808 MB | Business workflows from the typed-decisions dataset |
Sizes are from the model card; all three together take about 2.3 GB. Most apps need one or two. Details on each: Laya models.
How much RAM does Laya need?
On CPU, Laya runs in 32-bit floating point, so the weights take about 4 bytes per parameter in memory. By our estimate:
laya-multilingual: about 1.3 GB of weights in RAM.layaorlaya-typed-decisions: about 1.7 GB each.- Python and PyTorch add their own overhead on top.
What has been measured:
- Upstream ran all its CPU benchmarks on a 4-core AWS server with 16 GiB of RAM. Peak memory for the whole run, with up to five checkpoints loaded at once, was 9.3 GiB.
- The community laya-pi project serves Laya through ONNX Runtime on a Raspberry Pi 4 with 8 GB.
The Router keeps two checkpoints in memory by default (English and multilingual). On a small machine, use Router(max_loaded=1) or load a single checkpoint directly; the trade-off is a reload of several seconds whenever the language switches.
Our recommendation: 4 GB of RAM is enough to try one checkpoint, 8 GB for the default Router, and 16 GB for a server that preloads all three.
Do you need a GPU for Laya?
No. Laya runs on CPU, and for many batch jobs that is enough. A GPU mainly helps when you need answers in tens of milliseconds, or ask many questions per call.
CPU (upstream measurements, p50 for one question):
| Hardware | laya (English) | laya-multilingual |
|---|---|---|
| AMD EPYC server, 4 cores | 580 ms | 193 ms |
A Ryzen 9 laptop with 8 threads answered one question in 329 ms in a separate upstream test.
On CPU, each extra question costs about the same as the first, so ten questions take about ten times as long. Thread settings matter a lot: pinning PyTorch's inter-op threads to 1 made one laptop 12x faster. See Laya on CPU for the tuning steps.
GPU (upstream measurements, p50):
| Hardware | 1 question | 10 questions |
|---|---|---|
NVIDIA T4, laya-multilingual | 32.8 ms | 72.3 ms |
NVIDIA T4, laya | 39.5 ms | 158.6 ms |
Intel Arc B390 (XPU), laya | 29.7 ms | 96.9 ms |
Even stored in 32-bit, one checkpoint's weights take at most about 1.7 GB of video memory. By our estimate, any GPU with 4 GB or more fits one checkpoint; the T4 used above has 16 GB.
Supported hardware
| Hardware | How to run Laya |
|---|---|
| Any x86 or ARM CPU | Default install; see Laya on CPU |
| NVIDIA GPU | Default install with a CUDA build of PyTorch |
| Apple Silicon Mac | PyTorch on MPS (automatic), or the native MLX runtime |
| Intel GPU (Arc, integrated) | Install the XPU build of PyTorch first; Laya picks the XPU automatically |
| Huawei Ascend NPU | Community laya-Ascend project |
| Raspberry Pi 4 and RK3588 boards | Community projects: laya-pi (CPU) and laya-rk3588-turingpi-rk1 (NPU, Mali GPU) |
You do not have to use Python at all. Laya also runs from Node.js through ONNX Runtime, in native C++ and in Rust. For a faster CPU deployment, the community laya-openvino backend ships an INT8 English checkpoint of 405 MB. More options: Laya runtimes.
Requirements for fine-tuning Laya
Running Laya needs no GPU; fine-tuning it practically does. The official notebook runs on Kaggle's free 2x T4 GPUs, and the project reports about 4 to 5 hours for 4 epochs over roughly 30,000 questions. See Fine-tune Laya.
Is Laya free?
Yes. The code and all three checkpoints are released under the Apache 2.0 license. You can use, modify and self-host Laya, including in commercial products, with no per-request fees. Your only cost is the hardware you run it on. By comparison, TypeSafe's hosted Jev is billed per token; see Laya vs Jev.
To try Laya without installing anything, use the Playground.
FAQ
What are the minimum requirements for Laya?
Python 3.10 or newer on macOS, Linux or Windows, about 2 GB of disk and about 3 GB of free RAM for one checkpoint. No GPU is required.
How big is the Laya model?
The English laya and laya-typed-decisions checkpoints have 421M parameters and download as about 808 MB each. laya-multilingual has 322M parameters and is about 647 MB.
Can Laya run on CPU only?
Yes. On a 4-core server, one question takes about 580 ms on the English checkpoint and about 193 ms on the multilingual one. Tune the thread settings as described in Laya on CPU.
Does Laya run on a Mac?
Yes. On Apple Silicon, Laya uses the GPU through PyTorch MPS automatically. The community laya-mlx runtime runs it natively on MLX without PyTorch.
Does Laya work on Windows?
Yes. Upstream documents a Windows PowerShell install, and Intel GPUs are supported on Windows through the XPU build of PyTorch.
Can Laya run offline?
Yes. After the first download, the checkpoints load from the local Hugging Face cache. For a machine without internet access, download the model folder on another machine and copy it over.
Is Laya free for commercial use?
Yes. The Apache 2.0 license allows commercial use.
Related
- Install Laya: install commands for every platform
- Laya on CPU: speed and thread tuning without a GPU
- Laya models: the three checkpoints compared
- Self-host Laya: the HTTP server, Docker and MCP
- Laya runtimes: MLX, Node.js, Rust and more
Last verified against the Laya README, BENCHMARKS.md, Hugging Face and PyPI: September 28, 2026.