Laya Requirements: RAM, GPU, Model Size and Supported Systems

What you need to run Laya: Python 3.10+, about 3 GB of RAM per checkpoint on CPU, an optional GPU, 647-808 MB model files and supported hardware. Free.

Last updated: Sep 28, 2026

What do you need to run Laya? Python 3.10 or newer, about 1 GB of disk per checkpoint, and about 3 GB of free RAM for one checkpoint on CPU. A GPU is optional: Laya runs on any modern CPU, and in upstream tests a GPU made it roughly 6 to 40 times faster, depending on the checkpoint and the number of questions. Laya is free and open source under Apache 2.0.

This page collects the requirements in one place. Figures come from the Laya README, the upstream BENCHMARKS.md, the Hugging Face model files and PyPI. Where a number is our own estimate, we say so.

Laya requirements at a glance

MinimumRecommended
Python3.103.11 or 3.12
Operating systemmacOS, Linux or WindowsAny of the three
RAM (CPU inference)About 3 GB free for one checkpoint8 GB for the default Router, 16 GB for a server with all three checkpoints loaded
GPUNot requiredAny NVIDIA, Apple Silicon or Intel GPU with 4 GB or more
DiskAbout 2 GB (PyTorch + one checkpoint; more on Linux, see below)3 to 5 GB (all three checkpoints, CUDA libraries on Linux)
InternetOnly for the first downloadWorks offline after that
CostFree (Apache 2.0)Free; you pay only for your own hardware

Software: Python, PyTorch and operating systems

  • Python 3.10 or newer. The dependencies set this floor: transformers 5.x, huggingface_hub 1.x and torch 2.14 all need 3.10.
  • Dependencies. pip install laya pulls in PyTorch, Transformers, safetensors, huggingface_hub and NumPy. Extras such as laya[serve] (HTTP server) and laya[mcp] (MCP server) add only what those features need.
  • Operating systems. Upstream documents installs on macOS, Linux and Windows (PowerShell). See Install Laya for the exact commands.

PyTorch is the largest part of the install. For torch 2.14 on Python 3.12, the PyPI wheel is about 127 MB on Apple Silicon Macs, about 124 MB on Windows and about 555 MB on Linux x86_64. On Linux, the default wheel also pulls in NVIDIA CUDA libraries (cuDNN alone is about 519 MB), even on a machine without a GPU. On a CPU-only Linux server, install the CPU build of PyTorch first to skip them:

pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install laya

Laya model size and disk space

Laya ships three checkpoints. Each is downloaded from Hugging Face the first time you load it and cached on disk (by default under ~/.cache/huggingface).

CheckpointEncoderParametersDownload sizeUse it for
layaModernBERT-large421M~808 MBEnglish
laya-multilingualmmBERT-base322M~647 MB100+ languages, 2 to 3 times faster
laya-typed-decisionsModernBERT-large421M~808 MBBusiness workflows from the typed-decisions dataset

Sizes are from the model card; all three together take about 2.3 GB. Most apps need one or two. Details on each: Laya models.

How much RAM does Laya need?

On CPU, Laya runs in 32-bit floating point, so the weights take about 4 bytes per parameter in memory. By our estimate:

  • laya-multilingual: about 1.3 GB of weights in RAM.
  • laya or laya-typed-decisions: about 1.7 GB each.
  • Python and PyTorch add their own overhead on top.

What has been measured:

  • Upstream ran all its CPU benchmarks on a 4-core AWS server with 16 GiB of RAM. Peak memory for the whole run, with up to five checkpoints loaded at once, was 9.3 GiB.
  • The community laya-pi project serves Laya through ONNX Runtime on a Raspberry Pi 4 with 8 GB.

The Router keeps two checkpoints in memory by default (English and multilingual). On a small machine, use Router(max_loaded=1) or load a single checkpoint directly; the trade-off is a reload of several seconds whenever the language switches.

Our recommendation: 4 GB of RAM is enough to try one checkpoint, 8 GB for the default Router, and 16 GB for a server that preloads all three.

Do you need a GPU for Laya?

No. Laya runs on CPU, and for many batch jobs that is enough. A GPU mainly helps when you need answers in tens of milliseconds, or ask many questions per call.

CPU (upstream measurements, p50 for one question):

Hardwarelaya (English)laya-multilingual
AMD EPYC server, 4 cores580 ms193 ms

A Ryzen 9 laptop with 8 threads answered one question in 329 ms in a separate upstream test.

On CPU, each extra question costs about the same as the first, so ten questions take about ten times as long. Thread settings matter a lot: pinning PyTorch's inter-op threads to 1 made one laptop 12x faster. See Laya on CPU for the tuning steps.

GPU (upstream measurements, p50):

Hardware1 question10 questions
NVIDIA T4, laya-multilingual32.8 ms72.3 ms
NVIDIA T4, laya39.5 ms158.6 ms
Intel Arc B390 (XPU), laya29.7 ms96.9 ms

Even stored in 32-bit, one checkpoint's weights take at most about 1.7 GB of video memory. By our estimate, any GPU with 4 GB or more fits one checkpoint; the T4 used above has 16 GB.

Supported hardware

HardwareHow to run Laya
Any x86 or ARM CPUDefault install; see Laya on CPU
NVIDIA GPUDefault install with a CUDA build of PyTorch
Apple Silicon MacPyTorch on MPS (automatic), or the native MLX runtime
Intel GPU (Arc, integrated)Install the XPU build of PyTorch first; Laya picks the XPU automatically
Huawei Ascend NPUCommunity laya-Ascend project
Raspberry Pi 4 and RK3588 boardsCommunity projects: laya-pi (CPU) and laya-rk3588-turingpi-rk1 (NPU, Mali GPU)

You do not have to use Python at all. Laya also runs from Node.js through ONNX Runtime, in native C++ and in Rust. For a faster CPU deployment, the community laya-openvino backend ships an INT8 English checkpoint of 405 MB. More options: Laya runtimes.

Requirements for fine-tuning Laya

Running Laya needs no GPU; fine-tuning it practically does. The official notebook runs on Kaggle's free 2x T4 GPUs, and the project reports about 4 to 5 hours for 4 epochs over roughly 30,000 questions. See Fine-tune Laya.

Is Laya free?

Yes. The code and all three checkpoints are released under the Apache 2.0 license. You can use, modify and self-host Laya, including in commercial products, with no per-request fees. Your only cost is the hardware you run it on. By comparison, TypeSafe's hosted Jev is billed per token; see Laya vs Jev.

To try Laya without installing anything, use the Playground.

FAQ

What are the minimum requirements for Laya?

Python 3.10 or newer on macOS, Linux or Windows, about 2 GB of disk and about 3 GB of free RAM for one checkpoint. No GPU is required.

How big is the Laya model?

The English laya and laya-typed-decisions checkpoints have 421M parameters and download as about 808 MB each. laya-multilingual has 322M parameters and is about 647 MB.

Can Laya run on CPU only?

Yes. On a 4-core server, one question takes about 580 ms on the English checkpoint and about 193 ms on the multilingual one. Tune the thread settings as described in Laya on CPU.

Does Laya run on a Mac?

Yes. On Apple Silicon, Laya uses the GPU through PyTorch MPS automatically. The community laya-mlx runtime runs it natively on MLX without PyTorch.

Does Laya work on Windows?

Yes. Upstream documents a Windows PowerShell install, and Intel GPUs are supported on Windows through the XPU build of PyTorch.

Can Laya run offline?

Yes. After the first download, the checkpoints load from the local Hugging Face cache. For a machine without internet access, download the model folder on another machine and copy it over.

Is Laya free for commercial use?

Yes. The Apache 2.0 license allows commercial use.

Last verified against the Laya README, BENCHMARKS.md, Hugging Face and PyPI: September 28, 2026.