Laya vs Kev: Open Weights, Accuracy, Speed & Tradeoffs
Compare Laya vs Kev across architecture, open weights, self-hosting, accuracy, calibration, latency, fine-tuning and choosing the right size.
Laya and Kev both give you an open, self-hostable System One model, but they get there from opposite directions: Laya is a small encoder with a decision head, and Kev is a family of fine-tuned decoder LLMs (Qwen) ranging from 0.8B to 27B parameters. This page compares them on architecture, accuracy, calibration, latency and hardware cost, and links every number to its source.
At a glance
| Laya | Kev | |
|---|---|---|
| Maker | ConvAI Innovations | Jared Palmer |
| Architecture | Encoder (ModernBERT, mmBERT) with decision heads | Fine-tuned decoder LLM (Qwen3.5 / Qwen3.8) |
| Weights | Open | Open |
| License | Apache 2.0 | Apache 2.0 |
| Sizes | 322M / 421M | 0.8B, 4B, 9B, 27B |
| Deployment | Local / self-hosted | Local / self-hosted |
| Jev-compatible API | Yes, through laya-serve | Yes, POST /v1/systemone |
| Fine-tuning | Upstream RLCD workflow | Upstream training scripts, --init_from a released checkpoint |
| Typed primitives | choice, score, noul | choice, score, noul |
| CPU inference | Supported, see Laya on CPU | Not the target; hundreds of ms on a Mac |
Our head-to-head test: Laya vs Kev-0.8B
Kev ships four sizes; our own System One benchmark only tested the smallest, Kev-0.8B, against Laya on the same questions and examples, with no fine-tuning:
| Test | Kev-0.8B | Laya (English) | Laya (multilingual) | Jev 1.13 |
|---|---|---|---|---|
| SST-2 sentiment, 2 labels | 0.920 | 0.920 | 0.857 | 0.963 |
| TREC question type, 6 labels | 0.953 † | 0.773 | 0.860 | 0.927 |
| Banking77 intent, 77 labels | 0.823 † | 0.363 | 0.367 | 0.813 |
| MASSIVE, average of 7 non-English languages | 0.656 | 0.307 | 0.544 | 0.872 |
| Calibration error (ECE), SST-2 | 0.032 | 0.015 | 0.101 | 0.020 |
| Calibration error (ECE), Banking77 | 0.128 | 0.506 | 0.439 | 0.083 |
| Median latency per request (local, Apple M5 Pro) | 22.6–68.9 ms | 18.5–54.2 ms | 10.9–24.5 ms | ~410 ms (hosted) |
† Kev's model card lists Banking77, TREC and SST-5 (the same sentences as SST-2) in its training data; we scored held-out test and validation examples, but Kev has seen this kind of text and these label sets. Full method and every raw prediction: System One benchmark.
What this shows:
- Kev-0.8B is the strongest open model we tested on English classification, ahead of Laya on TREC and Banking77 — though the training-data overlap above means that gap is partly earned on Kev's own turf.
- Kev-0.8B also beat Laya multilingual on non-English input (0.656 vs 0.544) in our test, even though Kev's model card does not document multilingual support and Laya explicitly ships a dedicated multilingual checkpoint for 100+ languages.
- Laya is faster locally and better calibrated at scale. Laya multilingual answers in 11–25 ms against Kev-0.8B's 23–69 ms, and Kev's calibration error is far lower than Laya's on Banking77's 77 labels (0.128 vs 0.439–0.506), so Kev's probabilities stay more trustworthy as the option count grows.
- On simple two-label sentiment, Laya English and Kev-0.8B tie exactly (0.920).
Beyond 0.8B: how the larger Kev checkpoints compare
Our benchmark only covers Kev-0.8B. The Kev project publishes its own numbers for the larger checkpoints against Jev, on questions held out from training:
| Model | Accuracy (dev / test) | Brier score (lower is better) |
|---|---|---|
| Kev-0.8B | 0.648 / 0.697 | 0.481 / 0.416 |
| Kev-4B | 0.817 / 0.838 | 0.269 / 0.242 |
| Kev-9B | 0.822 / 0.852 | 0.286 / 0.237 |
| Kev-27B | 0.848 / 0.896 | 0.236 / 0.164 |
| Jev (hosted) | 0.857 / – | 0.211 / – |
Source: Kev model card. These are Kev's own measurements, on a different question set than our benchmark above, so do not mix the two tables directly. The pattern is consistent with what you would expect: bigger Kev checkpoints close the gap to Jev, at the cost of needing much larger hardware (see below). Kev also trails on knowledge-heavy questions: Kev-9B scores 0.74 on MMLU against Jev's 0.90, since Kev decides only from the text you give it while Jev can draw on more general knowledge.
Hardware: small encoder vs decoder LLM
Laya's two checkpoints (322M and 421M) run comfortably on a CPU. Kev's range is wider and shifts the floor upward:
| Laya | Kev | |
|---|---|---|
| Smallest | 322M (multilingual), ~1.3 GB RAM | 0.8B, Apple Silicon Mac or an NVIDIA L4 |
| CPU-only | Yes, see Laya requirements | Not the target |
| Largest | 421M | 27B, needs an 80 GB-class GPU (H100, H200 or B200) |
| Mac support | All checkpoints | Kev-0.8B only; no Mac path for 27B |
If you need the smallest possible footprint or CPU-only serving, Laya is the only one of the two that fits. If you can dedicate a GPU and want accuracy closer to Jev, a larger Kev checkpoint is the more direct path.
Fine-tuning
Both projects expect you to fine-tune for your own workflow rather than rely on zero-shot accuracy:
- Laya publishes an RLCD notebook that runs on Kaggle's free 2×T4 GPUs; its own benchmark shows the fine-tuned
laya-typed-decisionscheckpoint jumping from 0.362 to 0.766 on its typed-decisions benchmark. See Fine-tune Laya. - Kev trains with a JSONL dataset and
--init_froma released checkpoint; the project reports that starting from the raw base model scored 0.33 on its evaluation set, against 0.84 starting from the released Kev-4B checkpoint. Akev-finetuneagent skill can run training on Modal for about $1 of H100 time.
Switching between them
Both serve the same Jev-compatible POST /v1/systemone endpoint, so most client code only needs a new base URL. What does not transfer automatically:
- Confidence thresholds. Laya's
answer_confidenceand Kev's confidence score are not calibrated the same way; re-fit your threshold on your own labelled data after switching. - Option-set size. Laya's decision head shares a fixed token budget across all options in one
choicequestion (192–256 tokens); Kev, built on a full LLM, does not have that constraint in the same way. - Hardware. Moving from Laya to any Kev checkpoint above 0.8B means provisioning a GPU you may not currently need.
Which should you evaluate?
Consider Laya first when:
- you need CPU-only or the smallest possible footprint;
- low local latency matters more than the last few points of accuracy;
- you want a dedicated multilingual checkpoint documented for 100+ languages;
- you expect to fine-tune on a modest GPU budget (2×T4).
Consider Kev first when:
- you have a GPU and want accuracy closer to Jev, especially in English;
- your questions lean on general or world knowledge (within a decoder LLM's reach);
- you want a family of sizes to trade off accuracy against cost as your workload grows.
For production decisions, test the specific Kev size you would deploy against Laya on your own representative dataset — our table above covers only Kev-0.8B.
FAQ
Is Kev better than Laya?
It depends on the size and the task. In our test, the smallest Kev (0.8B) matched or beat Laya on English classification and non-English input, and Kev is far better calibrated on tasks with many options. Laya is faster locally, runs on CPU, and its multilingual checkpoint is explicitly documented for 100+ languages. Larger Kev checkpoints (4B–27B) score higher still, approaching Jev, but need much more GPU memory.
Can I run Kev on a CPU like Laya?
Not really. Kev's project targets GPU and Apple Silicon; on a Mac, answers take hundreds of milliseconds rather than Laya's tens of milliseconds, and Kev-27B needs an 80 GB-class GPU with no Mac path at all. If CPU-only serving matters, Laya is the better fit.
Which Kev size should I compare against Laya?
Start with Kev-0.8B, the size closest to Laya's footprint and the one in our own benchmark above. If you have GPU budget to spare, Kev-4B already scores much closer to Jev on the project's own numbers.
Is there another project called "Kev"?
Yes, and it is unrelated to the Kev covered on this page. NotJev : Kev by developer arjun988 is a separate, independently named open-source project: a TypeScript SDK, server and CLI that turn an LLM you already have (Ollama, vLLM or any OpenAI-compatible API) into a Jev-style typed-decision endpoint. It ships no trained model of its own — the "intelligence is whatever model you point it at," in the project's own words — so it is closer in kind to AnyJev than to the fine-tuned Kev family this page benchmarks. Check the GitHub URL if you are not sure which "Kev" a link points to.
Are Laya and Kev both free?
Yes. Both are released under Apache 2.0: no per-request fees, only the cost of the hardware you run them on. Compare that to TypeSafe's hosted Jev, billed per token.
Related
- Kev model guide: sizes, setup and fine-tuning
- System One benchmark: Jev, Kev, Laya, Von and GLiNER on the same questions
- Laya vs Jev: open weights, speed, accuracy and tradeoffs
- System One models compared: every model side by side
- Jev alternatives: open System One models you can self-host
Last verified: September 29, 2026.