Laya API: Hosted Endpoint, Pricing and HTTP Guide
Call Laya through a Jev-compatible API: request and response format, curl, Python and TypeScript examples, limits, errors, and hosted Laya API pricing vs Jev.
Hosted Laya API · Use Laya in production without a server: Jev-compatible API, $0.05 per 1M input tokens, 20 free calls to start.
The Laya API is the HTTP interface to Laya, the open-source decision model by Convai Innovations. You send a piece of text (the state) and typed questions, and the Laya API returns an answer and a probability for every option in one call. It uses the same POST /v1/systemone format as TypeSafe's Jev, so code written for Jev works against a Laya API by changing the base URL.
Convai Innovations publishes Laya as open weights and does not run an official hosted Laya API. You can get one in three ways:
- Use a hosted Laya API. laya-ai.com is opening one; planned pricing is below.
- Run the Laya API yourself with
laya-serve, the server in the officiallayapackage. - Skip HTTP and call Laya from Python with the
layalibrary.
Laya API at a glance
| Endpoint | POST /v1/systemone (single state), POST /v1/systemone/batch (up to 64 states, laya 0.3.22+) |
| Auth | Authorization: Bearer <key> |
| Question types | choice (pick one option), score (place on an ordered scale), noul (probability of yes) |
| Output | the answer, a probability per option, confidence values and token usage; no generated text |
| Models | english (421M), multilingual (322M, 100+ languages), typed-decisions (421M, fine-tuned) |
| Compatible clients | TypeSafe SDK, any Jev client, plain HTTP |
| Billing unit | input tokens; output tokens are free |
Hosted Laya API from laya-ai.com
We are opening a hosted Laya API so you can call Laya without running a server or a GPU. It speaks the standard Laya API above, runs the open Laya checkpoints, and bills per input token.
| Planned pricing | |
|---|---|
| Price | $0.05 per 1M input tokens; output tokens are free |
| Free trial | 20 calls when you sign up |
| Top-up | card, Alipay or WeChat Pay |
| Access | API key right after sign-up, no approval queue |
One decision with a short ticket and three questions is 200 to 400 input tokens, so $1 covers 50,000 to 100,000 decisions. Pricing is planned and will be confirmed at launch. Join the waitlist below and we will email you once when the hosted Laya API opens.
Laya API request format
curl https://YOUR_LAYA_API/v1/systemone \
-H "Authorization: Bearer $LAYA_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": "I was charged twice this month, I want my money back",
"questions": {
"queue": {"type": "choice", "instructions": "Which team?",
"criteria": {"billing": "billing and refunds",
"tech": "login and app issues",
"other": "everything else"}},
"urgency": {"type": "score", "instructions": "How urgent?",
"criteria": ["calm", "firm", "angry", "furious"]},
"refund": {"type": "noul", "instructions": "Does the customer ask for a refund?"}
}
}'
| Field | Required | What it does |
|---|---|---|
state | yes | the text, email, ticket or JSON document to decide on |
questions | yes | an object keyed by question id; each has a type, instructions and, for choice and score, criteria |
model | no | picks a checkpoint: english, multilingual or typed-decisions. Any other value, such as a Jev model id, lets the router choose |
lang | no | a language code (de, pt-BR) that skips language detection |
min_confidence | no | a threshold from 0 to 1; answers below it come back marked low_confidence |
max_len, head_max_len | no | token windows for long states or long option lists |
task | no | forces a checkpoint by workflow name |
Put all questions about one state in one request: they share the same read of the text, so three questions cost far less than three separate calls.
Laya API response format
{
"model": "laya-rl-agent",
"answers": {
"queue": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9281, "tech": 0.0412, "other": 0.0307},
"confidence": 0.4534, "answer_confidence": 0.9281},
"urgency": {"type": "score", "score": 2.6389,
"legend": {"0": "calm", "1": "firm", "2": "angry", "3": "furious"},
"probabilities": {"0": 0.0099, "1": 0.0713, "2": 0.536, "3": 0.3828}},
"refund": {"type": "noul", "noul": 0.97}
},
"usage": {"input_tokens": 74, "output_tokens": 0},
"routing": {"model": "english", "reason": "English Latin text"}
}
| Answer type | What you read |
|---|---|
choice | choice is the top option; probabilities has one value per option |
score | score is the expected level index and can fall between levels; legend maps each index to its text |
noul | noul is the probability of yes |
| every type | answer_confidence (probability of the reported answer) and confidence |
Gate automatic actions on answer_confidence, not on confidence: the two are computed differently. If you move from Jev, re-fit your thresholds, because Jev defines confidence another way. routing shows which checkpoint answered and why, and usage.input_tokens is what a hosted Laya API bills.
Call the Laya API from Python and TypeScript
Python, with requests:
import os, requests
resp = requests.post(
"https://YOUR_LAYA_API/v1/systemone",
headers={"Authorization": f"Bearer {os.environ['LAYA_API_KEY']}"},
json={
"state": "Shipment arrived damaged, please send a new one",
"questions": {
"queue": {"type": "choice", "instructions": "Which team?",
"criteria": {"shipping": "delivery and damage",
"billing": "payments and refunds"}},
},
},
timeout=10,
)
resp.raise_for_status()
answer = resp.json()["answers"]["queue"]
print(answer["choice"], answer["answer_confidence"])
TypeScript, with fetch:
const res = await fetch("https://YOUR_LAYA_API/v1/systemone", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LAYA_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
state: "Can I change the delivery address on order 1182?",
questions: {
intent: { type: "choice", instructions: "What does the user want?",
criteria: { change_order: "edit an order", track: "track a parcel",
other: "anything else" } },
},
}),
});
const { answers } = await res.json();
console.log(answers.intent.choice, answers.intent.probabilities);
Already using the TypeSafe SDK for Jev? Keep your code and set its base URL to the Laya API. See How to use Laya for writing good questions and choosing thresholds.
Batch requests
With laya 0.3.22 or later, POST /v1/systemone/batch takes a states list with one shared questions object and returns one result per state plus a total_usage sum. A batch holds up to 64 states; a larger one is refused with 413. Use it for backfills and queues; for live traffic, one request per state keeps latency low.
Laya API limits and errors
| Limit | Value |
|---|---|
| Request body | 2 MiB |
state | 50,000 characters |
| Questions per request | 64 |
Options per choice question | 100 |
Levels per score question | 32 |
| Options across all questions | 512 |
| Status | Meaning |
|---|---|
400 | the body is not valid JSON, or state or questions is missing |
401 | the API key is missing or wrong |
413 | a limit above was exceeded; the message names it |
422 | a question is invalid for Laya, for example too much option text for the model's option window |
503 | the server is busy; retry after the Retry-After delay |
Laya reads option text through a 192-token window by default. A choice question with dozens of long options hits a 422 or loses accuracy: keep option lists short, or split them into two questions.
Hosted Laya API, self-hosting or Jev?
| Hosted Laya API (laya-ai.com, opening soon) | Self-hosted laya-serve | TypeSafe Jev | |
|---|---|---|---|
| Model | open Laya checkpoints | open Laya checkpoints | closed Jev model |
| Price | $0.05 per 1M input tokens | your hardware only (Apache 2.0) | $0.042 per 1M input tokens |
| To start | API key right after sign-up, 20 free calls | install the package, download weights, run a server | a TypeSafe account |
| Servers to run | none | yours | none |
| Where requests run | our hosted servers | your servers | TypeSafe's cloud |
| Fine-tuning | not offered | yes, on your own labels | no |
| Accuracy out of the box | Laya | Laya | higher, most of all with many labels and outside English |
/v1/systemone format | yes | yes | yes |
Pick the hosted Laya API when you want Laya's price and open model without operating anything, and want to pay by card, Alipay or WeChat Pay. Pick self-hosting when data must stay on your servers or you plan to fine-tune. Pick Jev when you need the highest accuracy out of the box and do not need open weights. All three speak the same format, so you can move between them by changing the base URL.
Run your own Laya API
The official package ships the Laya API server:
pip install "laya[serve]"
LAYA_API_KEY=change-me laya-serve # http://0.0.0.0:8000/v1/systemone
It runs on CPU: on a 4-core server upstream measured 193 ms per question on laya-multilingual (Laya on CPU). For Docker, NixOS, sizing and security settings, follow Self-host Laya. ollaya and Unsloth serve the same API from a local app.
Self-hosting wins when data must stay on your servers or volume is in the millions of decisions per day. A hosted Laya API wins when you want no servers, no model downloads and a bill that follows usage.
Laya API vs Jev API
The request and response shape is the same, so switching is a base URL change. The models behind them are not the same:
- Accuracy: Jev is more accurate out of the box, and the gap grows with many labels and outside English. On Banking77's 77 intents, base Laya scored 0.363 against 0.813 for Jev in our System One benchmark.
- Speed: Laya answers in tens of milliseconds next to your application; a hosted call adds network time on top.
- Control: Laya's weights are open, so you can fine-tune it on your own labels; Jev cannot be fine-tuned.
Test both on a sample of your own data before you move production traffic. Details: Laya vs Jev.
Laya Python API
If your application is in Python and runs next to the model, you do not need HTTP at all:
import laya
agent = laya.load("convaiinnovations/laya")
result = agent.predict(state, questions)
The Python API returns the same answers object as the Laya API. See Laya in Python and Install Laya.
FAQ
Is there an official Laya API?
No. Convai Innovations releases Laya as open weights and a Python package, and does not operate a hosted API. Every hosted Laya API, including ours, is run by an independent provider on the open checkpoints.
Is the Laya API free?
Running it yourself is free: Laya is Apache 2.0 and you pay only for your hardware. Hosted Laya APIs charge per input token or per decision; ours plans $0.05 per 1M input tokens with 20 free calls to start.
Does the Laya API work with the TypeSafe SDK and Jev code?
Yes. The Laya API uses the same POST /v1/systemone request and response format, so a Jev client works after you change the base URL. Re-check accuracy and thresholds on your own data after switching.
Which languages does the Laya API support?
The multilingual checkpoint covers 100+ languages, and the router sends non-English text to it automatically. Accuracy outside English is lower than in English, so test your language before relying on it.
Can I try Laya before using the API?
Yes. The Laya Playground runs the real Laya checkpoints on ready-made examples, with no sign-up.
Last verified against laya 0.3.22 and the providers' public pages: September 30, 2026.