Laya AI Model Explained: Open-Weight Decision Model for Apps
Updated 2026-10-10
Short answer: Laya is an open-weight, non-autoregressive "decision model" family from Convai Innovations. You give it a state (a support ticket, an email, a JSON object) plus typed questions, and it returns structured answers with probabilities – a chosen label, a score on a scale, or the probability that a statement is true. It does not write text. Code and weights are Apache 2.0, it runs locally in Python, and the checkpoints are roughly 322M–421M parameters – tiny next to chat LLMs.
Searches for "laya ai" are still small but rising: Google's Keyword Planner (US) shows the term growing from about 10 searches a month in late 2025 to 70 in August 2026, and search interest climbed further in October 2026, after the first guides to the model appeared in late September. Below is what Laya is, how it works, and when it is the right tool.
What Laya does (and doesn't)
Most AI features in apps are really small decisions: which team should get this ticket, how urgent is it, does this message threaten to cancel? Teams usually ask a chat model and parse its JSON – which brings invalid JSON, unexpected labels and extra text. Laya replaces that step with a model whose allowed answers are defined before inference:
- choice – pick one option from a named set (e.g. billing / technical / other). Returns the option plus a distribution over all options.
- score – rate on an ordered rubric (e.g. routine / soon / blocking). Returns a distribution over levels and an expected score.
- noul – estimate a yes/no proposition as P(true) ("Does the customer explicitly threaten to cancel?").
Several questions can be asked about the same state in one call. What Laya does not do: draft replies, summarise, explain its reasoning in prose. Keep a generative LLM for those parts.
How it works
Instead of generating tokens one by one, Laya encodes the state and each question with a bidirectional encoder and scores the permitted answers with decision heads. The English checkpoint uses a ModernBERT encoder; the multilingual one uses mmBERT. A built-in Router picks the right checkpoint for the input's language and script.
The three checkpoints
| Checkpoint | Encoder, size | Default context | Use it for |
|---|---|---|---|
laya |
ModernBERT-large, ~421M params | 512 tokens | English decision workloads |
laya-multilingual |
mmBERT-base, ~322M params | 1,024 tokens | Non-English or mixed-language traffic |
laya-typed-decisions |
ModernBERT-large, ~421M params | 1,024 tokens | Workflows like its typed-decisions training tasks |
The multilingual checkpoint can be configured up to 8,192 tokens, but the project itself reports more variable accuracy on long documents (especially beyond roughly 4,000 tokens). Treat long context as something to test, not a given.
Run Laya locally
You need Python 3.10 or newer. Install the package in a virtual environment:
python3 -m venv .venv
.venv/bin/python -m pip install laya
Then pass a state and your questions to the Router:
from laya import Router
router = Router()
state = {"subject": "Duplicate charge",
"body": "We were billed twice for March. Please refund today or we will cancel."}
questions = {
"department": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, errors",
"other": "everything else"}},
"urgency": {"type": "score", "instructions": "How urgent is it?",
"criteria": ["routine", "soon", "blocking"]},
"churn_risk": {"type": "noul", "instructions": "Does the customer explicitly threaten to cancel?"},
}
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"])
The first call downloads the checkpoint from Hugging Face, so cache it in production. The project also documents an HTTP server (laya[serve]).
Where Laya fits – and its limits
Good fits: support-ticket routing, lead qualification, content review flags, and routing requests between agents or models. All cases where the answer space is small and known.
Limits worth knowing before you ship:
- Probabilities are not guarantees. A noul of 0.83 is an estimate, not an 83% certainty. Check calibration on your own labelled examples before using confidence thresholds.
- Scores are its weaker area in the project's own evaluation – check adjacent levels (soon vs blocking) carefully.
- Many options hurt. Option text shares a fixed budget; with dozens of categories, try a coarse first choice, then a smaller second one.
- Model output is input to your policy. A high churn-risk value should open a review task, not issue a refund on its own.
Not to be confused with
- Laya (notification app) – an open-source, local-first desktop app that merges Slack, Gmail, GitHub, Jira, Notion, Outlook and calendar notifications into one AI feed. Same name, different project.
- Layla – an AI trip planner (layla.ai) and a separate offline AI assistant app; one letter apart, often mixed up in search.
If you need a model that writes – prompts, stories, captions – a generative model is still the right tool; see our AI story generator for an example.
FAQ
Is the Laya AI model open source?
Yes. The code and published weights use the Apache 2.0 license, and the weights are on Hugging Face under Convai Innovations.
Can Laya run offline?
Yes, once the checkpoint is downloaded the Python package runs inference locally on your own hardware.
Does Laya generate text or JSON explanations?
No. It returns typed answers with probabilities. Pair it with a generative model if you need explanations or replies.
Which Laya checkpoint should I use?
laya for English, laya-multilingual for other or mixed languages (or let the Router decide), laya-typed-decisions for workloads like its training tasks. Pin the package and model version for reproducible results.
Is Laya a chatbot like ChatGPT?
No. It is a small decision model for bounded questions inside software, not an assistant you chat with.