Laya, an open-source decision engine from Convai Innovations, became one of the most-starred machine-learning repositories of September 2026. This non-autoregressive System 1 model uses a 421-million-parameter encoder to read text and typed questions, returning a probability for every option in a single forward pass with zero output tokens. It prioritises speed and calibrated probabilities, offering an alternative to TypeSafe’s Jev. This guide tests those promises on the banking domain of the CLINC150 intent dataset, measuring zero-shot accuracy against trained classifiers, the impact of option wording and order, the reliability of shipped probabilities, the effects of temperature calibration, an abstention gate fitted to an error budget, out-of-scope traffic, and typed outputs from a pydantic schema.
In this article
Setup and reproducibility
The test installs the released package, laya 0.3.27, and loads the English checkpoint. Two choices keep the run reproducible. By default laya.load follows the Hugging Face main branch, so the revision the library’s own authors reviewed gets pinned via laya.PINNED_REVISIONS. On CUDA Laya autocasts to half precision, so the switch is off to keep every device in fp32 and let a GPU run reproduce the CPU numbers shown here.
Printing the checkpoint’s shipped temperatures reveals the first finding before any prediction. The entry for choice questions with eleven or more options is 0.10, outside the valid range, so the loader clamps it to 0.5 and warns. A temperature below one sharpens probabilities, so every answer to a question with that many options will look twice as certain as the raw model is.
One forward pass, three typed questions
A single call to predict answers three typed questions about a support ticket in a single forward pass. The ticket reads: “Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan.” The questions ask for the department as a choice, the urgency as a score from 0 to 2, and the churn risk as a yes/no.
The result carries a probability for every option and two confidence fields that are easy to confuse. answer_confidence is the probability of the reported answer, and it is the quantity that calibration, the abstention gate and every metric later in this tutorial use. confidence is one minus the normalized entropy, whose scale depends on how many options a question has. The usage block shows zero output tokens, because Laya scores the options it is given and never generates text.
Cost model
Before building on Laya, the cost of a forward pass on a single message gets measured. Each question becomes its own row, paired with the message, so sixteen yes/no questions take about eight times as long as one. All the options of a choice question share one row and its head budget, so a forty-option choice costs barely twice a three-option choice.
The design rule is clear: ask one choice question with many options, not many yes/no questions. A 16 yes/no query takes roughly 16 rows and incurs a higher latency cost compared to a single choice question with 40 options.
What it means
Developers building routing or triage systems can expect faster inference than traditional generation models. The architecture allows multiple decisions per request without paying the token generation cost. However, the model requires careful handling of option counts due to temperature clamping and the need to group choices to maintain efficiency.



