innFactory AI Consulting from Rosenheim adds Laya to the model overview because it makes the System One idea we explained in Jev by TypeSafe: the AI model that writes no text available as an open model. For companies in the DACH region that is the decisive difference: Laya can run on your own infrastructure in the EU. This page is based on the vendor’s model card on Hugging Face, the GitHub repository and the documentation. As of 30 September 2026.
What is Laya?
Laya is a family of non-autoregressive decision models from Convai Innovations Pvt. Ltd. in Kasaragod (Kerala, India), a company that by its own description builds sovereign on-device AI for application security, compliance and clinical research. The models were released in September 2026 under the Apache 2.0 licence on Hugging Face.
The principle is the same as with Jev: a state (text or JSON) and typed questions go in, typed answers with calibrated probabilities come out, all questions in a single forward pass. Three question types are available: choice (one option from a list), score (a level on an ordered scale) and noul (probability of yes). Laya generates no text.
Technically, Laya consists of an encoder backbone and its own decision head: two transformer layers, a scorer for option markers at [MASK] positions and an act/escalate head. Per the model card it is trained with Reinforcement Learning for Calibrated Decisions (RLCD) against strictly proper scoring rules.
Three checkpoints
| Checkpoint | Backbone | Parameters | Context | Use |
|---|---|---|---|---|
laya | ModernBERT-large | 421M | 512 tokens | English, base for your own fine-tuning |
laya-multilingual | mmBERT-base | 322M | 1,024 tokens (up to 8,192 in windows) | over 100 languages |
laya-typed-decisions | ModernBERT-large | 421M | 1,024 tokens | fine-tuned for typed decisions |
A built-in router detects the script and language of the input and picks the matching checkpoint with under a millisecond of overhead.
The vendor’s comparison figures
Convai publishes its own measurements against Jev in the model card. For context: these are vendor figures; there is no paper and no independent evaluation yet.
| Measurement | Laya | Jev (per Convai) |
|---|---|---|
| typed-decisions, accuracy (fine-tuned) | 0.766 | 0.727 |
| AG News, accuracy | 0.950 | 0.910 |
| ECE after temperature calibration | 0.081 | 0.144 |
| Latency, one question (Tesla T4) | 32.8 to 39.5 ms | 236 to 276 ms |
| Choice with 77 options, accuracy | 0.425 | 0.870 |
The last row shows the limit: with many options Laya drops off markedly and needs a larger token budget or embedding-based shortlisting. Also documented: without fine-tuning the base checkpoints sit at 0.362 on typed-decisions, below the majority class (0.461); score questions are the weakest primitive (SST-5: 0.372); without temperature calibration the model is over-confident.
Operation: open, local, Jev-compatible
- Installation:
pip install laya, weights are loaded from Hugging Face; use via the SDK (laya.load,decide()with JSON schema) or as a Transformers pipeline. - HTTP:
laya-serveexposes the Jev-compatible endpointPOST /v1/systemone; anyone who built for Jev can switch the base URL. Optionally an MCP server with tools for single and batch predictions. - Hardware: CPU, GPU (CUDA, MPS, XPU) and TPU; Docker containers for CPU, NVIDIA GPU, ARM64 and DGX Spark. Batched, a Tesla T4 handles 103 to 332 questions per second per the vendor.
- Edge and browser: ONNX export with INT8 quantisation; the TypeScript SDK runs inference via ONNX directly in the browser.
- Fine-tuning: the SDK includes training and RLCD; for production accuracy, adapting the model to your own task is usually required per the model card.
Data protection and sovereignty
Laya never leaves your infrastructure: no hosted API, no telemetry to the vendor, weights under Apache 2.0. That meets the requirement many of our customers have for classifiers in the inbox, the ticket system or in guardrails: personal data stays in the company’s own data centre or with an EU cloud provider. In return, responsibility for operation, calibration and evaluation lies entirely with the company.
Positioning versus Jev
| Laya | Jev | |
|---|---|---|
| Licence and operation | Apache 2.0, self-hosting | proprietary API in the US |
| Model size | 322 to 421 million parameters (encoder) | not published |
| Language understanding | encoder class, fine-tuning usually needed | near frontier per TypeSafe, customised via prompt |
| Context | 512 to 1,024 tokens, windows up to 8,192 | 64,000 tokens combined |
| Cost | infrastructure | USD 0.042 per 1M input tokens |
| Data protection | data stays in-house | DPA with EU SCCs, zero data retention for enterprise |
For narrowly defined, trainable tasks with data sovereignty Laya is the natural choice; for broad, changing questions without training data and without an EU requirement Jev remains stronger. Since 1 October 2026 there is a third option with Clef and Clef-flash by Cloudflare: open models with 27 and 9 billion parameters, image and video input and a Jev-compatible interface, hosted on Workers AI or self-hosted. Neither replaces an LLM: for text, summaries and code, self-hosted models such as Mistral, Qwen or GPT-OSS remain in charge.
Integration with CompanyGPT and AI Gateway
As a self-hosted model, Laya can run next to the LLMs in your own EU cloud, for example as a front stage that classifies requests and hands them to the right language model. We assess its use per project, for instance for routing and guardrails in CompanyGPT workflows. For an assessment of whether a System One model fits your architecture, contact innFactory AI Consulting.
