↓ Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
SPECIALIZED Convai Innovations India

Laya (Convai Innovations)

Laya by Convai Innovations is an open, non-autoregressive decision model (Apache 2.0, 322 to 421 million parameters) with a Jev-compatible API: choice, score and noul with calibrated probabilities in roughly 33 to 40 ms per question. Weights on Hugging Face, self-hosting on your own EU infrastructure, CPU, GPU, ONNX or in the browser. As of 30 September 2026.

License Apache 2.0
GDPR Hosting Available
Context 8192 Tokens
Modality Text, JSON → Typed decisions (choice, score, noul) with probabilities

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
laya-typed-decisions (ModernBERT-large, 421M)
September 2026 (first upload to Hugging Face mid-September 2026, package laya 0.3.x)
Fine-tuned for typed decisions: per the model card 0.766 accuracy on typed-decisions (Jev: 0.727) and 0.950 on AG News (Jev: 0.910) Calibrated probabilities, ECE 0.081 after temperature calibration (vendor figure) 1,024 token context, long documents via predict_long in windows of up to 8,192 tokens Jev-compatible endpoint POST /v1/systemone via laya-serve, optional MCP server Runs on CPU, GPU (CUDA, MPS, XPU) and TPU; ONNX export with INT8 quantisation, TypeScript SDK for in-browser inference Apache 2.0 licence, fine-tuning and RLCD training via the SDK
On choice questions with many options (77 labels) 0.425 versus 0.870 for Jev per the model card; mitigated by a larger token budget or embedding-based shortlisting Score questions are the weakest primitive (SST-5: 0.372 per the model card) Encoder with 421 million parameters: no frontier-level language understanding, instructions in the request have limited effect All comparison figures from the vendor, no paper, no independent evaluation; young project (since September 2026)
Current
laya-multilingual (mmBERT-base, 322M)
September 2026
Reads over 100 languages per the vendor; 45 of 51 tested languages well above chance (model card) Smallest checkpoint: 32.8 ms per question and 72.3 ms for ten questions on a Tesla T4 (vendor figure) 1,024 token context, up to 8,192 tokens in windows; router picks the checkpoint automatically by script and language
Ships uncalibrated, temperature calibration on your own data required Position bias on score questions per the project's issue tracker No separate evaluation for German published
Current
laya (English, ModernBERT-large, 421M)
September 2026
English base checkpoint, 39.5 ms per question on a Tesla T4 (vendor figure) Starting point for your own fine-tuning on choice, score and noul tasks
Only 512 tokens of context Near chance on typed-decisions without fine-tuning (0.362 versus 0.461 majority class, per the model card) Noul questions sometimes follow the option labels instead of the state, per the model card
Current

Use Cases

Typical applications for this model

Ticket and email routing on-premises
Classification with confidence in regulated environments
Guardrail and moderation decisions without the cloud
Scoring on rubrics at high volume
Decision nodes in AI agents with data sovereignty
Inference in the browser or on edge devices

Technical Details

API, features and capabilities

API & Availability
Availability No hosted API from the vendor; self-hosting via pip install laya, laya-serve (HTTP, Jev-compatible), Docker containers (CPU, NVIDIA GPU, ARM64, DGX Spark), ONNX Runtime, TypeScript SDK in the browser
Latency (TTFT) 32.8 to 39.5 ms per question on a Tesla T4 (vendor figure)
Throughput 103 to 332 questions/s batched on a Tesla T4 (vendor figure) Tokens/Sec
Features & Capabilities
Structured Output
Training & Knowledge
Knowledge Cutoff not documented (encoder backbones ModernBERT and mmBERT)
Fine-Tuning Available (Full fine-tuning (SDK), RLCD (Reinforcement Learning for Calibrated Decisions))
Language Support
Best Quality English
Supported Over 100 languages in the multilingual checkpoint (per the vendor); 45 of 51 tested languages well above chance
No separate evaluation for German published; evaluate on your own German data before deployment

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-hosted
Your own EU infrastructure or EU cloud
Recommended: open weights, full data control, no transfer to the vendor
License & Hosting
License Apache 2.0
Security Filters None (decision model without text generation)
On-Premise Edge-capable

Benchmarks

Performance comparison with standardized tests

typed-decisions (laya-typed-decisions, vendor figure)
76.6
AG News (vendor figure)
95

innFactory AI Consulting from Rosenheim adds Laya to the model overview because it makes the System One idea we explained in Jev by TypeSafe: the AI model that writes no text available as an open model. For companies in the DACH region that is the decisive difference: Laya can run on your own infrastructure in the EU. This page is based on the vendor’s model card on Hugging Face, the GitHub repository and the documentation. As of 30 September 2026.

What is Laya?

Laya is a family of non-autoregressive decision models from Convai Innovations Pvt. Ltd. in Kasaragod (Kerala, India), a company that by its own description builds sovereign on-device AI for application security, compliance and clinical research. The models were released in September 2026 under the Apache 2.0 licence on Hugging Face.

The principle is the same as with Jev: a state (text or JSON) and typed questions go in, typed answers with calibrated probabilities come out, all questions in a single forward pass. Three question types are available: choice (one option from a list), score (a level on an ordered scale) and noul (probability of yes). Laya generates no text.

Technically, Laya consists of an encoder backbone and its own decision head: two transformer layers, a scorer for option markers at [MASK] positions and an act/escalate head. Per the model card it is trained with Reinforcement Learning for Calibrated Decisions (RLCD) against strictly proper scoring rules.

Three checkpoints

CheckpointBackboneParametersContextUse
layaModernBERT-large421M512 tokensEnglish, base for your own fine-tuning
laya-multilingualmmBERT-base322M1,024 tokens (up to 8,192 in windows)over 100 languages
laya-typed-decisionsModernBERT-large421M1,024 tokensfine-tuned for typed decisions

A built-in router detects the script and language of the input and picks the matching checkpoint with under a millisecond of overhead.

The vendor’s comparison figures

Convai publishes its own measurements against Jev in the model card. For context: these are vendor figures; there is no paper and no independent evaluation yet.

MeasurementLayaJev (per Convai)
typed-decisions, accuracy (fine-tuned)0.7660.727
AG News, accuracy0.9500.910
ECE after temperature calibration0.0810.144
Latency, one question (Tesla T4)32.8 to 39.5 ms236 to 276 ms
Choice with 77 options, accuracy0.4250.870

The last row shows the limit: with many options Laya drops off markedly and needs a larger token budget or embedding-based shortlisting. Also documented: without fine-tuning the base checkpoints sit at 0.362 on typed-decisions, below the majority class (0.461); score questions are the weakest primitive (SST-5: 0.372); without temperature calibration the model is over-confident.

Operation: open, local, Jev-compatible

  • Installation: pip install laya, weights are loaded from Hugging Face; use via the SDK (laya.load, decide() with JSON schema) or as a Transformers pipeline.
  • HTTP: laya-serve exposes the Jev-compatible endpoint POST /v1/systemone; anyone who built for Jev can switch the base URL. Optionally an MCP server with tools for single and batch predictions.
  • Hardware: CPU, GPU (CUDA, MPS, XPU) and TPU; Docker containers for CPU, NVIDIA GPU, ARM64 and DGX Spark. Batched, a Tesla T4 handles 103 to 332 questions per second per the vendor.
  • Edge and browser: ONNX export with INT8 quantisation; the TypeScript SDK runs inference via ONNX directly in the browser.
  • Fine-tuning: the SDK includes training and RLCD; for production accuracy, adapting the model to your own task is usually required per the model card.

Data protection and sovereignty

Laya never leaves your infrastructure: no hosted API, no telemetry to the vendor, weights under Apache 2.0. That meets the requirement many of our customers have for classifiers in the inbox, the ticket system or in guardrails: personal data stays in the company’s own data centre or with an EU cloud provider. In return, responsibility for operation, calibration and evaluation lies entirely with the company.

Positioning versus Jev

LayaJev
Licence and operationApache 2.0, self-hostingproprietary API in the US
Model size322 to 421 million parameters (encoder)not published
Language understandingencoder class, fine-tuning usually needednear frontier per TypeSafe, customised via prompt
Context512 to 1,024 tokens, windows up to 8,19264,000 tokens combined
CostinfrastructureUSD 0.042 per 1M input tokens
Data protectiondata stays in-houseDPA with EU SCCs, zero data retention for enterprise

For narrowly defined, trainable tasks with data sovereignty Laya is the natural choice; for broad, changing questions without training data and without an EU requirement Jev remains stronger. Since 1 October 2026 there is a third option with Clef and Clef-flash by Cloudflare: open models with 27 and 9 billion parameters, image and video input and a Jev-compatible interface, hosted on Workers AI or self-hosted. Neither replaces an LLM: for text, summaries and code, self-hosted models such as Mistral, Qwen or GPT-OSS remain in charge.

Integration with CompanyGPT and AI Gateway

As a self-hosted model, Laya can run next to the LLMs in your own EU cloud, for example as a front stage that classifies requests and hands them to the right language model. We assess its use per project, for instance for routing and guardrails in CompanyGPT workflows. For an assessment of whether a System One model fits your architecture, contact innFactory AI Consulting.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is Laya?

Laya is a family of open decision models from Convai Innovations in Kerala, India, released in September 2026 under the Apache 2.0 licence on Hugging Face. Like Jev by TypeSafe, Laya answers typed questions (choice, score, noul) about a state with calibrated probabilities in a single forward pass, without generating text. It is built on encoder models (ModernBERT-large with 421 million and mmBERT-base with 322 million parameters) plus a decision head trained with RLCD.

Can I run Laya in the EU?

Yes. The weights are on Hugging Face under Apache 2.0, and the laya Python package ships the SDK, the laya-serve HTTP server with a Jev-compatible /v1/systemone endpoint, Docker containers for CPU and NVIDIA GPU, and ONNX export. Laya therefore runs on your own infrastructure or with an EU cloud provider without data leaving the company. The vendor offers no hosted API.

How does Laya compare to Jev?

Per the vendor's model card (as of September 2026), the fine-tuned laya-typed-decisions checkpoint reaches 0.766 accuracy on its own typed-decisions benchmark versus 0.727 for Jev, and 0.950 on AG News versus 0.910; calibration (ECE) is 0.081. For a single question Convai measures 32.8 ms (multilingual) to 39.5 ms (English) on a Tesla T4. On choice questions with many options (77 labels) Jev leads clearly at 0.870 versus 0.425. All figures come from the vendor; there is no independent evaluation yet.

What are Laya's limits?

Without fine-tuning the base checkpoints reach only chance level on the typed-decisions benchmark (0.362); good results require adapting the model to your task. Score questions are the weakest primitive per the model card, the model is over-confident without temperature calibration, and the multilingual checkpoint ships uncalibrated. Context is 512 or 1,024 tokens per checkpoint; longer documents are scanned in windows of up to 8,192 tokens. With 0.3 to 0.4 billion parameters Laya does not have the language understanding of a frontier model.

Which languages does Laya support?

The multilingual checkpoint on an mmBERT base reads over 100 languages per the vendor; in the model card's tests 45 of 51 languages scored well above chance, 23 for the English checkpoint. Convai publishes no separate evaluation for German, so an evaluation on your own German data is essential.

Consultation for this model?

We help you select and integrate the right AI model for your use case.