innFactory AI Consulting from Rosenheim adds GLiNER2.5-Decide because Fastino has released another open model in the System One category we explained in Jev by TypeSafe: the AI model that writes no text. This page is based on Fastino’s announcement of 24 September 2026, the model cards on Hugging Face (GLiNER2.5-Decide, GLiNER2.5-multi-Decide, GLiNER2.5-Decide-1B) and the GLiNER2 paper. This overview reflects the state as of October 2026.
What is GLiNER2.5-Decide?
Decision models, also called System One models, generate no text. They take a state and a list of typed questions and return probabilities for predefined answers. That suits routing, triage and checking steps in front of or alongside a language model. Known examples are Jev, Cloudflare Clef and Laya.
GLiNER2.5-Decide by Fastino is such a model with 340 million parameters under Apache 2.0. Technically it is an encoder based on DeBERTa-v3-large, not a generative model. The lineage goes back to GLiNER2 (Zaratiana et al., arXiv, 24 July 2025), an encoder-based system that unifies entity recognition, text classification and structured extraction in one model through a schema-driven interface. Per the model card GLiNER2.5-Decide is a “specialist classifier for operational decisions” for intent routing, sentiment, document classification and yes/no questions over passages; it has “no reasoning, no explanations, no open-ended answers”.
Inputs, schema and output
The model receives a text and a set of typed questions. Per Fastino each question declares its permitted answers and whether one answer, several answers or an ordered value is expected. The encoder computes a compatibility score for every permitted answer; a constrained decoder then searches for the highest-scoring joint assignment that satisfies the declared rules. Fastino illustrates this with a safety classification: decided independently, “safe” (0.52) appears next to “prompt_injection” (0.82); joint decoding returns a coherent “unsafe” and “prompt_injection”.
In practice you load the model through the Python library gliner2 (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and pass text and label sets; the output is a dictionary with one prediction per label head. Documentation is at docs.fastino.ai.
Variants
| GLiNER2.5-Decide | GLiNER2.5-multi-Decide | GLiNER2.5-Decide-1B | |
|---|---|---|---|
| Parameters | 340 million | 287 million | about 1 billion |
| Encoder | DeBERTa-v3-large | mDeBERTa-v3-base | Ettin encoder (1B) |
| Languages | English | multilingual | English |
| Accuracy (model card, internal benchmark, 17 domains) | 60.2% | 56.7% | 59.6% |
| Licence | Apache 2.0 | Apache 2.0 | Apache 2.0 |
Important for the German market: per the model card the base model is English. For German text you need the multilingual variant, which per the model card is to be used for multilingual input. Fastino publishes no German evaluation, so evaluate on your own data.
Benchmarks: Fastino’s figures
Fastino compares on the internal benchmark “Fast Decisions” (17 datasets, 5,100 examples). These are vendor figures; per Fastino it is an internal benchmark and not JevBench.
| Model | Average (per Fastino, internal benchmark) |
|---|---|
| GLiNER2.5-Decide | 60.1% (leads on 9 of 17 datasets per Fastino) |
| JevK5 | 57.5% |
| SemIf | 56.4% |
| GLiFormer | 49.0% |
| Laya | 46.6% |
Individual values for GLiNER2.5-Decide: support intent 75.3%, banking intent 64.3%. The model card states 60.2% exact-match accuracy, the blog 60.1%; we use the blog value and point out the discrepancy.
Latency and hardware
Per Fastino the median latency (p50, end to end, 64 tokens) is:
| Hardware | p50 |
|---|---|
| Intel Xeon Platinum 8581C (48 vCPU) | 167.3 ms |
| NVIDIA V100 | 38.3 ms |
| NVIDIA T4 | 43.6 ms |
| NVIDIA L4 | 43.4 ms |
| NVIDIA A100 | 47.3 ms |
At 1,024 tokens Fastino states 52.6 ms (A100), 75.6 ms (V100) and 131.4 ms (L4). Per Fastino the model runs locally on CPUs and can be deployed in air-gapped environments. For Fastino’s hosted inference the blog names the API https://agent.fastino.ai.
Data protection, regions and self-hosting
The weights of all three variants are on Hugging Face under Apache 2.0. Because the 340-million-parameter model runs on CPU or GPU, it can be operated on your own infrastructure or with an EU cloud provider; processing then stays in the EU and data does not leave your environment. For Fastino’s hosted API EU data residency is not documented, nor is a company location on the Fastino website. For personal data we recommend self-hosting or a prior contractual clarification with the vendor.
Positioning: GLiNER2.5-Decide, Jev, Clef-flash and Laya
| GLiNER2.5-Decide | Jev | Clef-flash | Laya | |
|---|---|---|---|---|
| Size | 340M (encoder) | not published | 9B (Qwen base) | 0.3 to 0.4B (encoder) |
| Licence | Apache 2.0, open weights | proprietary API | Apache 2.0, open weights | Apache 2.0, open weights |
| Operation | self-hosting, Fastino API | API only (US) | Workers AI or self-hosting | self-hosting only |
| Inputs | text | text, JSON | text, JSON, images, video | text, JSON |
| Context | not stated | 64,000 tokens | 65,536 tokens | 512 to 1,024 tokens |
| Languages | base English, multi variant multilingual | not stated | not stated | not stated |
| EU residency | via self-hosting | not documented | via self-hosting | via self-hosting |
In size GLiNER2.5-Decide sits between Laya and Clef-flash; like Laya it remains a small encoder and needs no GPU. The vendors’ benchmark figures come from their own measurements and are not directly comparable. No model in this group replaces an LLM; for text, summaries and code, models such as Qwen or Mistral remain in charge.
Our recommendation
For companies in the DACH region GLiNER2.5-Decide is interesting above all as a self-hosted, CPU-capable classifier, for example for intent routing and ticket triage. Check two points before adopting it: for German text you need the multilingual variant, which you should evaluate on your own data, and the vendor figures come from an internal benchmark. We assess its use as a front stage to language models in CompanyGPT or as a router in the AI Gateway per project. For an assessment of whether a decision model fits your architecture, contact innFactory AI Consulting.
Related decision models
The decision model (System One) category also includes Microsoft-Decision-1, Cloudflare Clef, Laya, Jev, Julia-1. OpenAI is moving in the same direction with the Decisions API based on GPT-6 Luna; details are on the OpenAI GPT page.
