↓ Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
DECISION Fastino

GLiNER2.5-Decide (Fastino)

GLiNER2.5-Decide by Fastino: open decision model (340M parameters, Apache 2.0) answering with probabilities instead of text, CPU or GPU, multilingual variant for German.

License Apache 2.0
GDPR Hosting Available
Modality Text → Decisions for typed questions (single, multiple or ordered answer) with probabilities

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
GLiNER2.5-Decide (340M)
24 September 2026
340 million parameters, DeBERTa-v3-large encoder, Apache 2.0; per the model card runs on CPU or GPU through the GLiNER2 library Typed questions with permitted answers; per Fastino a constrained decoder searches for the highest-scoring joint assignment that satisfies the declared rules Per Fastino (internal benchmark, 17 datasets, 5,100 examples): 60.1% on average, support intent 75.3%, banking intent 64.3% Median latency per Fastino: 38.3 ms (V100), 43.6 ms (T4), 43.4 ms (L4), 47.3 ms (A100), 167.3 ms on a 48-vCPU Xeon Can be deployed in air-gapped environments (per Fastino)
Per the model card for English text; the multilingual variant is intended for German Not a general-purpose model: no reasoning, no explanations, no open-ended answers (model card) All comparison figures come from the vendor's internal benchmark, not JevBench Text input only
Current
GLiNER2.5-multi-Decide (287M, multilingual) Recommended
not stated
287 million parameters on mDeBERTa-v3-base, multilingual; per the model card to be used when the input is multilingual Apache 2.0, runs through the GLiNER2 library
Per the model card 56.7% exact-match accuracy in the internal benchmark across 17 domains (base model: 60.2% per the model card) No separate German evaluation published
Current
GLiNER2.5-Decide-1B
not stated
About 1 billion parameters (Ettin encoder), Apache 2.0, English Per the model card 59.6% exact-match accuracy on fast-decisions across 17 domains
English; larger than the 340M base model
Current

Use Cases

Typical applications for this model

Intent routing for support and banking requests
Sentiment and document type classification
Ticket and email routing
Yes/no questions over text passages
Safety classification with consistent multi-labels
Front stage in front of expensive language models

Technical Details

API, features and capabilities

API & Availability
Availability Weights on Hugging Face (GLiNER2 library, CPU or GPU); hosted inference API from Fastino at agent.fastino.ai
Latency (TTFT) 167.3 ms (48-vCPU Xeon) to 38.3 ms (V100) median at 64 tokens (Fastino measurement)
Features & Capabilities
Structured Output
Training & Knowledge
Knowledge Cutoff not documented
Fine-Tuning Available (Your own fine-tuning of the open weights, Fastino API as an inference and fine-tuning platform (per the Fastino website))
Language Support
Best Quality English
Supported Base model English; multilingual variant GLiNER2.5-multi-Decide (mDeBERTa-v3-base)
For German use the multilingual variant and evaluate it on your own German data before deployment; Fastino publishes no German evaluation

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-hosted
Your own EU infrastructure or EU cloud
Recommended for EU residency: Apache 2.0 weights, CPU-capable, GLiNER2 library
Fastino API
not documented
Hosted API at agent.fastino.ai; no documented EU data residency
License & Hosting
License Apache 2.0
Security Filters None (decision model without text generation)
On-Premise

Benchmarks

Performance comparison with standardized tests

Fast Decisions, average (GLiNER2.5-Decide, Fastino internal)
60.1
Support intent (GLiNER2.5-Decide, Fastino internal)
75.3
Banking intent (GLiNER2.5-Decide, Fastino internal)
64.3

innFactory AI Consulting from Rosenheim adds GLiNER2.5-Decide because Fastino has released another open model in the System One category we explained in Jev by TypeSafe: the AI model that writes no text. This page is based on Fastino’s announcement of 24 September 2026, the model cards on Hugging Face (GLiNER2.5-Decide, GLiNER2.5-multi-Decide, GLiNER2.5-Decide-1B) and the GLiNER2 paper. This overview reflects the state as of October 2026.

What is GLiNER2.5-Decide?

Decision models, also called System One models, generate no text. They take a state and a list of typed questions and return probabilities for predefined answers. That suits routing, triage and checking steps in front of or alongside a language model. Known examples are Jev, Cloudflare Clef and Laya.

GLiNER2.5-Decide by Fastino is such a model with 340 million parameters under Apache 2.0. Technically it is an encoder based on DeBERTa-v3-large, not a generative model. The lineage goes back to GLiNER2 (Zaratiana et al., arXiv, 24 July 2025), an encoder-based system that unifies entity recognition, text classification and structured extraction in one model through a schema-driven interface. Per the model card GLiNER2.5-Decide is a “specialist classifier for operational decisions” for intent routing, sentiment, document classification and yes/no questions over passages; it has “no reasoning, no explanations, no open-ended answers”.

Inputs, schema and output

The model receives a text and a set of typed questions. Per Fastino each question declares its permitted answers and whether one answer, several answers or an ordered value is expected. The encoder computes a compatibility score for every permitted answer; a constrained decoder then searches for the highest-scoring joint assignment that satisfies the declared rules. Fastino illustrates this with a safety classification: decided independently, “safe” (0.52) appears next to “prompt_injection” (0.82); joint decoding returns a coherent “unsafe” and “prompt_injection”.

In practice you load the model through the Python library gliner2 (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and pass text and label sets; the output is a dictionary with one prediction per label head. Documentation is at docs.fastino.ai.

Variants

GLiNER2.5-DecideGLiNER2.5-multi-DecideGLiNER2.5-Decide-1B
Parameters340 million287 millionabout 1 billion
EncoderDeBERTa-v3-largemDeBERTa-v3-baseEttin encoder (1B)
LanguagesEnglishmultilingualEnglish
Accuracy (model card, internal benchmark, 17 domains)60.2%56.7%59.6%
LicenceApache 2.0Apache 2.0Apache 2.0

Important for the German market: per the model card the base model is English. For German text you need the multilingual variant, which per the model card is to be used for multilingual input. Fastino publishes no German evaluation, so evaluate on your own data.

Benchmarks: Fastino’s figures

Fastino compares on the internal benchmark “Fast Decisions” (17 datasets, 5,100 examples). These are vendor figures; per Fastino it is an internal benchmark and not JevBench.

ModelAverage (per Fastino, internal benchmark)
GLiNER2.5-Decide60.1% (leads on 9 of 17 datasets per Fastino)
JevK557.5%
SemIf56.4%
GLiFormer49.0%
Laya46.6%

Individual values for GLiNER2.5-Decide: support intent 75.3%, banking intent 64.3%. The model card states 60.2% exact-match accuracy, the blog 60.1%; we use the blog value and point out the discrepancy.

Latency and hardware

Per Fastino the median latency (p50, end to end, 64 tokens) is:

Hardwarep50
Intel Xeon Platinum 8581C (48 vCPU)167.3 ms
NVIDIA V10038.3 ms
NVIDIA T443.6 ms
NVIDIA L443.4 ms
NVIDIA A10047.3 ms

At 1,024 tokens Fastino states 52.6 ms (A100), 75.6 ms (V100) and 131.4 ms (L4). Per Fastino the model runs locally on CPUs and can be deployed in air-gapped environments. For Fastino’s hosted inference the blog names the API https://agent.fastino.ai.

Data protection, regions and self-hosting

The weights of all three variants are on Hugging Face under Apache 2.0. Because the 340-million-parameter model runs on CPU or GPU, it can be operated on your own infrastructure or with an EU cloud provider; processing then stays in the EU and data does not leave your environment. For Fastino’s hosted API EU data residency is not documented, nor is a company location on the Fastino website. For personal data we recommend self-hosting or a prior contractual clarification with the vendor.

Positioning: GLiNER2.5-Decide, Jev, Clef-flash and Laya

GLiNER2.5-DecideJevClef-flashLaya
Size340M (encoder)not published9B (Qwen base)0.3 to 0.4B (encoder)
LicenceApache 2.0, open weightsproprietary APIApache 2.0, open weightsApache 2.0, open weights
Operationself-hosting, Fastino APIAPI only (US)Workers AI or self-hostingself-hosting only
Inputstexttext, JSONtext, JSON, images, videotext, JSON
Contextnot stated64,000 tokens65,536 tokens512 to 1,024 tokens
Languagesbase English, multi variant multilingualnot statednot statednot stated
EU residencyvia self-hostingnot documentedvia self-hostingvia self-hosting

In size GLiNER2.5-Decide sits between Laya and Clef-flash; like Laya it remains a small encoder and needs no GPU. The vendors’ benchmark figures come from their own measurements and are not directly comparable. No model in this group replaces an LLM; for text, summaries and code, models such as Qwen or Mistral remain in charge.

Our recommendation

For companies in the DACH region GLiNER2.5-Decide is interesting above all as a self-hosted, CPU-capable classifier, for example for intent routing and ticket triage. Check two points before adopting it: for German text you need the multilingual variant, which you should evaluate on your own data, and the vendor figures come from an internal benchmark. We assess its use as a front stage to language models in CompanyGPT or as a router in the AI Gateway per project. For an assessment of whether a decision model fits your architecture, contact innFactory AI Consulting.

Related decision models

The decision model (System One) category also includes Microsoft-Decision-1, Cloudflare Clef, Laya, Jev, Julia-1. OpenAI is moving in the same direction with the Decisions API based on GPT-6 Luna; details are on the OpenAI GPT page.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is GLiNER2.5-Decide?

GLiNER2.5-Decide is an open decision model with 340 million parameters under Apache 2.0, released by Fastino on 24 September 2026. It is an encoder (DeBERTa-v3-large) and generates no text: it takes a text and typed questions with permitted answers and returns a compatibility score for each answer. Per the model card it is a specialist for operational decisions such as intent routing, sentiment, document classification and yes/no questions over passages, not a general-purpose model.

Does GLiNER2.5-Decide understand German?

Per the model card the base model GLiNER2.5-Decide (340M) is intended for English text. For multilingual input Fastino points to the variant GLiNER2.5-multi-Decide (287 million parameters, mDeBERTa-v3-base). Fastino publishes no separate evaluation for German; test the multilingual variant on your own German data before deployment.

How does GLiNER2.5-Decide perform in benchmarks?

Per Fastino (internal benchmark "Fast Decisions", 17 datasets, 5,100 examples) GLiNER2.5-Decide reaches 60.1 percent on average, followed by JevK5 at 57.5, SemIf at 56.4, GLiFormer at 49.0 and Laya at 46.6 percent. Fastino itself notes that this is an internal benchmark and not JevBench. These are vendor figures without independent verification.

How fast is GLiNER2.5-Decide and what hardware does it need?

Per Fastino the median latency (p50, end to end, 64 tokens) is 167.3 ms on a 48-vCPU Intel Xeon CPU, 38.3 ms on a V100, 43.6 ms on a T4, 43.4 ms on an L4 and 47.3 ms on an A100. At 1,024 tokens it is 52.6 ms on the A100. According to the model card it runs on CPU or GPU through the GLiNER2 library, and per Fastino it can be deployed in air-gapped environments.

Can GLiNER2.5-Decide be used in a GDPR-compliant way?

The weights are on Hugging Face under Apache 2.0, the model is CPU-capable and can be self-hosted on your own infrastructure or with an EU provider, which keeps processing in the EU. Fastino also offers a hosted API (agent.fastino.ai); EU data residency is not documented there. For personal data we therefore recommend self-hosting.

How does GLiNER2.5-Decide differ from Jev, Clef and Laya?

All four are decision models (System One) that return probabilities instead of text. GLiNER2.5-Decide is an open encoder with 340 million parameters, Jev a proprietary API, Clef-flash an open 9B model on a Qwen base with image and video input, Laya an open encoder with 0.3 to 0.4 billion parameters. The base model of GLiNER2.5-Decide is English; Clef and Jev state no language coverage.

Consultation for this model?

We help you select and integrate the right AI model for your use case.