innFactory AI Consulting from Rosenheim adds Julia-1 because the model extends the category of System One models we explained in Jev by TypeSafe: the AI model that writes no text with a very small, open variant that runs on CPU. This page is based on the model card SupersonicLabs/Julia-1, the ONNX variant Julia-1-ONNX, the organisation page on Hugging Face and the Supersonic Labs website. In addition we cite figures from media reports (MarkTechPost, 26 September 2026) and mark them as such. This overview reflects the state as of October 2026.
What is Julia-1?
Julia-1 is a decision model, also called a System One model. A classic language model (LLM) generates text token by token. A decision model generates no text: it receives a state (for example a support message), a question and a list of predefined answers and returns probabilities over those answers. That is fast, easy to audit and suited to classification, routing and decision nodes in workflows. We cover the same model class for Jev, Cloudflare Clef and Laya.
On its website Supersonic Labs describes itself as “independent research into useful, local artificial intelligence” and Julia-1 as the first model of the Julia family. Per the model card Julia-1 has 144.3 million parameters and is based on mmBERT-small (a multilingual ModernBERT encoder from JHU CLSP). The licence is Apache 2.0, the weights are in Safetensors format, and the training pipeline is not part of the repository. Julia-1 is not a chat or text generation model.
Inputs and question types
The interface is typed, and the question types follow the pattern of the other decision models:
- choice: one option from 2 to 20 entries (ID with description).
- score: a level on an ordered rubric with 2 to 20 descriptions.
- noul: boolean decision, optionally with descriptions for true and false.
- Context: up to 8,192 combined tokens for state, question and options; the model card gives an example configuration with a 512-token budget for question and options and at most 48 tokens per option.
- More than 20 options: per the model card a hierarchical router narrows the selection via narrowing and reranking. The grouped result is explicitly not a global probability distribution.
- Input: text only. The output is typed decisions with probabilities.
Benchmarks: figures per the model card
All figures are vendor figures, measured per the model card on 24 September 2026 on an H200 in BF16.
| Benchmark | Julia-1 (per model card) | Comparison (Jev reference per model card) |
|---|---|---|
| Typed decisions, overall (2,000 questions) | 73.15% | not stated |
| of which choice / score / noul | 71.33% / 68.88% / 80.67% | not stated |
| AG News (4 labels, 100 examples) | 94% | 91% |
| DAIR Emotion (6 labels, 100 examples) | 86% | 48% |
| Banking77 (72 labels via 16-item shortlist, 100 examples) | 64% | 87% |
| MASSIVE, 18 labels, 52 locales | 71.5% | not stated |
On MASSIVE the model card gives 86.75 percent for English (en-US) and 86.25 percent for Portuguese (pt-PT). The pilots cover only 100 examples each and should be read with caution. The weak Banking77 result (64 versus 87 percent) is already in the model card; per media reports (MarkTechPost, 26 September 2026) it is highlighted as well. For tasks with many similar labels, use Julia-1 only with a shortlist and your own evaluation.
On reproducibility: the model card gives, for CPU FP32 (26 September), 426 of 600 choice, 542 of 800 score and 483 of 600 noul questions correct, close to the GPU figures.
The model card names limits: Julia-1 cannot supply missing facts, solve algebra or carry long chains of calculation; it struggles with ambiguous wording, unfamiliar domains and long label lists. The benchmarks do not establish accuracy for a new domain, every language or high-stakes use.
Hardware and operation
- CPU: inference works per the model card with a standard PyTorch installation (Python 3.11 or newer); no GPU is required.
- GPU: optional with CUDA and BF16 via
device="cuda". - Size: 550.5 MiB for the FP32 weights.
- ONNX: Julia-1-ONNX contains an export with a WebGPU adapter (WGSL kernels) and a Rust WebAssembly tokenizer, plus a Node N-API binding. Per the model card the median latency in the browser was 75.47 ms per decision, and 100 of 100 predictions matched the original. Accuracy was not re-validated for WebGPU.
- Latency per media reports: MarkTechPost cites about 33 ms on an Apple M4 and about 203 ms on the CPU of a Samsung tablet. We could not verify these figures in the primary sources.
Data protection and self-hosting
The weights are Apache 2.0 and the model is small enough for CPU operation. That makes self-hosting in the EU the obvious route: on your own infrastructure, with an EU cloud provider or even in the browser via the ONNX variant, without inputs leaving your system. We found no hosted API documented by the vendor in the primary sources we reviewed; per media reports an API is planned. The model card, the Hugging Face organisation page and the website do not state the company’s location explicitly, so we make no statement on it. For processing of personal data the usual applies: clarify data processing agreements, legal basis and evaluation before deployment.
Positioning: Julia-1, Jev, Clef-flash and Laya
| Julia-1 | Jev | Clef-flash | Laya | |
|---|---|---|---|---|
| Size | 144.3M parameters (mmBERT-small) | not published | 9B (Qwen base) | 0.3 to 0.4B (encoder) |
| Licence | Apache 2.0, open weights | proprietary API | Apache 2.0, open weights | Apache 2.0, open weights |
| Operation | self-hosting (CPU possible), no hosted API documented | API only (US) | Workers AI or self-hosting | self-hosting only |
| Inputs | text | text, JSON | text, JSON, images, video | text, JSON |
| Hosted price | no hosted API documented | USD 0.042 per 1M input | USD 0.09 per 1M input | no hosted API |
| EU residency | via self-hosting | not documented | via self-hosting | via self-hosting |
Julia-1 is the smallest and easiest-to-operate model in this group with a multilingual encoder. It replaces neither an LLM nor a larger decision model; the comparison figures come from the vendor.
Our recommendation
For companies in the DACH region Julia-1 is interesting as a very small, self-hosted decision model for routing, triage and pre-filtering where data stays on your own infrastructure. Before adopting it, check accuracy on your German data (the model card gives no German figures), handling of long label lists and the maturity of a community release from a young vendor. We assess its use as a front stage to language models in CompanyGPT or as a router in the AI Gateway per project. For an assessment of whether a decision model fits your architecture, contact innFactory AI Consulting.
Related decision models
The decision model (System One) category also includes Microsoft-Decision-1, Cloudflare Clef, Laya, Jev, GLiNER2.5-Decide. OpenAI is moving in the same direction with the Decisions API based on GPT-6 Luna; details are on the OpenAI GPT page.
