Ollama 0.35: Jev-Style Decision Models for Local AI

⚡ Quick Take
Ollama’s release of version 0.35 fundamentally changes local AI infrastructure by introducing native "Jev-style decision models"-bypassing generative text in favor of lightning-fast, structured logic for classification, scoring, and routing.
Summary:
Ollama has officially introduced support for Jev-style decision models through a dedicated /v1/systemone endpoint, enabling developers to run typed classification and scoring tasks entirely locally. This shifts local AI capabilities from purely generative text to structured, programmatic decision-making.
What happened:
With the rollout of version 0.35, Ollama partnered with AI labs to host specialized models—like Bespoke Labs’ Nimble (9B) and Together AI’s Tev1 (4B and 0.8B)—that take a state and question schema and return exact choices, probabilities, and scores rather than prose.
Why it matters now:
Developers have spent the last two years fighting brittle JSON outputs and complex prompt engineering just to get large language models to categorize data reliably. This release formalizes a critical architectural split in the AI ecosystem: separating generative tasks from deterministic decision tasks at the local compute layer.
Who is most affected:
AI engineers, MLOps teams, and enterprise developers building real-time triage systems, content moderation pipelines, and intelligent model routers.
The under-reported angle:
This is a direct threat to cloud-based LLM classification APIs. By allowing developers to run lightning-fast, 0.8B to 9B decision models locally with zero per-call token costs, enterprises can build highly scalable, privacy-first routing layers that never touch the public internet.
🧠 Deep Dive
Have you ever tried forcing a chatty generative model to give you a clean yes-or-no answer for a production system? It rarely goes smoothly. Generative AI models are inherently chatty, but enterprise infrastructure requires deterministic logic. Until now, developers building local AI workflows had to rely on text-generation models to classify data, triage tickets, or route prompts. This meant wrestling with complex prompt engineering, forcing JSON-mode constraints, and enduring high latency just to get a model to output a simple "yes" or "no." Ollama’s version 0.35 dismantles this bottleneck by integrating TypeSafe’s Jev API format, creating a specialized pipeline specifically for AI logic.
Through the new /v1/systemone endpoint, Ollama has effectively bifurcated its platform. Instead of processing a prompt and streaming text tokens, decision models digest a state (context) and a rigid schema of questions, returning exact choices, confidence probabilities, and rubric scores. This allows developers to construct highly reliable local decision pipelines—such as intent classification or content moderation—without the operational complexity of parsing unstructured text outputs or relying on traditional, inflexible ML classifiers.
The hardware and latency implications are significant, primarily due to the specialized models launching alongside this capability. Bespoke Labs has introduced Nimble, a 9B parameter model fine-tuned from Qwen3.5, designed as a robust heavyweight for complex probability mapping across a 256K context window. In contrast, Together AI has released the experimental Tev1 family, which includes an ultra-lightweight 0.8B variant. The availability of a sub-1B parameter decision model means developers can run continuous, real-time ticket triage or model routing on consumer-grade edge devices or tightly constrained server environments without thermal throttling or GPU bottlenecks.
By replacing prompt-based classification with the /v1/systemone endpoint, the ecosystem is solving a massive MLOps pain point: probability calibration. Traditional LLMs are notoriously overconfident and poorly calibrated when generating text-based classifications. By outputting raw probabilities mapped to predefined allowed answers, Ollama’s decision models give MLOps teams the numerical visibility needed to set threshold cutoffs, build confusion matrices, and reliably measure latency—essential steps for deploying autonomous AI agents in production.
From what I've seen, this rollout reflects the industry's rapid shift toward "Compound AI Systems." Developers are no longer looking for a single massive model to do everything. Instead, they want a microservices architecture for intelligence. An application can now use a hyper-fast 0.8B local Tev1 model to classify user intent for free, and only wake up a heavier cloud LLM or a local 70B generative model if the probability scores dictate that complex reasoning is required.
📊 Stakeholders & Impact
Stakeholder / Aspect | Impact | Insight |
|---|---|---|
AI Developers & MLOps | High | Eliminates brittle JSON parsing and prompt-engineering hacks in favor of deterministic APIs and calibrated probability scores. |
Cloud AI Providers | Negative | Local, cost-free triage endpoints threaten API query volumes, as simple classification and routing tasks migrate from the cloud to the edge. |
Enterprise Infrastructure | High | Enables localized, privacy-compliant, zero-latency moderation and routing pipelines without external data leakage. |
Hardware / Silicon (Edge) | Significant | The introduction of 0.8B logic-focused models allows sophisticated AI routing on consumer-grade chips and constrained edge servers. |
✍️ About the analysis
This independent analysis synthesizes official Ollama documentation, GitHub version-release data, and model library specifications to provide technical leaders, CTOs, and developers with an architectural view of the shift toward local decision models.
🔭 i10x Perspective
The introduction of native decision models signals the maturation of the intelligence infrastructure stack. We are moving past the era where a monolithic 70B parameter model is wasted on answering a simple categorical question. As open-source providers like Together AI and Bespoke Labs optimize fine-tunes strictly for logic over language, lightweight local routers will become the mandatory "front door" for all enterprise AI applications. Watch for this architecture—where local decision models act as the high-speed traffic controllers for heavier generative tasks—to become the default standard for building reliable, autonomous agents over the next 18 months.
Related News

DeepSeek TileLang: Open-Source CUDA Alternative for Huawei Ascend
DeepSeek open-sources TileLang and compute libraries to program Huawei Ascend chips, challenging Nvidia CUDA dominance. Learn how this toolkit enables efficient LLM training without Nvidia GPUs.

OpenAI Agents: API, SDK, AgentKit & Dots Explained
OpenAI launched a fragmented suite of agent tools including the Agents API, SDK, AgentKit, and consumer Dots. Discover how developers and enterprises can choose the right runtime while managing security risks. Explore the analysis.

OpenAI Dots: Persistent AI Agents on GPT-6 Astra
OpenAI Dots mark a shift to always-on AI agents powered by GPT-6 Astra. Each agent runs on its own cloud computer across 4,000+ apps. Discover the move from reactive chatbots to ambient intelligence.