Open Source · On-Premise · Switzerland

    Ollama for Enterprise
    AI local. Data sovereign.

    CNEXT implements Ollama for Swiss organisations: run open-source language models securely on your own infrastructure — without cloud dependency, without token costs, fully nDSG-compliant.

    Why Ollama?

    Ollama makes running large language models as simple as a package manager — and completely locally.

    100% On-Premise

    Your data never leaves your infrastructure. Full control, no cloud risk.

    Open Models

    Llama 3, Mistral, Phi-3, Qwen, Gemma and many more — freely selectable.

    Fast Inference

    Optimised for CPU and GPU. Locally faster than many cloud APIs for smaller models.

    nDSG Compliant

    Comply with Swiss data protection law without cloud processing of personal data.

    Technical Capabilities

    What Ollama can do

    From the model library to OpenAI-compatible APIs and flexible deployment — Ollama provides a complete local AI platform.

    Local Model Library

    Ollama manages open-weight models like a package manager: simple pull, run and update.

    • Llama 3.1 / 3.3 (8B, 70B)
    • Mistral / Mixtral
    • Phi-4, Qwen 2.5, Gemma 3
    • Code models: DeepSeek-Coder, CodeLlama

    OpenAI-Compatible API

    Ollama speaks the same API format as OpenAI — redirect existing applications without code changes.

    • Drop-in replacement for OpenAI SDK
    • Chat completions endpoint
    • Embeddings endpoint
    • Streaming support

    Flexible Deployment Options

    From a single server to a cluster — Ollama adapts to your infrastructure.

    • On-premise (own servers)
    • Private cloud (Azure / Swiss data centre)
    • GPU servers for large models
    • Raspberry Pi / edge devices for small models
    Advantages over Cloud AI

    Ollama vs. Cloud LLMs

    For privacy-critical scenarios, high volumes or regulated industries, Ollama offers decisive advantages over cloud models.

    Data Sovereignty

    • No data leaving to US cloud
    • Meets nDSG and industry-specific requirements
    • Ideal for healthcare, banking, public sector
    • Operation in Swiss data centre possible

    Cost Efficiency

    • No token costs at high volume
    • One-time hardware investment instead of ongoing API fees
    • Smaller models run on standard hardware too
    • No vendor lock-in

    Customisability

    • Fine-tuning on your own data
    • Custom system prompts (Modelfiles)
    • Integration with your own vector databases
    • Open source — full transparency

    Use Cases in Switzerland

    Industries with high data protection requirements benefit most from local AI.

    Healthcare & Clinics

    Patient data must never leave your servers. Swiss hospitals run AI assistants for documentation, coding and research with Ollama — fully locally.

    AI-assisted medical letters locally
    ICD coding assistant
    Offline medical literature search
    No data sent to US cloud

    Banks & Insurance

    Financial institutions with strict compliance requirements benefit from local AI without external data processing.

    Contract review and summarisation
    Risk documentation locally
    Customer email drafts internally
    Regulator-compliant (FINMA)

    Industry & SME

    Manufacturing companies use local models for technical documentation, fault analysis and internal knowledge bases.

    Search technical manuals
    Fault analysis and troubleshooting
    Internal knowledge chatbots
    Offline-capable in production environments

    Government & Public Administration

    Federal agencies and cantons that want to use AI without sending data abroad — Ollama is the sovereign solution.

    AI assistant for internal documents
    No US cloud dependency
    Open-source transparency
    Operation on federal infrastructure
    CNEXT Services

    What CNEXT implements for you

    From the first installation to production operation — CNEXT accompanies you the entire way.

    Ollama Setup & Deployment

    Installation, configuration and model selection — production-ready in days.

    • Hardware sizing (CPU vs. GPU)
    • Model evaluation and selection
    • API configuration and security setup
    • Monitoring & alerting

    RAG with Local Models

    Semantic search and document chat over your internal knowledge bases — fully local.

    • SharePoint content as knowledge base
    • Local vector database (Qdrant, Chroma)
    • Run embeddings model locally
    • Hybrid search (keyword + vector)

    App Integration

    Switch existing Microsoft 365 applications to local models with a drop-in.

    • Power Automate → Ollama endpoint
    • SharePoint bot with local LLM
    • Custom chat interface (Open WebUI)
    • OpenAI SDK drop-in

    Fine-Tuning & Modelfiles

    Adapt models with your company's language and context.

    • Creating Modelfiles (system prompts)
    • LoRA fine-tuning on your data
    • GGUF conversion of custom models
    • Benchmark tests before production release

    Training & Enablement

    Your team learns how to effectively use and operate local AI models.

    • Prompting workshops for local models
    • Technical deep-dive for IT teams
    • Best practices for model selection
    • Hands-on labs with your own infrastructure

    Strategic Consulting

    When does on-premise AI pay off — and when is cloud the better choice? CNEXT helps you decide.

    • TCO analysis (on-premise vs. cloud)
    • Model roadmap for your organisation
    • Compliance check for regulated industries
    • Hybrid strategy (local + Azure)

    Local AI in Numbers

    Why more and more Swiss organisations are turning to on-premise AI.

    100%
    Data Control
    All data stays on your infrastructure
    0 CHF
    Token Costs
    No ongoing API fees per request
    70+
    Open-Source Models
    Immediately available in the Ollama library
    CH
    Data Residency
    Operation in Swiss data centre possible

    What is Ollama?

    Ollama is an open-source framework that allows large language models (LLMs) to run entirely locally — on your own hardware or a private server. Unlike cloud AI services such as ChatGPT, Ollama sends no data to external servers. CNEXT deploys Ollama for Swiss organisations that want AI capabilities without sacrificing data sovereignty and compliance with the Swiss DPA.

    70+

    open-source models supported, including Llama 3, Mistral, Qwen and Phi-4

    Ollama Library, 2025
    CHF 0

    ongoing API costs with self-hosted deployment — no token consumption, no subscription

    CNEXT
    100%

    data control: all requests and responses remain within your own network

    CNEXT

    The most common misconception about local AI is that you need powerful servers. For many Swiss SME use cases — internal search, document summarisation, HR assistants — a modern workstation is entirely sufficient.

    Marcel HaasCEO & AI Architect, CNEXT

    Last updated: July 2026

    Ready for Local AI?

    CNEXT analyses your use case and shows you whether and how Ollama delivers the most value in your environment.