Ollama for Enterprise
AI local. Data sovereign.
CNEXT implements Ollama for Swiss organisations: run open-source language models securely on your own infrastructure — without cloud dependency, without token costs, fully nDSG-compliant.
Why Ollama?
Ollama makes running large language models as simple as a package manager — and completely locally.
100% On-Premise
Your data never leaves your infrastructure. Full control, no cloud risk.
Open Models
Llama 3, Mistral, Phi-3, Qwen, Gemma and many more — freely selectable.
Fast Inference
Optimised for CPU and GPU. Locally faster than many cloud APIs for smaller models.
nDSG Compliant
Comply with Swiss data protection law without cloud processing of personal data.
What Ollama can do
From the model library to OpenAI-compatible APIs and flexible deployment — Ollama provides a complete local AI platform.
Local Model Library
Ollama manages open-weight models like a package manager: simple pull, run and update.
- Llama 3.1 / 3.3 (8B, 70B)
- Mistral / Mixtral
- Phi-4, Qwen 2.5, Gemma 3
- Code models: DeepSeek-Coder, CodeLlama
OpenAI-Compatible API
Ollama speaks the same API format as OpenAI — redirect existing applications without code changes.
- Drop-in replacement for OpenAI SDK
- Chat completions endpoint
- Embeddings endpoint
- Streaming support
Flexible Deployment Options
From a single server to a cluster — Ollama adapts to your infrastructure.
- On-premise (own servers)
- Private cloud (Azure / Swiss data centre)
- GPU servers for large models
- Raspberry Pi / edge devices for small models
Ollama vs. Cloud LLMs
For privacy-critical scenarios, high volumes or regulated industries, Ollama offers decisive advantages over cloud models.
Data Sovereignty
- No data leaving to US cloud
- Meets nDSG and industry-specific requirements
- Ideal for healthcare, banking, public sector
- Operation in Swiss data centre possible
Cost Efficiency
- No token costs at high volume
- One-time hardware investment instead of ongoing API fees
- Smaller models run on standard hardware too
- No vendor lock-in
Customisability
- Fine-tuning on your own data
- Custom system prompts (Modelfiles)
- Integration with your own vector databases
- Open source — full transparency
Use Cases in Switzerland
Industries with high data protection requirements benefit most from local AI.
Healthcare & Clinics
Patient data must never leave your servers. Swiss hospitals run AI assistants for documentation, coding and research with Ollama — fully locally.
Banks & Insurance
Financial institutions with strict compliance requirements benefit from local AI without external data processing.
Industry & SME
Manufacturing companies use local models for technical documentation, fault analysis and internal knowledge bases.
Government & Public Administration
Federal agencies and cantons that want to use AI without sending data abroad — Ollama is the sovereign solution.
What CNEXT implements for you
From the first installation to production operation — CNEXT accompanies you the entire way.
Ollama Setup & Deployment
Installation, configuration and model selection — production-ready in days.
- Hardware sizing (CPU vs. GPU)
- Model evaluation and selection
- API configuration and security setup
- Monitoring & alerting
RAG with Local Models
Semantic search and document chat over your internal knowledge bases — fully local.
- SharePoint content as knowledge base
- Local vector database (Qdrant, Chroma)
- Run embeddings model locally
- Hybrid search (keyword + vector)
App Integration
Switch existing Microsoft 365 applications to local models with a drop-in.
- Power Automate → Ollama endpoint
- SharePoint bot with local LLM
- Custom chat interface (Open WebUI)
- OpenAI SDK drop-in
Fine-Tuning & Modelfiles
Adapt models with your company's language and context.
- Creating Modelfiles (system prompts)
- LoRA fine-tuning on your data
- GGUF conversion of custom models
- Benchmark tests before production release
Training & Enablement
Your team learns how to effectively use and operate local AI models.
- Prompting workshops for local models
- Technical deep-dive for IT teams
- Best practices for model selection
- Hands-on labs with your own infrastructure
Strategic Consulting
When does on-premise AI pay off — and when is cloud the better choice? CNEXT helps you decide.
- TCO analysis (on-premise vs. cloud)
- Model roadmap for your organisation
- Compliance check for regulated industries
- Hybrid strategy (local + Azure)
Local AI in Numbers
Why more and more Swiss organisations are turning to on-premise AI.
What is Ollama?
Ollama is an open-source framework that allows large language models (LLMs) to run entirely locally — on your own hardware or a private server. Unlike cloud AI services such as ChatGPT, Ollama sends no data to external servers. CNEXT deploys Ollama for Swiss organisations that want AI capabilities without sacrificing data sovereignty and compliance with the Swiss DPA.
ongoing API costs with self-hosted deployment — no token consumption, no subscription
CNEXTdata control: all requests and responses remain within your own network
CNEXT“The most common misconception about local AI is that you need powerful servers. For many Swiss SME use cases — internal search, document summarisation, HR assistants — a modern workstation is entirely sufficient.”
Marcel Haas— CEO & AI Architect, CNEXT
Last updated: July 2026
Ready for Local AI?
CNEXT analyses your use case and shows you whether and how Ollama delivers the most value in your environment.