The Real Blocker
The problem isn't AI. It's where your data goes.
Law firms, medical practices, financial and professional businesses are sitting on exactly the material AI is brilliant with — contracts, records, case history, institutional knowledge. But sending that into a public AI tool means handing confidential and personal information to a third party, which POPIA and professional duty simply don't allow. So most firms either ban AI and fall behind, or use it quietly and take on real exposure. There's a third option: keep the AI, keep the data, and run the sensitive part yourself.
What We Build
Your knowledge, your hardware, the right AI for each job.
One system, engineered so the private stays private and the heavy thinking still happens — with a clear line between the two.
A private knowledge base
Your SOPs, contracts, product docs and institutional know-how, organised into a searchable, owned knowledge base that lives with you — not in someone else's cloud.
Local AI for anything sensitive
Open models running on your own hardware (Ollama-based), so queries over confidential or personal information are answered on your premises — nothing leaves the building. POPIA-safe by design.
Frontier AI for the heavy thinking
For complex, non-sensitive work, we route to a frontier model (like Claude) under governed access and guardrails — so you get top-tier reasoning where it's safe to use it.
Grounded in your documents
The AI answers over your material and cites where it found it (retrieval-augmented), so staff get accurate, checkable answers — not confident guesses.
What Teams Use It For
Institutional knowledge, on tap.
Straight Talk
We tell you the trade-offs up front.
Local models are private and free of per-message cloud fees, but they are not as powerful as the best frontier models — so we use each for what it's good at: local for the private and the routine, frontier for the genuinely hard reasoning. Good local performance needs decent hardware, which we spec honestly. And an AI over your documents still needs guardrails and oversight — which is exactly what we do. Anyone promising a free local model that matches Claude is selling you something. We won't.
Which Models
Named, not hidden behind “AI-powered”.
You are entitled to know what is processing your information, and your procurement team will ask. So we say it plainly rather than making you request it.
Open-weight models — running on your hardware
These are the models whose weights are published, so they can run entirely inside your environment. Nothing leaves the building, which is the whole point for privileged, clinical or personal information.
- Qwen (Alibaba) — strong general and multilingual performance, and several sizes are Apache-2.0 licensed, which is the cleanest commercial position of the major families.
- Llama (Meta) — widely supported with the largest tooling ecosystem. Released under Meta's own community licence, not a standard open-source one, and that licence carries conditions worth reading before you commit.
- Mistral and Mixtral (Mistral AI) — efficient for their size, so they run well on more modest hardware. Several are Apache-2.0.
- Gemma (Google) — small and capable, under Google's own terms of use rather than a standard open-source licence.
Served locally through Ollama or a comparable runtime. Which one we recommend depends on your hardware, your languages and the work — not on which name is fashionable this quarter.
Frontier models — for the genuinely hard reasoning
For complex work that carries no sensitive information, we route to a hosted frontier model — typically Claude (Anthropic), and GPT (OpenAI) or Gemini (Google) where a client already has an account or a preference. These are materially stronger on the hardest tasks and will remain so for the foreseeable future.
The trade is explicit: data sent to a hosted model leaves your environment, and usually leaves the country. That is a section 72 question under POPIA, not merely a preference. So the routing rule is written down, agreed with you, and enforced in the system rather than left to whoever is typing.
On licences. “Open” does not mean “unrestricted”. Apache-2.0 is genuinely permissive; Meta's and Google's licences carry conditions that can matter to a regulated business or a resale arrangement. We confirm the licence for the specific model and version at selection, and record it in the handover pack — because the answer changes between releases and you should not have to take ours on faith.
Model families move quickly. We name them here because you deserve a straight answer, and we re-check this page against what we actually deploy rather than leaving a list to rot.
The Deliverable
Set up, trained, and yours.
- The system designed and built around your actual documents and workflows
- Honest hardware selection and setup — sized to your needs and budget
- Software and model selection, installation and configuration
- Guardrails, access control and a POPIA-aware data position
- Prompt and usage training so your team actually gets value from it
- Full ownership — it runs on your hardware, in your control, with an optional maintenance retainer
Prefer AI training for your team first? See AI Fluency & Consulting. Need existing AI work reviewed or secured? That's AI Consulting & Oversight.
Common Questions
Straight answers.
Is a private AI knowledge system actually POPIA-safe?
That's the whole point. The sensitive work runs on local models on your own hardware, so confidential and personal information isn't sent to a third-party cloud AI. We build it with access control, a documented data-retention position and clear rules for what may and may not use the frontier model — so it stands up to a compliance conversation, not just a demo.
Is a local AI model as good as ChatGPT or Claude?
Honestly, not for the hardest reasoning — frontier models are still ahead. That's why we use both: local for anything private and for routine work, and a frontier model (under guardrails) for the genuinely complex, non-sensitive tasks. You get privacy where it matters and power where it's allowed.
Which AI models do you actually use, and who owns them?
For anything sensitive we run open-weight models on your own hardware: Qwen, Llama, Mistral or Gemma, served locally, so the data never leaves your environment. For hard reasoning on non-sensitive work we route to a hosted frontier model, usually Claude, or GPT or Gemini where you already have an account. The licences differ and it matters commercially: several Qwen and Mistral releases are Apache-2.0, while Meta and Google use their own terms with conditions attached. We confirm the licence for the specific model and version at selection and record it in your handover pack.
What hardware is needed for on-premise AI?
It depends on how many people use it and how heavy the local work is — from a single capable workstation to a small server with a GPU. We spec it honestly to your needs and budget rather than overselling; the goal is the right machine, not the biggest invoice.
Do we own the private AI system, or are we locked in?
You own it. It runs on your hardware, in your control, built on open tooling — no rented black box you can't leave. We can maintain and update it on a retainer if you want, but you're never held hostage.