Private AI Knowledge Systems

An AI that knows your business — without your data leaving the building.

Your team wants the productivity of AI. But you can't paste client files into public ChatGPT — not under POPIA, not with confidential work. So we build you a private one you own.

Direct answer: A private AI knowledge system connects your company's documents and know-how to AI — with the sensitive work running locally on your own hardware so nothing is sent to the cloud, and heavy reasoning handled by frontier models under strict guardrails. It answers your team's questions over your own material, and you own the whole thing.
Get a Quote Chat on WhatsApp

The Real Blocker

The problem isn't AI. It's where your data goes.

Law firms, medical practices, financial and professional businesses are sitting on exactly the material AI is brilliant with — contracts, records, case history, institutional knowledge. But sending that into a public AI tool means handing confidential and personal information to a third party, which POPIA and professional duty simply don't allow. So most firms either ban AI and fall behind, or use it quietly and take on real exposure. There's a third option: keep the AI, keep the data, and run the sensitive part yourself.

What We Build

Your knowledge, your hardware, the right AI for each job.

One system, engineered so the private stays private and the heavy thinking still happens — with a clear line between the two.

Your knowledge

A private knowledge base

Your SOPs, contracts, product docs and institutional know-how, organised into a searchable, owned knowledge base that lives with you — not in someone else's cloud.

Runs local

Local AI for anything sensitive

Open models running on your own hardware (Ollama-based), so queries over confidential or personal information are answered on your premises — nothing leaves the building. POPIA-safe by design.

Frontier reasoning

Frontier AI for the heavy thinking

For complex, non-sensitive work, we route to a frontier model (like Claude) under governed access and guardrails — so you get top-tier reasoning where it's safe to use it.

Answers with sources

Grounded in your documents

The AI answers over your material and cites where it found it (retrieval-augmented), so staff get accurate, checkable answers — not confident guesses.

What Teams Use It For

Institutional knowledge, on tap.

Instant internal answers"What's our policy on X?" answered from your own documents in seconds, for every staff member.
Faster onboardingNew hires ask the system instead of interrupting senior staff — knowledge that used to live in one head is now shared.
Drafting & researchFirst drafts, summaries and research grounded in your templates, precedents and prior work.
Confidential reviewAnalyse sensitive contracts, records or case files locally — the exact work you can't send to a public tool.

Straight Talk

We tell you the trade-offs up front.

Local models are private and free of per-message cloud fees, but they are not as powerful as the best frontier models — so we use each for what it's good at: local for the private and the routine, frontier for the genuinely hard reasoning. Good local performance needs decent hardware, which we spec honestly. And an AI over your documents still needs guardrails and oversight — which is exactly what we do. Anyone promising a free local model that matches Claude is selling you something. We won't.

Which Models

Named, not hidden behind “AI-powered”.

You are entitled to know what is processing your information, and your procurement team will ask. So we say it plainly rather than making you request it.

Open-weight models — running on your hardware

These are the models whose weights are published, so they can run entirely inside your environment. Nothing leaves the building, which is the whole point for privileged, clinical or personal information.

  • Qwen (Alibaba) — strong general and multilingual performance, and several sizes are Apache-2.0 licensed, which is the cleanest commercial position of the major families.
  • Llama (Meta) — widely supported with the largest tooling ecosystem. Released under Meta's own community licence, not a standard open-source one, and that licence carries conditions worth reading before you commit.
  • Mistral and Mixtral (Mistral AI) — efficient for their size, so they run well on more modest hardware. Several are Apache-2.0.
  • Gemma (Google) — small and capable, under Google's own terms of use rather than a standard open-source licence.

Served locally through Ollama or a comparable runtime. Which one we recommend depends on your hardware, your languages and the work — not on which name is fashionable this quarter.

Frontier models — for the genuinely hard reasoning

For complex work that carries no sensitive information, we route to a hosted frontier model — typically Claude (Anthropic), and GPT (OpenAI) or Gemini (Google) where a client already has an account or a preference. These are materially stronger on the hardest tasks and will remain so for the foreseeable future.

The trade is explicit: data sent to a hosted model leaves your environment, and usually leaves the country. That is a section 72 question under POPIA, not merely a preference. So the routing rule is written down, agreed with you, and enforced in the system rather than left to whoever is typing.

On licences. “Open” does not mean “unrestricted”. Apache-2.0 is genuinely permissive; Meta's and Google's licences carry conditions that can matter to a regulated business or a resale arrangement. We confirm the licence for the specific model and version at selection, and record it in the handover pack — because the answer changes between releases and you should not have to take ours on faith.

Model families move quickly. We name them here because you deserve a straight answer, and we re-check this page against what we actually deploy rather than leaving a list to rot.

The Deliverable

Set up, trained, and yours.

Every engagement includes:
  • The system designed and built around your actual documents and workflows
  • Honest hardware selection and setup — sized to your needs and budget
  • Software and model selection, installation and configuration
  • Guardrails, access control and a POPIA-aware data position
  • Prompt and usage training so your team actually gets value from it
  • Full ownership — it runs on your hardware, in your control, with an optional maintenance retainer

Prefer AI training for your team first? See AI Fluency & Consulting. Need existing AI work reviewed or secured? That's AI Consulting & Oversight.

Common Questions

Straight answers.

Is a private AI knowledge system actually POPIA-safe?

That's the whole point. The sensitive work runs on local models on your own hardware, so confidential and personal information isn't sent to a third-party cloud AI. We build it with access control, a documented data-retention position and clear rules for what may and may not use the frontier model — so it stands up to a compliance conversation, not just a demo.

Is a local AI model as good as ChatGPT or Claude?

Honestly, not for the hardest reasoning — frontier models are still ahead. That's why we use both: local for anything private and for routine work, and a frontier model (under guardrails) for the genuinely complex, non-sensitive tasks. You get privacy where it matters and power where it's allowed.

Which AI models do you actually use, and who owns them?

For anything sensitive we run open-weight models on your own hardware: Qwen, Llama, Mistral or Gemma, served locally, so the data never leaves your environment. For hard reasoning on non-sensitive work we route to a hosted frontier model, usually Claude, or GPT or Gemini where you already have an account. The licences differ and it matters commercially: several Qwen and Mistral releases are Apache-2.0, while Meta and Google use their own terms with conditions attached. We confirm the licence for the specific model and version at selection and record it in your handover pack.

What hardware is needed for on-premise AI?

It depends on how many people use it and how heavy the local work is — from a single capable workstation to a small server with a GPU. We spec it honestly to your needs and budget rather than overselling; the goal is the right machine, not the biggest invoice.

Do we own the private AI system, or are we locked in?

You own it. It runs on your hardware, in your control, built on open tooling — no rented black box you can't leave. We can maintain and update it on a retainer if you want, but you're never held hostage.