Stark Chat Logo

Is it safe to put company documents into an AI knowledge base?

The security, data and quality control questions to ask when setting up an AI knowledge base.

Is it safe to put company documents into an AI knowledge base?

Your team are probably already using AI without thinking about security.

It's 9:40 on a Tuesday. Your ops manager has a 40-page supplier contract open and one question: does the termination clause let us walk in November or do we owe another quarter? She's good at her job. She also has eleven other things to do before lunch. So she opens a free AI tool in a second tab, drags the PDF in, and types "when can we terminate this without penalty?"

Answer in nine seconds. Contract now sitting on somebody else's infrastructure.

This is the awkward part of the procurement conversation. While you're asking "is an AI knowledge base safe enough to put our documents into?", the documents are already leaving. Okta's 2026 research found 52% of employees admit to using AI tools without approval, and of those knowledge workers, 39% share confidential company documents [1].

So the honest comparison isn't the vendor versus nothing. It's the vendor versus the tab your team already has open. What follows is the check we'd run before uploading a single file — six questions, each one answerable this week, most of them by email.

52% of employees admit to using AI tools without approval. Of knowledge workers using unapproved AI tools, 39% share confidential company documents, 54% share internal messages and emails and 45% share HR-related information. Meanwhile 58% of executives reported that their organization experienced an AI-related security issue or close call in the past 12 months — against executive confidence of 90% in their visibility into AI tools.

Okta

AI Agents at Work 2026: Securing the agentic enterprise

Step one: get the training answer in writing, not on the pricing page

"Does AI train on my company data?" is the question everyone asks and almost nobody verifies. Kiteworks surveyed 459 security, compliance and technology professionals and found 27% of organisations have not evaluated or technically verified whether the AI vendors they rely on use their data for model training [2].

A marketing page saying "we don't train on your data" is not verification. The contract is. Send one email and ask for three things: the specific clause in the terms or data processing agreement that says your content is excluded from model training; the name of the underlying model provider; and confirmation that zero data retention is enabled on that provider's API. That last one matters, because a vendor can honestly say they don't train on your data while their upstream model provider retains prompts for 30 days by default.

While you're there, ask for the sub-processor list. If a vendor can't tell you who else touches your handbook, you've learned something useful for the price of an email.

DO THIS WEEK: Reply to your shortlisted vendors asking for the anti-training clause, the model provider, and the sub-processor list. Any that take more than two working days to answer, park.

Step two: treat encryption and ISO 27001 as documents you read, not badges you see

Everyone says "encrypted". Fewer do both ends of it. IBM's 2026 Cost of a Data Breach Report found only 37% of breached organisations stated that they encrypt sensitive data both at rest and in transit, and that more than 20% of organisations reported a breach targeting AI models or applications — with compromised APIs, applications or plug-ins and cloud misconfigurations affecting AI workloads tied as the leading root causes at 27% each [3].

So when you're comparing on AI knowledge base security, ask for the artefacts. An ISO 27001 AI knowledge base claim should come with a certificate you can read: the certificate number, the accreditation body that issued it, the expiry date, and — the bit people skip — the scope statement. Scope is where the interesting detail lives. A certificate covering the corporate office and not the product platform is a certificate covering the wrong thing.

One caveat we'd rather say out loud: ISO 27001 tells you the company runs a managed information security process. It tells you nothing about whether the assistant gives good answers. Those are two separate reviews and step four is the other one.

DO THIS WEEK: Request the certificate PDF and read the scope statement. If the platform isn't named in scope, ask why not.

If you're mid-evaluation, the fastest thing you can do is read someone's actual security and accreditation detail rather than their claims about it — then hold every other vendor to the same standard.

View security

Step three: decide who gets in before you decide what goes in

Most data leaks inside a knowledge base aren't dramatic. They're the sales team asking a question and getting back a paragraph from the redundancy consultation deck, because someone connected the whole of Google Drive to one assistant and moved on.

The fix is boring and it works: one knowledge base per audience, each pointed only at the sources that audience should see. Support gets the help docs, policies and past ticket exports. Sales gets case studies and won proposals. HR gets the handbook. Nobody gets a folder "just in case".

Then set the door. There are only three sensible access modes and you should be able to name which one you want before the demo: a public link for anything genuinely public, an @company.com domain suffix so everyone with a work email joins themselves, or an email allowlist where you invite named people one at a time. For a GDPR AI knowledge base you'll also want to know where the data is hosted, whether there's a signed DPA, and how deletion works when someone leaves.

Ask one more thing, because it separates the serious from the shiny. Can the vendor produce a record of who asked what, and when? Kiteworks found half of the organisations surveyed could not produce a complete AI access record within one business day, and only 33% have tamper-evident audit trails in place [2].

DO THIS WEEK: Sketch a three-column grid — audience, sources they may see, access mode. Twenty minutes on paper. This is the document your DPO will actually want.

27% of organizations have not evaluated or technically verified whether the AI vendors they rely on use their data for model training. 65% of organizations discovered shadow AI usage in the past 12 months. Half of the organizations we surveyed could not produce a complete AI access record within one business day, and only 33% have tamper-evident audit trails in place.

Kiteworks

Data Security and Compliance Risk: 2026 Annual Survey Report

Step four: run the ten-question test, including the questions with no answer

Safety isn't only about who can read your files. It's whether the answers can be checked, because a confidently wrong answer about statutory sick pay causes its own kind of incident.

Microsoft's 2026 Work Trend Index, based on 20,000 knowledge workers using AI, found 86% say they treat AI output as a starting point, not a final answer, and 50% now name quality control of AI output as a human skill becoming more important [4]. Which is exactly why cited AI answers matter — a citation is what turns "trust me" into "go and look".

Here's the test, and it takes an hour. Write ten questions before the trial starts. Three with answers in a single document. Three that need two documents combined. Two where the only source is an out-of-date policy you've deliberately left in. And two where the answer genuinely isn't in your files at all.

Score the last two hardest. An assistant that says "I can't find that in your sources" has just told you it's grounded in your documents. One that invents a plausible-sounding holiday allowance has told you something more expensive. Then click every citation. The failure mode to watch for isn't a missing link, it's a link that opens a page which doesn't contain the claim.

DO THIS WEEK: Draft your ten questions now, while you're not being demoed to. Send the same ten to every vendor.

The watch-outs nobody puts in the security review

Three things a compliance checklist won't catch.

Imports are a snapshot, not a live feed. When you import a policy, the assistant knows the version you imported. If you rewrite the expenses policy in March, the March version isn't in there until you import it again. So decide who owns re-importing and when — quarterly, or whenever a policy changes. An answer sourced from a document you retired last year is technically accurate and practically wrong.

Cost is a governance problem too. DoiT surveyed 500 finance leaders in the US and UK and found 79% experienced cost overruns on AI in the past 12 months, with only 15% able to calculate AI ROI without significant bottlenecks [5]. Ask what happens when usage spikes: is there a query allowance and rate limiting, or an invoice surprise? Predictable is a security property when it stops someone quietly shifting work to an unmonitored free tool.

And scope drifts. The knowledge base you carefully limited in July has four extra folders connected by October, because it was easier than saying no. Put a recurring 30-minute review on the calendar and check what's connected.

Is an AI knowledge base safe? It depends what you can prove by Friday

None of the six checks needs a security consultant. An email about training and sub-processors. A certificate you read to the scope statement. A grid of who sees what. Ten questions, two of them unanswerable.

The version of this that used to require a six-month project and a data science team is now something you can stand up in an afternoon, scoped to one team's documents, gated to one email domain. Which means the honest risk in 2026 isn't moving too fast on a controlled, cited AI knowledge base. It's the eleven months you spend deciding while the supplier contracts go into the second tab.

Pick one team. Give them ten documents and ten questions. See what comes back, and click the citations.

Run the six checks against a live trial rather than a slide deck — one team, one set of sources, and every answer traceable to the document it came from.

Start free trial

Sources

[1] Okta, AI Agents at Work 2026: Securing the agentic enterprise, 2026 — https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/

[2] Kiteworks, Data Security and Compliance Risk: 2026 Annual Survey Report, 2026 — https://www.kiteworks.com/cybersecurity-risk-management/ai-governance-gap-widens-2026/

[3] IBM / Ponemon Institute, 2026 Cost of a Data Breach Report, 2026 — https://newsroom.ibm.com/2026-07-29-ibm-study-one-in-four-malicious-breaches-are-ai-enabled,-costing-companies-6-million-on-average

[4] Microsoft, 2026 Work Trend Index Annual Report: Agents, human agency, and opportunity, 2026 — https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization

[5] DoiT, AI Spend Reality Check, 2026 — https://www.doit.com/blog/ai-spending-survey

Stark Chat - Bespoke AI knowledge bases

Connect your sources, brand it, set who gets access, and publish — a bespoke AI knowledge base your team can use in ten minutes. No code, no dev.

Stark Chat interface