Stark Chat Logo

Lessons from Cerebras' internal AI knowledge base that answers 15,000 questions a day

An AI knowledge base that handles 15,000 questions a day, effectively a companies brain. We break it down to 4 lessons we can learn and how to apply to your own business.

Lessons from Cerebras' internal AI knowledge base that answers 15,000 questions a day

First day of work jitters, no idea where to find ANYTHING

It's her first day. She's got a question about a product specification. There are four places it might be in. She picks the wrong one. Two hours later a kind senior technician replies with a link to a document that she'd never have found, because it doesn't contain a single word she typed.

That's the problem an AI driven internal knowledge base is supposed to solve, and it's the reason we read the Cerebras write-up twice. Cerebras builds enormous chips, employs people across data centre operations, chip design, hardware, training, inference and cloud platform, and hires hundreds of new staff a year.

Three months after launch, the KB now takes more than 15,000 a day. What's interesting isn't the number. It's the mindset.

It's unlikely you're reading this and putting yourself in the same scale as 15,000 questions a day. However, from small business and up, an AI driven knowledge base can offer a massive productivity boost. Observing not what Cerebras did but how they approached the problem can teach us good lessons about how to build a knowledge base that eliminates human bottlenecks on information transfer.

They didn't expect an organisational shift

Asking contributors to change their everyday workloads to accommodate the data that a knowledge base needs to work, is like running up a hill, in the rain, with a broken ankle. Habits are ingrained and data is siloed, procedures are either in place or they aren't but it's likely that staff process information in the environment that works best for their job.

Cerebras realised that in order to have a knowledge base that truly worked they needed to go to the data rather than bring the data to them. Their tooling was where their data lived so they built integrations to directly fetch that data. They also made the API available so any time a team wanted to add a new integration they could.

Discuss an engineering ticket in Github? No problem - we'll listen to a Github hook on the issue and fetch the idea.

Have a customer support ticket? Easy - we can periodically fetch them, dedupe and load into them in.

Invoices in Xero? Simple. A custom Xero connections can make them available.

LESSON ONE: Don't expect the data to come to you, go to the data.

Finding information inside an organization is hard. The data is scattered across tools, and every quarter or so someone proposes the same brilliant fix: let's record everything in one platform so that all information is in a single place. The dream of a single source of truth, of course, rarely works in practice.

Cerebras

How we built our knowledge base, 2026

Searching by keyword misses intention

The problem with modern search is you don't know what you don't know. The data is labelled and indexed against a set of keywords that the labeller understands, the problem is the new employee has no idea how the labeller decided how to label their documents. This leads to the keyword miss, intention hit problem.

For example, a customer asks "What's the best battery for my solar panels?", so the employee starts searching for "battery", "solar panels", "best battery" which sure will provide the keyword matches, but the intention of the employees query was to find the best battery for the customers specific solar panels.

This is where intention based mixed with keyword matching (or in technical speak, hybrid search - combining keyword with semantic) bridges the gap and is the foundation of every AI knowledge base.

Cerebras used this foundational understanding to move based either intention based or keyword base and mixed the two. A hybrid search approach led to employees finding answers to questions without knowing the exact way the labeller of the content decided to write them.

LESSON TWO: Employees who use your knowledge base will have wildly different ways of asking a question, listen for both intention and keyword matches

Overwhelmed by technicalities? Don't worry - we have you covered. Stark Chat requires no technical knowledge at all.

See how

They distill rather than store verbatim

Not all data is equal and treating it as much not only dilutes the signal of the important data, it risks polluting your knowledge base with information that is objectively incorrect. Cerebras reads that data source and distills the most important aspects, leaning in to the context surrounding the data and making accurate assertions around what is helpful.

The distilled version is what goes into the knowledge base.

Take a Slack thread, you can have hundreds of messages back and forth with only about 20% of them that are important. Cerebras introduced several patterns including evaluating "bursts" of message, where the same engineers write consecutively on a problem area - they take that burst of messages and look for clear signal such as unique word patterns (technical term is IDF score), or messages that have been reacted to by a colleague.

This is the core of filtering signal from noise, looking for social signals and signs of rarity and deciding on how to rank messages on indexing. Of course, a message that looks like an answer doesn't always been the answer is correct. This is where citations help, allowing you to cross reference the answers to the documentation. Anthropic, writing about its own internal analytics assistant, is refreshingly blunt that a source footer "doesn't make the answer more correct, but it does help the consumer judge how much they can trust the response" [2].

LESSON THREE: Distill before indexing, capture the important data. Ensure answers cite where their data came from.

Scope knowledge bases to teams

Just because you can, doesn't mean you should. Indexing everyone's data into one giant chat make for cool technical reading but with too much noise you dilute the signal that good AI depends on. This is a bedrock of successful AI integration, understanding who needs what data to do their work.

In fact, this is where organisations frequently fail, 69% of knowledge workers told Atlassian's State of Teams research that their data and knowledge foundations are "not optimized for AI" [3]. It's a big problem to solve and not one that you can solve easily without at least some manual intervention and data labelling. LLM's are fantastic at being specific as long as you give them specific documentation. As the maxim goes, "trash in, trash out."

LESSON FOUR: Understand who needs what data and build your knowledge bases around it.

What does this mean for the mere mortal?

Exciting engineering is always popular and nothing is more exciting than a company brain answering 15,000 question a day across a human and agentic surface. No matter how hard we try, most of us don't have that problem and probably can't afford to spend what it takes to build a system of this structure.

Fortunately, you don't have to. You can achieve 80% of what Cerebras has with zero technical knowledge and about 10 minutes of your time. If you don't believe us check out our demo and give it a try.

You can build and publish an internal knowledge base in 10 minutes, don't just take our word for it.

Try it out for free

Sources

[1] Cerebras, How we built our knowledge base, 2026 — https://www.cerebras.ai/blog/how-we-built-our-knowledge-base

[2] Anthropic, How Anthropic enables self-service data analytics with Claude, 2026 — https://claude.com/blog/how-anthropic-enables-self-service-data-analytics-with-claude

[3] Atlassian, Why individual AI speed isn't delivering the ROI CIOs expected — https://www.atlassian.com/blog/ai-at-work/why-ai-speed-isnt-delivering-roi-for-cios

Stark Chat - Bespoke AI knowledge bases

Connect your sources, brand it, set who gets access, and publish — a bespoke AI knowledge base your team can use in ten minutes. No code, no dev.

Stark Chat interface