We Built 3 AI Assistant Personas. Zero RAG.
For an in-product help experience on a B2B SaaS tool, we skipped retrieval-augmented generation entirely — and it was the right call, for now.
When my team scoped an in-product AI help experience for a client — 3 configurable AI assistant personas, each scoped to a persona × feature area × task — the first question in our planning thread was some version of "so, RAG, right?" It's become the reflexive answer whenever "AI assistant that knows things" comes up.
We didn't use it. Here's the actual architecture, and why.
What we built instead
No RAG, no embeddings, no vector store. Each persona's knowledge is written directly into the prompt on every turn, hard-capped at 6,000 characters — roughly 1,000-1,500 tokens. Add the last 12 messages of conversation history, plus up to 5 auto-generated cross-session memory summaries (100-150 words each), and that's the whole context window. Model is Groq's Llama 3.3 70B.
Why it held up
Each persona's knowledge is narrow by design — one feature area, one task, not "everything about the product." Narrow scope stays small. When a scope's reference material comfortably fits under a few thousand characters, retrieval adds infrastructure, latency, and another thing that can silently fail — without adding accuracy. Prompt stuffing wins on simplicity when the content actually fits, and in this case, it fit with room to spare.
Where this breaks
This isn't the architecture forever — it's the right one for the current scope. Two things will force a change: a scope's knowledge outgrowing that 6,000-character budget (right now it's silently truncated, no relevance ranking — that's the line for "you need RAG now"), and richer source material like PDF or Word uploads, which the system already stores but doesn't parse yet.
We haven't hit either limit. When we do, we'll know exactly what to build next — which is sort of the point. We didn't build for the ceiling before we needed it.