Back to Blog
AI Integration

Our Production RAG Implementation: 5 Lessons Learned

7 min read
Engineering Leaders
AI Engineers

Our Production RAG Implementation: 5 Lessons Learned

AI conversations in Dubai and across the GCC are getting serious. The era of flashy demos is over. CTOs and founders are now under pressure to deliver real business value, not just proof-of-concepts. One of the most promising, practical patterns to emerge is Retrieval-Augmented Generation, or RAG. It lets you ground powerful Large Language Models (LLMs) in your company’s private data. But moving from a simple notebook experiment to a production-grade RAG implementation is a huge leap.

At SouzaLabs, we’ve been in the trenches building these systems for our clients. We recently shipped Trajex, a sophisticated RAG-based knowledge management tool for a major enterprise in the region. The journey taught us some critical, hard-won lessons that you won’t find in most tutorials.

This isn't another high-level overview. This is a practical guide for engineering leaders in the UAE who need to ship robust AI features. We’re sharing the five biggest lessons we learned building a production RAG implementation, so you can avoid costly mistakes. You can read more about the final outcome in our Trajex case study.

1. Chunking Is Not 'Set It and Forget It'

Your RAG system is only as good as the information it can retrieve. The first step, breaking down your documents into searchable “chunks,” is deceptively critical. It’s tempting to grab a standard library, use a `RecursiveCharacterTextSplitter`, and call it a day. In production, this falls apart quickly.

Naive chunking leads to two major problems:

  • Loss of Context: A chunk might start or end mid-sentence, cutting off the very information needed to answer a question correctly.
  • Irrelevant Retrieval: Large, unfocused chunks can be semantically similar to a query without containing the specific answer, leading to noise.
  • What to do instead:

  • Experiment with Overlap: Start with a character splitter but experiment heavily with the `chunk_size` and `chunk_overlap` parameters. A larger overlap can help preserve context across chunk boundaries, but it also increases the size of your vector database.
  • Use Semantic Chunking: Instead of splitting by a fixed character count, consider splitting based on semantic boundaries. This could mean splitting by paragraphs, sections, or even using smaller language models to determine the most logical break points.
  • Build Document-Specific Logic: Do you have structured data like tables in PDFs or formatted reports? A generic text splitter will mangle them. You need to write custom parsers for these documents to preserve their structure, perhaps by converting tables to Markdown format before chunking and ingestion.
  • 2. Your Embedding Model Is a Critical Choice

    The embedding model is the engine that turns your text chunks into numerical vectors for searching. The default choice for many is OpenAI's `text-embedding-ada-002`, but it's not always the best tool for the job, especially for businesses in the GCC.

    When evaluating models, consider:

  • Cost: At scale, embedding millions of chunks can become a significant operational expense. Cheaper or self-hosted open-source models can offer huge savings.
  • Performance: Don't just trust the leaderboards. A model that performs well on a generic English benchmark might struggle with your specific domain jargon, like legal documents or engineering reports.
  • Multilingual Needs: For any business operating in the UAE and wider Middle East, support for Arabic is non-negotiable. You need a model that performs exceptionally well in both English and Arabic. Test this rigorously.
  • We benchmarked three different embedding models for Trajex before settling on one. We created a small, high-quality dataset of questions and answers relevant to the client's business and measured which model retrieved the correct documents most often. The winner wasn't the most famous model, but the one that best understood their specific terminology.

    3. Naive Retrieval Will Only Get You 80% There

    The core of RAG is retrieving relevant chunks. A basic implementation performs a vector similarity search: find the chunks whose vectors are closest to the query's vector. This works surprisingly well and will get you a decent demo.

    In production, it’s not enough. You’ll quickly find the system retrieving chunks that are thematically related but factually incorrect or irrelevant. To deliver a reliable product, you need more advanced retrieval strategies.

    Two techniques to level up your retrieval:

  • Hybrid Search: This combines keyword-based search (like BM25) with vector search. Vector search is great for understanding conceptual meaning, but it can miss specific keywords, product codes, or names. Keyword search excels at this. By combining them, you get the best of both worlds. Most modern vector databases offer hybrid search as a feature.
  • Re-ranking: This adds a second stage to your retrieval process. First, retrieve a larger number of potential chunks (e.g., the top 20). Then, use a more powerful but slower model (like a cross-encoder) to re-rank those 20 chunks for relevance to the query. You then pass only the top 3-5 re-ranked chunks to the LLM. This significantly improves the quality of the context and reduces the chances of the LLM seeing irrelevant information.
  • 4. Evaluation Is Everything, and It's Hard

    How do you know if your changes to chunking, embedding, or retrieval are actually making the system better? You have to measure it. A robust evaluation framework is non-negotiable for a successful RAG implementation.

    Don't rely on random spot-checks. You need a systematic approach. We recommend building a “golden dataset” of `(question, context, answer)` triplets that represent real-world use cases.

    With this dataset, you can measure a few key metrics:

  • Context Precision & Recall: Of the chunks you retrieved, how many were actually relevant? Did you miss any relevant chunks that were in the knowledge base?
  • Answer Faithfulness: Does the LLM's answer stick to the facts provided in the retrieved context? This is key to preventing hallucinations.
  • Answer Relevancy: Does the generated answer actually address the user's question?
  • Frameworks like RAGAs and TruLens can help automate this, but even a simple script that runs your golden dataset through the pipeline and flags regressions is a massive step up from manual testing. This evaluation step is what separates a toy project from a production system.

    5. The LLM Is a Generation Layer, Not a Magic Brain

    It’s a common mistake to assume the LLM at the end of the chain will magically figure things out if you just throw enough context at it. Garbage in, garbage out. If your retrieval is poor, the LLM will either generate a poor answer or hallucinate. The LLM’s job is to synthesize and present information, not to find it.

    The key to controlling the LLM is meticulous prompt engineering. Your prompt is the contract between you and the model. A good RAG prompt should be explicit:

  • Define the Role: `You are an expert assistant for XYZ Company.`
  • State the Core Instruction: `Answer the user's question based ONLY on the provided context below. Do not use any other information.`
  • Provide a Failure Condition: `If the answer is not found in the context, state clearly 'I do not have enough information to answer that question.'`
  • Structure the Data: Clearly demarcate the context and the user's question. For example:
  • `--- CONTEXT ---`

    `{retrieved_chunks}`

    `--- QUESTION ---`

    `{user_question}`

    This level of instruction dramatically reduces hallucinations and ensures the LLM behaves as a grounded, factual agent based on your data.

    Work with SouzaLabs

    Building a high-performing RAG implementation is a serious engineering challenge that goes far beyond gluing a few APIs together. It requires careful architecture, rigorous testing, and a deep understanding of the full stack, from data ingestion to prompt engineering.

    If you're a CTO or engineering leader in the UAE or GCC looking to build real, value-driving AI solutions without the guesswork, our team can help. We specialize in taking complex AI projects from concept to production. You can learn more about our approach by reading about our AI integration services.

    Ready to see how we can apply these lessons to your business? Book a free consultation with our team today.