Almost every team building with large language models starts the exact same way. You load a batch of PDFs into LangChain, run a recursive character text splitter, push five-hundred-token chunks into a vector database and wire up top-four cosine similarity retrieval. In a local notebook or during an internal slide demo, it looks like magic.
Then you deploy it to real users and everything begins to fall apart. Someone asks for the total revenue across three different quarters and the vector search returns random paragraphs from the introduction. Another user searches for a specific error code and the model hallucinates an apology because pure semantic embeddings smoothed away the exact keyword match. If you want RAG to work reliably in production, you have to look past naive vector similarity.
1. Document Structure Dictates Chunking Strategy
Splitting text blindly by character counts is the quickest way to destroy context. A sentence split mid-thought or a table chopped across two arbitrary chunks leaves the embedding model with fragmented noise. Contextual chunking respects the natural boundaries of your source material:
Parent Document Retrieval: Split documents into small granular chunks for precise vector matching, but store a pointer to the larger parent section. When a small chunk matches the query, feed the broader parent context to the LLM so it sees the complete thought.
Markdown and Table Parsing: Never feed raw PDF text streams into an embedding model. Convert tables to clean markdown structures or JSON representations before embedding so row and column relationships survive.
Document Metadata Enrichment: Inject document titles, section headers and revision dates directly into the header of every chunk before generating embeddings.
2. Hybrid Retrieval Beats Semantic Search Alone
Vector embeddings are great at capturing conceptual themes, but they are notoriously clumsy with precise strings, part numbers, SKUs and acronyms. A user asking for "policy clause 4.2b" will often get irrelevant clauses that discuss similar legal themes rather than the exact numbered clause they requested.
Production pipelines rely on hybrid search. Run dense vector retrieval (using pgvector or Qdrant) in parallel with sparse lexical keyword search (such as BM25). Combine the candidates using Reciprocal Rank Fusion (RRF) and pass the top twenty candidates through a cross-encoder reranker like Cohere or BGE. The reranker scores each passage against the user query directly, stripping out false positives before they hit your model prompt window.
3. Enforcing Strict JSON Schemas for Tool Execution
Free-form conversational responses are fine for a generic chatbot, but enterprise workflows require deterministic execution. If an AI agent extracts customer claims or triggers an internal database query, an ambiguous paragraph is useless to your downstream API.
Never ask a model to "output valid JSON" inside a raw prompt and hope for the best. Use structured output mechanisms (such as Pydantic models or OpenAI structured outputs) that constrain token generation at the decoding level. If the output does not conform to your expected schema, catch it before it reaches your user or backend services.
4. Continuous Evaluation Over Manual Spot Checks
You cannot improve what you do not measure. Instead of manually inspecting twenty sample prompts whenever you tweak your chunk size or system prompt, build a continuous evaluation test suite. Track three concrete metrics on every deployment:
Faithfulness: Did the model derive every single claim in its answer strictly from the retrieved context?
Context Recall: Did the retrieval step gather all the reference materials required to answer the prompt?
Answer Relevance: Did the final response directly address the user query without rambling or leaking internal prompt instructions?
Conclusion
Building a RAG prototype takes an afternoon. Engineering a production-grade system that answers customer queries accurately without hallucinating requires discipline around data parsing, hybrid search and automated evaluation. Respect the source document structure, verify your retrieval quality and treat your prompts like version-controlled software contracts.
