---
title: "How to move a RAG prototype into production"
canonical: "https://amartripathi.com/guides/rag-prototype-to-production"
---

# How to move a RAG prototype into production

Turn a retrieval demo into a maintainable product with evaluation, data controls, observability, and clear user states.

Updated: September 2026

Canonical page: https://amartripathi.com/guides/rag-prototype-to-production

## Quick answer

A production RAG system needs representative evaluation cases, document access controls, reliable ingestion, source references, monitoring, and failure recovery. Retrieval quality alone does not make the product safe or useful.

## Evidence and decision points

### Evaluation set

Use representative questions, source documents, expected evidence, and failure cases.

### Data boundaries

Retrieval must respect document access, retention, deletion, and provider boundaries.

### Visible states

Users need clear ingestion, citation, uncertainty, failure, and retry states.

## Create evaluation before expansion

Collect real questions and expected source evidence before changing chunk size, embeddings, prompts, or models. Without a stable evaluation set, each change becomes a subjective demo.

Include answerable questions, unanswerable questions, conflicting sources, and access-boundary cases.

## Make ingestion an operating workflow

Document parsing can fail, stall, or produce incomplete content. Track each stage and give users a clear state instead of hiding processing behind one spinner.

Store document identity and processing version so the team can trace which content supported an answer.

- Validate file type, size, and access before processing.

- Track parsing, segmentation, indexing, completion, and failure.

- Support safe retry without duplicate document state.

- Remove derived data when the source document is deleted.

## Add product controls

Monitor retrieval latency, model latency, failure rates, token use, and expensive document paths. Set practical limits before traffic grows.

Show source references where users need verification. State uncertainty when the available evidence does not support a complete answer.

## Prototype versus production RAG

| Area | Prototype | Production |

| --- | --- | --- |

| Evaluation | A few manual questions | Versioned representative cases |

| Documents | Sample files | Access, status, retention, and deletion |

| Answers | Plausible text | Evidence, citations, and uncertainty |

| Operations | Local logs | Latency, failures, cost, and alerts |

## Common questions

### Which embedding model should I use?

Choose through representative evaluation, language needs, latency, cost, and deployment constraints.

### Do I need citations?

Use citations when users must verify claims or inspect the original source. The interface should link to permitted source material.

### Can I improve RAG without changing the model?

Yes. Ingestion, segmentation, metadata, filtering, retrieval, context assembly, and interface states often matter greatly.

## Sources and further reading

- [NIST AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): NIST provides a framework for measuring and managing AI risks.

- [OWASP Top 10 for LLM applications](https://genai.owasp.org/llm-top-10/): OWASP covers prompt injection, sensitive disclosure, supply chain, and excessive agency risks.

## Related resources

- [RAG development](https://amartripathi.com/services/rag-development)

- [RAG document processing](https://amartripathi.com/case-studies/rag-document-processing)

- [AI SaaS development](https://amartripathi.com/services/ai-saas-development)

## Contact

Email Amar Tripathi at [theamartripathi@gmail.com](mailto:theamartripathi@gmail.com) with the product context, current stack, expected outcome, constraints, and target timeline.