Building a RAG PDF Assistant with LangChain, FAISS and Gemini

How Chatur answers questions about PDFs with RAG: PyPDF2 extraction, recursive chunking, Gemini embeddings in FAISS, a grounded prompt, and lessons on chunk size.

Satyam KesharwaniSatyam Kesharwani 6 min read
  • AI
  • Python
  • RAG

TL;DR. Chatur is a small retrieval-augmented generation (RAG) app I built in Python: upload one or more PDFs, ask questions in plain English, and get answers grounded in the documents. It extracts text with PyPDF2, splits it with LangChain's recursive splitter (10,000-character chunks with 1,000 characters of overlap), embeds the chunks with Google's embedding-001 model into a local FAISS index, and has Gemini Pro answer from the retrieved chunks at a temperature of essentially zero, under a prompt that forbids answering beyond the context. The UI is Streamlit. This post explains how each step works, and what I would change now that I understand the trade-offs better.

Code: GitHub

Why RAG instead of just asking the model

A language model answers from what it learned during training. It has never seen your lecture notes, a company's internal report or yesterday's research paper, and when asked about them it may confidently invent an answer. You could paste the whole PDF into the prompt, but long documents overflow the context window, and even when they fit, cost and latency grow with every page.

Retrieval-augmented generation splits the job in two:

  1. Retrieve: find the few passages of the document that are most relevant to the question.
  2. Generate: give the model only those passages, and instruct it to answer from them.

The model's job shrinks from "know everything" to "read this and answer", which is something it does well.

The pipeline

 indexing (once per upload)                    answering (per question)

 PDFs ──▶ PyPDF2 text ──▶ chunks ──▶ embeddings     question ──▶ embedding
                         (10,000 chars,     │                         │
                          1,000 overlap)    ▼                         ▼
                                       FAISS index ◀── similarity search (top chunks)
                                                                      │
                                              prompt = context + question ──▶ Gemini Pro ──▶ answer

1. Extracting text

PyPDF2 reads each uploaded file page by page and concatenates the text:

def get_pdf_text(pdf_docs):
    text = ""
    for pdf in pdf_docs:
        pdf_reader = PdfReader(pdf)
        for page in pdf_reader.pages:
            text += page.extract_text()
    return text

This works for PDFs with a real text layer. Scanned documents are just images of text, so they need OCR first, which Chatur does not do yet.

2. Chunking

Embedding models and retrieval both work on passages, not whole books, so the text is split into chunks:

def get_text_chunks(text):
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=10000, chunk_overlap=1000)
    return text_splitter.split_text(text)

The recursive splitter tries to break on paragraph boundaries first, then lines, then sentences, and only cuts mid-sentence as a last resort, which keeps chunks coherent. The 1,000-character overlap means an idea that straddles a boundary appears whole in at least one chunk.

3. Embedding and indexing with FAISS

Each chunk is turned into a vector by Google's embedding model, so that passages with similar meaning end up close together in vector space. The vectors go into FAISS, Meta's library for fast similarity search, which is saved to disk:

def get_vector_store(text_chunks):
    embeddings = GoogleGenerativeAIEmbeddings(model="models/embedding-001")
    vector_store = FAISS.from_texts(text_chunks, embedding=embeddings)
    vector_store.save_local("faiss_index")

For a few documents, FAISS's exact search over every vector is effectively instant. Its approximate indexes only start to matter at millions of vectors.

4. Retrieval and a grounded prompt

At question time, the question is embedded with the same model, FAISS returns the closest chunks, and they are "stuffed" into a prompt together with the question:

prompt_template = """
# Answer the question as detailed as possible from the provided context, make sure to provide all the details, if the answer is not in
# provided context just say, "answer is not available in the context", don't provide the wrong answer\n\n
Context:\n {context}?\n
Question: \n{question}\n

Answer:
"""

model = ChatGoogleGenerativeAI(model="gemini-pro", temperature=0.00001)
prompt = PromptTemplate(template=prompt_template, input_variables=["context", "question"])
chain = load_qa_chain(model, chain_type="stuff", prompt=prompt)

Two choices here do most of the work against hallucination. The instruction gives the model an explicit, acceptable way out ("answer is not available in the context") instead of forcing it to produce something. And a temperature of essentially zero makes the model pick its most likely answer instead of sampling creatively, which is what you want for factual questions about a document.

5. A Streamlit front end

Streamlit turns the script into a web app with a few lines: a sidebar to upload PDFs and build the index, and a text box for questions. It was the fastest way to put a usable interface on the pipeline.

Handling API quotas

The free tier of the Gemini API limits requests per minute, and indexing a long PDF sends one embedding request per chunk. I wrote a wrapper that caps calls at 150 per minute with the ratelimit package and, on a ResourceExhausted error, waits 10 seconds and retries:

@sleep_and_retry
@limits(calls=150, period=ONE_MINUTE)
def safe_embed_request(text_chunk, embeddings):
    try:
        return embeddings.embed(text_chunk)
    except ResourceExhausted:
        time.sleep(10)
        return safe_embed_request(text_chunk, embeddings)

To be honest about it: the indexing path currently calls FAISS.from_texts directly, so this wrapper is not wired in yet. Making index building go through it (or using exponential backoff with a retry limit) is the first fix on my list.

What I would change

  • Smaller chunks. 10,000 characters per chunk is large. Retrieval returns several chunks per question (LangChain's default is four), so every question sends tens of thousands of characters to the model. That is slower and more expensive, and it dilutes the relevant sentences with surrounding text. Chunks of around 1,000 characters with a little overlap usually retrieve more precisely; the right size is worth measuring on real questions.
  • Citations. Keeping the page number with each chunk would let every answer point to where it came from, which is the quickest way for a user to trust or check it.
  • Conversation memory. Every question is answered independently today, so a follow-up like "and what about the second method?" has no context.
  • Per-session indexes. The index is a single folder on disk, so a second upload replaces the first, and two users of a shared deployment would see each other's documents.
  • Hybrid search and re-ranking. Combining keyword search (BM25) with vector search helps with exact terms like names and error codes, and a re-ranking step can pick the best few chunks out of a larger candidate set.
  • Evaluation. A small set of questions with known answers would turn "it seems to work" into a number, and would show whether a change, such as a new chunk size, actually helps.
  • Current models. The code uses gemini-pro and embedding-001, which Google has since superseded; LangChain makes swapping models a one-line change.

Building it end to end taught me that RAG quality is mostly decided before the model is ever called: by how documents are extracted, chunked and retrieved.

Frequently asked questions

What is retrieval-augmented generation (RAG)?

RAG answers a question in two steps: it first retrieves the passages of a document collection that are most relevant to the question, then gives only those passages to a language model with an instruction to answer from them. This grounds the answer in the documents, including documents the model never saw in training.

How does Chatur choose which parts of a PDF to use?

It splits the PDF text into overlapping chunks, embeds each chunk with Google's embedding model into a FAISS index, embeds the question the same way, and passes the closest chunks to Gemini Pro.

What chunk size should a RAG system use?

It depends on the documents and the questions, so it is worth measuring. Chatur uses 10,000-character chunks with 1,000 characters of overlap, which is large; chunks of around 1,000 characters usually retrieve more precisely and send less text to the model per question.

How does Chatur reduce hallucinations?

Its prompt tells the model to answer only from the retrieved context and to reply 'answer is not available in the context' otherwise, and the model runs at a temperature of essentially zero.

Source code on GitHub