ChatPDF
Gen AI PDF chat — upload a document, get an instant summary, and ask follow-ups grounded in your file. Multi-agent RAG + FAISS vector search, Groq LLM with multi-key failover, JWT + Google OAuth, and Neon PostgreSQL chat history.

←Use arrow keys or swipe to navigate→
The Problem
Students, researchers, and professionals routinely face 50–500 page PDFs — research papers, resumes, contracts, lecture notes — but reading every page to find one fact is slow and error-prone. Pasting the whole document into a generic chatbot blows token limits, hallucinates when context is missing, and cannot scope answers to *your* file. Existing PDF tools often stop at keyword search or a one-shot summary with no follow-up thread, no account persistence, and no guarantee that answers cite document content instead of the open internet.
The product gap is a **document-grounded Gen AI app**: sign in, upload once, get an automatic summary and starter questions, then chat in plain English with retrieval-augmented answers — all sessions saved per user, not lost on refresh.
My Role & Constraints
Solo Full-Stack Engineer, Architect & Product Owner — I designed and shipped ChatPDF end-to-end.
**Frontend:** React 19 SPA (Vite 6, TypeScript, Tailwind CSS v4, Framer Motion) with cinematic video hero landing, glass-style UI, features/how-it-works/FAQ/testimonials marquee, responsive navbar, email + Google auth screens, drag-and-drop PDF upload (up to 15 MB), sidebar session history, chat workspace with summary header, suggested starter questions, replace-PDF flow, and always-dark theme.
**Backend:** Node.js 20+ Express server (TypeScript, npm workspaces) — JWT + bcrypt auth, Google OAuth 2.0 callback, pdf-parse text extraction, chunking + FAISS (faiss-node) vector index, multi-agent RAG pipeline (query optimization → retrieval → context analysis → answer generation), Groq LLM integration with multi-key pool and automatic failover on rate limits, input truncation safeguards, and Neon PostgreSQL persistence.
**Ops:** Split Vercel (frontend) + Render (backend — native faiss-node requires full Node host, not serverless) + Neon DB; render.yaml Blueprint; schema auto-created from backend/db/schema.sql on boot.
System Design / Architecture
ChatPDF is an **npm workspaces monorepo** — frontend/ (React + Vite) and backend/ (Express + RAG).
1$React 19 + Vite (Vercel)2$│ JWT3$▼4$Express (Render) ── native faiss-node needs a full Node host5$│6$├─ pdf-parse ──▶ chunks ──▶ FAISS index (in-process)7$│8$├─ multi-agent RAG9$│ query optimiser ─▶ retrieval ─▶ context analyser ─▶ answer10$│ │11$│ ▼12$│ Groq (multi-key failover)13$│14$└──▶ Neon PostgreSQL15$users · chat_sessions · pdf_documents · chat_messages
**Request flow:**
1. User signs in → JWT issued, stored client-side
2. User uploads PDF → backend extracts text via pdf-parse, chunks content, builds FAISS index, generates AI summary
3. User asks a question → multi-agent RAG retrieves relevant chunks → Groq generates grounded answer
4. Messages + session metadata persisted to Neon PostgreSQL
**Frontend** (frontend/): React 19, TypeScript, Vite 6, Tailwind v4, Framer Motion. Landing: cinematic hero video, glass cards, smooth-scroll sections, CTA to app. App shell: sidebar (+ New Chat, session list, user avatar from Google photo or generated fallback), main chat pane (PDF metadata, word/page count, summary block, message thread, composer). Dev: empty VITE_API_URL — Vite proxies to backend :3002.
**Backend** (backend/): Express REST API. Auth routes: email register/login (bcrypt + JWT), Google OAuth redirect + callback. PDF routes: multipart upload, text extraction, chunking, FAISS index build per session, summary generation. Chat routes: question → RAG pipeline → Groq completion. Groq layer supports GROQ_API_KEY + _2, _3, _4 failover when rate-limited.
**RAG pipeline (multi-agent):** Query optimizer rewrites user question for retrieval → FAISS semantic search over document chunks → context analyzer ranks/filters passages → answer generator produces final response with truncation guards for token limits.
**Database** (Neon PostgreSQL, auto-migrated on start):
| Table | Purpose |
|-------|--------|
| users | Accounts (email + Google) |
| chat_sessions | Threads per user |
| pdf_documents | PDF metadata + summaries |
| chat_messages | Full conversation history |
**Deployment:** Frontend → Vercel (frontend/ root, dist output, VITE_API_URL → Render backend). Backend → Render (npm install && npm run build -w backend, npm run start -w backend). Database → Neon. CORS locked via CLIENT_URL env.
Key Engineering Decisions
- •Chose FAISS (`faiss-node`) for in-process vector search instead of hosted pgvector — keeps RAG self-contained per upload, no embedding table migrations, and fast semantic retrieval over chunked PDF text.
- •Built a multi-agent RAG pipeline (query optimization → retrieval → context analysis → answer generation) instead of single-shot prompt stuffing — improves relevance on long documents and reduces hallucination when only partial context fits the token window.
- •Groq multi-key pool with automatic failover (`GROQ_API_KEY` + `_2`/`_3`/`_4`) — free-tier rate limits don't kill the session; backend rotates keys transparently.
- •Split Vercel frontend + Render backend — `faiss-node` uses native modules and long AI workloads; serverless (Vercel Functions) cannot host the vector index build reliably.
- •Neon PostgreSQL for sessions, PDF metadata, and message history — serverless Postgres with connection pooling; chats survive refresh and work across devices when logged in.
- •Dual auth: email/password (bcrypt + JWT) and Google OAuth 2.0 — lowers signup friction for students already on Google Workspace while keeping email-only option.
- •Always-dark glass UI with cinematic landing hero — positions the product as a premium Gen AI tool, not a utilitarian upload form.
- •Input truncation safeguards before Groq calls — prevents token overflow on large PDFs while still retrieving the most relevant chunks via FAISS first.
Business / Product Thinking
ChatPDF targets the **document Q&A wedge** — students with lecture PDFs, job seekers with long resumes, researchers skimming papers, professionals reviewing contracts. Positioning: *Chat with your PDF* — upload, summarize, ask follow-ups, answers grounded in your file.
**Product-led loop:** landing hero CTA → sign up (Google one-click) → upload PDF → instant summary + suggested questions → natural chat → return via sidebar history.
**Go-to-market:** live at chat-with-pdf-ashen.vercel.app → open-source GitHub → portfolio + LinkedIn demos with resume-PDF chat screenshots (strong recruiter hook).
**Monetization paths (not shipped):** free tier with daily upload/message caps, Pro for larger files + longer retention, team workspaces with shared document libraries.
**Trust signals:** JWT-scoped sessions, per-user PDF isolation, no API keys in client bundle, CORS locked to production frontend URL.
Results & Impact
Live at chat-with-pdf-ashen.vercel.app — frontend on Vercel, Express API on Render, Neon PostgreSQL.
**Shipped (product):** Cinematic landing with glass UI · email + Google OAuth auth · drag-and-drop PDF upload (15 MB) · automatic AI summary · suggested starter questions · natural-language Q&A grounded in document · sidebar session history · replace PDF mid-thread · user avatar (Google photo or fallback) · always-dark theme.
**Shipped (platform):** Multi-agent RAG pipeline · FAISS vector search over chunks · Groq LLM with multi-key failover · pdf-parse extraction · JWT + bcrypt auth · 4-table Neon schema (users, chat_sessions, pdf_documents, chat_messages) · npm workspaces monorepo · Vite dev proxy · render.yaml Blueprint · schema auto-init on boot.
Open source at github.com/subhm2004/Chat_with_pdf with full README, architecture diagram, env templates, Google OAuth setup guide, and Vercel + Render deployment checklist.
What I'd Do Differently
Add streaming token delivery so long answers feel live. Move PDF file storage to S3 or persistent Render disk — local uploads are lost on backend redeploy. Redis adapter if horizontal scaling is needed. Page-level citations (click chunk → highlight in PDF viewer) would strengthen trust. Rate limiting per user tier on upload and chat endpoints. Integration tests covering upload → FAISS index → RAG → Groq answer chain.