PrepTalk
Voice-first AI mock interview platform — Ava, an animated interviewer, speaks TTS questions, listens, and cross-questions weak answers live. Monaco code execution, an Excalidraw whiteboard read by vision, Whisper speech analytics, and a shareable Report Card PDF. Groq (Llama 3.3 70B + Scout Vision + Whisper + Orpheus TTS) with Gemini fallback on Neon PostgreSQL & Upstash Redis.

←Use arrow keys or swipe to navigate→
The Problem
Reading interview questions is not the same as answering them out loud. Most prep tools hand you a static question bank or a text chatbot that grades a pasted answer — neither recreates the real thing: a human across the table who *speaks*, listens, hears you stumble, and immediately probes the gap in what you just said. Candidates walk into real rounds having never rehearsed the actual muscle — thinking out loud, defending a design, writing code while being watched, and recovering from a follow-up they never saw coming.
The harder gap is technical breadth in one product: real text-to-speech with a lip-synced avatar, speech-to-text scoring of spoken answers, live code execution, vision that can *read a hand-drawn system-design diagram*, adaptive cross-questioning when an answer is weak, and a shareable report at the end — all on free-tier infrastructure without hitting a rate-limit wall. PrepTalk exists to close that gap: a voice-first mock interview where Ava talks, listens, pushes back, and hands you a Report Card PDF.
My Role & Constraints
Solo Full-Stack Engineer, Architect & Product Owner — I designed and shipped PrepTalk end-to-end across three services.
**Frontend (frontend/):** React 19 + Vite + TypeScript + Tailwind v4 + Redux Toolkit + Framer Motion. Built the landing page (Ava talking live in the hero card), OTP-verified signup, Google + email login, password reset, the live Interview Runner (Monaco code editor, Excalidraw whiteboard, per-question audio recorder persisted in IndexedDB), the animated **Ava avatar with real-time lip-sync** (WebAudio AnalyserNode drives mouth amplitude, plus idle blink / breathing sway / eyebrow raises), the session-review Question Narrative screen, Report Card PDF export, the ATS Resume Analyzer, and the analytics dashboard.
**Backend (backend/):** Node.js + Express 5 on Neon PostgreSQL (primary) and Upstash Redis (queues / cache). Auth (bcrypt + JWT in HttpOnly cookies + rotating refresh tokens, Google OAuth, email OTP), session orchestration, the **server-side TTS proxy** (Ava's voice is served by question index so clients can never synthesize arbitrary text through the Groq quota), follow-up cross-question logic, the BullMQ resume pipeline, and gamification (XP, levels, streaks, 16 badges, leaderboard).
**AI service (ai-service/):** Python + FastAPI wrapping Groq — Llama 3.3 70B (question generation, evaluation, follow-ups), Llama 4 Scout Vision (diagram evaluation), Whisper v3 Turbo (transcription + speech analytics), and Orpheus TTS (Ava's voice) — with multi-key failover and smart cooldowns, an automatic Gemini fallback for the resume analyzer, and PyMuPDF / Pytesseract resume parsing.
System Design / Architecture
PrepTalk is a **three-service system over one serverless data tier** — a React frontend, an Express orchestrator, and a Python / FastAPI AI microservice, with Neon PostgreSQL as the primary datastore and Upstash Redis for queues, caches, and the leaderboard buffer.
1$React 19 + Vite (Vercel)2$│ REST + Socket.io3$▼4$Express 5 backend (Render)5$│ │ X-API-Key (internal only)6$│ ▼7$│ FastAPI ai-service (Render)8$│ │9$│ ├──▶ Groq · Llama 3.3 70B questions · eval · follow-ups10$│ ├──▶ Groq · Llama 4 Scout whiteboard vision11$│ ├──▶ Groq · Whisper v3 transcription12$│ ├──▶ Groq · Orpheus TTS Ava's voice13$│ └──▶ Gemini résumé-analyzer fallback14$│15$├──▶ Neon PostgreSQL users · sessions (JSONB) · resumes · gamification16$└──▶ Upstash Redis BullMQ · caches · atomic XP buffer
**Topology:**
- **Frontend** (React 19 / Vite) ↔ **Backend** (Express 5) over REST + Socket.io
- **Backend** ↔ **Neon PostgreSQL** (pg) for all durable data; ↔ **Upstash Redis** (ioredis + BullMQ) for jobs and caches
- **Backend** ↔ **AI service** (FastAPI) over an internal X-API-Key — the AI service is never exposed to the browser
- **AI service** ↔ Groq (Llama 3.3 70B, Llama 4 Scout Vision, Whisper v3 Turbo, Orpheus TTS) with Gemini fallback for the resume analyzer
**Interview lifecycle (Socket.io):** create session → AI_GENERATING → QUESTIONS_READY (Groq tailors questions to role, seniority, interview type, and optionally the uploaded resume) → per question the candidate records audio / writes code / draws a diagram → submit → AI_TRANSCRIBING (Whisper) → AI_EVALUATING (Llama scores technical + confidence and extracts speech analytics) → if the score is below 60, AI_FOLLOWUP generates a targeted probe (FOLLOW_UP_ADDED, capped at 2 per session, tagged followUpOf) → session completed computes final scores.
**Ava's voice path:** the client calls POST /api/sessions/:id/speak with a question *index* (not text); the backend looks the text up server-side, proxies to Groq Orpheus TTS, and streams back WAV — the browser plays it and feeds amplitude into the lip-sync AnalyserNode. If TTS is unavailable it degrades to browser speechSynthesis.
**Data tier (Neon Postgres):** self-bootstrapping schema (CREATE TABLE IF NOT EXISTS, no migration tooling). Flexible AI output — questions with evaluations, speech metrics, and follow-ups — lives in JSONB; relational data gets real constraints and indexes. sessions carries a (user_id, created_at DESC) index for dashboard pagination; gamification a partial index on xp DESC WHERE leaderboard_opt_in so the leaderboard is a single ORDER BY. Concurrent per-question evaluations are serialized with a per-session in-process lock so whole-blob JSONB writes never lose updates.
**Redis (Upstash):** bull:resume-processing:* job queue, resume-cache:* (same file → skip Groq), a 24h resume-view:* API cache, and user:{id}:xp_buffer — an atomic XP counter incremented during the interview and flushed to Postgres on completion.
**Resume analyzer:** upload PDF / DOCX → BullMQ background job → PyMuPDF / Pytesseract extraction → ATS score across 6 sections (/100), entity extraction, deep insights with role-alignment bars, JD matching, and streamed AI bullet rewrites + tailored cover letters.
**Deploy:** frontend on Vercel; backend + AI service on Render via render.yaml; Neon + Upstash are serverless and TLS from anywhere.
Key Engineering Decisions
- •Split into three services (React frontend, Express orchestrator, Python / FastAPI AI) instead of one app — the AI workloads (Groq LLM, Vision, Whisper, TTS) live behind an internal API key in Python where the ML tooling is native, while Express owns auth, sessions, and real-time; each scales and deploys independently.
- •Served Ava's TTS by question index, never by client-supplied text — the browser sends `POST /speak` a question index and the backend resolves the text server-side before hitting Groq Orpheus. Clients can never synthesize arbitrary speech through the Groq quota, closing an obvious cost and abuse hole.
- •Drove lip-sync from real audio amplitude (WebAudio `AnalyserNode`), not a canned loop — Ava's mouth tracks the actual speech waveform, with idle blink / breathing / eyebrow motion, so the avatar reads as present rather than a looping GIF; browser `speechSynthesis` is the graceful fallback when TTS is unavailable.
- •Made cross-questioning score-gated (below 60 triggers a probe, capped at 2 per session) — a targeted follow-up attacks the exact gap ("you mentioned X, but what happens when Y?") instead of interrogating every answer; the cap keeps sessions bounded and the `followUpOf` tag preserves the thread in the final report.
- •Chose Neon Postgres as primary with JSONB for AI output plus real indexes for relational data — questions, evaluations, and speech metrics are schemaless blobs, but users, sessions, and gamification get constraints and targeted indexes (dashboard pagination, single-`ORDER BY` leaderboard). The schema self-bootstraps with `CREATE TABLE IF NOT EXISTS`, so there is no migration tooling to operate.
- •Serialized per-question evaluations with a per-session in-process lock — because a whole session (all questions + evaluations) is one JSONB blob, concurrent writes would clobber each other; the lock guarantees no evaluation is lost when answers submit in quick succession.
- •Buffered XP in Redis during interviews and flushed to Postgres on completion — atomic `user:{id}:xp_buffer` increments keep the hot gamification path off the primary DB, while the leaderboard reads a Postgres partial index so live XP never thrashes the durable store.
- •Ran the resume analyzer as a BullMQ background pipeline with a content-hash cache — uploads don't block the request, identical files skip Groq entirely (`resume-cache:*`), responses cache for 24h, and Socket.io `resume:status` streams progress to the UI.
- •Made Groq multi-key with cooldowns and an automatic Gemini fallback for the analyzer only — interviews and voice stay on Groq for latency, but when every Groq key is rate-limited the resume analyzer transparently fails over to Gemini instead of erroring.
- •Used Llama 4 Scout Vision to actually read the Excalidraw whiteboard — the diagram is uploaded to Cloudinary and visually evaluated, so system-design answers are graded on the boxes-and-arrows the candidate drew, not just what they said out loud.
Business / Product Thinking
PrepTalk sits in the **interview-readiness wedge** — candidates who have read a hundred questions but never rehearsed answering out loud, defending a design, or recovering from a follow-up. Positioning: *Skip the Nerves. Ace the Interview.* — a live AI interviewer (Ava) who speaks, listens, and pushes back, not another static question bank.
**Product-led loop:** landing (Ava talking in the hero) → OTP signup / Google → Interview Now → pick role + seniority + type and optionally attach your resume → talk to Ava, write code, draw the diagram → get scored with live cross-questions → **Report Card PDF** (overall / technical / confidence, speech analytics, per-question feedback) that's LinkedIn-share-worthy → return for the streak, XP, and badges.
**Second surface (top of funnel):** the free ATS Resume Analyzer — upload a resume for an instant /100 ATS score, entity extraction, role alignment, JD match, and AI rewrites — pulls in job seekers who then try a mock interview.
**Go-to-market:** live at prep-talk-eight.vercel.app → GitHub with SETUP / ARCHITECTURE / DEPLOYMENT docs → placement cells, r/cscareerquestions, and LinkedIn posts sharing a Report Card.
**Monetization paths (not shipped):** free tier with N interviews per month, Pro for unlimited sessions + company-specific interview packs, and institution workspaces with cohort analytics.
Results & Impact
Live at prep-talk-eight.vercel.app with source at github.com/subhm2004/PrepTalk.
**Shipped (interview):** talking **Ava** avatar with real-time lip-sync (Groq Orpheus TTS, speechSynthesis fallback) · role / seniority / type + resume-tailored questions (Llama 3.3 70B) · Monaco live code execution (JDoodle, 20+ languages) · Excalidraw whiteboard evaluated by Llama 4 Scout Vision · per-question audio recording in IndexedDB · Whisper speech analytics (pace / WPM, filler words, pauses, clarity) · technical + confidence scoring with ideal answers · **score-gated cross-questioning** (follow-up probes, 2 per session) · session-review Question Narrative (Neural Analytics vs Execution Reference) · one-click **Report Card PDF**.
**Shipped (resume analyzer):** PDF / DOCX upload → BullMQ pipeline → ATS score across 6 sections (/100) · entity extraction + profile summary · deep insights with role-alignment bars · JD match scoring · streamed AI bullet rewrites + tailored cover letters · downloadable ATS PDF.
**Shipped (platform):** email OTP + Google OAuth + password reset · JWT in HttpOnly cookies + rotating refresh tokens · Socket.io interview event stream (AI_GENERATING → QUESTIONS_READY → AI_TRANSCRIBING → AI_EVALUATING → AI_FOLLOWUP → session completed) · Neon PostgreSQL (JSONB + targeted / partial indexes, self-bootstrapping schema) · Upstash Redis (BullMQ, caches, atomic XP buffer) · gamification (XP, levels, daily streaks, 16 badges, Postgres-indexed leaderboard) · analytics dashboard (score trends, performance by role, speech charts) · Groq multi-key failover + Gemini analyzer fallback · Jest / Supertest + Vitest + Pytest suites in CI · one-click Render + Vercel deploy.
What I'd Do Differently
Stream evaluation tokens so long feedback feels live instead of arriving as a block. Replace the per-session in-process lock with a distributed lock (Redis) before running the backend on multiple instances — the current lock is single-node. Move the Socket.io fan-out to a Redis adapter for horizontal scale. Add company-specific interview packs and a coding round with test-case grading beyond single execution. Cache TTS audio per question so replays and repeat sessions skip Groq. Persist whiteboard evaluations with region-level feedback that highlights the part of the diagram that lost points. Grow speech analytics into cross-session trend coaching, not just per-question metrics.