Back to Projects

PrepTalk

GenAI + Full Stack10 min read

Voice-first AI mock interview platform — Ava, an animated interviewer, speaks TTS questions, listens, and cross-questions weak answers live. Monaco code execution, an Excalidraw whiteboard read by vision, Whisper speech analytics, and a shareable Report Card PDF. Groq (Llama 3.3 70B + Scout Vision + Whisper + Orpheus TTS) with Gemini fallback on Neon PostgreSQL & Upstash Redis.

PrepTalk - Image 1
GenAI + Full Stack
1 / 6

←Use arrow keys or swipe to navigate→

The Problem

Reading interview questions is not the same as answering them out loud. Most prep tools hand you a static question bank or a text chatbot that grades a pasted answer — neither recreates the real thing: a human across the table who *speaks*, listens, hears you stumble, and immediately probes the gap in what you just said. Candidates walk into real rounds having never rehearsed the actual muscle — thinking out loud, defending a design, writing code while being watched, and recovering from a follow-up they never saw coming.

The harder gap is technical breadth in one product: real text-to-speech with a lip-synced avatar, speech-to-text scoring of spoken answers, live code execution, vision that can *read a hand-drawn system-design diagram*, adaptive cross-questioning when an answer is weak, and a shareable report at the end — all on free-tier infrastructure without hitting a rate-limit wall. PrepTalk exists to close that gap: a voice-first mock interview where Ava talks, listens, pushes back, and hands you a Report Card PDF.

My Role & Constraints

Solo Full-Stack Engineer, Architect & Product Owner — I designed and shipped PrepTalk end-to-end across three services.

**Frontend (frontend/):** React 19 + Vite + TypeScript + Tailwind v4 + Redux Toolkit + Framer Motion. Built the landing page (Ava talking live in the hero card), OTP-verified signup, Google + email login, password reset, the live Interview Runner (Monaco code editor, Excalidraw whiteboard, per-question audio recorder persisted in IndexedDB), the animated **Ava avatar with real-time lip-sync** (WebAudio AnalyserNode drives mouth amplitude, plus idle blink / breathing sway / eyebrow raises), the session-review Question Narrative screen, Report Card PDF export, the ATS Resume Analyzer, and the analytics dashboard.

**Backend (backend/):** Node.js + Express 5 on Neon PostgreSQL (primary) and Upstash Redis (queues / cache). Auth (bcrypt + JWT in HttpOnly cookies + rotating refresh tokens, Google OAuth, email OTP), session orchestration, the **server-side TTS proxy** (Ava's voice is served by question index so clients can never synthesize arbitrary text through the Groq quota), follow-up cross-question logic, the BullMQ resume pipeline, and gamification (XP, levels, streaks, 16 badges, leaderboard).

**AI service (ai-service/):** Python + FastAPI wrapping Groq — Llama 3.3 70B (question generation, evaluation, follow-ups), Llama 4 Scout Vision (diagram evaluation), Whisper v3 Turbo (transcription + speech analytics), and Orpheus TTS (Ava's voice) — with multi-key failover and smart cooldowns, an automatic Gemini fallback for the resume analyzer, and PyMuPDF / Pytesseract resume parsing.

System Design / Architecture

PrepTalk is a **three-service system over one serverless data tier** — a React frontend, an Express orchestrator, and a Python / FastAPI AI microservice, with Neon PostgreSQL as the primary datastore and Upstash Redis for queues, caches, and the leaderboard buffer.

bash
1$ React 19 + Vite (Vercel)
2$ │ REST + Socket.io
3$ ▼
4$ Express 5 backend (Render)
5$ │ │ X-API-Key (internal only)
6$ │ ▼
7$ │ FastAPI ai-service (Render)
8$ │ │
9$ │ ├──▶ Groq · Llama 3.3 70B questions · eval · follow-ups
10$ │ ├──▶ Groq · Llama 4 Scout whiteboard vision
11$ │ ├──▶ Groq · Whisper v3 transcription
12$ │ ├──▶ Groq · Orpheus TTS Ava's voice
13$ │ └──▶ Gemini résumé-analyzer fallback
14$ │
15$ ├──▶ Neon PostgreSQL users · sessions (JSONB) · resumes · gamification
16$ └──▶ Upstash Redis BullMQ · caches · atomic XP buffer

**Topology:** - **Frontend** (React 19 / Vite) ↔ **Backend** (Express 5) over REST + Socket.io - **Backend** ↔ **Neon PostgreSQL** (pg) for all durable data; ↔ **Upstash Redis** (ioredis + BullMQ) for jobs and caches - **Backend** ↔ **AI service** (FastAPI) over an internal X-API-Key — the AI service is never exposed to the browser - **AI service** ↔ Groq (Llama 3.3 70B, Llama 4 Scout Vision, Whisper v3 Turbo, Orpheus TTS) with Gemini fallback for the resume analyzer

**Interview lifecycle (Socket.io):** create session → AI_GENERATING → QUESTIONS_READY (Groq tailors questions to role, seniority, interview type, and optionally the uploaded resume) → per question the candidate records audio / writes code / draws a diagram → submit → AI_TRANSCRIBING (Whisper) → AI_EVALUATING (Llama scores technical + confidence and extracts speech analytics) → if the score is below 60, AI_FOLLOWUP generates a targeted probe (FOLLOW_UP_ADDED, capped at 2 per session, tagged followUpOf) → session completed computes final scores.

**Ava's voice path:** the client calls POST /api/sessions/:id/speak with a question *index* (not text); the backend looks the text up server-side, proxies to Groq Orpheus TTS, and streams back WAV — the browser plays it and feeds amplitude into the lip-sync AnalyserNode. If TTS is unavailable it degrades to browser speechSynthesis.

**Data tier (Neon Postgres):** self-bootstrapping schema (CREATE TABLE IF NOT EXISTS, no migration tooling). Flexible AI output — questions with evaluations, speech metrics, and follow-ups — lives in JSONB; relational data gets real constraints and indexes. sessions carries a (user_id, created_at DESC) index for dashboard pagination; gamification a partial index on xp DESC WHERE leaderboard_opt_in so the leaderboard is a single ORDER BY. Concurrent per-question evaluations are serialized with a per-session in-process lock so whole-blob JSONB writes never lose updates.

**Redis (Upstash):** bull:resume-processing:* job queue, resume-cache:* (same file → skip Groq), a 24h resume-view:* API cache, and user:{id}:xp_buffer — an atomic XP counter incremented during the interview and flushed to Postgres on completion.

**Resume analyzer:** upload PDF / DOCX → BullMQ background job → PyMuPDF / Pytesseract extraction → ATS score across 6 sections (/100), entity extraction, deep insights with role-alignment bars, JD matching, and streamed AI bullet rewrites + tailored cover letters.

**Deploy:** frontend on Vercel; backend + AI service on Render via render.yaml; Neon + Upstash are serverless and TLS from anywhere.

Key Engineering Decisions

  • •Split into three services (React frontend, Express orchestrator, Python / FastAPI AI) instead of one app — the AI workloads (Groq LLM, Vision, Whisper, TTS) live behind an internal API key in Python where the ML tooling is native, while Express owns auth, sessions, and real-time; each scales and deploys independently.
  • •Served Ava's TTS by question index, never by client-supplied text — the browser sends `POST /speak` a question index and the backend resolves the text server-side before hitting Groq Orpheus. Clients can never synthesize arbitrary speech through the Groq quota, closing an obvious cost and abuse hole.
  • •Drove lip-sync from real audio amplitude (WebAudio `AnalyserNode`), not a canned loop — Ava's mouth tracks the actual speech waveform, with idle blink / breathing / eyebrow motion, so the avatar reads as present rather than a looping GIF; browser `speechSynthesis` is the graceful fallback when TTS is unavailable.
  • •Made cross-questioning score-gated (below 60 triggers a probe, capped at 2 per session) — a targeted follow-up attacks the exact gap ("you mentioned X, but what happens when Y?") instead of interrogating every answer; the cap keeps sessions bounded and the `followUpOf` tag preserves the thread in the final report.
  • •Chose Neon Postgres as primary with JSONB for AI output plus real indexes for relational data — questions, evaluations, and speech metrics are schemaless blobs, but users, sessions, and gamification get constraints and targeted indexes (dashboard pagination, single-`ORDER BY` leaderboard). The schema self-bootstraps with `CREATE TABLE IF NOT EXISTS`, so there is no migration tooling to operate.
  • •Serialized per-question evaluations with a per-session in-process lock — because a whole session (all questions + evaluations) is one JSONB blob, concurrent writes would clobber each other; the lock guarantees no evaluation is lost when answers submit in quick succession.
  • •Buffered XP in Redis during interviews and flushed to Postgres on completion — atomic `user:{id}:xp_buffer` increments keep the hot gamification path off the primary DB, while the leaderboard reads a Postgres partial index so live XP never thrashes the durable store.
  • •Ran the resume analyzer as a BullMQ background pipeline with a content-hash cache — uploads don't block the request, identical files skip Groq entirely (`resume-cache:*`), responses cache for 24h, and Socket.io `resume:status` streams progress to the UI.
  • •Made Groq multi-key with cooldowns and an automatic Gemini fallback for the analyzer only — interviews and voice stay on Groq for latency, but when every Groq key is rate-limited the resume analyzer transparently fails over to Gemini instead of erroring.
  • •Used Llama 4 Scout Vision to actually read the Excalidraw whiteboard — the diagram is uploaded to Cloudinary and visually evaluated, so system-design answers are graded on the boxes-and-arrows the candidate drew, not just what they said out loud.

Business / Product Thinking

PrepTalk sits in the **interview-readiness wedge** — candidates who have read a hundred questions but never rehearsed answering out loud, defending a design, or recovering from a follow-up. Positioning: *Skip the Nerves. Ace the Interview.* — a live AI interviewer (Ava) who speaks, listens, and pushes back, not another static question bank.

**Product-led loop:** landing (Ava talking in the hero) → OTP signup / Google → Interview Now → pick role + seniority + type and optionally attach your resume → talk to Ava, write code, draw the diagram → get scored with live cross-questions → **Report Card PDF** (overall / technical / confidence, speech analytics, per-question feedback) that's LinkedIn-share-worthy → return for the streak, XP, and badges.

**Second surface (top of funnel):** the free ATS Resume Analyzer — upload a resume for an instant /100 ATS score, entity extraction, role alignment, JD match, and AI rewrites — pulls in job seekers who then try a mock interview.

**Go-to-market:** live at prep-talk-eight.vercel.app → GitHub with SETUP / ARCHITECTURE / DEPLOYMENT docs → placement cells, r/cscareerquestions, and LinkedIn posts sharing a Report Card.

**Monetization paths (not shipped):** free tier with N interviews per month, Pro for unlimited sessions + company-specific interview packs, and institution workspaces with cohort analytics.

Results & Impact

Live at prep-talk-eight.vercel.app with source at github.com/subhm2004/PrepTalk.

**Shipped (interview):** talking **Ava** avatar with real-time lip-sync (Groq Orpheus TTS, speechSynthesis fallback) · role / seniority / type + resume-tailored questions (Llama 3.3 70B) · Monaco live code execution (JDoodle, 20+ languages) · Excalidraw whiteboard evaluated by Llama 4 Scout Vision · per-question audio recording in IndexedDB · Whisper speech analytics (pace / WPM, filler words, pauses, clarity) · technical + confidence scoring with ideal answers · **score-gated cross-questioning** (follow-up probes, 2 per session) · session-review Question Narrative (Neural Analytics vs Execution Reference) · one-click **Report Card PDF**.

**Shipped (resume analyzer):** PDF / DOCX upload → BullMQ pipeline → ATS score across 6 sections (/100) · entity extraction + profile summary · deep insights with role-alignment bars · JD match scoring · streamed AI bullet rewrites + tailored cover letters · downloadable ATS PDF.

**Shipped (platform):** email OTP + Google OAuth + password reset · JWT in HttpOnly cookies + rotating refresh tokens · Socket.io interview event stream (AI_GENERATING → QUESTIONS_READY → AI_TRANSCRIBING → AI_EVALUATING → AI_FOLLOWUP → session completed) · Neon PostgreSQL (JSONB + targeted / partial indexes, self-bootstrapping schema) · Upstash Redis (BullMQ, caches, atomic XP buffer) · gamification (XP, levels, daily streaks, 16 badges, Postgres-indexed leaderboard) · analytics dashboard (score trends, performance by role, speech charts) · Groq multi-key failover + Gemini analyzer fallback · Jest / Supertest + Vitest + Pytest suites in CI · one-click Render + Vercel deploy.

What I'd Do Differently

Stream evaluation tokens so long feedback feels live instead of arriving as a block. Replace the per-session in-process lock with a distributed lock (Redis) before running the backend on multiple instances — the current lock is single-node. Move the Socket.io fan-out to a Redis adapter for horizontal scale. Add company-specific interview packs and a coding round with test-case grading beyond single execution. Cache TTS audio per question so replays and repeat sessions skip Groq. Persist whiteboard evaluations with region-level feedback that highlights the part of the diagram that lost points. Grow speech analytics into cross-session trend coaching, not just per-question metrics.

Tech Stack

React
Vite
TypeScript
Tailwind CSS
FastAPI
Python
Node.js
Express
Groq
Gemini
Neon PostgreSQL
Redis
BullMQ
Socket.IO
JWT
Redux Toolkit
Framer Motion
Monaco Editor
Excalidraw
Whisper
Cloudinary
JDoodle
Vercel
Render
GitHub

Want to see more?

View All Projects