Serve TTS by Index, Not by Text: A Tiny API Decision That Closes a Cost Hole
In PrepTalk, an AI interviewer named Ava speaks her questions out loud with real text-to-speech, lip-synced to the audio. The obvious way to build that is an endpoint that takes text and returns audio: POST /speak { text }. It works instantly — and it hands every visitor a button that spends your paid TTS quota on any string they like.
The abuse the naive endpoint invites
A POST /speak { text } endpoint is a public text-to-speech proxy wearing your API key. Someone can script it to synthesize paragraphs of arbitrary text, burn through your Groq quota, and run up the bill — all through a legitimate route. The client is deciding what you pay to generate, which is exactly backwards.
Let the server decide what gets spoken
The fix is to stop trusting the client with the text. The browser already knows which question it's on, so it sends an index, not a string. The backend looks the text up server-side from the session and only then calls TTS:
1$// backend — client sends an index, never free text2$app.post("/api/sessions/:id/speak", async (req, res) => {3$const { questionIndex } = req.body;4$const session = await getSession(req.params.id);5$const text = session.questions[questionIndex]?.prompt;6$if (!text) return res.status(400).json({ error: "bad index" });7$8$const wav = await groqOrpheusTTS(text); // our key, our text9$res.type("audio/wav").send(wav);10$});
Now the only thing a client can synthesize is a question that already exists in their own session. There is no arbitrary-text path to abuse.
Lip-sync stays on the client, cheaply
Serving audio server-side doesn't cost you the animation. The browser plays the returned WAV and feeds its amplitude into a WebAudio AnalyserNode, which drives Ava's mouth in real time — so the avatar tracks the actual waveform without any extra API calls. If TTS is unavailable, it falls back to the browser's built-in speechSynthesis, and the interview keeps working.
The general rule
This is a small change with a wider principle behind it: never let a client decide what a paid, privileged resource does on your behalf. Pass a reference the server can resolve — an index, an ID, an enum — not the raw payload. The same shape shows up everywhere: signed asset keys instead of arbitrary URLs, template IDs instead of raw email bodies, product SKUs instead of client-sent prices. Give the client a pointer; keep the spending decision on the server.
From the project
PrepTalk
GenAI + Full Stack