Back to Projects

Splitr

Microservices + Full Stack10 min read

A production Splitwise rebuilt as five independent microservices behind one gateway — auth, expense, AI and notification — each owning its own Postgres schema with zero cross-schema queries and its own migrations. Money never touches a float (integer paise, largest-remainder splits); a min-cash-flow algorithm collapses a group's tangled debts into the fewest transfers; one email logs in with Google or a password to the very same account; Groq turns "dinner 1200 split with rahul" into a ready expense; and idempotent cron reminders physically cannot double-mail. Every service re-verifies the JWT itself — the gateway is convenience, not security.

Splitr - Image 1
Microservices + Full Stack
1 / 4

←Use arrow keys or swipe to navigate→

The Problem

Splitwise is a genuinely useful app, and rebuilding it as a monolith would teach nothing. The interesting problems in a shared-expense app are not the CRUD screens — they are the three parts almost everyone gets subtly wrong.

**Money as a float.** 0.1 + 0.2 !== 0.3 in IEEE-754, and a bill-splitting app that stores rupees as floats will, given enough expenses, drift a balance off by a paisa that then never settles to zero. Most tutorials paper over this with a tolerance = 0.01 check; that is not a fix, it is a decision to be quietly wrong.

**Tangled group debts.** Six friends on a trip rack up eleven overlapping IOUs — A owes B, B owes C, C owes A — and paying every one of them back individually is absurd. Collapsing that into the fewest possible transfers is a real algorithm, not a SUM().

**One person, two logins.** Sign in with Google today and a password tomorrow, on the same email, and you expect the same data — which quietly requires separating *who you are* from *how you prove it*, and getting the account-linking security right so an attacker cannot pre-register your email and wait to inherit your Google login.

Splitr is a ground-up rewrite of a Next.js + Convex + Clerk tutorial into **five independent microservices** — same product, redesigned identity model, exact money math, and an architecture built to survive being poked.

My Role & Constraints

Solo Engineer & Architect — five backend services, the gateway, the Next.js frontend, the migrations, the Docker fleet and the CI.

**gateway (:4000):** the only public entry point — CORS, rate limiting (a tight budget on /api/auth), routing to the right service, and scrubbing the x-internal-key header off any outside traffic. /health is a pure liveness probe; /health/fleet aggregates every service.

**auth-service (:4001, auth schema):** Google OAuth (server-side code flow) and email+password (bcrypt), JWT issuing, refresh-token rotation, and the user directory (search + batch hydrate). Owns the account-linking logic that makes one email one account regardless of login method.

**expense-service (:4002, expense schema):** groups, expenses, the three split strategies, settlements, all balance math, debt simplification and dashboard aggregation. Where every rupee is computed.

**ai-service (:4003, stateless):** a Groq adapter — natural language into a structured expense draft, and monthly spending insights. Carries no database code at all.

**notification-service (:4004, notify schema):** Brevo transactional email, a daily payment-reminder cron and a monthly insights cron, with an idempotent send log. Carries no JWT code at all.

**Frontend (Next.js 15 + React 19):** the SPA — route groups, a client-side auth guard, a REST client with single-flight token refresh, shadcn/ui, Recharts dashboards, AI quick-add and one-click settle.

**Ops:** an npm-workspaces monorepo (one npm run dev boots the fleet), per-service checksummed SQL migrations, one Docker image for the whole backend, and GitHub Actions that typechecks, builds, runs migrations against a live database and smoke-tests the API.

System Design / Architecture

Splitr is **microservices-first**: five independent services behind one gateway, each owning its own data, its own dependencies and its own platform code. The dependency graph has no cycles, and identity is a leaf that depends on nothing.

bash
1$ Next.js SPA · :3000
2$ REST client · JWT context
3$ │ Bearer JWT
4$ ┌─────▼─────┐
5$ │ gateway │ :4000 CORS · rate-limit · route
6$ └─────┬─────┘ scrub x-internal-key at edge
7$ ┌─────────────┬───┴───────┬──────────────┐
8$ ▼ ▼ ▼ ▼
9$ auth :4001 expense :4002 ai :4003 notify :4004
10$ OAuth + pwd splits·ledger Groq NL Brevo · cron
11$ JWT issue simplify insights idempotent
12$ directory dashboard (no db) (no jwt)
13$ │ │ │
14$ [ auth ] [ expense ] [ notify ] one Neon Postgres —
15$ schema schema schema schema-per-service,
16$ no cross-schema FK
17$
18$ internal calls (x-internal-key, constant-time compared):
19$ expense → auth ai → expense notify → expense, ai

**One public door.** Only the gateway faces the internet; /internal/* paths are 404'd from outside and the x-internal-key header is scrubbed at the edge, so nobody can forge an internal call.

**Database-per-service.** Each service owns its Postgres schema and never reads another's tables — **zero cross-schema foreign keys, zero cross-schema queries**. expense.expenses.paid_by_user_id holds an auth.users.id but carries no FK; identity is resolved through the auth service's batch API, cached 60 seconds in the expense service.

**Zero trust between processes.** Every service re-verifies the JWT itself rather than trusting a gateway-set header — if a service port ever leaks onto a network, nobody can forge x-user-id: <victim>. Internal calls additionally require a constant-time-compared secret. The gateway is convenience, not security.

**Share-nothing code.** There is no shared library. Each service vendors exactly the platform code it uses in its own src/core/ — env, logging, errors, db, guards — so the AI service carries no db code and the notification service carries no JWT code. Any service could be lifted into its own repository without dragging anything along.

**Independent evolution.** Every service ships and applies its own checksummed migrations; npm run migrate simply runs all three in sequence. One Docker image runs any service via its start command, so one can be scaled or redeployed without touching the rest.

Key Engineering Decisions

  • •Rebuilt a Convex + Clerk tutorial as five hand-rolled microservices — the point was to own the parts a BaaS hides: the identity model, the money math, the service boundaries and the migrations. A managed backend would have left every interesting decision made by someone else.
  • •Database-per-service (schema-per-service) on one Neon instance. Each service owns its schema and never touches another's tables; cross-service data moves over HTTP through `/internal/*` only. It keeps the services genuinely decoupled while staying on a single free-tier database — the boundary is enforced in code, not paid for as three databases.
  • •No cross-schema foreign keys. `expense` stores an `auth.users.id` with no FK to it; identity is hydrated through auth's batch API and cached 60 seconds. Real referential integrity across a service boundary is a distributed-systems trap — dropping it is what keeps each schema independently migratable.
  • •Every service re-verifies the JWT itself instead of trusting a gateway-set `x-user-id`. If any internal port is ever exposed, a forged header buys nothing. The gateway does CORS, routing and rate limits — it is convenience, never the security boundary.
  • •Money is integer paise everywhere, never a float. `0.1 + 0.2 !== 0.3`, so the legacy code hid behind a `tolerance = 0.01` check; the rewrite removes the problem instead. Equal splits use largest-remainder distribution (₹100 across 3 is `[3334, 3333, 3333]` paise, summing to exactly 10000), and exact splits must equal the whole with no tolerance at all.
  • •Debt simplification by min-cash-flow: repeatedly match the largest debtor with the largest creditor, emit the smaller amount, repeat. For n non-zero balances that is at most n−1 transfers instead of up to n(n−1)/2. It is a pure, deterministic function — sorted with id tiebreaks — and because paise make zero mean exactly zero, termination needs no epsilon.
  • •Stayed honest about optimality: the true minimum-transaction-count problem is NP-hard. Greedy largest-vs-largest gives ≤ n−1 transfers and optimal total volume, which is what production apps actually ship — so the write-up says NP-hard out loud rather than claiming a false optimum.
  • •One account, two doors. `auth.users` is keyed on `lower(email)`; login *methods* live in a separate `auth.identities` table. Google or password on the same email resolve to the same user id, so every other service sees identical data regardless of the door used.
  • •Closed the account-linking pre-hijack: a Google login is accepted only when Google reports `email_verified: true`. So an attacker who registered your email with a password cannot wait to inherit your Google sign-in — your verified Google login owns the row, and password access still needs the password they set.
  • •Refresh tokens are opaque 384-bit values, stored SHA-256 hashed and single-use; the consume-and-revoke is one atomic SQL statement, so a token replayed after rotation is dead on arrival. Access tokens are 15-minute HS256 with a pinned algorithm, so an `alg: none` token is rejected outright.
  • •Idempotent cron by construction: every successful send claims an idempotency key (`payment_reminder:<userId>:<date>`) under a partial unique index. Overlapping crons or a restarted container physically cannot double-mail — the database refuses the second insert.
  • •Share-nothing services with no common package: each vendors its own `src/core/` and owns its migrations, so the AI service ships zero db code and the notification service ships zero JWT code. Twelve classic design patterns — Strategy for splits, Adapter for the email and AI ports, Proxy for the cached user directory, and so on — each do a real job rather than decorate.

Business / Product Thinking

Splitr is a **systems-credibility project**: the product is Splitwise, but the argument is the architecture. It is aimed squarely at the reader who asks *can you actually design a service boundary, or only call one?*

**Who it is for:** engineers preparing for system-design interviews — service decomposition, idempotency, money handling and auth are all on the menu here — and anyone evaluating backend depth beyond CRUD.

**The hooks are the three hard parts.** Money that settles to exactly zero because it never was a float; a knot of eleven debts collapsing to four payments in front of you; and one email that logs in two ways to the same account. Each is a claim most apps get subtly wrong, demonstrated rather than asserted.

**Go-to-market:** live at splitrrr-ten.vercel.app → GitHub with an architecture deep-dive, a full database reference and an ER diagram → system-design and backend communities. MIT-licensed, and every service is documented well enough to read as a reference.

**What it signals:** most portfolios prove you can wire a frontend to a database. This one proves the harder claims — a defensible service boundary, zero trust between processes, exact money arithmetic, and an auth model that resists a real attack — the things that separate "it works on my machine" from "it survives being poked."

Results & Impact

Live at splitrrr-ten.vercel.app with source at github.com/subhm2004/Splitwise.

**Shipped (architecture):** **five independent services** — gateway, auth, expense, ai, notification — behind one public gateway · **schema-per-service** on one Neon Postgres with zero cross-schema foreign keys and zero cross-schema queries · **zero-trust**: every service re-verifies the JWT, internal calls need a constant-time-compared x-internal-key · **share-nothing code**: each service vendors its own src/core/ and owns its checksummed migrations · one Docker image runs any service · an npm-workspaces monorepo booted by a single npm run dev.

**Shipped (money + debts):** integer-paise arithmetic throughout · three split strategies — **equal** (largest-remainder), **percentage** (leftovers to the largest remainders), **exact** (must equal the whole, no tolerance) · pairwise-netted running balances · settlements with partial support · **min-cash-flow debt simplification** to at most n−1 transfers, deterministic and epsilon-free.

**Shipped (auth):** **one account, two doors** — Google OAuth and email+password on the same lower(email), login methods in a separate identities table · email_verified-gated linking that blocks the pre-hijack · 15-minute pinned-algorithm access tokens · opaque, SHA-256-hashed, single-use refresh tokens rotated in one atomic statement · uniform login errors with a dummy bcrypt compare so timing never reveals whether an account exists · the tightest rate limits on /api/auth.

**Shipped (AI + email):** Groq (llama-3.3-70b-versatile, JSON mode) parses *"dinner 1200 split with rahul and priya"* into a matched-contact expense draft, and writes monthly spending insights · Brevo transactional email · a daily payment-reminder cron and a monthly insights cron, both **idempotent by a unique index** so a restart cannot double-mail.

**Shipped (product + ops):** 1:1 and group expenses, dashboard charts, invite links with rotation, a settle-guard on leaving a group, dark mode · GitHub Actions that typechecks, builds, runs migrations against a live database, smoke-tests the API and publishes images to GHCR.

What I'd Do Differently

The honest limitations are the ones any microservices build eventually meets. Cross-service reads go over HTTP with a 60-second cache, so a renamed user can be briefly stale in the expense service — acceptable here, but a real event bus (or at least cache invalidation) is the grown-up answer. There is no distributed tracing yet, so a request that fans out across services is harder to follow than it should be; a correlation id threaded end to end is the obvious next step.

The services share one Postgres instance by schema, which is the right call for cost and simplicity but is not true physical isolation — a genuinely independent deployment would give each its own database. And the AI service trusts Groq's JSON mode to return well-formed output; it validates the shape, but a stricter schema-guard with a repair pass would harden the natural-language path.

None of these are load-bearing for what the project sets out to prove. They are the distance between a convincing architecture and a production one — and I stopped where the boundaries were demonstrably real.

Tech Stack

Next.js
React
TypeScript
Tailwind CSS
Express
Node.js
PostgreSQL
Neon
Groq
Brevo
JWT
Google OAuth
Docker
GitHub Actions
Vercel
GitHub

Want to see more?

View All Projects