Banking transaction ETL & fraud detection
A pipeline that turns messy, multi-format financial data into a clean database — and scores every transaction for fraud risk.
2.2M+
records processed
3
input formats
5
fraud rules
30
automated tests
The problem
Financial data rarely arrives clean. The same transaction might appear as a CSV row, a JSON object or a fixed-width line from a legacy system, with different field names and different date formats. Before you can detect fraud, you have to make the data trustworthy.
What I built
- Multi-format ingestion (Java). Readers for CSV, JSON and fixed-width text, with validation and error logging.
- Normalization. Maps inconsistent field names to one schema and handles seven different date/timestamp formats, plus duplicate detection.
- Storage (PostgreSQL). Batch inserts in 1,000-record chunks through a HikariCP connection pool, Flyway migrations for versioned schema changes, and a JSONB column that preserves the raw source record alongside the normalized fields.
- Fraud detection (Scala). Immutable, composable rules: high-value transactions (over $5,000), velocity (many transactions in a short window), statistical anomalies via z-score, unusual hours (2–5 AM), and first-time merchants. Each transaction receives a score and a risk level.
- Cloud-ready. AWS SDK integration for S3 as a raw-data lake and RDS for managed PostgreSQL.
- Quality. JUnit, Mockito, ScalaTest and Testcontainers for real-database integration tests; Docker for local development; shell scripts for pre-processing and validation.
Data
The project runs on public research datasets rather than synthetic toy data: a credit-card transaction set (24,319 records), the Lending Club peer-to-peer loan set (2,260,668 records, managed with Git LFS) and the German Credit set (1,000 records) — about 2.29 million records in total.
Architecture
Sources (CSV · JSON · fixed-width)
↓
Ingestion layer (Java): readers + validation
↓
Normalization (Java): schema mapping, dates, dedupe
↓
PostgreSQL (RDS or local): Flyway migrations, indexes, JSONB
↓
Scala fraud engine: rules + z-score → score & risk level
Why it matters for GEO work
Good GEO is a data problem as much as a content problem: you collect noisy outputs from many engines, normalize them into comparable records and look for patterns over time. The same discipline — clean ingestion, a stable schema, tested logic — underpins how I approach AI-visibility measurement.