GeoRanker
← All work
Data engineeringJava 21Scala 2.13PostgreSQL 15AWS S3 / RDS

Banking transaction ETL & fraud detection

A pipeline that turns messy, multi-format financial data into a clean database — and scores every transaction for fraud risk.

2.2M+

records processed

3

input formats

5

fraud rules

30

automated tests

The problem

Financial data rarely arrives clean. The same transaction might appear as a CSV row, a JSON object or a fixed-width line from a legacy system, with different field names and different date formats. Before you can detect fraud, you have to make the data trustworthy.

What I built

  • Multi-format ingestion (Java). Readers for CSV, JSON and fixed-width text, with validation and error logging.
  • Normalization. Maps inconsistent field names to one schema and handles seven different date/timestamp formats, plus duplicate detection.
  • Storage (PostgreSQL). Batch inserts in 1,000-record chunks through a HikariCP connection pool, Flyway migrations for versioned schema changes, and a JSONB column that preserves the raw source record alongside the normalized fields.
  • Fraud detection (Scala). Immutable, composable rules: high-value transactions (over $5,000), velocity (many transactions in a short window), statistical anomalies via z-score, unusual hours (2–5 AM), and first-time merchants. Each transaction receives a score and a risk level.
  • Cloud-ready. AWS SDK integration for S3 as a raw-data lake and RDS for managed PostgreSQL.
  • Quality. JUnit, Mockito, ScalaTest and Testcontainers for real-database integration tests; Docker for local development; shell scripts for pre-processing and validation.

Data

The project runs on public research datasets rather than synthetic toy data: a credit-card transaction set (24,319 records), the Lending Club peer-to-peer loan set (2,260,668 records, managed with Git LFS) and the German Credit set (1,000 records) — about 2.29 million records in total.

Architecture

Sources (CSV · JSON · fixed-width)
        ↓
Ingestion layer (Java): readers + validation
        ↓
Normalization (Java): schema mapping, dates, dedupe
        ↓
PostgreSQL (RDS or local): Flyway migrations, indexes, JSONB
        ↓
Scala fraud engine: rules + z-score → score & risk level

Why it matters for GEO work

Good GEO is a data problem as much as a content problem: you collect noisy outputs from many engines, normalize them into comparable records and look for patterns over time. The same discipline — clean ingestion, a stable schema, tested logic — underpins how I approach AI-visibility measurement.