Live roles / sunday
Compatibility brief · Discovered by NoBoards 1d ago

Senior Data Engineer

sunday · bangkok, TH
Sponsorship not statedMode not stated5+ years requestedIndustry: insurtech
Extracted role summary · not eligibility evidence

Design, build, and operate batch and incremental data pipelines end-to-end for a global insurance group using a Databricks lakehouse on AWS.

DatabricksAWSUnity CatalogDelta LakePySparkSpark ConnectPostgreSQLMySQLMSSQLDB2

Data Engineering is a small, high-leverage platform team serving every Sunday entity (insurance, broker, care, technology) across two countries. We run one lakehouse and a set of frameworks — not a pile of one-off pipelines. You will have unusual ownership: the systems you design run production for the whole group.

Key Areas of Responsibility

The lakehouse: Databricks on AWS (Unity Catalog, Delta) across Thailand and Indonesia, prod and non-prod workspaces. Ingestion & export frameworks: our Python/PySpark (Spark Connect) batch framework pulling from operational databases (Postgres, MySQL, MSSQL, DB2, MongoDB) into Delta — and pushing back out: hash-diff incremental reverse-ETL from the lakehouse into operational Postgres serving live systems, plus Kafka (Avro/Confluent) streams. dbt warehouses: layered medallion architecture for Thailand and Indonesia, tested and CI-gated.

Data Governance

End to End data governance across Thailand and Indonesia.

Everything as code

Terraform/Terragrunt for infrastructure, Databricks Asset Bundles for jobs and schedules, Jenkins for CI/CD. If it isn't in git, it doesn't exist.

Responsibilities

Design, build, and operate batch and incremental pipelines end-to-end — from source system to the table an underwriter's dashboard reads. Extend frameworks rather than write one-offs: when you solve a problem, the next ten sources get the solution for free. Own data quality and reliability: tests, monitoring, SLAs, and the root-cause analysis when something breaks at 7am. Model insurance data (policies, claims, members, payments) so analysts and data scientists can trust and reuse it.

Partner directly with data scientists, analysts, and business stakeholders across four business entities and two countries. Review more code than you write — including code written by AI. We work AI-assisted by default. Mentor other engineers and raise the team's bar for design and operational discipline.

Requirement

5+ years building production data systems, with at least one system you owned end-to-end (design → build → operate → evolve). Expert SQL and solid data modeling — dimensional and lakehouse/medallion patterns, and the judgement to know when each applies. Strong Python engineering: typed, tested, reviewable code — not just notebooks. Production experience with Spark and a lakehouse platform (Databricks strongly preferred) or equivalent scale elsewhere.

Real understanding of incremental processing

CDC, merge/upsert semantics, idempotency, late-arriving data, backfills. Operational depth: you've debugged pipelines under pressure and can tell the story of a root cause you found. Fluency with AI coding tools and a track record of catching their mistakes. We don't screen AI out of our hiring process — we screen for people who use it well. Clear written and spoken English — it's our working language, and much of our design work happens in documents and PRs.