About the Role
OdoCore is looking for an experienced Data Engineer to join a large-scale data platform modernization engagement. This is a hands-on role focused on migrating a legacy Hadoop-based data warehouse to a modern lakehouse architecture. You'll be working directly on production-critical pipelines that power core business reporting and analytics, so real-world experience with the tools below — not just theoretical familiarity — is essential.
The Engagement
You'll be part of a team modernizing a large-scale data platform, covering:
- Migrating Spark 2 → Spark 3
- Migrating Hive → Iceberg
- Migrating Oozie → Apache Airflow
- Offloading IBM Netezza and IBM DataStage workloads onto a Spark 3 / Iceberg Lakehouse
Required Hard Skills
- Strong SQL, with real experience reading and re-optimizing complex, poorly-written legacy queries
- Apache Spark (Spark 2 and Spark 3) — PySpark or Scala
- Apache Hive and Apache Iceberg — table formats, partitioning, schema evolution
- ETL / data warehousing fundamentals — medallion (bronze/silver/gold) architecture, dimensional modeling
- Orchestration tools — Oozie and/or Apache Airflow
- Linux/Unix comfort, Git
- Cloudera CDP platform exposure (Impala, Ranger) — or a fast ability to ramp up on it
Experience Level
3+ years in data engineering, with at least one prior migration, ETL modernization, or lakehouse build under your belt. Mid-to-senior individual contributors preferred.