Senior Data Engineer – Apache Spark / Iceberg / AWS S3
Location: Remote, Finland
Contract: 12 months, extendable
Working Model: Remote
We are looking for an experienced Data Engineer to join a data engineering team responsible for building scalable data pipelines and curated datasets for analytics and machine learning use cases.
The role requires strong hands-on experience with Apache Spark, SQL, AWS S3 and data lakes, with Apache Iceberg and Dremio experience being highly desirable.
Key Responsibilities
- Develop scalable data pipelines using Apache Spark to transform raw data into curated datasets.
- Analyse raw data structures through Dremio and identify appropriate source tables, fields, joins, filters and business keys.
- Build consistent offline training datasets and batch-scoring datasets using governed feature definitions.
- Work with Data Scientists to ensure reliable and reusable data for machine learning use cases.
- Analyse existing data flows and identify improvements across data quality, performance, lineage and business logic.
- Collaborate with Data Scientists, Data Architects, Business Analysts, Product Owners and source-system experts.
- Validate data definitions, transformation logic and business requirements.
- Work with data lakes and data marts, ensuring data is structured and optimised for downstream consumption.
- Contribute to data-platform, feature-store, MLOps and lakehouse initiatives.
- Continuously improve data engineering practices, automation and pipeline performance.
Essential Skills & Experience
- 6–12 years of Data Engineering experience.
- Strong hands-on experience with Apache Spark.
- Strong SQL skills, preferably including Oracle SQL.
- Experience working with AWS S3 and cloud-based data platforms.
- Strong understanding of Data Lakes and Data Marts.
- Experience with Apache Iceberg.
- Good understanding of Unix/Linux environments and command-line operations.
- Experience with AWS DevOps and automation.
- Strong understanding of data transformation, ETL/ELT and scalable data pipelines.
- Ability to work with complex data structures and identify appropriate joins, filters and business keys.
- Strong communication and collaboration skills.
Desirable Skills
- Hands-on experience with Dremio.
- Experience with Feature Stores and ML data pipelines.
- Knowledge of MLOps.
- Experience with Lakehouse architectures.
- Exposure to batch scoring and machine learning feature engineering.
- Experience with data lineage and data-quality frameworks.





