
Ref. AO5648
We are looking for an Azure Data Engineer to join our team.
If you consider yourself a flexible and proactive person and want to face new professional challenges, send us your application! We look forward to being part of your growth and we will certainly build a successful future together!
Technical skills:
– Databricks & PySpark Developer Databricks Ecosystem: Proficiency in building, managing, and optimising Databricks Workspaces and Notebooks
– Experience with Databricks Unity Catalog for data governance, access control, and metadata management
– Knowledge of Databricks Workflows (Jobs, Orchestration)
– Distributed Computing (PySpark): Strong hands-on experience using PySpark (Spark DataFrames, Datasets, and RDDs) for large-scale data processing
– Deep understanding of Spark Architecture: Drivers, Executors, Workers, Partitions, Memory Management, and Shuffle operations
– Ability to tune and optimise Spark applications (handling data skew, broadcast joins, caching strategies, and memory tuning)
– Languages & Data Warehousing Python: Familiarity with data analysis libraries (e.g., Pandas, NumPy) and standard library automation scripts. SQL: Expert-level SQL skills (complex JOINs, Window functions, CTEs, Aggregations, and Query Tuning)
– Experience with Spark SQL for data transformation and analytics
– Data Architecture & Pipeline Engineering Data Pipelines & ETL/ELT: Designing, building, and maintaining high-volume batch and real-time streaming data pipelines (Structured Streaming)
– Implementation of Medallion Architecture (Bronze, Silver, and Gold layers) using Delta Lake
– Incremental data loading patterns (Change Data Capture - CDC, MERGE INTO, Auto Loader)
– Data Warehousing Concepts: Knowledge of data modelling techniques (Dimensional Modelling, Star/Snowflake Schemas, Data Vault)
– Cloud Infrastructure, CI/CD & DevOps (Nice to Have / Preferred) Cloud Providers: Hands-on experience with Azure, integration with Databricks (e.g., AWS S3, Azure ADLS Gen2, Snowflake)
– DevOps & Software Engineering Practices: Version control using Git (GitHub, Azure DevOps, or GitLab)
– Experience with CI/CD pipelines for deploying Databricks code (using Databricks Asset Bundles, Terraform, or REST APIs)
– Automated testing for PySpark jobs (pytest) and data validation tools (e.g., Great Expectations)
Personal skills:
– Good communication and interpersonal skills
– Proactive
– Team player
Are you interested in this opportunity?
Fill in the form.
Share this opportunity: