Data Engineer | Azure
🇵🇹 Portugal
NumPy
Unity
Management
Java
Python
Angular
PHP
AWS
Azure
Git
Terraform
Snowflake
GitHub
Design
Blockchain
Recruitment
Data Science
Devops
SQL
Analyst
Testing
Data Engineer | Azure
from 🇵🇹 Portugal
Make an impact by working for sectors where technology is the enabler, everything is ground-breaking and there’s a constant need to be innovative.
Create and enhance projects in Java, Python, Angular, PHP, .NET and so much more while diving in the world of Blockchain, Artificial Intelligence, Data Science, Security and Internet of Things.
Be part of the team that combines business knowledge, technological edge and a design experience. Our different backgrounds and know-how are key in developing solutions and experiences for digital clients.
Face challenges and learn other ways of thinking and seeing the world - there’s always room for your energy and creativity.
About the role
Data Engineer is responsible for building and maintaining Data Platforms. Recognizes the importance of data for the organization in the areas where it is the key to success. Maintaining an eye on the big picture and knowing the details of the business are decisive for this role.
This role focuses on designing, developing, and maintaining the data platform for data storage, processing, orchestration, and analysis.
The mission involves implementing scalable, high-performance data pipelines and data integration solutions.
Agnostic of data sources and technologies to ensure efficient data flow and high data quality, enabling data scientists, analysts, and other stakeholders to access and analyze data effectively.
As a part of your job, you will:
Design, build, and maintain scalable data platforms;
Collect, process, and analyze large and complex data sets from various sources;
Develop and implement data processing workflows using data processing framework technologies such as Spark and Apache Beam;
Collaborate with cross-functional teams to ensure data accuracy and integrity;
•Ensure data security and privacy through proper implementation of access controls and data encryption;
Extraction of data from various sources, including databases, file systems, and APIs;
Monitor system performance and optimize for high availability and scalability.
What are we looking for?
Technical Skills Profile: Databricks & PySpark Developer
Databricks Ecosystem:
Proficiency in building, managing, and optimisingDatabricks Workspaces andNotebooks.
Experience withDatabricks Unity Catalog for data governance, access control, and metadata management.
Knowledge ofDatabricks Workflows (Jobs, Orchestration).
Distributed Computing (PySpark):
Strong hands-on experience usingPySpark (Spark DataFrames, Datasets, and RDDs) for large-scale data processing.
Deep understanding ofSpark Architecture: Drivers, Executors, Workers, Partitions, Memory Management, and Shuffle operations.
Ability to tune and optimise Spark applications (handling data skew, broadcast joins, caching strategies, and memory tuning).
Languages & Data Warehousing
Python:
Familiarity with data analysis libraries (e.g., Pandas, NumPy) and standard library automation scripts.
SQL:
Expert-level SQL skills (complex JOINs, Window functions, CTEs, Aggregations, and Query Tuning).
Experience withSpark SQL for data transformation and analytics.
Data Architecture & Pipeline Engineering
Data Pipelines & ETL/ELT:
Designing, building, and maintaining high-volume batch and real-time streaming data pipelines (Structured Streaming).
Implementation ofMedallion Architecture (Bronze, Silver, and Gold layers) using Delta Lake.
Incremental data loading patterns (Change Data Capture - CDC, MERGE INTO, Auto Loader).
Data Warehousing Concepts:
Knowledge of data modelling techniques (Dimensional Modelling, Star/Snowflake Schemas, Data Vault).
Cloud Infrastructure, CI/CD & DevOps (Nice to Have / Preferred)
Cloud Providers:
Hands-on experience withAzure, integration with Databricks (e.g., AWS S3, Azure ADLS Gen2, Snowflake).
DevOps & Software Engineering Practices:
Version control usingGit (GitHub, Azure DevOps, or GitLab).
Experience withCI/CD pipelines for deploying Databricks code (using Databricks Asset Bundles, Terraform, or REST APIs).
Automated testing for PySpark jobs (pytest) and data validation tools (e.g., Great Expectations)
Personal traits
Ability to adapt to different contexts, teams, and Clients;
Teamwork skills but also a sense of autonomy;
Motivation for international projects and ok if travel is included;
Willingness to collaborate with other players;
Strong communication skills.
At Celfocus, we are committed to cultivate a diverse and inclusive workplace. As an equal-opportunity employer, we welcome applicants of all backgrounds, gender identities, and abilities. We are dedicated to providing reasonable accommodations for candidates with specific needs. If you require any adjustments during the selection process, please inform our Talent Acquisition Team.
Come join the Team!






