Senior Data Engineer
Xebia
5 - 10 years
Pune City
Posted: 28/06/2026
Job Description
Streaming & Real-Time Data Processing
- Design and develop real-time streaming pipelines using Databricks Structured Streaming.
- Build and maintain Kafka-based ingestion frameworks.
- Handle late-arriving events using watermarks, event-time processing, and stateful streaming concepts.
- Implement exactly-once processing and checkpointing mechanisms.
- Monitor and optimize streaming workloads for performance and reliability.
Spark Declarative Pipelines (SDP) & Delta Live Tables (DLT):
- Design and implement Spark Declarative Pipelines using Databricks.
- Develop Delta Live Table (DLT) pipelines for scalable data transformations.
- Implement data quality expectations, validations, retries, and failure handling within DLT pipelines.
- Manage pipeline dependencies and orchestration using declarative approaches.
- Understand advantages of SDP over traditional ETL pipelines, including:
- Simplified pipeline development
- Reduced operational overhead
- Automated lineage tracking
- Improved maintainability
- Enhanced observability
Databricks Asset Bundles:
- Develop and deploy Databricks Asset Bundles (DAB) for CI/CD.
- Configure mandatory bundle components including:
- databricks.yml
- Resource definitions
- Job definitions
- Environment configurations
- Source code artifacts
- Manage deployment across development, testing, and production environments.
Data Ingestion & Auto Loader
- Build scalable ingestion frameworks using Databricks Auto Loader (CloudFiles).
- Configure schema inference and schema evolution strategies.
- Handle duplicate records and implement deduplication mechanisms.
- Design robust ingestion pipelines from ADLS Gen2 and other cloud storage systems.
- Work with CloudFiles architecture and incremental processing patterns.
Lakehouse & Medallion Architecture
- Design and implement Bronze, Silver, and Gold layer architectures.
- Build enterprise-grade data products using Lakehouse principles.
- Implement Delta Lake optimization techniques.
- Ensure data quality, governance, and lineage across layers.
Azure Data Engineering
- Work extensively with:
- Azure Data Lake Storage Gen2 (ADLS Gen2)
- Azure Service Principals
- Azure Key Vault
- Azure Data Factory (ADF)
- Azure Databricks
- Implement secure authentication and authorization mechanisms.
- Troubleshoot and debug ADF pipeline failures.
Spark Optimization & Performance Tuning
- Optimize Spark jobs using:
- Partitioning strategies
- Adaptive Query Execution (AQE)
- Broadcast joins
- Caching and persistence
- Z-Ordering
- File compaction techniques
- Analyze Spark UI for job failures and performance bottlenecks.
- Troubleshoot executor failures, stage failures, skew issues, memory problems, and shuffle bottlenecks.
Unity Catalog & Governance
- Implement and manage Unity Catalog.
- Configure data access controls and governance policies.
- Establish lineage tracking and data security standards.
- Manage catalog, schema, and table-level permissions.
CI/CD & DevOps
- Implement CI/CD pipelines for Databricks projects.
- Work with Git branching strategies:
- Feature Branches
- Pull Requests
- Code Reviews
- Merge Processes
- Integrate VS Code with Databricks.
- Automate deployments using Databricks Asset Bundles and DevOps pipelines.
- Experience with Azure DevOps, GitHub Actions, Jenkins, or equivalent CI/CD platforms.
Data Modeling & SQL
- Develop complex SQL transformations.
- Build optimized analytical data models.
- Write performant SQL queries for large-scale datasets.
- Design dimensional and Lakehouse data models.
Services you might be interested in
We Search & Apply Jobs for You!
Our team scans through 1000s of opportunities and applies to roles best suited to your profile
Save 100+ hours and focus on what matters - cracking interviews and landing offers.
