Login Sign Up

Senior Data Engineer

Xebia

5 - 10 years

Pune City

Posted: 28/06/2026

Job Description

Streaming & Real-Time Data Processing

  • Design and develop real-time streaming pipelines using Databricks Structured Streaming.
  • Build and maintain Kafka-based ingestion frameworks.
  • Handle late-arriving events using watermarks, event-time processing, and stateful streaming concepts.
  • Implement exactly-once processing and checkpointing mechanisms.
  • Monitor and optimize streaming workloads for performance and reliability.

Spark Declarative Pipelines (SDP) & Delta Live Tables (DLT):

  • Design and implement Spark Declarative Pipelines using Databricks.
  • Develop Delta Live Table (DLT) pipelines for scalable data transformations.
  • Implement data quality expectations, validations, retries, and failure handling within DLT pipelines.
  • Manage pipeline dependencies and orchestration using declarative approaches.
  • Understand advantages of SDP over traditional ETL pipelines, including:
  • Simplified pipeline development
  • Reduced operational overhead
  • Automated lineage tracking
  • Improved maintainability
  • Enhanced observability

Databricks Asset Bundles:

  • Develop and deploy Databricks Asset Bundles (DAB) for CI/CD.
  • Configure mandatory bundle components including:
  • databricks.yml
  • Resource definitions
  • Job definitions
  • Environment configurations
  • Source code artifacts
  • Manage deployment across development, testing, and production environments.

Data Ingestion & Auto Loader

  • Build scalable ingestion frameworks using Databricks Auto Loader (CloudFiles).
  • Configure schema inference and schema evolution strategies.
  • Handle duplicate records and implement deduplication mechanisms.
  • Design robust ingestion pipelines from ADLS Gen2 and other cloud storage systems.
  • Work with CloudFiles architecture and incremental processing patterns.

Lakehouse & Medallion Architecture

  • Design and implement Bronze, Silver, and Gold layer architectures.
  • Build enterprise-grade data products using Lakehouse principles.
  • Implement Delta Lake optimization techniques.
  • Ensure data quality, governance, and lineage across layers.

Azure Data Engineering

  • Work extensively with:
  • Azure Data Lake Storage Gen2 (ADLS Gen2)
  • Azure Service Principals
  • Azure Key Vault
  • Azure Data Factory (ADF)
  • Azure Databricks
  • Implement secure authentication and authorization mechanisms.
  • Troubleshoot and debug ADF pipeline failures.

Spark Optimization & Performance Tuning

  • Optimize Spark jobs using:
  • Partitioning strategies
  • Adaptive Query Execution (AQE)
  • Broadcast joins
  • Caching and persistence
  • Z-Ordering
  • File compaction techniques
  • Analyze Spark UI for job failures and performance bottlenecks.
  • Troubleshoot executor failures, stage failures, skew issues, memory problems, and shuffle bottlenecks.

Unity Catalog & Governance

  • Implement and manage Unity Catalog.
  • Configure data access controls and governance policies.
  • Establish lineage tracking and data security standards.
  • Manage catalog, schema, and table-level permissions.

CI/CD & DevOps

  • Implement CI/CD pipelines for Databricks projects.
  • Work with Git branching strategies:
  • Feature Branches
  • Pull Requests
  • Code Reviews
  • Merge Processes
  • Integrate VS Code with Databricks.
  • Automate deployments using Databricks Asset Bundles and DevOps pipelines.
  • Experience with Azure DevOps, GitHub Actions, Jenkins, or equivalent CI/CD platforms.

Data Modeling & SQL

  • Develop complex SQL transformations.
  • Build optimized analytical data models.
  • Write performant SQL queries for large-scale datasets.
  • Design dimensional and Lakehouse data models.

Services you might be interested in

We Search & Apply Jobs for You!

Our team scans through 1000s of opportunities and applies to roles best suited to your profile

Save 100+ hours and focus on what matters - cracking interviews and landing offers.