Data Engineer- GCP/Databricks

AuxoAI Engineering Pvt. Ltd.Bengaluru, KarnatakaOn-siteFull-timeSenior, 5–8 yearsListed 1 hour ago

Apply now

About this role

AuxoAI is seeking a Senior Data Engineer to lead the design, development, and optimisation of modern data pipelines and cloud-native platforms. This role is ideal for someone with deep experience building scalable batch and streaming data workflows across cloud and lakehouse environments, strong hands-on engineering skills, and a drive to mentor junior engineers.

You will work closely with AI engineers, solution architects, and cross-functional teams to build production-grade pipelines spanning ingestion, transformation, and curated data delivery — enabling high-quality data for AI and analytics use cases at scale.

Location: Bangalore / Mumbai / Hyderabad / Gurgaon (Hybrid — 3 days in office)

Responsibilities

- Design and build scalable batch and streaming data pipelines across bronze, silver, and gold medallion layers.
- Build historical and incremental ingestion using Auto Loader/Spark Structured Streaming/Kafka feeds, with GCS/Azure/AWS storage and Databricks Jobs orchestration.
- Develop and maintain Databricks-based pipelines using Spark and Delta Lake for lakehouse architecture, including migration of legacy or on-premises data sources.
- Design and maintain analytical data layers in BigQuery or Databricks SQL, applying best practices in partitioning, clustering, and performance tuning.
- Implement SQL/PySpark transformations for wide and semi-structured data, including wide-to-long processing and typed or hybrid models suited to consumer requirements.
- Collaborate with AI engineers and data scientists to build pipelines that feed ML models, AI agents, and analytical systems.
- Implement data governance, quality controls, and security best practices including schema enforcement, lineage tracking, and access controls.
- Drive engineering best practices across CI/CD, testing, monitoring, and pipeline observability.
- Partner with solution architects to translate data requirements into technical designs.
- Mentor junior data engineers and contribute to documentation, code reviews, and agile ceremonies.

Requirements

- 5+ years of hands-on experience in data engineering, building and operating production-grade pipelines.
- Hands-on experience with Databricks on GCP, including BigQuery, GCS, Databricks, Spark, Delta Lake, and structured streaming.
- Hands-on experience with Databricks and Apache Spark, including Delta Lake and end-to-end lakehouse implementations.
- Strong programming skills in Python and/or Scala, with solid SQL for modelling and transformation.
- Experience with data modelling, ETL/ELT, pipeline orchestration, and data warehousing concepts, including experience working with large, evolving JSON/map/array payloads, wide-to-long transformations, event-time context joins and schema-change handling.
- Familiarity with Git, CI/CD pipelines, and data quality monitoring frameworks.
- Solid understanding of data architecture, schema design, and performance tuning.
- Experience with Unity Catalog, source reconciliation, schema evolution, correction handling and replay/recovery testing.
- Strong problem-solving and collaboration skills.

Bonus Skills

- GCP Professional Data Engineer certification.
- Experience with Vertex AI, Cloud Functions, Dataproc, or real-time streaming architectures.
- Experience with factory or industrial data sources — MES systems, IoT sensor streams, or operational telemetry.
- Familiarity with data governance and cataloguing tools such as Dataplex, Unity Catalog, Atlan, or Collibra.
- Exposure to Docker, Kubernetes, API integration, and infrastructure-as-code (Terraform).