Data Engineer

TechdomeHyderabad, TelanganaOn-siteFull-timeMid level, 2–5 yearsListed 3 hours ago

Apply now

About this role

Role summary: The Data Engineer builds the data foundation that AI agents, scorecards, and dashboards run on: ingestion pipelines, connectors, metadata and search indexes, vector stores, data marts, and quality instrumentation.
Experience: 3 to 6 years in data engineering. Senior Data Engineer: 2+ years, including ownership of data foundation architecture and connector frameworks.
### Key responsibilities

- Design and build ingestion pipelines for structured data, unstructured content (documents, PDFs), and metadata.
- Build reusable connectors to enterprise systems, catalogs, content repositories, and third-party or licensed sources via APIs.
- Model and build raw-to-mart data layers that serve analytics and AI use cases.
- Implement metadata extraction, enrichment, and search indexing, including semantic and vector indexes for RAG.
- Set up and manage vector databases and knowledge repositories used by LLM agents.
- Implement data quality rules, profiling, scoring outputs, and exception handling.
- Register lineage and maintain source registries, version tracking, and refresh controls.
- Design data stores for signals, findings, audit trails, and user feedback loops.
- Apply security and governance controls: RBAC, PII/sensitivity flagging, and approved data handling.
- Support SIT/UAT data validation, defect fixes, and production release activities.

Required skills

- Strong Python and SQL; solid data modeling (dimensional and normalized).
- Hands-on experience with a modern data platform such as Databricks, Snowflake, or Azure/AWS data services.
- Pipeline orchestration and transformation (Spark, Airflow, ADF, dbt, or similar).
- API-based integration (REST), JSON handling, and incremental/CDC ingestion patterns.
- Data quality frameworks and testing practices for pipelines.
- Version control (Git) and CI/CD for data workloads.

Preferred skills

- RAG data preparation: chunking, embeddings, vector databases (Azure AI Search, pgvector, Pinecone, or similar).
- Unstructured content processing: text extraction, OCR, document parsing.
- Metadata management, data catalogs, ontologies, or knowledge graphs (for example Neptune or other graph databases).
- Experience supporting LLM or agentic applications with grounded, traceable data.
- Life sciences data exposure (commercial, medical, regulatory, or launch data) and regulated-data handling.
- Cloud certification (Azure Data Engineer, Databricks, AWS, or Snowflake).