About this role
Position name - Data architect Experience- 8+ years Location - Bangalore employment type - Full time
Job Description: We are looking for an experienced Data Architect to lead the architecture and technical direction of a new real-time data streaming, processing, and integration platform.The role will be responsible for designing scalable, resilient, secure, and low-latency data architectures using modern data engineering and streaming technologies. The ideal candidate should have strong hands-on experience with Apache Flink, Apache Kafka, Debezium, Change Data Capture (CDC), ClickHouse, and data orchestration frameworks such as Apache Airflow. The platform will ingest real-time data from multiple source systems, process and transform streaming data, and make it available through analytical engines, APIs, dashboards, and future downstream integrations. The Data Architect will work closely with engineering, application, infrastructure, product, and business teams to establish architecture standards, technical patterns, and engineering practices for the platform.
Key Responsibilities Data Architecture & Solution Design
- Define the overall architecture for the real-time data streaming and integration platform.
- Design end-to-end data architectures covering ingestion, CDC, streaming, processing, transformation, storage, analytics, APIs, and downstream integrations.
- Define scalable, fault-tolerant, highly available, and low-latency data processing architectures.
- Establish architectural patterns for batch, streaming, event-driven, and hybrid data workloads.
- Evaluate technical requirements and translate them into scalable data architecture and solution designs.
- Define data flow, system interactions, interfaces, dependencies, and technology boundaries.
- Conduct architecture reviews and provide technical recommendations and trade-offs.
- Create architecture diagrams, technical design documents, data flow diagrams, and implementation guidelines.
Real-Time Streaming & Event Processing
- Design real-time streaming architectures using Apache Kafka and Apache Flink.
- Define Kafka architecture including:
- Topics
- Partitions
- Replication
- Consumer groups
- Retention
- Message delivery semantics
- Schema management
- Define Flink architecture and patterns for:
- Stream processing
- Transformations
- Filtering
- Enrichment
- Aggregation
- Joins
- Windowing
- Event-time processing
- State management
- Checkpointing
- Fault tolerance Define strategies for managing high-volume and low-latency streaming workloads.
- Evaluate and optimize streaming architecture for throughput, latency, scalability, and reliability.
CDC & Data Integration
- Design Change Data Capture (CDC) architecture for ingesting data from transactional/source systems.
- Define and implement architectural patterns using Debezium.
- Design reliable CDC pipelines for capturing inserts, updates, deletes, and schema changes.
- Define strategies for handling:
- Initial loads
- Incremental loads
- Schema evolution
- Data consistency
- Ordering
- Duplicate events
- Replay and recovery Design integration patterns between source systems, Kafka, Flink, analytical platforms, APIs, and downstream consumers. Data Orchestration
- Define the orchestration and workflow architecture for batch, CDC, streaming, and downstream data processing.
- Establish standards and reusable patterns using Apache Airflow or equivalent orchestration frameworks.
- Define workflows for scheduling, dependency management, retries, backfills, monitoring, and alerting.
- Design integration between orchestration frameworks and Kafka, Debezium, Flink, ClickHouse, APIs, and other data services.
- Define appropriate approaches for coordinating batch workflows and real-time/eventdriven processing.
- Evaluate orchestration technologies such as Apache Airflow, Dagster, Prefect, or Apache NiFi based on use-case requirements. Analytical Data Platform
- Define the architecture for ClickHouse and other analytical data platforms.
- Design data models optimized for high-volume analytical workloads.
- Define ingestion patterns from Kafka/Flink into ClickHouse.
- Establish strategies for partitioning, indexing, retention, aggregation, and query performance.
- Define data lifecycle and storage strategies across real-time and historical data.
Data Quality, Governance & Security
- Establish data quality standards and validation frameworks.
- Define data contracts, schemas, metadata, lineage, and data ownership.
- Establish standards for schema evolution and compatibility.
- Define data governance, security, access control, and compliance requirements.
- Establish data observability standards covering freshness, completeness, accuracy, availability, and consistency.
- Define monitoring and alerting standards for critical data pipelines.
Performance & Reliability
- Evaluate system performance across Kafka, Flink, CDC, orchestration, and analytical platforms.
- Identify architectural bottlenecks and recommend improvements.
- Define strategies for high availability, fault tolerance, disaster recovery, and scalability.
- Establish SLAs/SLOs for critical data pipelines and services.
- Drive performance tuning and capacity planning.
Technical Leadership
- Provide technical leadership and mentoring to Data Engineers.
- Conduct design and code reviews where appropriate.
- Establish development standards and reusable engineering patterns.
- Lead technical POCs and evaluate new technologies.
- Collaborate with engineering and infrastructure teams on deployment and operational architecture.
- Support production troubleshooting and complex technical issues.
- Drive continuous modernization and improvement of the data platform.
Required Skills & Experience
- 8+ years of experience in Data Engineering, Data Architecture, or related roles.
- Strong experience designing real-time and streaming data platforms.
- Strong hands-on expertise with Apache Flink.
- Strong experience with Apache Kafka.
- Hands-on experience with Debezium and Change Data Capture (CDC).
- Strong experience with analytical databases; ClickHouse experience is highly preferred.
- Strong experience with Apache Airflow or equivalent orchestration technologies.
- Strong understanding of:
- Distributed systems
- Event-driven architecture
- Streaming architecture
- Batch and real-time processing
- Data integration patterns
- Strong SQL and data modeling skills.
- Experience designing cloud-based data platforms.
- Experience with data quality, observability, monitoring, governance, and security.
- Experience with REST APIs and downstream integrations.
- Strong architecture documentation and communication skills.
- Ability to translate business and technical requirements into scalable architecture.
Preferred Skills
- Experience with Apache Airflow, Dagster, Prefect, Apache NiFi, or similar technologies.
- Strong knowledge of Flink:
- State management
- Checkpoints
- Savepoints
- Watermarks
- Event time
- Windows
- State backends
- Experience with Kafka Schema Registry.
- Experience with Avro, Protobuf, or JSON.
- Experience with Kubernetes and containerized workloads.
- Experience with AWS, Azure, or GCP.
- Experience with CI/CD and DevOps practices.
- Experience with Infrastructure as Code such as Terraform.
- Experience with data observability and monitoring platforms.
- Experience building high-volume, low-latency data platforms.
- Experience with API and microservices-based integration architectures.
