About this role
### Data Architect – GCP / BigQuery / Power BI
Experience: 10+ years in data engineering and architecture, including 4+ years designing data platforms on GCP
#### Key Responsibilities
Discovery and assessment
- Assess the current client data landscape: the PostgreSQL schemas, volumes and growth; the third-party spreadsheets; the REST API integrations; and the HTML sources.
- Run workshops with business owners, product, IT, security and report consumers. Capture KPIs, reporting needs, data freshness SLAs and pain points.
- Document current data flows, dependencies, data ownership and data quality issues.
Target architecture design
- Design the end-to-end data platform on GCP: ingestion, landing zone (GCS), BigQuery warehouse, transformation, semantic layer and Power BI consumption.
- Define an ingestion pattern for each source type: PostgreSQL: CDC or incremental loads (for example, Datastream)
- REST APIs: Cloud Run or Cloud Functions with Pub/Sub or Cloud Scheduler
- Spreadsheets: GCS or Google Sheets with schema validation
- HTML files: parsing and extraction pipelines
- Define the BigQuery layers (Raw → Curated → Marts), naming conventions and how datasets and projects are organized.
- Design dimensional data models (star schemas and SCD handling) optimized for Power BI.
- Pick the orchestration and transformation stack (Cloud Composer, Dataform/dbt, Dataflow) and record each choice in Architecture Decision Records.
Scalability, performance and cost
- Plan capacity for 10–20% monthly growth. That rate takes ~1 TB to roughly 3–9 TB within 12 months.
- Set standards for partitioning, clustering, materialized views, BI Engine and query optimization.
- Recommend a BigQuery pricing model (on-demand or Editions/slot reservations), storage lifecycle policies and cost guardrails.
Power BI integration
- Define the connectivity approach (Import, Direct Query or Composite), gateway needs and refresh strategy, including incremental refresh.
- Guide semantic model design, row-level security (RLS) and workspace/deployment strategy.
Security, governance and quality
- Design IAM, VPC Service Controls, CMEK encryption, and column- and row-level security (policy tags).
- Set up cataloguing, lineage and metadata management (Dataplex), plus PII classification and masking.
- Define a data quality and observability framework: validation, reconciliation and alerting.
Delivery enablement
- Produce the architecture blueprint, HLD/LLD, NFRs (SLA, HA, DR, RPO/RTO), a phased roadmap and cost and effort estimates.
- Define environment strategy (Dev/QA/Prod), CI/CD and Terraform (Infrastructure-as-Code) standards.
- Hand over to the Tech Lead and act as design authority during implementation.
- Present architecture options and trade-offs to client leadership.
#### Technical Skills Required
Must have
- Expert in BigQuery: modelling, partitioning/clustering, performance tuning, cost and slot management
- GCP data services: GCS, Dataflow, Datastream, Pub/Sub, Cloud Composer, Cloud Run/Functions
- Deep PostgreSQL knowledge: CDC/logical replication, migrating TB-scale datasets
- Data warehousing and dimensional modelling (Kimball), SCD, medallion architecture
- ELT with Dataform or dbt, plus strong SQL
- Integrating REST APIs, spreadsheets and semi-structured and HTML data
- Power BI architecture: semantic models, DirectQuery vs Import, gateways, RLS, performance with BigQuery
- GCP security and governance: IAM, VPC-SC, KMS, policy tags, Dataplex
- Terraform and CI/CD
Good to have
- GCP Professional Cloud Architect or Professional Data Engineer certification
- Python
- Data observability tools
- Streaming analytics experience
- Telecom domain exposure
Soft skills
- Stakeholder management and workshop facilitation
- Can present trade-offs to executives