About this role
Work Schedule
Standard (Mon-Fri)
Environmental Conditions
Office
Job Description
Join Thermo Fisher Scientific, the world leader in serving science, as a Staff Engineer, Software to make a meaningful impact. In this role, you will provide technical leadership and architectural guidance while developing innovative software solutions that enable our customers to make the world healthier, cleaner, and safer. Working in a supportive, multi-functional environment, you will design and implement sophisticated solutions across our product portfolio, from cloud platforms to scientific instrumentation. You will support the growth of other engineers, drive adoption of best practices, and help shape the technical direction of critical projects. This position offers the opportunity to work with advanced technologies while contributing to groundbreaking scientific discoveries.
Description
Experience: 10+ years of professional experience building analytics platforms, including 5+ years of hands-on Databricks experience
Role: Senior Data Platform Engineer
Primary Skills: Databricks, Apache Spark/PySpark, Python, SQL, Delta Lake, Unity Catalog, MLflow, GenAI, BI, Cloud, CI/CD, Terraform
Role Overview
We are seeking an experienced Sr. Data Platform Engineer to build and operate secure, scalable data platforms and deliver reliable data, AI, generative AI, and analytics solutions using Databricks, Apache Spark, and modern cloud technologies.
The role combines platform engineering, data engineering, AI/ML enablement, business intelligence, and operational support, working closely with engineering, data science, analytics, security, DevOps, architecture, and business stakeholders.
Key Responsibilities
Outcome 1: Secure, scalable, and governed Databricks platform
- Configure and manage Databricks workspaces, clusters, policies, runtimes, and separate Development, QA/UAT, and Production environments.
- Implement platform security and governance using Unity Catalog, IAM/RBAC, service principals, secrets, private connectivity, and enterprise access controls.
- Automate Databricks infrastructure and deployments using Terraform, source control, and CI/CD practices.
Outcome 2: Reliable and high-performing data products
- Build and maintain scalable ETL/ELT pipelines and Lakehouse solutions using Apache Spark, PySpark, Databricks workflows/jobs, and Delta Lake.
- Integrate Databricks with enterprise data sources, cloud storage, databases, APIs, and downstream applications.
- Optimize Spark workloads, queries, clusters, and resource usage for performance, reliability, scalability, and cost.
Outcome 3: Production-ready AI, machine learning, and generative AI solutions
- Prepare data and build ML workflows covering feature engineering, model training, evaluation, deployment, and lifecycle management using Databricks and MLflow.
- Develop generative AI solutions using LLMs, prompt engineering, embeddings, vector search, retrieval-augmented generation, and model serving.
- Evaluate and monitor AI solutions for accuracy, relevance, safety, bias, latency, cost, privacy, and governance.
Outcome 4: Trusted analytics and decision support
- Create analytical datasets, semantic models, KPIs, dashboards, and reporting layers using Databricks SQL and tools such as Power BI, Tableau, or Looker.
- Translate business requirements into accurate, performant, user-friendly dashboards and self-service analytics solutions.
- Ensure analytics solutions follow applicable data governance, security, quality, and accessibility standards.
Outcome 5: Operable and reusable platform capabilities
- Establish monitoring, logging, alerting, troubleshooting, and operational support for data, AI, generative AI, and analytics workloads.
- Collaborate across architecture, security, infrastructure, engineering, data science, analytics, and business teams to define practical Databricks standards and best practices.
- Maintain concise technical documentation for platform configuration, deployment, data pipelines, AI workflows, dashboards, and operational procedures.
Required Technical Skills
- At least 10 years of professional experience building analytics platforms, including at least 5 years of hands-on Databricks experience; a Databricks certification is required.
- Strong production experience with Databricks, Apache Spark/PySpark, Python, SQL, Delta Lake, Lakehouse architecture, workspaces, clusters, jobs/workflows, and scalable ETL/ELT pipelines.
- Experience with Azure, AWS, or Google Cloud, including cloud storage, Unity Catalog, IAM/RBAC, secrets management, networking, and secure connectivity.
- Experience with Terraform or similar Infrastructure-as-Code tools, CI/CD, and source-control practices.
- Working experience with MLflow and AI/ML/GenAI delivery, including model lifecycle, LLMs, embeddings, vector search, RAG, prompt engineering, and model serving.
- Experience with Databricks SQL and BI tools such as Power BI, Tableau, or Looker, including semantic models and KPIs.
- Strong troubleshooting, performance optimization, monitoring, and production-support skills.
Preferred Skills
- Experience designing enterprise-scale Databricks Lakehouse platforms and standardized multi-environment deployments.
- Experience with Databricks Asset Bundles, Apache Airflow, Azure Data Factory, or similar deployment and orchestration frameworks.
- Experience with streaming technologies such as Spark Structured Streaming, Kafka, or Event Hubs.
- Experience with advanced MLOps and GenAI tooling such as LangChain, LlamaIndex, Hugging Face, Azure OpenAI, Amazon Bedrock, fine-tuning, or model evaluation.
- Knowledge of data governance, metadata management, data-quality frameworks, BI governance, privacy controls, and responsible AI practices.
- Additional Databricks certifications or relevant cloud certifications, and experience in regulated or large enterprise environments, are a plus.
Behavioral Competencies
- Strong analytical, troubleshooting, and outcome-oriented problem-solving skills.
- Strong ownership of reliable, secure, scalable, and maintainable solutions.
- Ability to collaborate effectively across engineering, data science, analytics, architecture, DevOps, security, and business teams.
- Ability to translate business needs into practical data, analytics, and AI solutions.
- Clear verbal and written communication, including concise technical documentation.
- Strong attention to data governance, security, privacy, responsible AI, and operational discipline.
Education
Bachelor’s or master’s degree in computer science, Information Technology, Engineering, or a related discipline.