Data Scientist

UpshopToronto, OntarioOn-siteFull-timeJunior, 1–2 yearsListed 8 hours ago

Apply now

About this role

Position Overview

We are seeking an engineering-focused Data Scientist to build, operationalize, and maintain production-grade retail forecasting and optimization models. In this role, model development goes hand-in-hand with operational reliability: success is measured by accurate forecasts and stable, low-latency, deterministic pipelines running natively on the Databricks Lakehouse.

You will bridge the gap between applied data science and machine learning engineering. Working in close partnership with Product and Data Engineering , you will own the operational lifecycle of item-store level demand forecasts and downstream replenishment/optimization engines—from distributed feature pipelines to automated Databricks workflows, model tracking, and runtime monitoring.

Key Responsibilities

- Production Forecasting & Optimization: Develop, calibrate, and tune time-series forecasting engines and downstream supply chain/inventory optimization logic using tree-based ensembles ( LightGBM , CatBoost ) and distributed Python/PySpark.

- Databricks-Native MLOps: Build, schedule, and maintain automated training and batch inference pipelines using Databricks Workflows and Jobs . Leverage MLflow for robust experiment tracking, artifact logging, and model registry management.

- Cross-Functional Partnership:

With Product: Translate business requirements into technical specs, define operational SLAs, and provide technical feasibility assessments for new forecasting features.

- With Data Engineering: Establish strict data contracts, define schema validations, and optimize data ingestion/consumption patterns from upstream Delta tables.

- Pipeline Quality & Stability: Treat ML pipelines as critical production software. Implement pre-inference data validation gates (e.g., schema checks, missingness thresholds, null checks) and automated alerting to prevent corrupted data from reaching scoring jobs.

- Model & Pipeline Observability: Track pipeline health, monitor runtime performance, and detect feature drift, target drift, and forecast degradation across high-cardinality retail catalogs.

- Software Excellence: Write modular, maintainable, and testable code. Champion version control best practices, unit/integration testing with pytest , and automated CI/CD checks within Git.

- Spec-Driven Execution: Embrace a Spec-Driven Development (SDD) mindset, leveraging modern agentic AI development workflows (e.g., Cursor, Claude Code) to move rapidly from research to reliable production code.

Required Technical Skills

- Tree-Based Ensembles: Hands-on experience developing, tuning, and deploying gradient boosted decision trees—specifically LightGBM and CatBoost —on high-cardinality, tabular, and time-series datasets.

- Time-Series Retail Forecasting: Deep practical understanding of demand forecasting challenges: trend, seasonality, calendar events, promotional uplifts, stockouts, and intermittent/sparse demand patterns.

- Databricks Platform: Proven experience building within the Databricks ecosystem, specifically authoring and managing multi-task Databricks Workflows/Jobs , navigating Delta Lake, and using MLflow across the model lifecycle.

- Data Manipulation & PySpark: Strong proficiency in Python and PySpark for distributed data processing, feature engineering, and memory-conscious transformations across massive retail datasets.

- Software Engineering Fundamentals: Solid understanding of clean code principles, modular package design, virtual environments, automated testing ( pytest ), and standard Git workflows (pull requests, branching, code reviews).

Preferred Qualifications

- Optimization & Supply Chain: Familiarity with inventory optimization mechanics (safety stock calculation, reorder point modeling, lead time variability, allocation constraints).

- Databricks Advanced Features: Experience leveraging Delta Live Tables (DLT), Unity Catalog for data and model governance, or Photon compute engine.

- Explainable AI ( XAI ): Experience implementing TreeSHAP or similar interpretability methods within production batch scoring jobs.

- Continuous Integration: Experience setting up or integrating with CI/CD pipelines (e.g., GitHub Actions) to automate testing and deployment into Databricks workspaces.

Cultural Alignment & Values

- Production Mindset: You believe a model is only finished when it is tested, automated, monitored, and running reliably in production.

- Ownership & Root-Cause Thinking: When a pipeline fails or a metric degrades, you dig into the logs, identify the root cause, write a regression test, and implement a durable fix.

- Collaborative Communicator: You easily speak the language of business trade-offs with Product managers and system architecture with Data Engineers.

The estimated pay ranges for this role are as follows:

- $120,000 - 140,000 CAD

The successful candidate’s starting salary will be determined based on permissible, non-discriminatory factors such as skills, experience, and geographic location.