Data Scientst - Student Position - Sales Business Analytics

AppleToronto, OntarioOn-siteFull-timeSenior, 5–8 yearsListed 19 hours ago

Apply now

About this role

The people here at Apple don’t just create products—they create the kind of wonder that’s revolutionized entire industries. It’s the diversity of those people and their ideas that inspires the innovation that runs through everything we do, from amazing technology to industry-leading environmental efforts. Join Apple, and help us leave the world better than we found it!

As a Data Scientist - Student Position, Sales Business Analytics in Apple's Sales organization, you'll play a key role in supporting our mission. Collaborating with Sales Professionals, Finance, Operations and Senior Leadership, you will deliver operational excellence by providing data insights, reporting, analytics, and tools to advance our strategic vision. We are searching for a student who can be flexible in the face of business ambiguity and eager to analyze/break down sophisticated datasets to derive clear decisions for our various business partners.

Please Note: This is a Limited Term Employment position (8-Month Co-op) from January to August 2027

This program offers mentorship and development to elevate your business acumen and give visibility to a wide spectrum of business projects at Apple. You will be embedded within a Sales Data and Analytics team and will own deliverables end to end. Depending on team fit, your work will center on a subset of the following:

1. Data Pipeline Modernization and Observability
- Upgrade and modernize mature production ETL/ELT pipelines that span multiple platforms to ensure stability and reliability.
- Build and maintain automated batch pipelines that ingest, clean, and transform multi-source data into PostgreSQL and Snowflake environments.
- Stand up pipeline monitoring and observability: freshness checks, volume and schema drift detection, data quality assertions.
- Optimize database schemas, warehouse models, and query performance for analytical and AI workloads.

2. Workflow Orchestration and Tooling Migration
- Migrate scheduled jobs and legacy workflows into Apache Airflow and Apple's internal orchestration platforms, with a focus on avoiding silent failures during and after cutover.
- Document existing job dependencies and build validation and parallel-run strategies to confirm parity before decommissioning.
- Contribute to a catalog of modern Canada tooling and automations.

3. AI Agent Development and Evaluation
- Architect and build agentic workflows, Retrieval-Augmented Generation (RAG) systems, and custom LLM applications against internal knowledge sources.
- Support the Canada AI Knowledge Hub and Skills Marketplace: knowledge base design, content structuring, retrieval quality, and the user-facing experience.
- Establish rigorous evaluation frameworks for AI systems, including ground-truth sets, benchmark suites, and methods for detecting confidently incorrect answers and hallucinations.
- Build AI agents for cross-source data reconciliation and monitoring, resolving discrepancies where multiple systems of record disagree.
- Implement structured logging, caching, and fallback mechanisms so AI features behave predictably in production.

4. Analytics, Modeling, and Tiering
- Conduct exploratory data analysis, feature engineering, and statistical modeling to uncover actionable insights.
- Develop and refine multi-factor scoring and tiering models (Business Tiering, POS Tiering, Carrier Analytics, and related frameworks) that combine several weighted inputs into a single classification.
- Build robust validation approaches and be prepared to explain and defend individual model outputs to business stakeholders as needed.
- Deliver reporting and visualization that makes model results usable, including Tableau dashboards and prepared datasets.

5. Data Governance and Enablement
- Contribute to data catalog and metadata efforts, including data lineage tracking and documentation of data requirements across sources.
- Curate and prepare datasets for business teams adopting AI tools, translating qualitative stakeholder needs into concrete metrics and data products.
- Support AI enablement activities such as office hours, documentation, and internal guidance.

6. Collaboration and Cross-Functional Impact
- Partner with business partners, peers, and IS&T to translate ambiguous problems into concrete data and AI deliverables.
- Write clean, well-documented, tested code; participate in code reviews; and present project milestones to both technical and executive stakeholders.

Minimum Qualifications

Enrolled in a Bachelor's or Master's program in Computer Science, Data Science, Statistics, Mathematics, Engineering, or related quantitative field, returning to studies after position’s term.
Strong proficiency in Python and standard data libraries (pandas, NumPy).
Strong command of SQL and relational database concepts, including experience with PostgreSQL & Snowflake or a comparable relational database.
Hands-on experience building end-to-end data pipelines or applications through coursework, personal projects, hackathons, or prior internships.
Experience with Git and collaborative version control workflows.
Foundational understanding of algorithms, data structures, and software engineering design.
Demonstrated AI literacy: familiarity with how LLMs work, their failure modes, and where they are and are not appropriate.
Strong communication and writing skills, critical to working across multiple teams and functions.
Flexibility to juggle multiple responsibilities independently, and the judgment to ask questions early when something is unclear.

Preferred Qualifications

Candidates are expected to have familiarity and/or experience in these areas, although deep expertise in all is not required.

Data Engineering and Orchestration
Apache Airflow (or Prefect, Dagster) for workflow orchestration
PostgreSQL, Snowflake, dbt
Pipeline testing, data quality frameworks (Great Expectations or similar), and monitoring or alerting tooling
Data lineage, metadata management, or data catalog tools

AI and LLM Application Development
LLM and agent frameworks: LangChain, LlamaIndex, or native tool-use and agent frameworks
RAG architecture, prompt engineering, and retrieval evaluation
Agent evaluation and benchmarking, including LLM-as-judge methods and hallucination detection
Vector search: pgvector, Chroma, Qdrant, or similar
FastAPI, Docker, and asynchronous Python

Data Science and Analytics
Statistical analysis, hypothesis testing, and feature engineering
scikit-learn, XGBoost, or LightGBM for classification and scoring problems
Composite scoring, weighting, and segmentation methodology
Tableau and Tableau Prep, or comparable BI tooling

Software Practices
CI/CD pipelines, containerization, and unit or integration testing
Clear technical documentation