Data Scientist

Ecolab Inc.Bengaluru, KarnatakaOn-siteFull-timeJunior, 1–2 yearsListed 1 hour ago

Apply now

About this role

Data Scientist

Ecolab Digital is seeking a Data Scientist to help turn raw operational data into prioritized, plain-language insights and recommendations that reach customers in the Institutional & Specialty segment.

You will join a team that pulls together diverse real-world signals, such as audits, inspections, sensor telemetry, checklists, and third-party sources, into one coherent view, then works across classical machine learning, LLM-based applications, and emerging agentic workflows to decide what matters and say why. Partnering with product, engineering, and business stakeholders, you will take problems from data exploration and modeling through deployment and monitoring, contributing production-grade components that customers and internal teams rely on to make decisions. You will apply established data science practice while learning and deploying the newer GenAI craft on real, customer-facing work.

What you will do

- Build, test, and deploy analytical and machine learning solutions such as classification, forecasting, summarization, Q&A, recommendations, and decision support, using data from diverse internal and external sources.
- Train, tune, and evaluate machine learning models, and defend your choices with sound error analysis rather than a single accuracy number.
- Build GenAI applications on LLM APIs: structured prompts with output schemas, retrieval-augmented generation over vector indexes, and GenAI reasoning combined with deterministic logic for reliable, auditable outputs.
- Contribute to multi-step AI workflows for task and goals based agents.
- Prepare, structure, and validate datasets so data quality and model reliability are earned, not assumed.
- Deploy solutions into cloud environments using established engineering and CI/CD practices.
- Monitor deployed models and AI systems, and support performance, reliability, latency, and cost efficiency after launch.
- Build dashboards and visualizations that turn model output into decisions stakeholders can act on.
- Translate requirements from product, engineering, and business partners into scalable solutions, and communicate insights and impact clearly to technical and non-technical audiences.

Minimum Qualifications

- Bachelor's degree in Data Science, Computer Science, Math, Statistics, or a related field with 3 years of hands-on data science experience, or a Master's degree in a related field with 1 or more years of hands-on experience.
- 3+ years of Python. You write clean, modular code and treat version control, testing, and code review as standard practice.
- 3+ years of SQL for querying and preparing data.
- Hands-on experience with PySpark and DataFrame APIs on a large-scale distributed platform.
- 2+ years building machine learning models, covering training, tuning, and evaluation. Classical machine learning and statistics are central to this team, not replaced by LLMs. They are how we keep the LLMs honest, from anomaly detection to composite scoring.
- 1+ year building GenAI or LLM solutions, personal projects included. This means building RAG pipelines, agents, or applications. A strong side project or open-source contribution counts here as much as work experience. You should be able to show what you built and problem it solved.
- Experience taking models and solutions all the way to production. We are looking for someone who has gone past proofs-of-concept and personal experiments and shipped something real that other people depended on, then stayed with it through the iteration cycle of fixing, improving, and re-releasing. Tell us what you put into production, who used it, and what you changed after launch.
- A working sense of evaluation: can you prove a change actually made things better? When you adjust a prompt, a model, or a pipeline, you can measure whether the output genuinely improved rather than just confirming it still runs. That means comparing results against a trusted answer key, watching quality after launch, and keeping errors and hallucinations in check, especially in anything customer-facing.
- Working knowledge of Git, agile practices, and CI/CD workflows as a normal part of how you ship.
- Experience building dashboards and data visualizations to communicate results, Power BI preferred.
- Strong analytical thinking, problem-solving, and the ability to explain your approach, assumptions, and trade-offs to a range of stakeholders.
- Solid data science foundation that carries into GenAI work: EDA, statistical reasoning, metric and evaluation design, sampling, and error analysis.

Preferred Qualifications

- Experience on the Databricks platform: Spark SQL, Delta Lake, Unity Catalog, MLflow, Vector Search, and Model Serving.
- Hands-on experience with GenAI frameworks and tools: LangChain, Anthropic, OpenAI or Hugging Face APIs, and vector databases.
- Exposure to agentic or multi-step AI workflows such as tool-calling and sequential handoffs.
- Proficiency across the Microsoft Azure suite and comfort with cloud APIs across multiple environments.
- Experience in Retail or Quick Service Restaurant businesses.

Sharing a public GitHub profile or project portfolio is encouraged; we would love to see examples of your hands-on work where available.

Our Commitment to a Culture of Inclusion & Belonging
Ecolab is committed to fair and equal treatment of associates and applicants and furthering the principles of Equal Opportunity to Employment. We will recruit, hire, promote, transfer and provide opportunities for advancement based on individual qualifications and job performance in all matters affecting employment, compensation, benefits, working conditions, and opportunities for advancement. Ecolab will not discriminate against any associate or applicant for employment because of race, religion, color, creed, national origin,citizenship status, sex, sexual orientation, gender identity and expressions, genetic information, marital status, age, or disability.