Senior Data Scientist

Providence Global CenterHyderabad, TelanganaHybridFull-timeSenior, 5–8 yearsListed 4 hours ago

Apply now

About this role

Job Description: Senior Data Scientist (Statistics, ML & Generative AI)

Location: Hyderabad, Telangana, India / Hybrid

Role Type: Full-Time

Position Overview

We are looking for a highly analytical and technical Senior Data Scientist with a strong mathematical foundation in statistics, probability, and classical machine learning, paired with cutting-edge experience in Large Language Models (LLMs) and Generative AI.

In this role, you will bridge the gap between rigorous statistical modeling and state-of-the-art AI technologies. You will design, develop, and deploy end-to-end data science solutions, RAG (Retrieval-Augmented Generation) architectures, and predictive models that turn complex data into strategic business value.

Key Responsibilities

- Core Statistical & ML Modeling:

Apply advanced probability theory, hypothesis testing, Bayesian methods, and statistical inference to validate data distributions, perform predictive analytics, and design robust experiments (A/B testing).

- Build, fine-tune, and evaluate classical machine learning algorithms (regression, classification, clustering, time-series forecasting, ensemble methods) for production systems.

- Generative AI & LLM Engineering:

Design and implement Generative AI architectures, including Retrieval-Augmented Generation (RAG) pipelines, multi-agent frameworks, and vector search systems.

- Fine-tune, evaluate, and prompt-engineer open-source and proprietary Large Language Models (LLMs) for specific domain applications.

- Implement guardrails, evaluation frameworks (e.g., RAGAS, DeepEval), and toxicity/bias detection for LLM applications.

- Data Engineering & System Architecture:

Collaborate with engineering teams to deploy ML/LLM models into production microservices via REST APIs, gRPC, or async worker queues.

- Optimize model inference latency, throughput, and GPU/memory consumption.

- Design data pipelines using Python, SQL, and modern data orchestration toolkits.

- Thought Leadership & Collaboration:

Partner with product and business stakeholders to translate complex business problems into scalable machine learning and AI frameworks.

- Stay at the forefront of emerging research in machine learning, deep learning, and generative AI.

Required Qualifications & Technical Skills

- Education: Bachelor's, Master's, or Ph.D. in Computer Science, Statistics, Applied Mathematics, Data Science, Quantitative Finance, or a related quantitative field.

- Core Mathematics & Statistics:

Strong expertise in probability distributions, stochastic processes, decision theory, regression analysis, and variance reduction techniques.

- Classical ML & Deep Learning:

Proficiency with ML frameworks (scikit-learn, XGBoost, LightGBM, PyTorch, or TensorFlow).

- Deep understanding of loss functions, optimization algorithms, evaluation metrics (ROC-AUC, RMSE, Precision/Recall, BLEU, ROUGE), and cross-validation techniques.

- Generative AI & LLM Ecosystem:

Practical experience with LLM frameworks like LangChain , LangGraph , LlamaIndex , or DSPy .

- Hands-on experience with Vector Databases ( ChromaDB , Pinecone , pgvector , Qdrant , or FAISS ).

- Knowledge of quantization techniques (LoRA, QLoRA), embedding models, and multi-agent coordination frameworks (e.g., Model Context Protocol / MCP servers).

- Programming & Tools:

High proficiency in Python (Pandas, NumPy, SciPy, PyTorch) and complex SQL .

- Familiarity with containerization ( Docker ) and modern development workflows (Git, CI/CD).

Preferred Qualifications

- Experience with time-series forecasting, quantitative modeling, or risk prediction algorithms.

- Knowledge of MLOps practices, model monitoring, and pipeline orchestration tools.

- Familiarity with cloud platforms (AWS, GCP, or Azure).