About this role
Job Description: Senior Data Scientist (Statistics, ML & Generative AI)
Location: Hyderabad, Telangana, India / Hybrid
Role Type: Full-Time
Position Overview
We are looking for a highly analytical and technical Senior Data Scientist with a strong mathematical foundation in statistics, probability, and classical machine learning, paired with cutting-edge experience in Large Language Models (LLMs) and Generative AI.
In this role, you will bridge the gap between rigorous statistical modeling and state-of-the-art AI technologies. You will design, develop, and deploy end-to-end data science solutions, RAG (Retrieval-Augmented Generation) architectures, and predictive models that turn complex data into strategic business value.
Key Responsibilities
- Core Statistical & ML Modeling:
Apply advanced probability theory, hypothesis testing, Bayesian methods, and statistical inference to validate data distributions, perform predictive analytics, and design robust experiments (A/B testing).
- Build, fine-tune, and evaluate classical machine learning algorithms (regression, classification, clustering, time-series forecasting, ensemble methods) for production systems.
- Generative AI & LLM Engineering:
Design and implement Generative AI architectures, including Retrieval-Augmented Generation (RAG) pipelines, multi-agent frameworks, and vector search systems.
- Fine-tune, evaluate, and prompt-engineer open-source and proprietary Large Language Models (LLMs) for specific domain applications.
- Implement guardrails, evaluation frameworks (e.g., RAGAS, DeepEval), and toxicity/bias detection for LLM applications.
- Data Engineering & System Architecture:
Collaborate with engineering teams to deploy ML/LLM models into production microservices via REST APIs, gRPC, or async worker queues.
- Optimize model inference latency, throughput, and GPU/memory consumption.
- Design data pipelines using Python, SQL, and modern data orchestration toolkits.
- Thought Leadership & Collaboration:
Partner with product and business stakeholders to translate complex business problems into scalable machine learning and AI frameworks.
- Stay at the forefront of emerging research in machine learning, deep learning, and generative AI.
Required Qualifications & Technical Skills
- Education: Bachelor's, Master's, or Ph.D. in Computer Science, Statistics, Applied Mathematics, Data Science, Quantitative Finance, or a related quantitative field.
- Core Mathematics & Statistics:
Strong expertise in probability distributions, stochastic processes, decision theory, regression analysis, and variance reduction techniques.
- Classical ML & Deep Learning:
Proficiency with ML frameworks (scikit-learn, XGBoost, LightGBM, PyTorch, or TensorFlow).
- Deep understanding of loss functions, optimization algorithms, evaluation metrics (ROC-AUC, RMSE, Precision/Recall, BLEU, ROUGE), and cross-validation techniques.
- Generative AI & LLM Ecosystem:
Practical experience with LLM frameworks like LangChain , LangGraph , LlamaIndex , or DSPy .
- Hands-on experience with Vector Databases ( ChromaDB , Pinecone , pgvector , Qdrant , or FAISS ).
- Knowledge of quantization techniques (LoRA, QLoRA), embedding models, and multi-agent coordination frameworks (e.g., Model Context Protocol / MCP servers).
- Programming & Tools:
High proficiency in Python (Pandas, NumPy, SciPy, PyTorch) and complex SQL .
- Familiarity with containerization ( Docker ) and modern development workflows (Git, CI/CD).
Preferred Qualifications
- Experience with time-series forecasting, quantitative modeling, or risk prediction algorithms.
- Knowledge of MLOps practices, model monitoring, and pipeline orchestration tools.
- Familiarity with cloud platforms (AWS, GCP, or Azure).