About this role
Devsinc is hiring a Senior AI/ML Engineer with 4+ years of experience , including development of LLM or generative AI applications . The ideal candidate brings strong machine learning fundamentals and expertise in Python backend services and APIs, RAG, semantic search, AI agents, and scalable ML infrastructure , using commercial and open-source models.
You will own the end-to-end AI lifecycle , from experimentation and evaluation to deployment, optimization, and monitoring, ensuring reliable, scalable, secure, and cost-efficient solutions. The role also involves architectural decisions, mentoring engineers, and collaborating with clients and cross-functional teams to deliver measurable business impact.
Responsibilities
- Design, develop, and deploy AI/ML and LLM-based applications , including AI agents, tool-using systems, and human-in-the-loop workflows , to solve business problems.
- Build scalable training, fine-tuning, evaluation, and inference pipelines , with experiment tracking and model versioning.
- Develop and optimize RAG and semantic-search systems using embeddings, document chunking, vector search, reranking, and grounding.
- Build backend APIs, microservices, and real-time inference services using Python, FastAPI, Flask, or Django .
- Improve model quality, latency, throughput, and cost through experimentation, hyperparameter tuning, quantization, batching, and caching .
- Implement MLOps, automated testing, CI/CD, and production monitoring , supported by evaluation datasets, quality criteria, regression tests, and human review.
- Guide architectural decisions and cloud deployment, ensuring scalability, reliability, security, and resource efficiency , with safeguards against prompt injection, data leakage, unauthorized access, and unsafe outputs.
- Evaluate emerging AI technologies and measure feature effectiveness through product analytics or A/B testing .
- Mentor engineers, collaborate with technical and non-technical stakeholders, and document designs, experiments, and outcomes.
Requirements
- Bachelor’s degree in Computer Science , Software Engineering , Data Science , or a related field.
- 4+ years of post-graduation professional experience in AI/ML engineering , with demonstrated ownership of production AI systems and hands-on experience developing LLM or generative AI applications .
- Strong production-level Python skills, with hands-on experience in PyTorch and/or TensorFlow and solid knowledge of machine learning, neural networks, NLP, feature engineering, and model optimization.
- Experience integrating commercial or open-source LLMs , including prompt design, structured outputs, tool calling, context management, and model limitations.
- Hands-on experience building RAG or semantic-search systems , including embeddings, chunking, retrieval, reranking, and grounding, using vector-search solutions such as pgvector, Pinecone, Weaviate, Qdrant, Milvus, or Elasticsearch .
- Experience developing and deploying APIs, microservices, or inference services using FastAPI, Flask, Django , or equivalent frameworks, with proficiency in SQL and PostgreSQL or MySQL .
- Experience deploying AI solutions on AWS, Azure, or Google Cloud , with working knowledge of Git, Docker, automated testing, CI/CD, and MLOps , including experiment tracking, model versioning, and monitoring.
- Understanding of AI evaluation, regression testing, human review, and security and privacy risks .
- Ability to own technical decisions, guide engineers, and communicate effectively with clients and cross-functional stakeholders.
Preferred Skills & Experience
- Experience with AI frameworks such as LangChain, LlamaIndex, or LangGraph , and evaluation tools such as LangSmith, Langfuse, or Ragas , is a plus.
- Experience with Hugging Face Transformers, vLLM, Ollama, LoRA/PEFT fine-tuning, or self-hosted models is preferred.
- Familiarity with MLflow, Kubeflow, Kubernetes, Terraform, distributed systems, or GPU acceleration is a plus.
- Experience with data orchestration, asynchronous processing, caching, or messaging , using tools such as Airflow, Redis, Celery, Kafka, or RabbitMQ , is preferred.
- Knowledge of advanced retrieval, knowledge graphs, recommendation systems, computer vision, multimodal AI, or A/B testing and product analytics is a plus.