Data Scientist Associate

JPMorgan Chase & Co.Bengaluru, KarnatakaOn-siteFull-timeNew grad, 0–1 yearsListed 1 day ago

Apply now

About this role

Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products.

As a Data Scientist Associate at JPMorganChase within the Asset and Wealth Management, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. Drive significant business impact through your capabilities and contributions, and apply deep technical expertise and problem-solving methodologies to tackle a diverse array of challenges that span multiple technologies and applications.

Build and productionize RAG and Agentic RAG applications for financial-services use cases (intelligent search, Q&A, summarization, and workflow assistants). This role blends core software engineering with applied data science skills—data cleaning, analytics, experimentation, and evaluation—to improve retrieval quality and model reliability.

Job Responsibilities

- Build end-to-end RAG applications: document ingestion → parsing → chunking → embeddings → indexing → retrieval → grounded generation (with citations/attribution where applicable).
- Implement Agentic RAG patterns (query planning, multi-hop retrieval, tool-based lookups, reranking, guardrails, and fallback behaviors) for complex user questions.
- Develop LLM-based NLP capabilities for classification, extraction, summarization, semantic search, and conversational flows tailored to financial domain needs.
- Perform data preparation and quality work: cleaning noisy text, de-duplication, normalization, metadata enrichment, labeling, and maintaining curated datasets for evaluation/training.
- Run applied data science experiments to improve relevance and answer quality: A/B tests, prompt/retrieval experiments, embedding model comparisons, chunking strategy tests, and reranker evaluations.
- Define and track quality metrics across retrieval and generation (e.g., recall@k, MRR, precision, groundedness, citation coverage, user satisfaction proxies) and create lightweight dashboards/regular reporting.
- Build basic analytics pipelines around usage and quality signals (feedback, clicks, escalation rates, latency/cost) to guide iteration.
- Implement testing and evaluation harnesses: golden question sets, automated regression tests, adversarial prompts, and safety checks to reduce hallucinations.
- Collaborate with product/design/stakeholders to translate requirements into shipped features and iterate quickly based on feedback.
- Ensure solutions follow security, privacy, and responsible AI requirements (safe handling of sensitive data, access control-aware retrieval, logging/audit needs).

Required qualifications, capabilities and skills

- 3+ years experience in software engineering, applied ML, data science engineering, or a related role building production systems.
- Strong programming in Python , with APIs and services.
- Working knowledge of applied data science fundamentals: data cleaning, exploratory data analysis (EDA), basic statistics, evaluation design, and communicating results.
- Experience with RAG development using frameworks such as LangChain/LlamaIndex (or equivalent),
- Comfortable with SQL and data tooling (e.g., pandas / Spark basics) to prepare datasets and run analyses.
- Experience with cloud (AWS or Azure) and standard SDLC practices (version control, CI/CD basics, testing).

Preferred qualifications, capabilities and skills

- Exposure to vector databases/search (e.g., OpenSearch/Elastic, Pinecone, Weaviate, FAISS) and reranking approaches.
- Experience with evaluation frameworks (offline relevance labeling, LLM-as-judge with guardrails, regression suites) and basic experiment design.
- Familiarity with agent frameworks (LangGraph/Semantic Kernel/etc.) and Agentic RAG workflows.
- Experience with Python.
- Familiarity with embeddings and retrieval concepts.