Research Intern - ML

InfrrdBengaluru, KarnatakaOn-siteInternshipListed 55 minutes ago

Apply now

About this role

About Infrrd

Infrrd (pronounced In-fur-d ) is an Enterprise AI company that automates document-heavy workflows for customers in mortgage, insurance, and finance. Our Research team works on the next generation of document intelligence: agentic systems that read, reason over, and audit complex documents with outputs that can be trusted and verified. We are looking for a Research Intern to join the team and contribute to experiments that shape what we ship.

About the Role

As a Research Intern, you will work on well-scoped research tasks under the guidance of senior researchers, across areas such as agentic document extraction, LLM-based auditing of mortgage documents, table extraction with calibrated trust scores, and verifiable evaluation of model outputs without ground truth. You will run experiments end to end: preparing and checking data, building prototypes, analysing errors and reasoning traces, and writing up what you found. This is a hands-on role for someone who enjoys rigorous experimentation and wants exposure to real enterprise-scale document AI problems.

Education details:  10+ 2/PUC mandatory (No Diploma), B.E/B.Tech/M.Tech students from all  Computer Science related backgrounds with a focus on machine learning, NLP, or computer vision.

Year of Graduation:  2027

Percentage criteria:  minimum 60% aggregate and higher throughout academics.

Internship duration:  1 year with an opportunity to convert to a full-time role based on performance.

What You Will Do

- Experimentation and Prototyping: Design and run experiments to validate document processing and agentic extraction approaches; build prototypes and proof-of-concept implementations using LLM and vision-language model APIs.

- Evaluation and Verification: Help build evaluation harnesses and verifier checks (cross-field consistency, structural invariants, multi-pass agreement) that measure whether an extraction or audit verdict can be trusted, including when no ground truth is available.

- Error and Trace Analysis: Conduct in-depth error analysis on model outputs and agent reasoning traces to identify failure modes, categorise them, and propose fixes.

- Data Quality and EDA: Verify the quality of datasets and synthetic document packages used in experiments; perform exploratory analysis to understand document characteristics and edge cases.

- Rule and Checklist Work: Assist in converting domain checklist rules into executable, testable checks and in measuring their precision and recall on real documents.

- Literature Tracking: Read and summarise recent papers on document AI, agent harnesses, RL post-training, and evaluation; present findings in internal paper discussions.

- Tooling and Workflow: Use AI coding assistants (Claude Code, Copilot, or similar) and internal tools effectively; track progress in Jira; participate actively in stand-ups and code reviews.

- Documentation and Communication: Document methodology, experiment setup, and results clearly so they are reproducible; contribute to technical reports, Confluence pages, and internal presentations.

Who You Are

- Strong mathematical, statistical, and probabilistic foundation with a solid grasp of core ML concepts.

- Strong Python skills, including writing clean, testable code within a larger codebase.

- Working knowledge of Transformer-based language models and how to use LLM APIs (prompting, structured outputs, tool or function calling).

- Familiarity with evaluation methodology: designing metrics, building test sets, and analysing results with rigour rather than anecdotes.

- Academic or project experience in NLP, computer vision, or document understanding (OCR, layout, tables, forms).

- Ability to run experiments scientifically, keep track of what was tried, and communicate outcomes clearly.

- Curiosity about agentic systems and initiative in learning new techniques and applying them to real problems.

Good to Have

- Experience with vision-language models or document-specific models for extraction and layout understanding.

- Exposure to agent frameworks, multi-agent orchestration, or harness design for LLM-based systems.

- Familiarity with RL post-training methods (GRPO, RLVR) or model fine-tuning.

- Experience with table extraction, PDF parsing, or synthetic data generation.

- Contributions to open source, published work, or a portfolio of research projects.

Pay Stubs Data Extraction

NLP-powered Table Extraction for Insurance Policy Data

Ally | Agentic AI for Mortgage

Automated Data Extraction from Engineering & Construction Drawings