About this role
Work Flexibility: Hybrid
What You will do:
- Analyze incoming order documents and map out their formats, fields, and edge cases.
- Evaluate OCR and document-extraction tools (e.g. Tesseract, PaddleOCR, layout-aware models like LayoutLM/LiLT) and help choose the right approach.
- Build a pipeline that extracts structured order data (customer, part numbers, quantities, prices, dates, PO references) from documents.
- Implement validation against master data and a review step that flags uncertain results for a human to check.
- Integrate the output with existing data infrastructure for downstream reporting and order entry.
- Document your work and support handover to the team.
What you will need:
Required Qualification:
- Working knowledge of Python (data processing, scripting).
- Masters degree in Engineering passing out in 2025 or 2026
- Interest in (or some exposure to) machine learning, computer vision, or NLP.
- Basic SQL and comfort with structured data.
- An analytical mindset and strong attention to data quality
- Basic knowledge of UI-based automations (eg. UiPath) and/or API-based automations (eg. PowerAutomate)
Preferred Qualification:
- Experience with OCR libraries or document-AI models.
- Familiarity with cloud data platforms and pipeline tooling.
- Version control (Git) and collaborative development experience.
Travel Percentage: 0%