About this role
About Impact Analytics Impact Analytics is an agentic-first AI software company transforming retail merchandising through cutting-edge AI, LLMs, and Generative AI technologies. As a fast-growing Series D company with deployments across five continents, it is building both industry-leading merchandising solutions and foundational AI agents that are redefining how retail decisions are made. What makes Impact Analytics unique is its combination of deep retail domain expertise, strong innovation culture, and global presence. It is one of the few India-born AI companies recognized globally by organizations like Fortune, Gartner, and the Inc. 5000. For candidates looking to work on next-generation AI products with global scale and real-world impact, Impact Analytics is an exciting place to build your career. Here’s a link to our website: www.impactanalytics.co. (http://impactanalytics.co)
The impact that you will be making
- The role is responsible for monitoring applications, investigating incidents, troubleshooting
- production issues, coordinating with engineering teams, and ensuring the availability and reliability of critical systems.
- This position requires working in rotational shifts, including night shifts, weekends, and holidays, to support a 24×7 production environment.
What this role entails Production & Application Support
- Monitor backend applications, services, APIs, databases, queues, and other critical production systems.
- Provide Level 2 and L3 production support and respond to incidents within defined SLAs.
- Troubleshoot application failures, performance issues, API errors, data issues, and service disruptions.
- Perform initial investigation, log analysis, and issue triaging before escalating to the appropriate engineering teams.
- Identify recurring production issues and work with development teams on permanent fixes.
Incident Management
- Respond promptly to production alerts and incidents during assigned shifts.
- Own incidents from identification through investigation, mitigation, resolution, and closure.
- Coordinate with Development, DevOps/SRE, QA, Database, and Infrastructure teams during critical incidents.
- Escalate issues based on defined severity levels and escalation procedures.
- Maintain clear and timely communication with relevant stakeholders during major incidents.
- Participate in Root Cause Analysis (RCA) and post-incident reviews.
Monitoring & Observability
- Monitor system health, application performance, availability, and error rates using observability and monitoring tools.
- Analyze application logs, metrics, traces, and alerts to identify issues and potential risks.
- Help improve alerting mechanisms to reduce false positives and improve incident response.
- Monitor scheduled jobs, batch processes, data pipelines, and integrations where applicable.
Operational Excellence
- Create and maintain runbooks, SOPs, troubleshooting guides, and knowledge-base documentation.
- Perform regular health checks and operational activities as per defined procedures.
- Identify opportunities for automation to reduce manual support effort and repetitive operational tasks.
- Participate in deployment validation and production readiness activities.
- Support release activities, including monitoring applications after production deployments.
- Ensure proper shift handovers, including open incidents, ongoing investigations, risks, and pending activities.
Continuous Improvement
- Analyze recurring incidents and operational trends.
- Work with engineering teams to improve application stability, reliability, performance, and supportability.
- Recommend improvements to monitoring, logging, alerting, automation, and operational processes.
- Contribute to reducing MTTR (Mean Time to Resolution) and improving service Availability.
What lands you in this Role
- Strong understanding of backend applications and distributed systems.
- Experience with one or more backend programming languages such as: Go (Golang), Java, Python, C++
- Good understanding of REST APIs and microservices architecture.
- Experience troubleshooting application and production issues using logs and monitoring tools.
- Knowledge of Linux/Unix commands and shell scripting.
- Understanding of databases such as PostgreSQL, MySQL, MongoDB, or similar technologies.
- Basic understanding of caching technologies such as Redis.
- Familiarity with messaging or streaming platforms such as Kafka, RabbitMQ, or similar technologies.
- Understanding of HTTP, networking fundamentals, and API troubleshooting.
- Experience with monitoring and observability tools such as Datadog, Grafana, Prometheus, ELK, Splunk, or similar tools.
- Familiarity with containerized environments and technologies such as Docker and Kubernetes is preferred.
- Understanding of CI/CD pipelines and release/deployment processes is an advantage.
Qualifications & Experience
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 4–7 years of experience in Backend Engineering, Application Support,
- Production Support, Site Reliability, or a similar role.
- Experience supporting business-critical or customer-facing production applications are preferred.
- Experience working with distributed teams and cross-functional stakeholders is desirable.
What we offer
- An opportunity to be part of some of the best enterprise SaaS products to be built out of India.
- Opportunities to quench your thirst for problem-solving, experimenting, learning, and implementing innovative solutions.
- A flat, collegial work environment, with a work hard, play hard attitude.
- A platform for rapid growth if you are willing to try new things without fear of failure.
- Remuneration with best-in-class industry standards with generous health insurance cover
Some of our accolades include
- Ranked as one of America's Fastest-Growing Companies by Financial Times for five consecutive years: 2020-2024.
- Ranked as one of America's Fastest-Growing Private Companies by Inc. 5000 for seven consecutive years: 2018-2024.
- Voted #1 by more than 300 retailers worldwide in the RIS Software LeaderBoard 2024 report.
- Ranked #72 in America’s Most Innovative Companies list in 2023 —by Fortune —alongside companies like Microsoft, Tesla, Apple, IBM, etc.
- Forged a strategic partnership with Google to equip retailers with cutting-edge generative AI tools.
- Recognized in multiple Gartner reports , including Market Guides and Hype Cycle , spanning assortments, merchandising, forecasting, algorithmic retailing, and Unified Price, Promotion, and Markdown Optimization Applications.
