About this role
##
Meet the team
Join us as we pursue our disruptive new vision to make machine data accessible, usable and valuable to everyone. We are a company filled with people who are passionate about our product and seek to deliver the best experience for our customers. At Cisco, we are building a more resilient digital world through our security and observability portfolio. Splunk, a Cisco company, helps organizations turn data into action and strengthen the resilience of their digital systems. Learn more about Cisco careers and how you can become a part of our journey!
Splunk Cloud is looking for an operations and automation engineer to support Network Operations Center (NOC) responsibilities and manage large-scale Splunk Cloud deployments. This role balances hands-on operational support with the development of reliable, maintainable automation.
Your impact
You will monitor and troubleshoot distributed cloud systems, respond to production incidents, and turn manual operational runbooks into safe, repeatable workflows. You will work with operations and engineering teams to improve how we detect, diagnose, and resolve service issues while reducing manual toil.
What you'll do
- Monitor and support large-scale Splunk Cloud deployments, ensuring service reliability, availability, and performance.
- Respond to monitoring alerts and production incidents according to defined playbooks and procedures.
- Participate in 24/7 operations and on-call support
- Investigate complex issues across cloud infrastructure, Linux systems, networking, distributed services, and application code.
- Participate in post-incident reviews and turn findings into operational and automation improvements.
- Identify repetitive, manual, or error-prone operational tasks that are suitable for automation.
- Convert operational runbooks into reliable automated workflows.
- Use Git-based development, peer review, documentation, and cross-functional collaboration to maintain and improve operational automation.
Minimum qualifications
- B.S. in a related field or equivalent work experience (5+ years).
- 3+ years of experience in systems or network administration in a cloud environment, including incident response and major incident management
- Proven experience developing maintainable python/go automation, operational tooling, or production support utilities.
- Experience using Unix or Linux systems and shell scripting.
- Experience using Git and participating in code review workflows.
- Working knowledge of software engineering practices, including modular design, testing, debugging, exception handling, logging, and documentation.
- Ability to translate a manual operational procedure into a safe, repeatable automated workflow.
- Understanding of production automation risks, including permissions, secrets management, validation, rollback, and human approval points.
- Strong troubleshooting, prioritization, collaboration, and communication skills, including the ability to remain effective during major service outages.
Why Cisco?
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.