About this role
Your impact
As a Staff Software Engineer within the Foundations organization, you will architect and deliver resilient, high-scale platform solutions that host Splunk's enterprise software across multi-cloud (AWS, GCP, Azure) and on-premises environments. You will balance hands-on distributed systems development with technical leadership-driving architectural decisions, establishing reliability standards, and mentoring engineers to power the next evolution of the Splunk Platform.
##
What you'll do
- Design, develop, and maintain Splunk Platform components across multiple cloud providers (AWS, Google Cloud Platform, Microsoft Azure) and on-premises infrastructure, focusing on scalable, secure, and maintainable software solutions.
- Oversee and participate in designing, implementing, testing, and deploying of distributed systems.
- Apply software engineering best practices in distributed systems programming, including debugging, performance tuning, and reliability engineering for complex systems like Kubernetes controllers, database platforms, and data replication.
- Define and improve service health indicators, observability, SLOs, RPO/RTO targets, alerting, runbooks, and end-to-end recovery testing.
- Lead incident diagnosis and learning, reducing operational risk and limiting customer impact through resilient design, automation, and recovery practices.
- Mentor and guide junior and mid-level engineers, promoting knowledge sharing and professional growth.
- Turn roadmap items into clear design, system boundaries, APIs, milestones, and engineering decisions.
- Lead design reviews and influence technical direction for the team’s roadmap and projects.
- Foster a culture of continuous learning and adaptability within an agile, fast-paced environment, encouraging innovation and exploration of new technologies.
Minimum Qualifications
- 8+ years of professional experience designing and operating large-scale distributed systems or SaaS platform infrastructure.
- Strong proficiency in Go and/or C++ for production systems development, with scripting experience in Python or Bash .
- Proven expertise building Kubernetes-native software (Custom Controllers, Operators) and deploying resilient applications across major cloud providers ( AWS, GCP, or Azure ).
- Deep understanding of concurrency, data replication, idempotency, partial failure modes, and eventual consistency.
- Proven track record of leading technical design reviews, author clear RFCs/design docs, and mentor engineering teams.
- Excellent verbal and written communication skills in English, with experience collaborating across cross-functional engineering teams.
Preferred Qualifications
- Hands-on experience scaling and running stateful workloads in Kubernetes (PostgreSQL, distributed NoSQL databases, caching layers, or object storage).
- Experience with Infrastructure as Code (Terraform), CI/CD pipelines (GitLab/Jenkins), and service mesh/networking infrastructure (load balancers, API gateways, DNS, TLS).
- Practical knowledge of SRE standards, distributed tracing, alerting frameworks, and chaos/recovery testing.
- Experience driving technical deliverables in fast-paced Agile (Scrum/Kanban) environments.
Why Cisco?
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.