Senior Site Reliability Engineer
US$165,000 – US$241,500 / yr base salary
Posted 31 Aug 2026
Advertised as “Senior Site Reliability Engineer, Production Engineer - ThousandEyes”
BranchFactor Summary
The original posting with the fluff stripped out
- Expected experience
- 5+ yrs
- Management level
- Individual contributor
- Employment
- Hybrid
- Contract
- Full-time
| Base salary | US$165,000 – US$241,500 / yr |
|---|
Study levelCertificate or higher
The application window is expected to close on: 09/01/2026
Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received.
This role follows a hybrid work model, with in-office attendance expected once a week in San Francisco, Seattle, Austin or New York.
Meet the Team
Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.
ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.
Your Impact
We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.
Responsibilities
- Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
- Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
- Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
- Participate in and improve our 24x7 incident response and on-call rotation.
- Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
- Automate production operations to provide guardrails and continuous platform operation.
- Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
- Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
- Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
- Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.