Site Reliability Engineer (SRE)
Remote
Job Id:
173592
Job Category:
Job Location:
Remote
Security Clearance:
No Clearance
Business Unit:
Piper Companies
Division:
Piper Enterprise Solutions
Position Owner:
Austin Richardson
Piper Companies is seeking a Site Reliability Engineer (SRE) – Kubernetes Platform to join a leading organization in the cloud networking and infrastructure industry. The Site Reliability Engineer (SRE) – Kubernetes Platform will be responsible for supporting and maintaining Kubernetes infrastructure across AWS and on-premises environments while driving reliability, automation, and operational excellence within highly available production systems.
Responsibilities of the Site Reliability Engineer (SRE) – Kubernetes Platform:
• Manage, build, and maintain Kubernetes clusters across AWS and on-premises environments.
• Monitor, troubleshoot, and resolve issues impacting Kubernetes infrastructure and production services.
• Administer and support Linux-based systems, including system performance tuning and root cause analysis.
• Implement and maintain Infrastructure as Code (IaC) solutions using Terraform, CloudFormation, or similar tools.
• Develop automation to improve operational efficiency, scalability, and platform reliability.
• Support Kubernetes cluster upgrades, patching, deployments, and ongoing maintenance activities.
• Partner with development and infrastructure teams to support application deployments and platform initiatives.
• Utilize monitoring and observability tools to proactively identify and resolve performance issues.
• Participate in on-call rotations and respond to production incidents as needed.
Qualifications of the Site Reliability Engineer (SRE) – Kubernetes Platform:
• 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or a related role.
• Strong hands-on experience managing Kubernetes clusters in production environments.
• Experience supporting Kubernetes infrastructure across both AWS and on-premises environments.
• Strong Linux administration, support, and troubleshooting experience.
• Hands-on experience with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or similar technologies.
• Experience with object-oriented programming or scripting languages such as Ruby, Python, Go, or similar.
• Knowledge of containerization, networking, and cloud infrastructure best practices.
• Experience with monitoring and observability tools including Prometheus, Grafana, or similar platforms.
• Exposure to CI/CD pipelines and deployment automation tools.
• Experience supporting regulated environments such as FedRAMP High, DoD IL5, or similar environments is preferred.
• Familiarity with ArgoCD and container security practices is a plus.
Compensation for the Site Reliability Engineer (SRE) – Kubernetes Platform includes:
• Salary range: $120,000 - $180,000 depending on experience
• Comprehensive benefits package including medical, dental, vision, 401(k), and PTO
• Fully remote work environment
This job opens for applications on 09/02/2026. Applications for this job will be accepted for at least 30 days from the posting date.
#LI-AR2
#REMOTE