Skip to content
All jobs

SRE Lead -- W2 Profiles -- Remote

Trebecon LLC

Job
29469
Posted
Location
Remote
Work type
Contract
Tax terms
W2, C2C, 1099
Experience
Experience open
Openings
1 opening

Opens your email app with a message to the employer, its subject naming this job. Attach your resume and send it from your own email.

Skills

  • Python
  • AWS
  • Terraform
  • Ansible
  • Kubernetes
  • Docker
  • GitLab
  • CI/CD
  • Prometheus
  • Grafana
  • Linux
  • Bash

About the job

Role: DevOps / SRE Cloud Lead

Location: 100% Remote

Key Responsibilities

First-Line Hands-On Troubleshooting

• Act as the primary technical escalation point and hands-on operational responder for infrastructure, network, database, and cloud workload issues across Google Cloud Platform (primary) / AWS.

• Troubleshoot GitOps deployment failures, synchronization issues, and application rollouts using Argo CD and GitLab CI/CD.

• Perform real-time diagnostic analysis using observability tools, cloud-native logs, and metrics to identify and resolve root causes.

• Execute standard operational procedures, automate repetitive maintenance tasks, and apply immediate hotfixes or operational patches to restore service availability.

• This is a heavy troubleshooting and project-based role requiring active participation in an on-call rotation for production support.

Incident Management

• Lead major incident management (MIM) processes, acting as the primary incident commander during outages, failed deployments, or high-severity events.

• Coordinate cross-functional technical teams to drive rapid resolution and minimize service impact (MTTR).

• Manage incident communication across stakeholders, leadership, and external partners.

• Own on-call rotations, incident escalation paths, and bridge management.

Problem Management

• Facilitate Post-Incident Reviews (PIRs) and root cause analysis (RCA) to uncover underlying systemic weaknesses.

• Track, prioritize, and manage problem records to prevent recurring pipeline, deployment, and infrastructure issues.

• Partner with SRE, DevOps, and Development teams to prioritize stability fixes, architectural remediations, and technical debt elimination.

Service Management & Operations Leadership

• Maintain operational SLAs, SLOs, and OLAs; continuously monitor and report on system uptime and operational metrics.

• Manage ITSM practices including Change Management, Event Management, Release Governance, and Capacity Planning.

• Lead, mentor, and build a team of cloud support engineers, fostering a continuous-improvement mindset around automated CI/CD and GitOps practices.

• Manage vendor relationships and cloud provider support agreements (Google Cloud Platform/AWS support tickets).

• Maintain strong security awareness across infrastructure and deployment practices, flagging and addressing risks proactively.

Qualifications & Skills

Must Have

• Kubernetes: Strong hands-on experience (deployments, troubleshooting, scaling, managed services like GKE/EKS).

• Argo CD: Expert-level, hands-on experience (application syncing, rollbacks, status monitoring, GitOps workflows).

• Database Knowledge: Solid working knowledge of databases (administration, performance troubleshooting, query-level debugging).

• Google Cloud Platform: Strong hands-on cloud experience (compute, networking, IAM, monitoring).

• Security Awareness: Working knowledge of cloud security best practices, IAM policies, and vulnerability awareness.

Technical Knowledge

• GitOps & Deployment Tools: Deep hands-on experience with Argo CD and GitLab (CI/CD pipelines, runner management, repository management).

• OS & Networking: Deep knowledge of Linux administration, networking protocols (TCP/IP, DNS, VPNs, Firewalls), and cloud networking (VPCs, route tables, load balancers).

• Monitoring & Observability: Hands-on experience with tools such as Datadog, CloudWatch, Google Cloud Monitoring, Prometheus, Grafana, Splunk, or New Relic.

• Automation & IaC: Proficiency with scripting languages (Python, Bash) and Infrastructure as Code (Terraform, Ansible).

• Containerization: Strong practical experience with Docker, Kubernetes, and managed container services (GKE, EKS).

Process & Management Experience

• ITSM Frameworks: Strong practical knowledge of ITIL guidelines (ITIL v4 certification is a plus).

• Leadership: 3+ years managing, leading, or mentoring cloud operations, support, or SRE teams.

• Troubleshooting: Proven ability to troubleshoot complex, distributed multi-tier architectures and automated CI/CD build/release failures under pressure.

• Communication: Excellent written and verbal communication skills for stakeholder updates and technical documentation.

• Availability: Willingness and ability to participate in on-call rotations, including off-hours incident response.

Preferred Requirements

• Bachelor's degree in Computer Science, Information Technology, or equivalent experience.

• Professional Cloud Certification (Google Cloud Associate Cloud Engineer / Professional Cloud Architect, or AWS equivalent).

• Kubernetes or GitOps certifications (CKA, CKAD, GitOps Certified Associate).

• Experience with incident management tools such as PagerDuty, Opsgenie, ServiceNow, or Jira Service Management.

Nice to Have

• FinOps skills - familiarity with cloud cost governance and reporting.

• Demonstrated cost-saving experience - track record of identifying and executing cloud cost optimization initiatives.

Similar jobs

See all jobs