Skip to content
All jobs

Site Reliability Engineer (SRE)

The Brixton Group

Job
33052
Posted
Location
Benton Harbor, MI
Work type
Full Time
Tax terms
W2, Yearly
Experience
Experience open
Openings
1 opening

Skills

  • Reliability Engineering
  • Akamai
  • Performance Engineering
  • Capacity Management
  • Stress Testing
  • Performance Analysis
  • Backup
  • Recovery
  • Disaster Recovery
  • Palo Alto
  • Firewall
  • Amazon DynamoDB
  • API
  • SaaS
  • Confluence
  • Production Engineering
  • DevOps
  • Amazon Web Services
  • Virtual Private Cloud
  • Amazon CloudFront

About the job

Duration: 12+ Months

Location: 100% Remote

We are seeking a Senior Site Reliability Engineer (SRE) with 10+ years of experience to own the reliability, availability, observability, and production operations of a complex, multi-region global e-commerce platform. The environment is primarily AWS-based, with extensive use of ECS/Docker, Lambda, API Gateway, DynamoDB, CloudFront/Akamai, and SaaS platforms such as VTEX. This is a highly hands-on role that requires deep expertise in production incident management, observability, performance engineering, AWS troubleshooting, and resilience.

Key Responsibilities:

  • Lead P1/P2 production incidents, drive recovery, and conduct RCA and blameless postmortems.
  • Own observability using New Relic, AWS CloudWatch, and Amazon Athena.
  • Define and monitor SLIs, SLOs, SLAs, reliability KPIs, and alerting strategies.
  • Perform capacity planning, load testing, stress testing, and performance analysis using tools such as k6 and JMeter.
  • Validate backup/restore, disaster recovery, resilience, patching, runtime upgrades, and SSL/TLS certificates.
  • Troubleshoot AWS networking, ALB traffic, CloudFront/CDN, DNS, WAF, DDoS/Bot attacks, and Palo Alto firewall interactions.
  • Troubleshoot production issues involving ECS, Lambda, S3, DynamoDB, API Gateway, and VTEX/SaaS integrations.
  • Read and assess Terraform/IaC, Docker containers, and production infrastructure for reliability risks.
  • Maintain operational runbooks and technical documentation in Confluence.

Required Skills:

  • 10+ years in SRE, Production Engineering, DevOps, or similar roles.
  • Strong hands-on experience with AWS: VPC, ALB, CloudFront, ECS, S3, Lambda, WAF, Secrets Manager.
  • Expert-level New Relic, CloudWatch, and Athena experience.
  • Strong incident management, RCA, postmortems, SLI/SLO/SLA, and reliability engineering experience.
  • Hands-on k6/JMeter performance and load testing.
  • Strong knowledge of Terraform, Docker, Git, Linux/Unix, DNS, TLS/SSL, and shell scripting.
  • Experience troubleshooting Node.js, PHP, Python, and Bash-based applications.
  • Experience with CDN/edge technologies, traffic analysis, security, and production troubleshooting.
  • Excellent communication and ability to operate independently in a mission-critical environment.

26-01122

#dice

Similar jobs

See all jobs