Site Reliability Engineer (SRE)
ASCII Group LLC
- Job
- 29300
- Posted
- Location
- Austin, TX
- Work type
- Contract
- Tax terms
- W2, C2C, 1099
- Experience
- Experience open
- Openings
- 1 opening
Opens your email app with a message to the employer, its subject naming this job. Attach your resume and send it from your own email.
Skills
- Python
- AWS
- Terraform
- Kubernetes
- CI/CD
- Linux
- Incident Response
About the job
Hi ,
The following requirement is open with our client.
Title : Site Reliability Engineer (SRE)
Location : Austin, TX (Onsite)
Duration : 12 Months
Relevant Experience : 8+
C2C or W2
Job Responsibilities:
- · We are looking for a skilled Site Reliability Engineer (SRE) to design, build, and maintain highly available, scalable, secure, and reliable cloud infrastructure and applications.
- · Must have Apple Exp.
- · Should be able to do coding independently in Python and GoLang, The ideal candidate will have strong hands-on experience with AWS, Kubernetes, Python, Linux, and cloud-native technologies.
- · You will work closely with Development, DevOps, Security, and Operations teams to improve system reliability, automation, observability, and operational efficiency.
- · Key Responsibilities
- · Design, implement, and maintain highly available and scalable infrastructure on AWS.
- · Deploy, manage, and troubleshoot containerized applications using Kubernetes.
- · Develop automation tools, scripts, and operational utilities using Python.
- · Build and maintain CI/CD pipelines for reliable and repeatable application deployments.
- · Monitor system health, availability, performance, and capacity. Implement observability using metrics, logs, traces, dashboards, and alerting.
- · Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews. Define and improve SLIs, SLOs, and SLAs. Automate repetitive operational tasks and reduce manual intervention.
- · Perform Kubernetes troubleshooting, including pods, deployments, services, ingress, networking, storage, and resource management. Optimize AWS infrastructure for performance, reliability, security, and cost.
- · Implement infrastructure as code using tools such as Terraform or CloudFormation.
- · Establish and maintain backup, disaster recovery, and business continuity mechanisms.
- · Work with development teams to improve application reliability and production readiness.
- · Participate in an on-call rotation and respond to production incidents when required.
- · Continuously identify opportunities to improve system resilience and operational processes.
- · Skills: Digital : Python~Digital : Kubernetes~Digital : Site Reliability Engineering (SRE)~AWS DevOps and Automation
- · Experience Required: 8-10
- · ** All submissions must have LinkedIn id of Candidate**
Must Have Skills:
- · Apple
- · AWS
- · Kubernetes
- · Python
- · SRE
Thanks and Regards,
Grace
Technical Recruiter |ASCII Group LLC.
Email: |Direct: