
Site Reliability Engineer of AI Infrastructure Operations (IMC)
at TSMC
Posted a day ago
No clicks
- Compensation
- Not specified
- City
- Country
- Not specified
Currency: Not specified
**Site Reliability Engineer - AI Infrastructure Operations** Apply your experienced SRE skills to ensure high availability and reliability of AI/ML infrastructure. Monitor and improve systems, collaborating with cross-functional teams. Key responsibilities involve designing scalable, fault-tolerant infrastructure, implementing CI/CD pipelines, and leveragingdata-driven insights for informed decisions. Required skills include proficiency in Kubernetes, Terraform, Prometheus, and Python. A bachelor's degree in Computer Science or a related field, along with 5+ years of experience in SRE or a similar role, is preferred. Familiarity with AI/ML workflows is a bonus.
No additional description provided.
