Site Reliability Senior Staff
at Synopsys
Posted 18 hours ago
No clicks
- Compensation
- Not specified
- City
- Country
- Vietnam
Currency: Not specified
**Site Reliability Senior Staff Engineer (Ho Chi Minh City, Vietnam)** As a Site Reliability Senior Staff Engineer, you'll ensure high availability and security of Synopsys' EDA engineering compute platforms and Synopsys.ai services. With senior-level experience in Unix/Linux infrastructure management, you'll team up with global IT, AE engineering, and product groups to provide reliable, secure compute environments for R&D and customer engineering teams, supporting HPC batch grids, containerized AI gateways, and vector data services. Key skills: Infrastructure as Code (IaC), continuous integration/continuous deployment (CI/CD), cloud and on-premises environments, customer-hosted and air-gapped deployments.
- Improves availability, operability, and utilization of business-critical EDA and Synopsys.ai infrastructure in Vietnam.
- Reduces operational risk for customer air-gapped and customer-hosted deployments through strong runbooks, monitoring, and incident practices.
- Connects customer HPC schedulers, shared storage, GPU compute, and platform gateways so Synopsys tools and GenAI services run predictably at scale.
- Extends APAC shared-support coverage with Japan, Korea, Taiwan, and China peers.
- Maintain Synopsys Vietnam engineering compute and data center environments per corporate IT, data center, and security standards.
- Support server hardware lifecycle activities: installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning.
- perform capacity planning, performance troubleshooting, and vendor/customer coordination at customer sites.
- Operate and troubleshoot Linux-based EDA compute: virtualization, engineering resource management, job schedulers, license connectivity, and remote access services.
- Run HPC operations: cluster health, scheduler integration (LSF, Slurm, or similar), InfiniBand where deployed, and performance-related incidents.
- Collaborate with corporate network and information security on data center connectivity, switching, firewalls, routing, and circuits.
- Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam: customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration (cloud or on-prem inference endpoints per customer policy), and customer-controlled vector database/GPU compute where applicable.
- Partner on install, upgrade, and configuration of platform components that connect user/agent workflows to the engineering grid—e.g., job services, unified MCP server patterns, and tool execution on LSF/Slurm-backed clusters (in coordination with products and global platform teams).
- Implement and maintain observability, alerting, authentication/authorization integration (e.g., customer IdP, OIDC/SSO/OAuth patterns), and operational documentation for platform services.
- Support agentic AI platform rollouts where components run locally, centrally, or on the grid; escalate design questions to global architecture and engineering teams.
- Own or co-own runbooks, monitoring/alerting policies, and automation (scripts, IaC, or equivalent) to reduce toil across the APAC team.
- Apply AI-assisted workflows to improve troubleshooting, knowledge management, and routine operations.
- On-call and incident response: shared rotation (weekday nights, weekends, and holidays); lead or support incident bridges; produce post-incident reports in English.
- Customer air-gapped and on-site support: planned upgrades, break-fix, and customer training for isolated HPC and on-prem AI sites. Travel within Vietnam and APAC as needed (typically 10–20% of time).
- Cross-regional collaboration: handoffs, documentation, and mentor junior staff where applicable.
- 5+ years in Linux system administration, infrastructure engineering, or production network/storage support.
- Enterprise data center experience: servers, storage, networking, monitoring, and operational processes.
- HPC or EDA engineering compute: Linux clusters, workload schedulers (LSF, Slurm, or similar), and production support.
- Ability to support GPU-enabled compute and containerized platform services in customer-controlled environments.
- Familiarity with virtualization, remote access, observability stacks, and infrastructure automation.
- Understanding of secure multi-tier deployments: gateways, service-to-service auth, and operational telemetry (logs, metrics, tracing) for distributed platform components.
Site Reliability Senior Staff
at Synopsys
Site Reliability Senior Staff
at Synopsys
Posted 18 hours ago
No clicks
- Compensation
- Not specified
- City
- Country
- Vietnam
Currency: Not specified
**Site Reliability Senior Staff Engineer (Ho Chi Minh City, Vietnam)** As a Site Reliability Senior Staff Engineer, you'll ensure high availability and security of Synopsys' EDA engineering compute platforms and Synopsys.ai services. With senior-level experience in Unix/Linux infrastructure management, you'll team up with global IT, AE engineering, and product groups to provide reliable, secure compute environments for R&D and customer engineering teams, supporting HPC batch grids, containerized AI gateways, and vector data services. Key skills: Infrastructure as Code (IaC), continuous integration/continuous deployment (CI/CD), cloud and on-premises environments, customer-hosted and air-gapped deployments.
- Improves availability, operability, and utilization of business-critical EDA and Synopsys.ai infrastructure in Vietnam.
- Reduces operational risk for customer air-gapped and customer-hosted deployments through strong runbooks, monitoring, and incident practices.
- Connects customer HPC schedulers, shared storage, GPU compute, and platform gateways so Synopsys tools and GenAI services run predictably at scale.
- Extends APAC shared-support coverage with Japan, Korea, Taiwan, and China peers.
- Maintain Synopsys Vietnam engineering compute and data center environments per corporate IT, data center, and security standards.
- Support server hardware lifecycle activities: installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning.
- perform capacity planning, performance troubleshooting, and vendor/customer coordination at customer sites.
- Operate and troubleshoot Linux-based EDA compute: virtualization, engineering resource management, job schedulers, license connectivity, and remote access services.
- Run HPC operations: cluster health, scheduler integration (LSF, Slurm, or similar), InfiniBand where deployed, and performance-related incidents.
- Collaborate with corporate network and information security on data center connectivity, switching, firewalls, routing, and circuits.
- Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam: customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration (cloud or on-prem inference endpoints per customer policy), and customer-controlled vector database/GPU compute where applicable.
- Partner on install, upgrade, and configuration of platform components that connect user/agent workflows to the engineering grid—e.g., job services, unified MCP server patterns, and tool execution on LSF/Slurm-backed clusters (in coordination with products and global platform teams).
- Implement and maintain observability, alerting, authentication/authorization integration (e.g., customer IdP, OIDC/SSO/OAuth patterns), and operational documentation for platform services.
- Support agentic AI platform rollouts where components run locally, centrally, or on the grid; escalate design questions to global architecture and engineering teams.
- Own or co-own runbooks, monitoring/alerting policies, and automation (scripts, IaC, or equivalent) to reduce toil across the APAC team.
- Apply AI-assisted workflows to improve troubleshooting, knowledge management, and routine operations.
- On-call and incident response: shared rotation (weekday nights, weekends, and holidays); lead or support incident bridges; produce post-incident reports in English.
- Customer air-gapped and on-site support: planned upgrades, break-fix, and customer training for isolated HPC and on-prem AI sites. Travel within Vietnam and APAC as needed (typically 10–20% of time).
- Cross-regional collaboration: handoffs, documentation, and mentor junior staff where applicable.
- 5+ years in Linux system administration, infrastructure engineering, or production network/storage support.
- Enterprise data center experience: servers, storage, networking, monitoring, and operational processes.
- HPC or EDA engineering compute: Linux clusters, workload schedulers (LSF, Slurm, or similar), and production support.
- Ability to support GPU-enabled compute and containerized platform services in customer-controlled environments.
- Familiarity with virtualization, remote access, observability stacks, and infrastructure automation.
- Understanding of secure multi-tier deployments: gateways, service-to-service auth, and operational telemetry (logs, metrics, tracing) for distributed platform components.
SIMILAR OPPORTUNITIES
No similar jobs available at the moment.
