About the Role
We are looking for SeniorMLOps& Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast’s ADAS and autonomous drivingprogramme-on-premiseGPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
What the Team Covers
You will contribute acrossall ofthe following over time. We do not expect one person to master every part on day one, but we do expect you to be willing to work in any of them.
GPU and HPC clusters running model training, large-scale inference and annotation workloads
MLOps: model serving, pipeline orchestration, experiment and model registry
Storage at multi-petabyte scale: tiering, lifecycle, backup and recovery
Ingestion of large volumes of recorded sensor data, with integrity verification and metadata extraction
CI/CD,GitOps, infrastructure-as-code, monitoring and observability
Access control, audit logging and data protection for internal and external users
Key Responsibilities
AI Compute & Serving
Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling,utilisationand capacity
Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
Manage multi-GPU distributed training and reproducible environment sandboxing
Scale pipeline orchestration on Kubernetes for large data processing jobs
Platform Automation & Delivery
DriveGitOps-based deployment and maintain infrastructure-as-code across the platform
Build and maintain CI pipelines, secure container builds and release automation
Build monitoring, logging and alerting so that failures are detected and actionable, never silent
Lead incident response and drive follow-up actions to closure
Data Infrastructure & Security
Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
Implement access control and single sign-on across platform services, with audit logging
Apply data protection measures to sensitive content before it reaches external users
Requirements
- 4+ years inMLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure
- Kubernetesadministration at production scale: Helm, ingress (Traefikor Envoy), CNI, andGitOps(Flux orArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- StrongLinux systems skillsand infrastructure-as-code (SaltStackor Ansible)
- Python and Bashfor automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack- compute, storage, networking, automation and security - rather than within a single specialty
- Good communication in English - technical documentation and working with international partners
Benefits
- Competitive salary
- Premium healthcare package, including PVI insurance & annual health check-ups
- 13th-month salary & performance bonuses to reward your contributions
- Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool...
- Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain
- Work Location:Technopark Tower, Gia Lam, Ha Noi
- Locations
- Gia Lâm, Hanoi, Hanoi, Vietnam
- Experience
- 4+ years
