Principal Observability Architect, AI and HPC puesto vacante

Vacancy caducado!

NVIDIA’s Hardware Infrastructure organization is seeking a Senior or Princip al Data and Observability Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world. What You’ll Be Doing:

C ollaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability.

Architect and lead teams to d evelop, test, and deploy data collectors, pipelines, visualization and retrieval services .

Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements.

Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency.

Continuously improve quality, workloads, and processes through better observability .

What We Need to See:

Experience designing and building large scale, distributed observability systems.

Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis.

Experience with turning raw data into actionable reports

Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools

Technical lead level Python programming experience and use of API calls

Passion for improving the productivity of others

Excellent planning and interpersonal skills

Flexibility/adaptability working in a dynamic environment with changing requirements

MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience

12 +yrs of relevant experience.

Ways To Stand Out from The Crowd:

Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology.

Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops

Experience in the management of datacenters and large-scale distributed computing

Experience in working with AI researchers and/or EDA developers

Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end.

The base salary range is 272,000 USD - 471,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.You will also be eligible for equity and benefits (https://www.nvidia.com/en-us/benefits/) . NVIDIA accepts applications on an ongoing basis. NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy caducado!

ID	#52810219
Estado	California
Ciudad	Santaclara
	Full-time
Salario	USD TBD TBD
Fuente	Nvidia
Showed	2024-11-01
Fecha	2024-11-02
Fecha tope	2024-12-31
Categoría	Etcétera
Crear un currículum vítae

Job Details

Principal Observability Architect, AI and HPC

Puestos de trabajo relacionados

»Architect/Sr. Principal Engineer, Backend - Cortex Cloud (Posture Security)

»Principal Security Researcher (Wildfire)

»Principal Software Engineer - JVM Ecosystem

»Principal Cloud Software Engineer (WildFire Cloud)

»Senior Principal IT Software Engineer

»Principal SDET Engineer (Cortex Cloud)

»Principal Product Security Researcher (Vulnerability Research)