Перейти к основному содержимому
Cerebras Systems

Hardware Analytics Engineer

УдалённоUnited States только
Опубликовано
Роль
Инженерия данных
Опыт
Мидл
Занятость
Полная занятость
$213.7k–$225k/yr
Проверьте доступность

Доступно для: US only. Укажите, откуда вы работаете, чтобы проверить доступность.

Коротко по делу

Hardware analytics/data engineer with a Master's degree and 3 years in hardware analytics, hardware engineering, data engineering, or related work. Must know large-scale ETL/data pipelines, Hive/Spark, Python, SQL, Tableau, Linux, ML modeling, anomaly detection, and hardware reliability/performance analytics. Sunnyvale, CA job site with telecommuting permitted.

Ключевые навыки

ETLSparkPython

Обязательные навыки

HiveSQLTableauLinuxAutomation scriptingMachine learningPredictive modelingStatistical analysisA/B testingAnomaly detectionData visualization

Чем предстоит заниматься

  • Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  • Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.
  • Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  • Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  • Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.
  • Define, operationalize, and maintain custom efficiency and reliability metrics; perform root cause analysis of systemic failures using large-scale statistical and machine learning methods; and deploy solutions that improve platform scalability, energy efficiency, and sustainability.
  • Collaborate with cross-functional engineering teams to troubleshoot complex failures, isolate defective components, and implement systemic fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Support the evolution and optimization of next-generation AI platforms and silicon products, including hardware subsystems to meet the performance, scalability, and efficiency demands of large language model training and inference workloads.

Что требуется

  • Master’s degree or foreign equivalent degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
  • 3 years of experience as Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or a related occupation required.
  • Experience with large-scale data pipeline architecture and ETL, distributed data processing, and dashboard development.
  • Experience with design, training, and deployment of machine learning models for hardware performance optimization and failure prediction.
  • Experience with predictive modeling, statistical analysis, A/B testing, anomaly detection, and data visualization in hardware reliability and performance.
  • Experience with hardware analytics for compute, storage, and AI servers; power and thermal optimization; GPU burn-in efficiency optimization; and reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD.

Преимущества

  • Build a breakthrough AI platform beyond the constraints of the GPU.
  • Publish and open source cutting-edge AI research.
  • Work on one of the fastest AI supercomputers in the world.
  • Enjoy job stability with startup vitality.
  • Simple, non-corporate work culture that respects individual beliefs.
  • Continuous learning, growth and support.

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

AI HardwareСтартап

Детали

Способ откликаDom
Также опубликовано ещё в 1 канале
$213.7k–$225k/yr