Skip to jobs
All jobs

Firmus Technologies · Platform Engineering

Principal Engineer, AI Cloud Software

Where
Singapore
Experience
7+ yrsstated in the description
Pay
Not stated
Posted
First seen by Unlisted 8 Oct, 02:22 UTC
Apply on GreenhouseOpens the employer's own posting

Checking Greenhouse for this posting…

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

ROLE SUMMARY 

Firmus Technologies is seeking a Senior AI Infrastructure Engineer, Observability, to join our Engineering and Technology team. You will establish how we measure, validate and communicate the health of GPU infrastructure used for customer and internal workloads. You will define trusted health signals and service-readiness criteria, and turn them into reusable dashboards, alerts, queries, diagnostic checks and operational guidance. Your work will help commissioning, infrastructure and operations teams bring capacity online safely, identify degradation early and recover from failures quickly. You will also make knowledge self-service by publishing clear reference implementations, runbooks and AI-ready operational knowledge that other teams can use and extend. 

 KEY RESPONSIBILITIES 

 

 

Rising correctable ECC counts, NVLink retries, thermal slowdown, XID patterns, wrong results with no error. Keep the knowledge current: what each signal means, what to do next on the machine, and where the operation team must make the final decision. 
 

 

SKILLS AND EXPERIENCE 

Highly Desirable Experiences

 

Location & Reporting

 

Employment Basis

Full-time

 

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.