UM Logo
DICC Logo
AboutGuidelinesServicesNewsContactRegister Now
UM LogoDICC Logo

© 2026 Data-Intensive Computing Centre, Universiti Malaya. All Rights Reserved.

UM Logo
DICC Logo
AboutGuidelinesServicesNewsContactRegister Now
Back to News
Incident
January 5, 2021

HPC CPU Node Crashed

Published by DICC

Dear all HPC Users,

There was an incident where one of the CPU compute nodes in the HPC pool, cpu06 crashed due to high CPU load caused by some processes stuck in the machine indefinitely. This incident has caused all the jobs running in cpu06 to fail as the worker daemon was not able to communicate with the scheduler due to high CPU load. 

The machine was rebooted physically this morning around 7.45am. If you were running some jobs in the affected machine, please verify and resubmit your job if necessary. We are still investigating the root causes of the incident and will take appropriate action to prevent this issue from happening again.

If you have any issue or question, please do not hesitate to contact us through the service desk. 

Thank you.

Categories: Incident

Related Posts

IncidentNews

Unauthorized Usage of HPC Resources

October 25, 2025

Dear HPC users, It has come to our attention that some users are purposely misusing the HPC resources that are available...

HPCIncident

Incidents on Scratch Storage Cleanup

December 10, 2024

Greetings HPC User, We are sorry to inform you that there was an oversight during the scheduled Scratch Storage cle...

HPCIncident

HPC Service Degradation

November 21, 2024

Greetings HPC Users, We are currently having issues with our virtualisation servers that are used to host all the suppor...

UM LogoDICC Logo

© 2026 Data-Intensive Computing Centre, Universiti Malaya. All Rights Reserved.