SDU eScience Center · status.cloud.sdu.dk

Ongoing

  • Instability causes jobs to be terminated

    We are currently experiencing some instability that causes jobs to be prematurely terminated. We are working on a solution.

    Aug 28, 08:00

August, 2026

  • Intermittent issues with storage

    We are experiencing intermittent issues with the storage system on Bitten, resulting in periodic issues when accessing files on the system. We are working on resolving the problem.

    Aug 24, 10:00 - Aug 26, 08:18

    We expect the problem is solved, and services resumed to a normal state.

    Aug 26, 08:17

  • Bitten unexpectedly rebooted

    Around 5.30 this morning, the entire Bitten system rebooted. Around one hour later, all services were back up and running again.

    Aug 15, 05:30 - Aug 15, 06:30
  • Problems during upgrade

    While rolling out an update to UCloud, we experienced some issues that made it impossible to access files and jobs.

    Aug 12, 15:30 - Aug 12, 17:15
  • Disruptions of network traffic to Bitten

    We experienced problems with network traffic to the Bitten cluster during the evening. The problem has been localized and traffic should be working again.

    Aug 6, 21:00 - Aug 6, 22:30
  • Reboot of all UCloud Nodes

    We are rolling out important patches on the system, which requires all machines to be rebooted. Running jobs will be terminated.

    Aug 6, 07:00 - Aug 6, 10:28

July, 2026

  • UCloud compute nodes are being rebooted

    We are rolling out important security patches on the system, which requires all machines to be rebooted. Running jobs will be terminated.

    Jul 10, 13:17 - Jul 12, 14:53

    The remaining compute nodes were rebooted during the afternoon.

    Jul 12, 14:49

June, 2026

  • Servicedesk maintenance

    We are updating the SDU eScience Center servicedesk so it is currently unavailable.

    Jun 30, 09:00 - Jun 30, 09:50
  • Bitten is unavailable

    The Bitten system is currently unavailable. During the night a transformer station in Aabenraa exploded and caused widespread problems for the surrounding power grid, including our data center. There has been no loss of data, but we are keeping the system offline during the weekend to assess the status and potential damages.

    Jun 26, 23:30 - Jun 29, 07:23

    Storage is now accessible, but compute nodes are still offline.

    Jun 28, 10:10

    Compute nodes are now back in production and everything should be in working order.

    Jun 29, 07:23

  • System was unexpectedly rebooted

    Around 5:30 this morning all machines at the Bitten provider were unexpectedly rebooted. We are working on bringing systems back online.

    Jun 18, 05:30 - Jun 18, 07:57
  • Issues with the Servicedesk

    We are currently experiencing issues with the SDU eScience Servicedesk. Tickets submitted via the UCloud support form and the SDU eScience support email are not reaching the Servicedesk. Longer response times are to be expected for now. It is still possible to open and reply to tickets, but only directly via support.escience.sdu.dk.

    Jun 3, 06:00 - Jun 4, 10:13

May, 2026

  • Extraordinary maintenance

    We have scheduled an extraordinary maintenance window with short notice, because we need to rollout some important changes to the system. All compute nodes will be rebooted during the maintenance window and they might be unavailable for up to two hours. All running jobs will be terminated.

    May 23, 07:00 - May 23, 08:29

    Maintenance has been completed and all nodes are now available again.

    May 23, 08:30

  • Maintenance window from our ISP

    Our ISP has announced a maintenance window on Friday (May 8th) that very likely will disrupt our internet connection to Bitten.

    May 8, 06:00 - May 8, 13:38
  • Power loss at SDU

    SDU is currently experiencing a power loss, which is also affecting services in our data center. This is also affecting the storage system at Bitten, which cannot talk to the second tier at the SDU site.

    May 6, 16:10 - May 6, 21:00
  • Issues with CPU nodes

    We are experiencing some stability issues with the new cpu-amd-zen5 machines. As a result, they sometimes need to be rebooted, affecting all jobs running on the node. We are working on finding a solution.

    May 6, 15:32 - Jun 3, 08:04
  • Internet connection down

    We temporarily lost our internet connection to the new data center. The problem has been identified by the ISP and we are waiting for them to solve it.

    May 5, 15:23 - May 6, 05:05

    The connection was finally restored again around 5am.

    May 6, 05:20

April, 2026

  • Extended UCloud downtime

    We are moving the UCloud platform to a new data center and this will require extended downtime. Downtime will start on Monday, April 27th, and continue for up to one week. The system is expected to be back online Monday, May 4th, at the latest.

    Apr 27, 08:00 - May 4, 07:00

    We are now taking UCloud offline for the data center migration.

    Apr 27, 07:59

    The migration has been completed and the system is back online.

    May 4, 07:01

March, 2026

  • Background tasks terminated for maintenance

    Unfortunately we have had to terminate background tasks (file transfers and copy operations). This was needed to deploy an updated which should improve the stability of the same feature. The background tasks have to be resubmitted to continue. We apologize for the inconvenience.

    Mar 17, 10:40 - Mar 17, 10:40

February, 2026

  • UCloud downtime

    UCloud was down from 01:00 until 08:55 due to a bug in the UCloud code. We have temporarily disabled the broken code and are working on deploying a permanent fix. UCloud is expected to work normally while we work on a fix.

    Feb 1, 01:00 - Feb 1, 08:55

January, 2026

  • Problems with Nvidia H100 cards

    After an update of the Nvidia drivers we are now experiencing problems with the Nvidia H100 (u3-gpu) machines. We will start downgrading the drivers again next week to restore full functionality.

    Jan 31, 11:32 - Feb 4, 07:47

    Drivers have been downgraded on all machines that experienced problems.

    Feb 4, 07:48

  • Unable to attach public IPs

    We are aware of an issue causing public IPs to not correctly be attached to machines. We are working on a fix.

    Jan 29, 07:04 - Jan 29, 09:30

    The issue has been resolved.

    Jan 29, 09:30

  • Downtime for all services

    UPDATE: This maintenance window has been expanded and it now covers ALL services offered by SDU eScience.

    We will be performing maintenance on January the 28th between 12:00 and 20:00. UCloud will be down during this period and jobs on SDU/K8s and AAU/K8s will be terminated at the start of the maintenance window.

    Jan 28, 12:00 - Jan 28, 17:27

    First part of the maintenance has been completed. We are now working on deploying the new version of UCloud.

    Jan 28, 15:20

    Update has completed. Please report errors that you may find.

    Jan 28, 17:28

December, 2025

  • Downtime for all services

    On December 1st there will be scheduled maintenance in the SDU data center, which requires a complete shutdown of all servers. For this reason UCloud and all other services offered by SDU eScience will be unavailable during the entire working today. We expect systems to be back online late in the afternoon.

    Dec 1, 07:00 - Dec 1, 15:16

    All services should be back up and running.

    Dec 1, 15:16

November, 2025

  • Hardware problems for u2-gpu

    The u2-gpu machines are currently experiencing hardware problems. User jobs are able to run, but they can be killed at any point due to maintenance.

    Nov 18, 09:00 - Dec 3, 11:32

    The machine has been powered off and hardware replacements should be performed later today.

    Dec 3, 09:21

    The hardware has been replaced and all GPUs are working again.

    Dec 3, 11:33

  • UCloud is experiencing issues

    UCloud is currently experiencing issues, we are working on fixing the problem.

    Nov 17, 19:26 - Nov 17, 20:21
  • UCloud partially unavailable

    During the night UCloud had an internal issue, which made it impossible to access jobs and files.

    Nov 14, 00:12 - Nov 14, 06:48
  • UCloud has been updated

    UCloud has been updated with the latest round of bug fixes and improvements to the UI. As always, this may have caused a few minutes of disruption to the service. Sorry for the inconvenience.

    Nov 11, 09:37 - Nov 11, 09:36

October, 2025

  • UCloud jobs will be terminated

    On Sunday, October 26th, several UCloud jobs that have been running for more than 14 days will be terminated due to hardware maintenance.

    Oct 26, 10:00 - Oct 26, 18:07

    The jobs have been terminated.

    Oct 26, 18:07

  • SDU/K8s unavailable

    The SDU/K8s provider was unavailable between 12/10/25 15:56:34 and 12/10/25 17:04:25 due to a software bug. A bug fix has been released now (13/10/25 07:00).

    Oct 12, 15:56 - Oct 12, 17:04

September, 2025

  • UCloud is being restarted for an update

    UCloud is being restarted for an update. The update will take a few minutes.

    Sep 30, 08:50 - Sep 30, 08:52
  • Two H100 nodes rebooted

    Two of the H100 nodes were rebooted this morning due to hardware maintenance, the first one around 9.00 and the second one around 10.30 A couple of jobs where stopped during the reboot.

    Sep 26, 08:00 - Sep 26, 11:05
  • Compute nodes unresponsive

    Around 8:50 this morning we started decommissioning an old storage system, which unfortunately affects the SDU/K8s compute nodes, making them partially unresponsive.

    Sep 22, 08:52 - Sep 22, 10:45

    Things should finally be returning to normal. If jobs are stuck for more than 10 minutes, start a new one.

    Sep 22, 10:31

  • Power fluctuation caused machines to reboot

    A fluctuation in the power grid caused around 30 machines to reboot in the UCloud server room. Jobs running on the machines were terminated during the reboot.

    Sep 9, 21:45 - Sep 10, 06:30
  • Issues with nodea0-19

    The machine 'nodea0-19' has been rebooted due to an error with one of the GPU cards. We are monitoring the node for a potential hardware issue.

    Sep 2, 07:35 - Sep 19, 13:19

    The error has reappeared, the card will most likely need to be replaced.

    Sep 4, 07:48

    A support case has been opened to get the card replaced.

    Sep 10, 07:49

    Hardware maintenance has been initiated.

    Sep 19, 12:46

    The faulty GPU has been replaced.

    Sep 19, 13:19

Read the original on status.cloud.sdu.dk ↗