FeM e.V. Infrastructure and System Status

FeM e.V. Infrastructure and System Status

This page shows an overview of the status of infrastructure, services and systems of FeM e.V.

Please keep in mind that all helpers do this voluntarily in their free time. Therefore, there are no guarantees for any given timelines or time estimations.

If you have any questions or specific problems, feel free to send us a ticket via our helpdesk: https://helpdesk.fem.tu-ilmenau.de.

All systems are operational.

Last updated 2 days ago
    • Internet Uplink

      Operational
      Last updated 2026-08-23 21:59:06
    • Wireless (Freescale system)

      House A

      Operational
      Last updated 2026-08-01 21:43:16
    • Wireless (Aruba system)

      House B, C, D, E, H, I, K, L, N, P, Q, CJD and FeM-Office

      Operational
      Last updated 2026-07-28 17:49:29
    • Wired

      Ethernet sockets

      Operational
      Last updated 2026-01-19 23:58:28
    • Network Management

      FeM AdminDB and network services (DHCP, DNS...)

      Operational
      Last updated 2026-04-23 21:48:54
    • Helpdesk

      Operational
      Last updated 2026-01-30 15:05:44
    • Storage systems

      Storage systems for servers (e.g. Ceph, ha-storage)

      Operational
      Last updated 2026-01-17 14:09:13
    • Monitoring

      All monitoring systems, e.g. Icinga, Grafana, ...

      Operational
      Last updated 2025-11-04 14:15:17
    • Databases

      Operational
      Last updated 2026-08-23 19:28:39
    • VM-Hosts

      Operational
      Last updated 2026-08-23 19:28:40
    • Mailserver

      Mail servers, inboxes & list systems

      Operational
      Last updated 2026-01-30 15:05:22
    • Webcluster

      Operational
      Last updated 2026-08-23 19:28:38
    • Mastodon

      Federated social network

      Operational
      Last updated 2026-05-10 14:39:06
    • Mattermost

      Internal Messenger

      Operational
      Last updated 2026-08-23 19:28:41
    • Pretix

      A Service to Sell Tickets and Manage Events

      Operational
      Last updated 2026-08-23 19:28:42
    • Dist-Mirror

      A Mirror for some selected Linux Distributions

      Operational
      Last updated 2025-10-31 12:35:48
    • XMPP

      XMPP / Jabber Instant Messaging

      Operational
      Last updated 2026-05-05 23:21:05
    • FeMCI

      FeM Content Infrastructure

      Operational
      Last updated 2026-08-23 19:28:43
    • Homepage

      Our homepage

      Operational
      Last updated 2026-05-22 15:00:39
    • Matrix

      Federated Instant Messaging

      Operational
      Last updated 2026-08-23 19:28:44

Past incidents

No incidents reported.

No incidents reported.

No incidents reported.

No incidents reported.

1 month ago 2026-07-15 11:15 UTC
Fixed
Affected Components: Wireless (Aruba system) and Wired

Fixed

1 month ago 2026-07-15 16:52 UTC

We replaced the broken switch module in the network switch of House H.

Identified

1 month ago 2026-07-15 12:20 UTC

We are currently in the process of preparing a replacement module and hope to replace it later today.

Reported

1 month ago 2026-07-15 11:15 UTC

We are seeing a failed switch module in the Switch of House H.

1 month ago 2026-07-14 17:43 UTC
Fixed
Affected Components: Wireless (Freescale system), Wireless (Aruba system) and Wired

Fixed

1 month ago 2026-07-14 21:13 UTC

After forcefully reenabling some of the Freescale access points, everything is working properly again.

Watching

1 month ago 2026-07-14 17:59 UTC

We are seeing system recovery and are monitoring the situation.

Reported

1 month ago 2026-07-14 17:43 UTC

We are seeing some network outages after the switches in Houses CJD and K and C have been forcefully rebooted. We suspect a power outage. The WiFi and network connection will suffer.

We are investigeting.

1 month ago 2026-07-13 19:00 UTC
Fixed
Affected Components: Wireless (Freescale system)

Fixed

1 month ago 2026-07-14 16:06 UTC

All access points are back online and working correctly.

This incident has been resolved.

Watching

1 month ago 2026-07-14 15:20 UTC

We were able to fix some configuration changes (made last Thursday) and to restart the faulty service and its networking interfaces.

We are now seeing recovering Freescale access points and are monitoring the situation.

Identified

1 month ago 2026-07-14 12:24 UTC

We have identified a faulty service, which handles orchestrating the access points and does the traffic tunneling. This service is currently unable to connect to the AP network.

We are currently in the process of reestablishing service operations.

Reported

1 month ago 2026-07-13 19:00 UTC

We are experiencing total connectivity and availability issues with our Freescale WiFi in Houses A and CJD. All access points are currently offline and do not allow network usage.

1 month ago 2026-07-13 17:50 UTC
Fixed
Affected Components: Internet Uplink, Wireless (Freescale system), Wireless (Aruba system), Network Management and Wired

Fixed

1 month ago 2026-07-14 07:20 UTC

We are happy to announce that all systems (except the Freescale WiFi of houses A and CJD) are functioning normally again.

Therefor, this incident has been resolved.

Watching

1 month ago 2026-07-13 21:30 UTC

We were able to pinpoint the (hopefully) final cause of this major outage.

Some technical details

The cause was the change of maximum packet size (MTU) between our core network hosts, which was deployed as part of regular maintenance on Thursday evening. This caused internal packet loss together with slow and steady builtup of query and logging backlogs, which caused a major disruption in one of the database master nodes, that handles all of the network management operations, IP assignments and DNS record storage. This database outage resulted in the DHCP issues and the partial inability to register new devices.

As we verified the of the MTU change on a testing network beforehand, we did not suspect it to be a major cause of this outage. After exhausting all other possible causes, we tested a rollback of this change today (Monday), which saw minor improvements. The full rollback was executed in early evening and caused another wave of network disruptions, as most central services needed to reboot.

But the issue still persisted, as one of the database nodes was still faulty. As a last ressort, we fully reset this database node and started a replication from other parts of the core system (which is currently in process). The query and logging backlog has been processed after the faulty database node was removed. As a result, network operations are currently returning back to normal operations (hopefully).

tl;dr;

The core database systems are currently recovering. We are monitoring the situation and hope to see a final full system recovery.

While most parts of the Aruba WiFi are already back online, the recovery of the Freescale WiFi (Houses CJD and A) ist still in progress.

Reported

1 month ago 2026-07-13 20:31 UTC

The NAT systems have recovered, DNS/DHCP issues are still persisting.

Investigating

1 month ago 2026-07-13 20:15 UTC

As mentioned beforehand: we are seeing some issues regarding our NAT systems.

Identified

1 month ago 2026-07-13 19:50 UTC

We are seeing partial recovery on DHCP and DNS systems.

We are expecting a short outage in a few minutes, as there will be a small disruption when restarting some services.

Investigating

1 month ago 2026-07-13 19:07 UTC

We are currently looking into the issue regarding the DNS (which seems to recover right now).

Reported

1 month ago 2026-07-13 18:39 UTC

After partially applying the aforementioned changes, we are seeing some more system failures, now regarding our DNS servers.

Identified

1 month ago 2026-07-13 18:21 UTC

After carefully rereading the changelogs of the maintenance of last Thrusday, we were able to pinpoint some potential causes of the widespread issues we are facing.

We are in the process of reverting one of these changes and are seeing partial recoveries in parts of our system.

Reported

1 month ago 2026-07-13 17:50 UTC

We are experiencing issues with IP address assignemnt.

This results in connection problems and outages.

Finalization of Core updates

1 month ago 2026-07-13 09:30 UTC
Complete
Affected Components: Mailserver, VM-Hosts, Mastodon, Matrix and Wired

We were not able to finish the maintenance of last Thursday. This will be done in this timeframe.

We hope to resolve the widespread outage of the weekend and reestablish normal network operations.