FeM e.V. Infrastructure and System Status
This page shows an overview of the status of infrastructure, services and systems of FeM e.V.
Please keep in mind that all helpers do this voluntarily in their free time. Therefore, there are no guarantees for any given timelines or time estimations.
If you have any questions or specific problems, feel free to send us a ticket via our helpdesk: https://helpdesk.fem.tu-ilmenau.de.
All systems are operational.
-
-
Internet Uplink
OperationalLast updated 2026-08-23 21:59:06 -
Wireless (Freescale system)
House A
OperationalLast updated 2026-08-01 21:43:16 -
Wireless (Aruba system)
House B, C, D, E, H, I, K, L, N, P, Q, CJD and FeM-Office
OperationalLast updated 2026-07-28 17:49:29 -
Wired
Ethernet sockets
OperationalLast updated 2026-01-19 23:58:28 -
Network Management
FeM AdminDB and network services (DHCP, DNS...)
OperationalLast updated 2026-04-23 21:48:54
-
-
-
Helpdesk
OperationalLast updated 2026-01-30 15:05:44 -
Storage systems
Storage systems for servers (e.g. Ceph, ha-storage)
OperationalLast updated 2026-01-17 14:09:13 -
Monitoring
All monitoring systems, e.g. Icinga, Grafana, ...
OperationalLast updated 2025-11-04 14:15:17 -
Databases
OperationalLast updated 2026-08-23 19:28:39 -
VM-Hosts
OperationalLast updated 2026-08-23 19:28:40 -
Mailserver
Mail servers, inboxes & list systems
OperationalLast updated 2026-01-30 15:05:22 -
Webcluster
OperationalLast updated 2026-08-23 19:28:38
-
-
-
Mastodon
Federated social network
OperationalLast updated 2026-05-10 14:39:06 -
Mattermost
Internal Messenger
OperationalLast updated 2026-08-23 19:28:41 -
Pretix
A Service to Sell Tickets and Manage Events
OperationalLast updated 2026-08-23 19:28:42 -
Dist-Mirror
A Mirror for some selected Linux Distributions
OperationalLast updated 2025-10-31 12:35:48 -
XMPP
XMPP / Jabber Instant Messaging
OperationalLast updated 2026-05-05 23:21:05 -
FeMCI
FeM Content Infrastructure
OperationalLast updated 2026-08-23 19:28:43 -
Homepage
Our homepage
OperationalLast updated 2026-05-22 15:00:39 -
Matrix
Federated Instant Messaging
OperationalLast updated 2026-08-23 19:28:44
-
Past incidents
No incidents reported.
No incidents reported.
No incidents reported.
No incidents reported.
Fixed
We replaced the broken switch module in the network switch of House H.
Identified
We are currently in the process of preparing a replacement module and hope to replace it later today.
Reported
We are seeing a failed switch module in the Switch of House H.
Fixed
After forcefully reenabling some of the Freescale access points, everything is working properly again.
Watching
We are seeing system recovery and are monitoring the situation.
Reported
We are seeing some network outages after the switches in Houses CJD and K and C have been forcefully rebooted. We suspect a power outage. The WiFi and network connection will suffer.
We are investigeting.
Fixed
All access points are back online and working correctly.
This incident has been resolved.
Watching
We were able to fix some configuration changes (made last Thursday) and to restart the faulty service and its networking interfaces.
We are now seeing recovering Freescale access points and are monitoring the situation.
Identified
We have identified a faulty service, which handles orchestrating the access points and does the traffic tunneling. This service is currently unable to connect to the AP network.
We are currently in the process of reestablishing service operations.
Reported
We are experiencing total connectivity and availability issues with our Freescale WiFi in Houses A and CJD. All access points are currently offline and do not allow network usage.
Fixed
We are happy to announce that all systems (except the Freescale WiFi of houses A and CJD) are functioning normally again.
Therefor, this incident has been resolved.
Watching
We were able to pinpoint the (hopefully) final cause of this major outage.
Some technical details
The cause was the change of maximum packet size (MTU) between our core network hosts, which was deployed as part of regular maintenance on Thursday evening. This caused internal packet loss together with slow and steady builtup of query and logging backlogs, which caused a major disruption in one of the database master nodes, that handles all of the network management operations, IP assignments and DNS record storage. This database outage resulted in the DHCP issues and the partial inability to register new devices.
As we verified the of the MTU change on a testing network beforehand, we did not suspect it to be a major cause of this outage. After exhausting all other possible causes, we tested a rollback of this change today (Monday), which saw minor improvements. The full rollback was executed in early evening and caused another wave of network disruptions, as most central services needed to reboot.
But the issue still persisted, as one of the database nodes was still faulty. As a last ressort, we fully reset this database node and started a replication from other parts of the core system (which is currently in process). The query and logging backlog has been processed after the faulty database node was removed. As a result, network operations are currently returning back to normal operations (hopefully).
tl;dr;
The core database systems are currently recovering. We are monitoring the situation and hope to see a final full system recovery.
While most parts of the Aruba WiFi are already back online, the recovery of the Freescale WiFi (Houses CJD and A) ist still in progress.
Reported
The NAT systems have recovered, DNS/DHCP issues are still persisting.
Investigating
As mentioned beforehand: we are seeing some issues regarding our NAT systems.
Identified
We are seeing partial recovery on DHCP and DNS systems.
We are expecting a short outage in a few minutes, as there will be a small disruption when restarting some services.
Investigating
We are currently looking into the issue regarding the DNS (which seems to recover right now).
Reported
After partially applying the aforementioned changes, we are seeing some more system failures, now regarding our DNS servers.
Identified
After carefully rereading the changelogs of the maintenance of last Thrusday, we were able to pinpoint some potential causes of the widespread issues we are facing.
We are in the process of reverting one of these changes and are seeing partial recoveries in parts of our system.
Reported
We are experiencing issues with IP address assignemnt.
This results in connection problems and outages.
Finalization of Core updates
We were not able to finish the maintenance of last Thursday. This will be done in this timeframe.
We hope to resolve the widespread outage of the weekend and reestablish normal network operations.