N11 - response problem

Incident Report for TostHost

Postmortem

ENG:

Postmortem - Dedicated Server Outage

Incident date: August 14–16, 2026
Incident type: Hardware / PCIe failure
Affected system: Dedicated server hosting customer services

Summary

Between August 14 and 16, 2026, a recurring hardware-related incident affected a dedicated server used to provide hosting services.

The server began experiencing extremely high system load and I/O wait. The Linux kernel repeatedly reported PCIe bus errors:

PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer

As a result, some services became unavailable or operated incorrectly. Websites and file downloads through Nginx remained partially functional, but SSH and other services became inaccessible. The KVM console also stopped accepting keyboard input.

Incident Timeline

August 14

The first major incident was observed. The load average increased to approximately 62.92, while CPU I/O wait reached approximately 67.45%.

The KVM console showed repeated PCIe errors related to:

0000:00:1b.0

The server provider performed an initial hardware check, but no clear hardware failure was detected during the brief diagnostic.

Following the incident, the system kernel was updated.

August 16

The same issue occurred again despite the kernel update.

The symptoms were similar to the previous incident: extremely high load and I/O wait, service instability, inability to connect via SSH, and recurring PCIe errors visible in the KVM console.

A hard reset was performed. The server initially recovered, but the same issue occurred again approximately 10 minutes after the reboot.

Hardware Replacement

Due to the recurring nature of the issue, the server provider replaced several major hardware components.

The following components were replaced:

  • CPU,
  • motherboard,
  • RAM,
  • network card.

The existing drives were retained.

After the hardware replacement, the server was brought back online and the affected services were restored.

Remediation

The following actions were taken:

  • hardware diagnostics were performed,
  • the Linux kernel was updated,
  • an emergency server reboot was performed,
  • the CPU was replaced,
  • the motherboard was replaced,
  • the RAM was replaced,
  • the network card was replaced,
  • affected services were restored,
  • additional monitoring was enabled.

Customer Compensation

Due to the service disruption, all customer services affected by the incident were automatically extended by 48 hours.

Status

Incident resolved.

Following the hardware replacement, the affected services have been restored and the server is operating normally. The server will continue to be monitored to ensure long-term stability.

PL:

Postmortem - Awaria serwera dedykowanego

Data incydentu: 14–16 sierpnia 2026
Typ incydentu: Awaria sprzętowa / PCIe
Dotknięty system: Serwer dedykowany używany do hostingu usług klientów

Podsumowanie

W dniach 14–16 sierpnia 2026 wystąpiła powtarzająca się awaria serwera dedykowanego wykorzystywanego do świadczenia usług hostingowych.

Serwer zaczął wykazywać bardzo wysokie obciążenie systemu oraz problemy z operacjami I/O. W systemie pojawiały się powtarzające się błędy magistrali PCIe:

PCIe Bus Error: severity=Uncorrectable (Non-Fatal), type=Transaction Layer

W wyniku awarii część usług przestała działać prawidłowo. Strony internetowe i pobieranie plików przez Nginx nadal działały częściowo, jednak SSH oraz inne usługi były niedostępne. Konsola KVM również przestała przyjmować dane wejściowe.

Przebieg incydentu

14 sierpnia

Po raz pierwszy zaobserwowano poważne problemy z serwerem. Load average wzrósł do około 62,92, a CPU I/O wait osiągnął około 67,45%.

W konsoli KVM pojawiały się powtarzające się błędy PCIe związane z portem:

0000:00:1b.0

Po przeprowadzeniu diagnostyki przez operatora nie wykryto jednoznacznych błędów podczas krótkiego testu sprzętowego.

Po incydencie systemowy kernel został zaktualizowany.

16 sierpnia

Problem wystąpił ponownie, pomimo aktualizacji kernela.

Objawy były bardzo podobne do tych z 14 sierpnia: wysokie obciążenie, bardzo wysoki I/O wait, problemy z usługami, brak możliwości połączenia przez SSH oraz powtarzające się błędy PCIe w KVM.

Po wykonaniu hard resetu serwer uruchomił się ponownie, jednak po około 10 minutach problem ponownie wystąpił.

Wymiana sprzętu

Ze względu na powtarzalność problemu oraz jego charakter operator serwera przeprowadził wymianę sprzętu.

W ramach wymiany zostały wymienione:

  • procesor,
  • płyta główna,
  • pamięć RAM,
  • karta sieciowa.

Dyski zostały zachowane.

Po wymianie serwer został ponownie uruchomiony i usługi zostały przywrócone.

Przyczyna

Najbardziej prawdopodobną przyczyną incydentu była usterka sprzętowa związana z platformą serwera i magistralą PCIe.

Ze względu na wymianę wielu kluczowych komponentów sprzętowych nie można jednoznacznie wskazać pojedynczego uszkodzonego elementu.

Aktualnie nie obserwujemy ponownego występowania opisanych problemów.

Działania naprawcze

  • przeprowadzono diagnostykę sprzętową,
  • zaktualizowano kernel systemu,
  • wykonano awaryjny restart serwera,
  • przeprowadzono wymianę procesora,
  • wymieniono płytę główną,
  • wymieniono pamięć RAM,
  • wymieniono kartę sieciową,
  • przywrócono działanie usług,
  • rozpoczęto dodatkowe monitorowanie serwera.

Rekompensata

W związku z niedostępnością usług wszystkie dotknięte awarią usługi klientów zostały automatycznie przedłużone o 48 godzin.

Status

Incydent zakończony.

Po wymianie sprzętu usługi zostały przywrócone i serwer działa prawidłowo. Będziemy monitorować jego stabilność, aby upewnić się, że problem nie wystąpi ponownie.

Posted Aug 16, 2026 - 17:30 CEST

Resolved

This incident has been resolved.
Posted Aug 16, 2026 - 17:12 CEST

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Aug 16, 2026 - 17:02 CEST

Identified

After contacting the data center, we decided to replace the server with different hardware while retaining the hard drives.
Posted Aug 16, 2026 - 14:39 CEST

Investigating

We are currently investigating this issue.
Posted Aug 16, 2026 - 14:18 CEST

Monitoring

A fix has been implemented and we are monitoring the results.
Posted Aug 16, 2026 - 14:05 CEST

Update

We are continuing to investigate this issue.
Posted Aug 16, 2026 - 14:05 CEST

Update

We are continuing to investigate this issue.
Posted Aug 16, 2026 - 13:16 CEST

Investigating

We are currently investigating this issue.
Posted Aug 16, 2026 - 13:10 CEST
This incident affected: Nodes (n11.tosthost.pl).