Process Terminated Abruptly: How to Confirm OOM in Ubuntu 24.04
Check the kernel log, systemd-oomd, and cgroup limits to distinguish memory exhaustion from similar failures and preserve facts before changes.

In this article
Процесс пропал, служба перезапустилась, а в приложении осталась только запись о неожиданном завершении. Это похоже на нехватку памяти, но по одному признаку выводить рано: тот же внешний симптом дают ручной SIGKILL, лимит контейнера, решение systemd-oomd и сбой самого приложения.
Главный ориентир: OOM подтверждает не слово «killed», а связка времени, сообщения механизма завершения и контекста памяти. Сначала проверим журнал ядра, затем отдельно systemd-oomd и cgroup. Только после этого есть смысл обсуждать лимит, нагрузку или утечку.
Границы инструкции
Команды рассчитаны на Ubuntu 24.04 LTS с systemd 255 и cgroup v2. Они читают состояние и журналы, не перезапускают службы и не меняют лимиты. Синтаксис journalctl и systemctl выполнен 29 сентября 2026 года в изолированной среде Ubuntu 24.04.3 LTS; в ней не было сохранённых журналов и события OOM. Утилита oomctl на стенде отсутствовала, поэтому её команда сверена с руководством Ubuntu для systemd 255, но фактически не выполнялась. Намеренно вызывать OOM ради проверки на рабочем сервере не нужно.
1. Зафиксируйте загрузку и время события
Если после сбоя сервер перезагрузили, поиск только в текущей загрузке даст пустой результат. Сначала посмотрите, какие загрузки сохранил журнал.
journalctl --list-boots --no-pager
Номер 0 обычно относится к текущей загрузке, -1 — к предыдущей. Отсутствие предыдущей загрузки в списке не доказывает, что события не было: журнал мог храниться только в памяти, быть ротирован или находиться на стороне хоста. Запишите часовой пояс и узкий интервал вокруг сбоя, например десять минут.
2. Проверьте сообщения OOM ядра
Для события в текущей загрузке ограничьте одновременно источник, время и характерные сообщения. Даты ниже — пример: замените их своим интервалом.
journalctl -k -b 0 --since "2026-09-29 02:10:00" --until "2026-09-29 02:20:00" --grep='oom-kill|Out of memory|Killed process' --no-pager
Формулировки зависят от версии ядра, поэтому отсутствие совпадения по шаблону не исключает просмотр всего журнала ядра в том же интервале. Если записи есть, сохраните несколько строк до и после события, а не только имя завершённого процесса.
oom-kill:помогает увидеть контекст события: нехватку на всём узле либо ограничение внутри cgroup.Killed processуказывает выбранный процесс и снимок его памяти в момент решения, но не доказывает, что именно он создал первопричину.Упоминание cgroup или memcg связывает событие с локальным лимитом. Свободная память вне этой группы могла ещё оставаться.

3. Отделите systemd-oomd от OOM ядра
В Ubuntu процесс может быть завершён и пользователем пространства systemd-oomd. Он следит за давлением памяти в cgroup v2 и действует до классического глобального OOM ядра. Поэтому его журнал проверяют отдельно.
systemctl is-active systemd-oomd
journalctl -u systemd-oomd -b 0 --since "2026-09-29 02:10:00" --until "2026-09-29 02:20:00" --no-pager
Если служба активна, её запись о выбранной cgroup и времени — прямое направление для проверки. Текущий снимок наблюдаемых групп можно запросить отдельно, когда установлен пакет с oomctl.
oomctl --no-pager dump
Вывод oomctl показывает текущее состояние, а не восстанавливает историю уже произошедшего события. Неактивный или отсутствующий systemd-oomd также ничего не говорит о решениях OOM ядра.
4. Свяжите событие со службой и лимитом cgroup
Для службы приложения соберите её текущее состояние и параметры памяти. Вместо example.service подставьте реальное имя юнита.
systemctl show example.service -p ActiveState -p Result -p ExecMainCode -p ExecMainStatus -p OOMPolicy -p ControlGroup -p MemoryCurrent -p MemoryPeak -p MemoryMax
Значение Result полезно, но не исчерпывает картину: рабочий процесс мог быть завершён, а мастер остаться активным. Поля MemoryCurrent и MemoryPeak после перезапуска описывают уже новый период и не являются доказательством прошлого пика. MemoryMax нужно сопоставлять с cgroup, указанной в ControlGroup.
Если используется cgroup v2, дополнительные счётчики OOM относятся к конкретной группе. Сначала получите ControlGroup из предыдущей команды и проверьте её счётчики штатными средствами вашей среды: не угадывайте каталог службы и не ослабляйте права доступа ради чтения.
Поля oom и oom_kill показывают накопленные события для этой cgroup, а oom_group_kill — групповые завершения, если они учитываются ядром. У счётчиков нет времени события; после пересоздания cgroup они могут начать новый период. Надёжнее сравнивать сохранённые снимки с журналом и мониторингом.

5. Если подтверждения нет
Пустой результат не превращает гипотезу в опровержение. Проверьте журнал приложения и контейнерного рантайма, консоль или события провайдера, сохранность journald и правильную загрузку. Код завершения, связанный с SIGKILL, сам по себе не называет инициатора: это мог быть OOM-механизм, администратор или оркестратор.
Сопоставляйте один и тот же интервал с доступной памятью, активностью swap, давлением PSI, лимитом cgroup и числом параллельных задач. График после перезапуска не восстанавливает состояние до сбоя, а свободная RAM на узле не исключает локальный лимит контейнера или службы.
Когда остановиться
Журнал недоступен, а событие уже прошло: сначала сохраните доступные факты и запросите данные у владельца хоста.
OOM повторяется, но неизвестно, какие процессы можно ограничивать или перезапускать.
Для продолжения нужно менять
MemoryMax, swap, параметры ядра или политикуsystemd-oomd.Сбой затрагивает базу данных, очередь заказов или другой сервис, где принудительный перезапуск может ухудшить восстановление.
Подтверждённый OOM отвечает на вопрос, кто и в каком контексте завершил процесс. Он ещё не доказывает утечку памяти и не обещает, что добавление RAM решит проблему. Сохраните время, строки журнала, cgroup, лимит и наблюдаемую нагрузку; изменение выбирают уже по этой связке, с резервной копией и планом отката там, где оно затрагивает данные или доступность.
The process vanished, the service restarted, and the application only shows a record of an unexpected termination. This looks like memory exhaustion, but drawing conclusions from a single symptom is premature: the same external sign can result from a manual SIGKILL, a container limit, a systemd-oomd decision, or an application crash.
Key Indicator: OOM is confirmed not by the word "killed" alone, but by the combination of timestamp, termination mechanism message, and memory context. First, check the kernel log, then examine systemd-oomd and cgroup separately. Only then is it meaningful to discuss limits, load, or leaks.
Instruction Boundaries
The commands are designed for Ubuntu 24.04 LTS with systemd 255 and cgroup v2. They read the state and logs without restarting services or changing limits. The syntax for journalctl and systemctl was executed on September 29, 2026, in an isolated Ubuntu 24.04.3 LTS environment; it contained no saved logs and no OOM events. The utility oomctl was absent from the test stand, so its command was cross-referenced with the Ubuntu documentation for systemd 255 but was not actually executed. There is no need to intentionally trigger an OOM event for testing on a production server.
1. Record the boot and event time
If the server was rebooted after the failure, searching only the current boot will yield no results. First, check which boots the log has preserved.
journalctl --list-boots --no-pager
The 0 entry usually refers to the current load, while -1 refers to the previous one. The absence of a previous load in the list does not prove that the event did not occur: the log might have been stored only in memory, rotated, or located on the host side. Record the time zone and a narrow interval around the failure, for example, ten minutes.
2. Check for kernel OOM messages
For an event in the current boot, constrain the source, time, and characteristic messages simultaneously. The dates below are an example: replace them with your own time range.
journalctl -k -b 0 --since "2026-09-29 02:10:00" --until "2026-09-29 02:20:00" --grep='oom-kill|Out of memory|Killed process' --no-pager
Wording depends on the kernel version, so a lack of pattern match does not rule out reviewing the entire kernel log within the same interval. If records exist, save several lines before and after the event, not just the name of the terminated process.
oom-kill:helps reveal the event context: memory shortage across the entire node or a limit within a cgroup.Killed processindicates the selected process and a snapshot of its memory at the moment of the decision, but it does not prove that this process created the root cause.Mentions of cgroup or memcg link the event to a local limit. Free memory outside this group might still have been available.

3. Distinguish systemd-oomd from kernel OOM
In Ubuntu, a process may be terminated by the systemd-oomd user-space process. It monitors memory pressure in cgroup v2 and acts before the classic global kernel OOM. Therefore, its log must be checked separately.
systemctl is-active systemd-oomd
journalctl -u systemd-oomd -b 0 --since "2026-09-29 02:10:00" --until "2026-09-29 02:20:00" --no-pager
If the service is active, its record of the selected cgroup and timestamp provides a direct path for verification. A current snapshot of observed groups can be queried separately when the package with oomctl is installed.
oomctl --no-pager dump
The output of oomctl shows the current state, not a recovery of history for an event that has already occurred. An inactive or missing systemd-oomd also provides no information about kernel OOM decisions.
4. Correlate the event with the service and cgroup limit
For the application service, collect its current state and memory parameters. Replace example.service with the actual unit name.
systemctl show example.service -p ActiveState -p Result -p ExecMainCode -p ExecMainStatus -p OOMPolicy -p ControlGroup -p MemoryCurrent -p MemoryPeak -p MemoryMax
The value of Result is useful but does not tell the whole story: the workflow might have completed while the master remained active. Fields MemoryCurrent and MemoryPeak after a restart describe a new period and do not prove a past peak. MemoryMax must be matched with the cgroup specified in ControlGroup.
If cgroup v2 is in use, additional OOM counters apply to the specific group. First, retrieve the ControlGroup from the previous command and check its counters using standard tools in your environment: do not guess the service directory and do not weaken access permissions for the sake of reading.
Fields oom and oom_kill show accumulated events for this cgroup, while oom_group_kill shows group completions if accounted for by the kernel. Counters have no event timestamp; after a cgroup is recreated, they may start a new period. It is more reliable to compare saved snapshots with the log and monitoring data.

5. If there is no confirmation
An empty result does not turn a hypothesis into a refutation. Check the application and container runtime logs, the console, or provider events, the integrity of journald, and the correct loading. A termination code associated with SIGKILL does not identify the initiator on its own: it could be the OOM mechanism, an administrator, or an orchestrator.
Correlate the same time interval with available memory, swap activity, PSI pressure, cgroup limits, and the number of parallel tasks. A graph after a restart does not restore the state prior to the failure, and free RAM on the node does not rule out a local limit for the container or service.
When to stop
The log is unavailable and the event has already passed: first save the available facts and request data from the host owner.
The OOM condition repeats, but it is unknown which processes can be limited or restarted.
To proceed, you must change
MemoryMax, swap, kernel parameters, or the policy insystemd-oomd.The failure affects the database, order queue, or another service where a forced restart could worsen recovery.
A confirmed OOM event answers who and in what context the process terminated. It does not yet prove a memory leak, nor does it promise that adding RAM will solve the issue. Save the timestamp, log lines, cgroup, limit, and observed load; the change is selected based on this combination, with a backup and rollback plan where it affects data or availability.





Discussion 0
Share your experience and ask questions. Comments without links appear after editorial review.
No comments yet. Start the discussion.