長時間実行されているmlx5dumpプロセスが原因でAzure CVOが再起動する
環境
- Cloud Volumes ONTAP(CVO)
- Microsoft Azure
問題
- EMSログには、mlx5dumpプロセスの実行時間が長くなっていることが示されています:
Thu Jun 15 06:38:25 -0400 [cluster1-01: sched_monitor: mgr.stack.longrun.proc:notice]: Long running process: mlx5dumpThu Jun 15 06:43:28 -0400 [cluster1-01: sched_monitor: sk.hog.runtime:notice]: Process mlx5dump ran for 15569 milliseconds
- 上記のプロセスにより、ノードの再起動がトリガーされます:
Thu Jun 15 06:43:36 -0400 [cluster1-01: pha_main000: kern.shutdown.initiator:debug]: SK reboot was initiated by "maytag.ko::fm_handleReserved+763".Thu Jun 15 06:59:16 -0400 [cluster1-01: sfo_status: callhome.reboot.giveback:notice]: Call home for REBOOT (after giveback)
- ノードが繰り返し再起動し、コンソールに以下のエラーが表示されます。
mlx5_core2: ERR: mlx5e_ioctl:4600:(pid 0): tso6 disabled due to -txcsum6.mlx5_core2: ERR: mlx5e_ioctl:4622:(pid 0): enable txcsum6 first.e0c: Forced delayed initialize mlx5_core2 before network ifconfig calle0c: mlx5_core2 SIOCGRSSKEY failed: 22[node-01:netif.init.failed:ALERT]: Initialization of network interface mlx5_core2 failed due to unexpected software error mlx5_core err=0xffffffc4:100.