Alta carga de CPU, baixo uso do núcleo, erro de memória (ECC) no kernel

1

Estou tendo um comportamento super estranho ... A carga da CPU do meu computador passa pelo telhado (> 4 em uma máquina 8core), mas não há processo que esteja levando muito CPU (consulte a imagem anexada) Embora o núcleo da máquina esteja com carga alta estar entre 30-70% de oscilação.

Esse comportamento aparece após X minutos de uso do computador (aleatório, variando de alguns minutos a algumas horas). Além disso, depois que isso aconteceu, o computador acabará por congelar.

Estou com perda aqui, tive esse problema em 15.04, atualizado para 15.10, mesmo.

A máquina tem essas partes: Placa-mãe: Asus Z10PE-D8WS CPU: CPU Intel (R) Xeon (R) E5-1620 v3 a 3,50 GHz RAM: 2x Kingston 16Go PC4-2133 CL15 - Registo ECC (KVR21R15D4 / 16) HDD: 2x 2Para ATA ST2000DM001-1ER1 no RAID 0

A única coisa estranha que encontrei foram aquelas linhas no log do kernel:

Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17386.894665] CMCI storm detected: switching to poll mode
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.299974] EDAC MC0: 4 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x1042 offset:0x100 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.299989] EDAC MC0: 4 CE memory read error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x85392b offset:0xa80 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:1 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.299999] EDAC MC0: 2 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x850da9 offset:0x580 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300009] EDAC MC0: 3 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x85f599 offset:0x100 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300018] EDAC MC0: 3 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x11b2 offset:0x780 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300022] EDAC MC0: 2 CE Error at MMIOH area, on addr 0x000000087fd43a40 on any memory ( page:0x0 offset:0x0 grain:32 syndrome:0x0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300032] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x8474e2 offset:0xf00 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300042] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x8476f8 offset:0xd80 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300051] EDAC MC0: 2 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x8466eb offset:0x500 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300060] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x846b23 offset:0x7c0 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300070] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x846b23 offset:0xcc0 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:1 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300080] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x846d32 offset:0xe40 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:1 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300089] EDAC MC0: 2 CE memory read error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x5c251b offset:0x640 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:1 rank:1)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300099] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x8474e3 offset:0x1c0 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:02 XXXX-Z10PE-D8-WS kernel: [17387.300108] EDAC MC0: 1 CE memory read error on CPU_SrcID#0_Ha#0_Chan#1_DIMM#0 (channel:1 slot:0 page:0x847711 offset:0xf40 grain:32 syndrome:0x0 -  area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:2 rank:0)
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17387.891537] EDAC sbridge MC0: HANDLING MCE MEMORY ERROR
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17387.891561] EDAC sbridge MC0: CPU 0: Machine Check Event: 0 Bank 7: cc08388000010090
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17387.891566] EDAC sbridge MC0: TSC 0 
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17387.891569] EDAC sbridge MC0: ADDR 87fc60500 EDAC sbridge MC0: MISC 14032b286 
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17387.891576] EDAC sbridge MC0: PROCESSOR 0:306f2 TIME 1455579963 SOCKET 0 APIC 0
Feb 15 18:46:03 XXXX-Z10PE-D8-WS kernel: [17388.299184] EDAC MC0: 8418 CE Error at MMIOH area, on addr 0x000000087fc60500 on any memory ( page:0x0 offset:0x0 grain:32 syndrome:0x0)
Feb 15 18:51:03 XXXX-Z10PE-D8-WS kernel: [17687.707744] CMCI storm subsided: switching to interrupt mode

com essas linhas repetindo muito

Feb 15 19:07:47 XXXX-Z10PE-D8-WS kernel: [18691.236569] EDAC sbridge MC0: HANDLING MCE MEMORY ERROR
Feb 15 19:07:47 XXXX-Z10PE-D8-WS kernel: [18691.236586] EDAC sbridge MC0: CPU 0: Machine Check Event: 0 Bank 7: cc00064000010090
Feb 15 19:07:47 XXXX-Z10PE-D8-WS kernel: [18691.236589] EDAC sbridge MC0: TSC 0 
Feb 15 19:07:47 XXXX-Z10PE-D8-WS kernel: [18691.236592] EDAC sbridge MC0: ADDR 103fb00 EDAC sbridge MC0: MISC 4062e286 
Feb 15 19:07:47 XXXX-Z10PE-D8-WS kernel: [18691.236597] EDAC sbridge MC0: PROCESSOR 0:306f2 TIME 1455581267 SOCKET 0 APIC 0

espaçados por alguns

Feb 15 19:07:48 XXXX-Z10PE-D8-WS kernel: [18692.381405] EDAC MC0: 26415 CE memory read error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x1042 offset:0xa00 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0001:0090 socket:0 ha:0 channel_mask:1 rank:0)
Feb 15 19:07:48 XXXX-Z10PE-D8-WS kernel: [18692.381481] EDAC MC0: 4 CE memory scrubbing error on CPU_SrcID#0_Ha#0_Chan#0_DIMM#0 (channel:0 slot:0 page:0x7c5acf offset:0x0 grain:32 syndrome:0x0 -  OVERFLOW area:DRAM err_code:0008:00c1 socket:0 ha:0 channel_mask:1 rank:1)

Ajuda!

    
por Xqua 16.02.2016 / 01:42

2 respostas

2

Obrigado por me lembrar de terminar isso!

De fato, depois de olhar as linhas, notei que: slot: 0 foi o problema. Assumindo que era uma memória ruim, eu peguei (Slots são alocados pela sua placa-mãe, ou pelo menos no meu, o slot zero era o slot 1 da placa-mãe)

Assim eu retirei, testei por 48 horas, e nenhum erro apareceu. Enviou a memória RAM para garantia, conseguiu um novo de volta.

Tudo é perfeito no país das maravilhas!

    
por Xqua 24.04.2016 / 21:54
4

Ainda está rastreando esse problema? Parece que você tem um módulo de memória ruim, a máquina faz uma pausa apenas aguardando o hardware para corrigir esse erro por si só. Pode ser necessário tentar remover ou substituir a memória na sua primeira CPU, segundo canal e primeiro slot. Consulte: link

Espero que ajude.

    
por Shisoft 23.04.2016 / 20:13