Contents
How to diagnose memory errors on AMD x86 _ 64 using EDAC?
This is a writeup I put together to help identify the defective DIMM from EDAC errors on linux x86_64. Over several years of managing a linux cluster I have occaisionally had systems with a bad memory DIMM. An early manifestation of these errors is EDAC errors (Error Detection and Correction kernel module) reported in the kernel ring buffer.
Why are there 2 channels in EDAC AMD64?
EDAC amd64: using x8 syndromes. This memory controller uses 8 chip select rows (MC 0-7) and with the current DIMM installation is showing 2 channels (DCT0 and DCT1). That is a confusing print out because the two characters, MC, are used in multiple places and seem to mean different things.
How is data accessed by the memory controller?
The data accessed by the memory controller is contained into one dimm only. E. g. if the data is 64 bits-wide, the data flows to the CPU using one 64 bits parallel access. Typically used with SDR, DDR, DDR2 and DDR3 memories.
How many DIMMs do I need for EDAC to work?
In order to make sure that you are interpreting the EDAC information correctly, you have to know the current actual DIMM setup. I have four 4GB DIMMS in the ‘A’ slots of each processor. That is a total of 16GB per processor and 64GB on the board. There are 16 DIMMS installed total.
How big is the memory controller in MC3?
This means that memory of one 4GB DIMM in slot 1A and one 4GB DIMM in slot 2A show up in two rows and two channels. For MC3, the csrow2 and csrow3 files contain the total size of the memory managed by this memory controller instance.
How is processor 2 served by MC2 and MC3?
Processor 2 is served by MC2 and MC3. Each of the DIMMS is ‘dual ranked’ which means that there are 2GB per ‘chip select row’ (csrow). A ‘rank’ corresponds to a populated csrow. Thus, these 4GB DIMMS show up in two csrows. The csrow2/ and csrow3/ directories contain the following files: managing.