Controller, firmware, flash
What is actually inside an SSD. And why none of it behaves like a hard disk.
An SSD has no motor, no heads and no platters, so nothing inside it can be repaired with a screwdriver or a clean bench. What it has is a controller, a small processor; the firmware that runs on it; and a bank of NAND flash chips. Everything you think of as your files is really a lookup table the controller keeps to itself. When an SSD fails, it is almost always that table, or the chip that reads it, that has gone. The flash underneath is usually still there. This page says how, plainly, for anyone whose drive has just stopped and who wants to understand what the bench is about to do to it.
Rather talk it through? An engineer answers the bench line
0800 6890668
The controller and its firmware.
The controller is a system-on-chip that does everything between the connector and the flash. It handles error correction, encryption, wear levelling and the translation of each address your computer asks for into the physical page where the data actually sits. The firmware it runs lives partly in a small ROM on the chip and mostly in a reserved service area on the NAND itself. If that service area becomes unreadable, the controller cannot finish booting and falls back to its ROM. That is why a failed drive can suddenly report a factory name such as SATAFIRM S11 or SM2258XT, or a capacity of 0GB, 2MB or 8MB, or nothing at all.
The flash translation layer.
Flash cannot be overwritten in place. It is written in pages and erased in much larger blocks, so the controller never puts your data where the computer thinks it is. The flash translation layer, the map, is the record of where each logical block really lives, and it changes constantly as the drive shuffles data about. Lose the map and the data is still physically present, but as a scatter of pages with no index. Rebuilding that map is the core of SSD recovery.
NAND types.
SLC stores one bit per cell, MLC two, TLC three and QLC four. More bits per cell means cheaper capacity, slower writes and fewer program-and-erase cycles before the cell wears. Planar flash gave way to 3D NAND in the mid-2010s, with cells stacked in layers now well past two hundred deep. Most consumer drives today are TLC or QLC 3D NAND. QLC drives lean heavily on a fast cache, and once it is full, write speed can drop sharply. That is a trait, not a fault.
The SLC cache, DRAM and drives without it.
Most TLC and QLC drives run part of the flash in single-bit mode as a write buffer. Many drives also carry a DRAM chip to hold the map while running. DRAM-less NVMe drives borrow a slice of the computer's own memory instead, through the host memory buffer; DRAM-less SATA drives have no such option and keep more of the map on the flash, which is one reason budget SATA drives are over-represented among firmware failures. A power cut in the middle of a map update is one of the classic ways a consumer SSD loses everything at once, and consumer drives have no capacitors to finish the write. Enterprise drives do.
Wear levelling, TRIM and garbage collection.
Because each block survives only a limited number of erases, the controller constantly moves data to spread the wear. When you delete a file, the operating system sends a TRIM command on SATA or a Deallocate on NVMe, telling the drive those addresses no longer matter, and the drive erases them in the background whenever it has a moment. The guidance from the makers of the recovery equipment we use is that this takes anywhere from ten minutes to twenty-four hours. That is why deleted-file recovery on an SSD is measured in hours, and why the answer to a deletion is a full shutdown, this minute, and nothing else.
Why the bench uses diagnostic modes.
A failed SSD cannot be read by plugging it into another computer, because its own firmware is what has broken, and the computer can only talk to it through that firmware. Laboratory equipment talks to the controller in its factory diagnostic state instead, using the controller family's own commands, and loads a temporary copy of working firmware into the controller's memory rather than onto the flash. The temporary code stops the drive's own background erasing, reads the service area and rebuilds the map, and it disappears when the power is removed. Nothing is written to the flash, so nothing is lost by trying. Every controller family, Phison, Silicon Motion, Marvell, Maxio, InnoGrit and the in-house designs, has its own diagnostic state and its own way of laying out the service area, which is why the model number on the form saves a day.
Chip-off, and why it is the last resort.
Where the controller itself is dead, the flash chips are removed and read directly at their pads. That gives a raw dump in which the controller's scrambling, its interleaving across dies and channels, its XOR patterns and its error correction all have to be undone in software before the map can even be looked for. Lifting the chips takes an hour. Reading what comes off them is the part that takes the day, and it has to be learned family by family. On drives that encrypt in hardware, chip-off yields only ciphertext, and there the original controller has to be revived instead. Samsung, WD and SanDisk, SK hynix, Kioxia, Micron's own lines and every Apple drive encrypt in hardware whether or not you ever set a password.
Where that leaves your drive.
Nearly every SSD that reaches us has a firmware or controller fault with intact flash behind it, and nearly every one of those is recovered through the diagnostic state without a chip being touched. A smaller number need board repair first: a dead regulator after a surge, corrosion after a spill, a snapped connector. A smaller number still need chip-off. The free look tells you which yours is, and what it will cost, before anything chargeable happens.
The questions that come up first.
Can data be recovered from a dead SSD?
Usually. Most dead SSDs have a firmware or controller fault with the flash intact behind it, and the map is rebuilt in the controller's diagnostic state without writing to the flash.
What is chip-off recovery?
Reading the flash chips directly when the controller is dead, then undoing the controller's scrambling, interleaving and error correction in software. It is slower than a firmware fix, which can move a figure within its band, and it is impossible on drives that encrypt in hardware.
Why can't I just swap the controller from an identical drive?
Unlike a hard disk's board, an SSD's controller holds nothing that can be moved, and the map on the flash is specific to the way this controller laid it out. On hardware-encrypted drives the keys are bound to this controller too.
Is an SSD harder to recover than a hard disk?
Different rather than harder. A hard disk fails at its moving parts; an SSD fails at the controller or in the tables the controller keeps. So long as TRIM has not already been through the blocks you want, the outlook is reasonable.
Does any of this cost me anything to find out?
No. The free look identifies the controller, the firmware state and the health of the flash, and one figure follows in writing. £300 + VAT for one drive.
Now you know what the bench is about to do.
Send the form with the model number, and the first look tells you which of these your drive needs, and what it would cost.