Floppy Emu Hardware Failure Analysis Results

Nobody enjoys troubleshooting non-working hardware, and I’m no exception. Every time I ask a contract manufacturer to assemble a batch of more Floppy Emus, there are always a few that don’t pass QA testing. What happens with those? For several years the accumulating QA failures have been sitting in a pile in the corner of my office, along with a few customer returns, all waiting for the day when I would dedicate time to their investigation. It was a long time coming, but “Failure Analysis Day” finally arrived, or more like Failure Analysis Month, and the results were pretty interesting.

Here’s a breakdown of the unique causes of failure that I identified, and the number of boards affected by each cause.

Hardware Issue Diagnosis Count
clock crystal / no clock 26
clock crystal / bad clock 16
bad CPLD chip 6
microcontroller not programmed 5
soldering problems 5
bad microcontroller chip 4
broken display header 1
missing component 1
cracked PCB 1
no issue, good board 1
unknown, couldn’t resolve 2


Clock Crystal

This was a strange failure that required a lot of time to track down initially, but once I learned to recognize the symptoms, I realized that most of my QA failures were due to this problem. I wrote about the clock crystal mysteries in more detail in a separate post last month. The short story is that during my most recent manufacturing batch, my normal supplier of clock crystals was out of stock. The contract manufacturer, with my approval, substituted a different crystal with the same specs. It shouldn’t have affected anything. But somehow it did.

26 of the QA failures appeared to have no functioning clock at all. The board utterly failed to do anything when the microcontroller clock source was changed to the external crystal. A further 16 displayed some level of function, but with erratic behavior or failures at higher clock speeds. Initially I wasn’t sure whether this was due to the newer crystals exposing some defect or fragility in the Floppy Emu design, or whether it was simply caused by defective or damaged crystals. Eventually I came to the opinion that the crystals were damaged by rough handling or overheating during assembly. Replacing the crystals and reprogramming the boards resolved all the issues.


CPLD Chip

The Xilinx CPLD chip on the Floppy Emu is a delicate flower that has long been a source of challenges. Prior to this year’s crystal-gate debacle, the CPLD was the single biggest source of failures that I’d observed. As the chip that’s directly connected to the Floppy Emu’s external interface, it bares the brunt of any static discharge or electrical stress. It’s a 5V-tolerant 3.3V part, but its 5V tolerance has sometimes seemed a bit questionable, at least in the way it’s used here.

Symptoms of a failed or bad CPLD can include disk emulation failures, overheating, or erratic behavior. Usually the device will still be functional and text appears on its display, but the disk features no longer work. In extreme cases the failed CPLD acts as a hard short-circuit from power to ground, and then nothing works. Replacing and reprogramming the CPLD resolved all of the problems with these boards.


Microcontroller not programmed

Amusingly, or depressingly depending on your perspective, the third leading cause of QA failures was that the microcontroller simply wasn’t programmed. Somebody fell asleep at the switch at the contract manufacturer, lost track of what they were doing, put a PCB in the wrong pile, or whatever. A board with an unprogrammed microcontroller will appear completely dead at first glance, but it still responds in the debugger and it only takes a few seconds to flash the chip and get everything working.


Soldering problems

Every component must be electrically bonded to the PCB with solder. Soldering problems can be tough to spot with the naked eye, but usually jump out under magnification, so one of my first troubleshooting steps is usually to look at a problematic board at 10x. I really should get a cool desktop microscope, but for the moment I’m using a cheap 10x jeweler’s loupe which works well enough.

A couple of boards had too much solder in places, resulting in a solder bridge that unintentionally connected two adjacent IC pins. But it was more common to find joints with insufficient solder or poor solder joints, where an IC pin was sort of resting on the PCB pad without actually bonding to it. Fortunately both problems were easy to fix with a soldering iron and a bit of flux.


Bad microcontroller chip

The onboard microcontroller trip is another potential source of failure. In my experience, these microcontrollers are pretty robust, and the only failures I have seen are caused in the field when customers accidentally connect the Floppy Emu cable backwards to their Apple II Disk II controller. This is distressingly easy to do, since the Disk II controller has bare pin connectors instead of a shrouded and keyed header. A backwards connection results in +12 and -12 volts applied to the mcu’s pins, killing it.

Unfortunately the symptoms of a bad microcontroller are nearly identical to the symptoms of a bad clock crystal: a completely unresponsive board, with no debugger activity. I had to review each non-responsive board’s history in order to guess which issue was at fault. In some cases I guessed wrong, and I ended up replacing the microcontroller, and then when that didn’t help, also replacing the clock crystal.


Other

The remaining issues were all one-offs. One board’s display header was physically broken and missing a pin, resulting in a blank display. Another board failed QA because the LED didn’t illuminate, except there was no LED! There was only a blank pad on the PCB. I also encountered a failure due to a PCB that was physically cracked, a long line running down the breadth of the PCB that could only be seen clearly under light from a specific angle. Two more boards defied my efforts to pinpoint the cause of their failures, and after spending too much time on them, I threw them into the scrap bin.

The very last board that I examined turned out to have no problems at all. I tested it extensively and it worked perfectly. This might have failed QA due to something external like a bad power supply or cable, or maybe it was simply miscategorized.


Final results

Of 68 Floppy Emus in the failure analysis heap, I managed to resuscitate 65 of them. That’s a pretty solid percentage! I learned to recognize the symptoms of certain failure causes, so I can address them faster if I see them again. More importantly, the failure analysis learnings (especially about clock crystals) will also help guide me in making design and assembly process changes, so I can reduce the number of future QA failures. Knowledge is power, as they say, and now I have lots of power.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论