Failed: VAVA Portable SSD Touch 1TB (Maxio MAS0902A + Micron 29F2T08EMLCE)

Back in 2022, I decided to buy a pair of 1TB VAVA Portable SSDs for a pretty decent price. I cared little for brands, when it came to solid-state drives, since most vendors weren’t making the controllers or NAND themselves, so what mattered most was what was inside. For my <AU$100, I was getting Micron TLC-based NAND with an average, but common and serviceable Maxio MAS0902A DRAM-less controller in a nice USB enclosure with a SATA bridge. I wished it was NVMe, but beggars can’t be choosers.

On the whole, the two drives served well, except when one drive developed one uncorrectable sector on a data recovery job, making my life more painful than it had to be as I was chasing a deadline. That drive has since become the scratch SSD on which all my blog content “in-progress” resides and some of my smaller backups do as well. The other has been a holiday SSD which I take along with me to record a secondary back-up of my photos and collected experimental data.

Just three days before my recent trip to Singapore, I was busy analysing some data from my London trip, using this latter “holiday SSD” to do some scratch work. It was filled to the brim when suddenly … it froze. 100% utilisation, no data movement and it wasn’t hot enough to be thermally throttled. I gave it five minutes and then, by its own accord, it dropped offline. I cycled the USB connection … the device was detected, but my data was nowhere to be found.

That was when I knew that one of my blog posts I had intended to write prior to my trip to Singapore would not make it before I flew out … because I’ve just lost a day’s worth of work.

Its Behaviour

It’s not a “failure”, but a “learning experience”, right? Let me see what I can learn from this. The disk was detected but the partition didn’t show. Disk management said there was no partitions at all – was this just a case of a hosed partition table?

I grabbed the trusty TestDisk and … let it run for a bit. No dice. It had one partition, but TestDisk wasn’t finding it.

Examining the disk in my trusty hex editor, it became apparent it was behaving strangely. Everything was zeroes. It looks like a fresh surface … but try writing anything to it and it simply gobbles it up with no error. Read it back and you’ll find it’s still all zeroes!

In essence, the drive had become a write-only memory. But unlike the joke, this was not very funny (to me, at the time).

According to CrystalDiskInfo, the drive still had its capacity and its ID intact. Firmware version and serial are there … but the hours blocks written/read have been zeroed. Some of the cycle data may still be there – I believe that’s 13 average P/E cycles and 16 maximum P/E cycles, but I’m not sure how accurate it was. It definitely is a lightly-used drive.

This lit a fire underneath my butt – the failure was at an inconvenient time with only days to my impending departure, but even more inconvenient would be if the other “brother” drive were to fail with all my work-in-progress and critical documents. So I ended up eating up another day just to archive all the data on that other drive and back it up just-in-case.

For reference, it’s running on an SBC running Ubuntu – it’s SMART data looks like this:

smartctl 7.4 2023-08-01 r5530 [x86_64-linux-7.0.0-28-generic] (local build)
Copyright (C) 2002-23, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Device Model:     1TB SSD
Serial Number:    K50021R002023
LU WWN Device Id: 5 000000 000000000
Firmware Version: V1.3.0
User Capacity:    1,024,209,543,168 bytes [1.02 TB]
Sector Size:      512 bytes logical/physical
Rotation Rate:    Solid State Device
Form Factor:      2.5 inches
TRIM Command:     Available, deterministic, zeroed
Device is:        Not in smartctl database 7.3/5528
ATA Version is:   ACS-2 (minor revision not indicated)
SATA Version is:  SATA 3.2, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Mon Sep 21 10:24:09 2026 AEST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
AAM feature is:   Unavailable
APM feature is:   Unavailable
Rd look-ahead is: Enabled
Write cache is:   Enabled
DSN feature is:   Unavailable
ATA Security is:  Disabled, NOT FROZEN [SEC1]
Wt Cache Reorder: Unknown

=== START OF READ SMART DATA SECTION ===
SMART Status not supported: Incomplete response, ATA output registers missing
SMART overall-health self-assessment test result: PASSED
Warning: This result is based on an Attribute check.

General SMART Values:
Offline data collection status:  (0x02) Offline data collection activity
                                        was completed without error.
                                        Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (   33) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        (   2) minutes.
Conveyance self-test routine
recommended polling time:        (   2) minutes.
SCT capabilities:              (0x0031) SCT Status supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAGS    VALUE WORST THRESH FAIL RAW_VALUE
  5 Reallocated_Sector_Ct   PO--C-   100   100   050    -    0
  9 Power_On_Hours          -O--C-   100   100   000    -    36344
 12 Power_Cycle_Count       -O--C-   100   100   000    -    146
192 Power-Off_Retract_Count -O--C-   100   100   000    -    92
194 Temperature_Celsius     -O---K   027   027   000    -    27
199 UDMA_CRC_Error_Count    -O--C-   100   100   000    -    0
206 Unknown_SSD_Attribute   -O--CK   200   200   000    -    11
207 Unknown_SSD_Attribute   -O--CK   200   200   000    -    52
208 Unknown_SSD_Attribute   -O--CK   200   200   000    -    22
241 Total_LBAs_Written      -O--CK   100   100   000    -    9712
242 Total_LBAs_Read         -O--CK   100   100   000    -    16791
                            ||||||_ K auto-keep
                            |||||__ C event count
                            ||||___ R error rate
                            |||____ S speed/performance
                            ||_____ O updated online
                            |______ P prefailure warning

Device Statistics (GP Log 0x04)
Page  Offset Size        Value Flags Description
0x01  =====  =               =  ===  == General Statistics (rev 1) ==
0x01  0x008  4             146  ---  Lifetime Power-On Resets
0x01  0x010  4           36344  ---  Power-on Hours
0x01  0x018  6     20368904311  ---  Logical Sectors Written
0x01  0x020  6        36402682  ---  Number of Write Commands
0x01  0x028  6     35214558578  ---  Logical Sectors Read
0x01  0x030  6        87219451  ---  Number of Read Commands
0x01  0x038  6  32722325135644  ---  Date and Time TimeStamp
0x07  =====  =               =  ===  == Solid State Device Statistics (rev 1) ==
0x07  0x008  1               6  N--  Percentage Used Endurance Indicator
                                |||_ C monitored condition met
                                ||__ D supports DSN
                                |___ N normalized value

SATA Phy Event Counters (GP Log 0x11)
ID      Size     Value  Description
0x0001  2            3  Command failed due to ICRC error
0x0003  2            0  R_ERR response for device-to-host data FIS
0x0004  2            0  R_ERR response for host-to-device data FIS
0x0006  2            0  R_ERR response for device-to-host non-data FIS
0x0007  2            0  R_ERR response for host-to-device non-data FIS
0x0008  2            0  Device-to-host non-data FIS retries
0x0009  4            0  Transition from drive PhyRdy to drive PhyNRdy
0x000a  4            2  Device-to-host register FISes sent due to a COMRESET
0x000f  2            0  R_ERR response for host-to-device data FIS, CRC
0x0010  2            0  R_ERR response for host-to-device data FIS, non-CRC
0x0012  2            0  R_ERR response for host-to-device non-data FIS, CRC
0x0013  2            0  R_ERR response for host-to-device non-data FIS, non-CRC

It’s a bit more used, but showed no data loss or signs of failure despite being the drive that previously caused me stress on a data recovery job.

The Drive

I disassembled the drive to get a closer look inside.

On the black PCB is the promised MAS0902A controller and Micron 29F2T08EMLCE B28B FortisFlash NAND. Based on a visual inspection, the two pins nearest the screw hole appear to be the ones used for invoking the bootloader.

On the back is a serial number label – the numbers match the SMART data. So I wonder if the drive’s inability to write or read data is simply a failure of a whole NAND package such that maybe the firmware which may be stored redundantly is still intact, but the user data is “broken” irretrievably?

Using jm_id, it seems not much had changed – somehow this drive wasn’t making a fuss about its condition:

v0.25a
Drive: 2(USB)
Model: 1TB SSD                                 
Fw   : V1.3.0  
Size : 976762 MB [1024.2 GB]
IOCtl: Unlk failed 0x0!
Firmware id string[2D0]: MKSSD_201000004037850100,Jun  9 2020,14:40:13,DM9343,EBWOCAPC
Project id string[280] : e:/rd/dm9343/svn/update_code/w_b27_noraid
Controller             : MAS0902
IOCtl: ID2 failed 0x0!
Ch0CE0: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch1CE0: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch4CE0: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch5CE0: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch0CE1: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch1CE1: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch4CE1: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch5CE1: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch0CE2: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch1CE2: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch4CE2: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch5CE2: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch0CE3: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch1CE3: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch4CE3: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die
Ch5CE3: 0x2c,0xc3,0x8,0x32,0xe6,0x0,0x0 - Micron 96L(B27B) TLC 512Gb/CE 512Gb/die

All the flash memory still seems to be there?

Saving the SSD?

I had realised that I probably shouldn’t worry about the data that was on this failed SSD as it was still backed up on another SSD and I could simply repeat the data analysis. That would take time I didn’t yet have, but given the exorbitant cost of SSDs, if I could rescue the drive, I probably should.

I went looking around for tools and the only source I could find was from the venerable usbdev.ru.

Using the diagnostic utilities first, it seems that I could see some of the parameters used by the original manufacturer. It’s a little more detailed than jm_id.

It seems that bad blocks are not the cause of the issue – or perhaps the firmware hides it well – as everything is zero.

Not even the SMART data is any more illuminating compared to CrystalDiskInfo.

Without gleaning any new information, I reasoned that the SSD is already broken, the data has already been willfully sacrificed, so let’s just try re-manufacturing the SSD using the MP tools. What I did not expect was the complexity and relatively undocumented nature of the tools – after all, an end-user like myself isn’t really supposed to have them.

This is where I ran into a problem. I’m not the first to encounter this, but it seems the other person who did was not able to find a solution. Put simply, the first major issue is that the exact B27B type NAND used in this drive is not in the NAND lists of any of the released tools.

As a compromise, I decided to short out the loader jumper and try running MP anyway with a smaller B27B relative selected. Perhaps I’d get a drive 1/16th the capacity, but if I got anywhere, I’d still be making progress.

But then, I ran into lots of configuration issues. Turns out I wasn’t using the tool correctly!

Hidden in the Device Setting tab, I needed to load the Para.ini file corresponding to the type of memory in use. I tried the generic B27B ini file which helped get rid of that error … but instead …

… I now have an error that the required system firmware is not present. Checking the folder, indeed, that particular .bin file is missing. It’s listed in the bottom of the main window, but no matter how I change things, it seems it wants this particular file. It didn’t seem connected to the choice of NAND either. I tried using a different version of the tool that used a different .bin file, but those tools had no support for B27B at all.

So I did what someone who had nothing to lose might have done – I chose the closest-named .bin file and renamed it to what it wanted.

At last, a little progress. Now I’m getting a different fault – a GPIO fault. I suppose the PCB design may not be the “reference” design or it may not match whichever product this MP tool was programmed up for. Moving over to the “Test Items” tab …

… one could sneakily untick the “Check GPIO” tickbox to make it forget all about this test. Let’s try again …

It spends a few minutes grinding away. I tried several versions of the tool … but from here, things don’t exactly succeed.

Despite using the .exe file that has “for MK” in it, the tool complains of needing to use a tool “MK MPT” as if this version is not the right one. I suppose this tool does have “MX” all over it.

In desperation, I looked at the MPTOOL.ini file and used the INI Setting system to toggle a few things and replace every “MX” with “MK” in the vain hope it would change something. In the end, no dice here, so I decided to try a different version of the tool with a slightly different approach.

I asked myself the question – would it work if I tried running MP as if it had a B27A flash instead and disabled all checks. Turns out, it doesn’t either. While the tool seemed much happier to do MP for B27A, it’s likely there is some difference in interface or flash geometry that makes “pretending” to be a B27A to be a failure.

In desperation, I leafed through some of the other MP tool pages, but nothing stood out to me. It would seem that the only way out would be to get a (presumably older) MK MP tool that had the appropriate flash parameters for the 29F2T08EMLCE programmed in to give this SSD a second chance. Until then, it’s pretty much a joke item that swallows every write you throw its way.

Conclusion

While traditional hard drives might show some warning signs of distress prior to failure, my experience with SSD failures has been interesting to say the least. Some of them do start throwing read errors, while USB sticks and SD cards may go into a read-only mode which makes recovery rather straightforward. In this case, the failure mode is sudden and seemingly complete loss of access to the data, despite the drive still retaining its identity, capacity and most-likely, firmware. This was not my expectation, but perhaps this suggests that a whole NAND device (but not all) may have failed, perhaps critical data structures are damaged beyond repair and internal backups were not viable, or perhaps there might be a firmware bug?

Recertifying the drive via manufacturing tools ran into a critical hurdle – simply the lack of tool versions out there with support and firmware for this particular type of Micron B28B flash memory and something specific about this particular controller that demands “MK” tools rather than “MX” tools. Had this not been the case, perhaps there may have been a chance to restore the drive to operation, perhaps with reduced capacity.

This comes at a bit of a bad time, as AI has pushed the price of both magnetic hard drives and flash memory to much higher levels, making them relatively less affordable. Considering I purchased the VAVA 1TB drive back in 2022 for AU$87~93 which seemed a bargain given its TLC nature (albeit, being SATA-based), even questionable-brand SSDs of 1TB run for AU$200+ nowadays. For technology, which is famous for its rapid depreciation and obsolescence, this sort of pricing is very unusual. But four years of service isn’t entirely bad, even if it was mostly light service.

For SSDs, I’ve not been as worried about brands – after all, it’s usually the choice and configuration of controller plus the NAND type and grade that matter most, as long as the PCB design is competent and the firmware is stable. The barrier to entry is much lower compared to building hard drives as it’s more of an assembly operation, meaning that there may be space for lower-end competitors to bring down prices, but vertical integration in larger companies and supply shortages have mostly quashed this dream. Right now, it would seem, that I’ll just have to live with a 1TB sized “hole” in my storage portfolio for now …

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论