Forum Discussion

pragsmike's avatar
pragsmike
Copper Contributor
Aug 27, 2026

KB5101650 / KB5121003: Thunderbolt 4 NVMe causes WHEA PCIe errors, bugcheck or hang, rollback fixes

Looking for anyone else seeing this since the July or August 2026 cumulative updates.

 

Hardware: ASUS ProArt Z790-CREATOR WIFI (BIOS 2801), Core i9-14900KF, Windows 11 25H2, Intel Thunderbolt 4 (JHL8540) with an OWC Thunderbolt Express 4M2 enclosure holding four NVMe SSDs. Intel Thunderbolt software 1.41.1412, which is ASUS's current release for this board.

 

Symptom: on build 26200.8875 (July update, KB5101650) and again on 26200.9168 (August update, KB5121003), within about a minute of connecting the enclosure the System log fills with WHEA-Logger event 17, corrected hardware error on the PCI Express root port behind the Thunderbolt controller. Shortly after, one of the enclosure SSDs is surprise-removed (disk event 157), and then the machine either bugchecks with IRQL_NOT_LESS_OR_EQUAL or hard-hangs with no bugcheck. In July it also produced DRIVER_POWER_STATE_FAILURE and a stornvme controller reset on an internal NVMe drive.

 

Rolling the cumulative back fixes it completely. Uninstalling KB5121003 with DISM (wusa refuses on this package) and returning to 26200.8655 with nothing else changed, the same enclosure on the same port runs clean: zero WHEA events over more than ten minutes including a 6 GB copy. I did the bad-then-good comparison on the same evening in one boot cycle, and got the same result in July with KB5101650.

 

Ruled out: the enclosure and SSDs (flawless on the older build, SMART clean), scheduled jobs, and any BIOS or driver change. The only variable is the cumulative update.

 

One possible connection: the resolved-issue note for the Dell / Intel IPF shutdown problem says it was caused by the new Windows USB-C Connection Manager interface introduced in the June 23 preview update (KB5095093). That change is inside my regression window and is not mentioned in that update's release notes. This looks like the same class of USB-C / Thunderbolt port-management regression on Intel Maple Ridge instead of Dell IPF.

 

I have filed this in Feedback Hub with minidumps and event log exports attached (I'll add the link in a reply so this post doesn't trip the spam filter). This is a production audio workstation with the working sample libraries in that enclosure, so I'm stuck on the June build with updates paused, without the August security fixes. The next preview (KB5120998) lists no related fix.

 

Reproducible on demand in under a minute. Has anyone else seen WHEA 17 or surprise removals on Thunderbolt-attached NVMe since July?

3 Replies

  • pragsmike's avatar
    pragsmike
    Copper Contributor

    Thanks both — and thank you Jamony for the careful reply.

     

    First, a housekeeping note. This thread is a shorter, link-free duplicate. My original post, with the full timeline, event tables, and evidence links, was auto-flagged as spam and has since been restored by the Tech Community team. Please treat that one as the canonical thread:

     

    https://techcommunity.microsoft.com/discussions/windows11/kb5101650kb5121003-hang-thunderbolt-4-nvme-%E2%86%92-whea-17%E2%86%92-bugcheck-0xa--z790-26200-8/4550187

     

    The Feedback Hub item, which carries the attachments (System/Setup event logs, the 233-event WHEA/disk/bugcheck export as CSV and XML, the July minidumps, and the sysinfo/msinfo export), is here — it is public and upvotes/"I have the same issue" help:

     

    https://aka.ms/AA139bx1

     

    On the technical point — agreed. WHEA-Logger 17 is a corrected PCIe error record, so the drip on root port 0000:00:1C.4 (8086:7A3C) is evidence of link instability, not proof of any particular cause. I named the USB-C Connection Manager change in KB5095093 as the best lead because it is the only documented change in the regression window (26200.8655 → .8875) that touches this subsystem and because Microsoft cited it for the acknowledged Dell/Intel-IPF issue; I do not claim it is the mechanism. What the A/B does establish is that the only variable between clean and failing is the cumulative update — same BIOS, drivers, enclosure, cables, and boot cycle.

     

    Versions, for the record (all already in the FH item): ASUS ProArt Z790-CREATOR WIFI, BIOS 2801; Intel JHL8540 TB4 controller (8086:1137) behind root port #5 (8086:7A3C, driver 10.1.46.3); Intel Thunderbolt software 1.41.1412 (ASUS's current WHQL release for this board); OWC Thunderbolt Express 4M2 with 2× Crucial P3 4 TB + 2× WD Blue SN5000 4 TB, all SMART-clean; Windows 11 25H2 — fails on 26200.8875 (KB5101650) and 26200.9168 (KB5121003), clean on 26200.8655.

     

    On the recommendations. I'll open a Microsoft support case referencing the Feedback Hub item and the good-build/bad-build comparison — that is the one channel I had not used. The rollback is containment, not a destination: updates are paused to a fixed date, and my plan for the September cumulative is an attended install with the enclosure detached, then a repeat of the same hot-attach reproduction — with PCIe Link State Power Management disabled first as a cheap discriminator, since a similar USB4/PCIe-tunnel freeze on another platform was worked around that way. Backups are on separate media and were re-verified after the rollback. The reproduction is repeatable on request if anyone at Microsoft or ASUS wants a specific trace captured.

  • This is a Thunderbolt 4 NVMe controller driver compatibility issue caused by the KB5101650/KB5121003 updates. You will need to roll back to the June version or wait for Microsoft to release a fix.

  • Your controlled rollback comparison is strong evidence of an update-sensitive regression, but WHEA event 17 is a corrected PCIe hardware-error report; it does not prove that the USB-C Connection Manager change is responsible. The surprise removals and bugchecks justify formal escalation.

     

    Keep the Feedback Hub submission and add the WHEA event XML, minidumps, PCI bus/device/function, msinfo32 export, BIOS version, Thunderbolt driver, and enclosure firmware. Install only ASUS-approved BIOS, chipset, and Thunderbolt updates plus current enclosure firmware, then repeat the same reproduction test. A Microsoft support case should reference the Feedback Hub item and bad-build/good-build comparison.

     

    Remaining on the known-good build with updates paused can be temporary containment, but it leaves security fixes unapplied. Avoid permanent rollback. Ask Microsoft or ASUS for a confirmed mitigation and retest the next servicing update before reconnecting production storage. Maintain a separate backup because surprise removals can corrupt active files even when SMART remains clean.