## Reassessing with the new report

This report changes the picture substantially. Two things from it are decisive:

1. **The mpipe threads are idle.** That kills my earlier "crypto buffer pool exhaustion" theory outright — there's no queued work waiting on a `dm_target_crypt` buffer. The dm-crypt layer isn't holding anything back; `dd` through the same mapper device gets full throughput during the hang.
2. **The stuck threads are sleeping on `"vmiopg"` and `"vmpfw"`, not on anything crypt-specific.** I traced both tags in the actual source:

- `"vmiopg"` → `vfs_vmio_release()` in `sys/kern/vfs_bio.c`, in `getnewbuf()`'s buffer-recycling path. Comment right above it: *"We need to own the page in order to safely unwire it."* This is buffer-cache code calling `vm_page_busy_wait()` before it can call `vm_page_unwire()` on a page it's about to detach from a `struct buf`.
- `"vmpfw"` → ordinary `vm_fault()`, waiting for a page that another thread has busied, then retrying (`KERN_TRY_AGAIN`).

So both stuck threads are blocked the exact same way: waiting for a `vm_page_t`'s busy bit to clear, and nobody ever clears it. This is a **generic VM busy/wire accounting bug**, not something inside `dm_target_crypt.c` — that file was a red herring in my earlier analysis. `sys/kern/vfs_bio.c` isn't even touched by `8ec33efba7563efab10f0d90b5684480d94daf0a`.

### What I verified in the actual diff of 8ec33ef

I re-pulled the real diff (rather than trusting my memory from the cleared session) — here's what's solid:

**A genuine, confirmed accounting irregularity in `pmap_enter()`:**

```c
-	if (wired && (origpte & pmap->pmap_bits[PG_W_IDX]) == 0)
-		atomic_add_long(&pmap->pm_stats.wired_count, 1);
+	if (wired) {
+		atomic_add_long(&pmap->pm_stats.wired_count, 1);
+		if ((m->flags & PG_FICTITIOUS) == 0)
+			vm_page_wire(m);
+	}
```

Old code only bumped the *pmap-level statistic* when the PTE wasn't already wired, and never touched `m->wire_count` at all — callers wired the page themselves, once. New code calls `vm_page_wire(m)` **unconditionally, every single time `pmap_enter()` is invoked with `wired=1`**, even if that exact page at that exact VA was already wired. Since `vm_page_wire()` is a bare `atomic_fetchadd_int(&m->wire_count, 1)`, this is a true refcount — and `vm_fault()`'s success path calls `pmap_enter()` on **every fault**, passing `wired = fs.wflags & FW_WIRED`, which stays true for the *entire lifetime* of a wired map entry, not just the fault that first wires it:

```c
for (n = 0; n < mextcount; ++n) {
    pmap_enter(fs.map->pmap, vaddr + (n << PAGE_SHIFT),
               fs.mary[n], fs.prot | inherit_prot,
               fs.wflags & FW_WIRED, fs.entry);
}
```

So any *re-fault* of an already-wired entry (COW resolution, a racing double-fault from two LWPs, etc.) adds an extra, unmatched `wire_count` increment. This is real, but it only fires for entries with `wired_count != 0` — i.e. `mlock()`'d regions, or (historically) vkernel `VPAGETABLE` regions. It does **not** explain plain `read()`/`write()`/ordinary-mmap traffic from `cc1`/`as`/`make`, which never sets `FW_WIRED`. I checked `vm_map_entry_unwire_all()` too — it loops `while (entry->wired_count) vm_fault_unwire(...)`, and `wired_count` is a signed `int`; if it's ever over-decremented (possible if unwind logic double-counts somewhere) that loop never terminates, which would be a real hang — but again scoped to wired entries.

**I also checked `pmap_remove_pages()`** (used at process-exit teardown) to see if the bulk path skips the wire-decrement — it doesn't; it funnels through the same `pmap_remove_callback()` → `pmap_removed_pte()` that correctly calls `vm_page_unwire(m, -1)` when `PG_W` is set. So that specific "leak on process exit" theory I was chasing doesn't hold up.

**Bottom line on my code reading:** I found one confirmed, real bug (surplus wiring on refault of `mlock()`'d/wired entries), but I could not pin down, from static reading alone, the exact mechanism that leaves an *ordinary* buffer-cache page's busy bit permanently set with a busy vs. wire interaction. `vm_page_wire()`/`vm_page_unwire()` themselves are tightly asserted (`KKASSERT` on bad `wire_count`, panics rather than silently corrupting) — which argues that if this were a clean over/under-wire, you'd likely get a panic, not a silent hang. The fact that you're getting a silent hang instead points at something more specific: either the kernel isn't built with `INVARIANTS` (so the assertions that would catch this are compiled out and the corruption is silent), or the bug is a genuine **lost wakeup** rather than a wire-count mismatch — i.e., some path busies a page and takes an error/retry branch that forgets to call `vm_page_wakeup()`, independent of wire_count being right or wrong.

--------------------------

## Why the virtual machine likely can't reproduce it

**1. Workload mismatch — this is probably the biggest factor.** ...

**2. CPU count / real parallelism.** ...

**3. A concrete, code-grounded reason fast real hardware might matter specifically.** I went back into `dmtc_bio_read_decrypt()`:

```c
/*
 * Unconditionally copy in data. Never decrypt in place!
 * For reads with bogus page we can't decrypt in place as stuff
 * can get ripped out from under us.
 */
memcpy(data_buf, bio->bio_buf->b_data, bytes);
bio->bio_buf->b_error = dmtc_bio_encdec(priv, data_buf, bytes,
    bio->bio_offset, CRYPTOAPI_CIPHER_DECRYPT);
...
struct bio *obio = pop_bio(bio);
biodone(obio);
```

That comment — *"stuff can get ripped out from under us"* — is the developers themselves acknowledging the exact class of race I've been hypothesizing (a buffer-cache page being concurrently touched by page-fault machinery while I/O is in flight). More importantly for your question: in the common case (mpipe buffer available immediately), this whole sequence — a `memcpy`, the AES-NI decrypt of up to ~256KB, a second `memcpy`, then `biodone(obio)` — runs **synchronously, inline, in whatever context the underlying disk driver's completion callback runs in**. It's not deferred to a separate thread unless the mpipe pool is exhausted.

On a real NVMe drive, hardware completion latency is tens of microseconds. Splicing in a synchronous, CPU-bound decrypt-plus-two-memcpys of comparable magnitude *directly into that completion chain* is a substantial, consistent perturbation to when `biodone()`/`bufdone()` actually runs relative to everything else happening on that core. On a VirtualBox emulated disk, completion latency is typically much larger and jittery (often sub-millisecond to multi-millisecond) — the crypto work becomes proportionally tiny noise on top of already-dominant emulated-disk latency, and won't shift relative thread interleaving nearly as much. So: same kernel, same bug, but the *relative* timing perturbation that makes the race "easily repeatable" on real hardware may simply not exist in the VM's timing profile.

## Re-checked: does this resolve "why only encrypted disks" despite the wait being in generic `vfs_bio.c`?

Yes — and I think this was the missing piece in my last answer. The wait site being in generic, encryption-agnostic code (`vfs_bio.c`) is fully consistent with a bug that's **not in `dm_target_crypt` at all**, but in the shared VM/pmap wiring code, which dm-crypt merely exposes by changing I/O completion timing. That resolves the apparent contradiction you flagged: the bug lives in generic code because it's a generic race; encryption is required only as a timing perturber, not as a causal participant.
