All posts

Subvisor: Reading an OpenVMM Guest Without Asking It Anything

12 min read

Every memory forensics tool I have written starts at the same wall: you have the bytes, but the bytes do not explain themselves. DumpIt acquires the RAM; something still has to find the kernel inside it and name what it is looking at. On Windows that something is KdDebuggerDataBlock, the obfuscated structure a debugger decodes to bootstrap an analysis. On Linux, for a long time, it was “hope you have the matching System.map.”

Subvisor is a different take on the acquisition side. It is built on OpenVMM, Microsoft’s open-source, cross-platform VMM written in Rust, and it reads a guest from outside, through the VMM, with nothing running inside the guest. This post is about the offline half: a standalone tool that takes an OpenVMM snapshot and turns it into a core dump that crash and drgn open directly, recovering the guest kernel’s full symbol table from memory without a single byte of help from the guest.

Reading a VM from outside is not new, and I have some history with it. Back in 2010 I added the ability to debug a running Hyper-V virtual machine to Sysinternals LiveKd — Mark Russinovich wrote it up at the time — and that work grew into LiveCloudKd, which reconstructs a live kernel-debugging view out of a Hyper-V guest’s memory with no agent inside the guest. As far as I know that was the origin of what people now call virtual machine introspection. For the last several years Arthur “Gerhart” Khudyaev has carried LiveCloudKd forward, doing deep, sustained Hyper-V internals work on it that goes well past where I left off. Subvisor is the same idea pointed at a new VMM: inspect the guest from outside, and make sense of it without trusting it.

The mechanism that made host-side introspection work on Hyper-V is a nice bit of history in its own right. Reaching a guest’s memory went through Vid.sys, and its IOCTL handler trusted only the process that had created the partition — vmwp.exe, the VM worker process. So the whole game was getting your code to act as that process, which for a long time you simply could. Arthur has documented this boundary and how it moved over the years:

The interesting part is not the container plumbing. It is that the kernel already tells you everything you need, and almost nobody uses it.

The snapshot is three files#

OpenVMM’s snapshot format is its own, not WinDbg’s .vmrs. A snapshot directory is three files:

  • memory.bin — flat guest RAM, with the architectural MMIO holes squeezed out.
  • state.bin — device and vCPU state, encoded with mesh, Microsoft’s protobuf-compatible serialization.
  • manifest.bin — the guest size, architecture, and vCPU count.

memory.bin is not a flat image of the address space. OpenVMM packs RAM ranges contiguously and skips the holes, so a guest-physical address is not a file offset. The manifest records only the total size, so Subvisor reconstructs the layout per architecture and then refines it. The vCPU registers it actually needs — CR3 or TTBR1 for the page tables, IDTR for the Windows path — are buried in a tree of mesh Any messages inside state.bin, each tagged with a type URL. The tool walks that tree looking for VpSavedState by its URL, which keeps it independent of the VMM’s own type registry and means it reads snapshots from any OpenVMM fork.

Anatomy of an OpenVMM snapshot. memory.bin is flat guest RAM with the MMIO hole removed, so guest-physical addresses map to file offsets through a per-architecture layout, not one-to-one. state.bin is a tree of mesh Any messages; Subvisor walks it by type URL to pull each vCPU’s CR3/TTBR1 and IDTR without linking the VMM.

None of this is confidential-computing territory. With SEV-SNP or TDX the host sees ciphertext and this whole approach stops at the door; that is a job for a paravisor inside the trust boundary, which is where the engine is designed to go next. For an ordinary guest, the snapshot is plaintext and the only question is how to make sense of it.

VMCOREINFO is the Linux KDBG#

Here is the part worth the post. The Linux kernel maintains a block of text in RAM, written at boot, that describes itself for the benefit of crash-dump tools. It is called VMCOREINFO, it exists so that kdump works, and it is the Linux counterpart of KdDebuggerDataBlock — except it is plain ASCII and it is not obfuscated.

Scan a snapshot’s memory.bin for OSRELEASE= and you find it. No registers, no page tables, no guest cooperation. On the arm64 Linux 6.18 guest I tested, the block reads, in part:

OSRELEASE=6.18.33
BUILD-ID=b89f5b19c7c1dec1509a816bef994a3ba6c8525b
KERNELOFFSET=0
SYMBOL(swapper_pg_dir)=ffffffc08103c000
SYMBOL(_stext)=ffffffc080010000
SYMBOL(kallsyms_names)=ffffffc080d8e7f0
SYMBOL(kallsyms_token_table)=ffffffc080e8f660
SYMBOL(kallsyms_offsets)=ffffffc080e8fbc0
SYMBOL(kallsyms_relative_base)=ffffffc080edc6d8
NUMBER(kimage_voffset)=0xffffffc03ee00000
NUMBER(PHYS_OFFSET)=0x40000000

That is the whole bootstrap, laid out for you:

  • KERNELOFFSET defeats KASLR.
  • PHYS_OFFSET tells you where RAM starts, which calibrates the file-offset mapping and lets the page-table walker work.
  • swapper_pg_dir, combined with kimage_voffset, gives the page-table root as a physical address with no walk required — the same value the masked TTBR1 from state.bin points at.
  • The kallsyms_* symbols give the exact locations of the kernel’s own symbol tables.
  • BUILD-ID is the key to full debug info: for a distribution kernel you hand it to debuginfod and get the matching vmlinux back, exactly as an RSDS GUID fetches a PDB from Microsoft’s symbol server.

The Linux symbol-recovery bootstrap. A raw scan for OSRELEASE finds VMCOREINFO, which hands over the KASLR slide, the RAM base, the page-table root, and the addresses of the kallsyms tables. Subvisor decodes those tables from memory, then validates the result by checking the decoded _stext against VMCOREINFO’s own value before trusting the encoding.

Decoding kallsyms from memory#

With the table addresses in hand, Subvisor decodes the kernel’s compressed symbol table the way the kernel itself does. kallsyms_names is a stream of length-prefixed, token-compressed names; kallsyms_token_table and kallsyms_token_index are the dictionary; kallsyms_offsets and kallsyms_relative_base give each symbol’s address. The first character of every decoded name is its nm-style type letter.

The one genuine ambiguity is the address encoding. Modern kernels are base-relative, but some — x86-64 in particular — use an “absolute percpu” scheme where a non-negative offset is an absolute address and a negative one is relative_base - 1 - offset. Rather than guess from the architecture, Subvisor decodes _stext under both interpretations and keeps the one whose result equals VMCOREINFO’s SYMBOL(_stext). The kernel’s self-description becomes the oracle that validates the decode.

On the test guest this recovers all 78,534 symbols, matching the guest’s own /proc/kallsyms byte for byte:

$ subvisor info ./snap
architecture:   aarch64
memory:         1024 MiB
guest os:       Linux
kernel:         6.18.33
build-id:       b89f5b19c7c1dec1509a816bef994a3ba6c8525b
vmcoreinfo:     15 symbols
kallsyms:       78534 symbols decoded

This also closes a real weakness in the live side of the project. The live introspection worker currently takes a supplied /proc/kallsyms, which a compromised kernel could falsify. Symbols recovered from the kallsyms tables by walking the page tables cannot be forged the same way, and folding this path into the live worker removes the need to trust the guest for its own symbol list.

A core dump the usual tools understand#

Recovering symbols is most of the battle; emitting a container is the easy part. Subvisor writes a kdump-style ELF core: one PT_LOAD program header per RAM range, an NT_PRSTATUS note per vCPU, and — the important detail — the raw VMCOREINFO block copied verbatim into a VMCOREINFO note. crash and drgn read that note to locate symbols exactly as they do for a real vmcore, so the dump is self-describing.

$ subvisor dump ./snap -o guest.core
wrote ELF core dump to guest.core
open with: crash <vmlinux> guest.core
       or: drgn -c guest.core

The output validates as ET_CORE, EM_AARCH64, with the VMCOREINFO and NT_PRSTATUS notes in a PT_NOTE segment and a page-aligned PT_LOAD covering the full gigabyte of guest RAM. Nothing special is required to open it.

Windows, without VMCOREINFO#

Windows has no VMCOREINFO, so the bootstrap is the classic one. Subvisor takes the saved IDTR base, scans backward page by page for the MZ/PE headers of ntoskrnl, walks the PE debug directory to the CodeView RSDS record, and reports the PDB name and the symbol-server signature:

guest os:       Windows
kernel base:    0xfffff80000000000
pdb:            ntkrnlmp.pdb
signature:      <GUID><age>
symbol url:     https://msdl.microsoft.com/download/symbols/ntkrnlmp.pdb/<GUID><age>/ntkrnlmp.pdb

That is the identity a debugger uses to pull symbols from Microsoft’s server. A native .dmp writer — decoding KdDebuggerDataBlock the way LiveCloudKd does, into the dump rather than the live VM — is the next step on that side. The ELF path is Linux-only by design: WinDbg will open the container but will not understand a Linux kernel inside it, and crash and drgn will not understand nt.

Why bother#

Agent sandboxes are the immediate reason, and they are why I started looking at OpenVMM snapshots at all. nvx — Microsoft’s cross-platform micro-VM sandbox for agentic workloads, out of its Systems Research Group and Azure Research, built on OpenVMM and the Nanvix lineage — runs untrusted code such as coding agents and OCI images in microVMs and snapshots them constantly for fast restore. My friend Ryan MacArthur maintains an Apple Silicon fork of it, and brainstorming with him around that work is what sent me down this path. Once you are snapshotting a guest every few milliseconds to clone it, that same snapshot is a perfect forensic artifact — a frozen, complete, tamper-evident image of a guest at an instant. subvisor info is itself an integrity signal, because a snapshot whose VMCOREINFO, page tables, and kallsyms do not agree with each other has been tampered with. Diffing a clean template against a later checkpoint shows exactly what a workload changed.

This is arriving at the right moment, because the ground under “just run it in a VM” is shifting. Trail of Bits made the point sharply in VMs won’t contain cyber-capable agents (Artem Dinaburg, August 2026): their cyber-capable model escaped a QEMU/KVM sandbox three separate times, chaining bugs across QEMU, the Linux kernel and libslirp, and the conclusion was blunt — “an off-the-shelf VM is not enough to contain a modern, cyber-capable AI agent. There is simply too much attack surface.” It is not a hypothetical, either: over the same weekend I was writing this, Vercel’s Guillermo Rauch confirmed a KVM 0day found through their Sandbox bounty program — “affecting the industry’s gold standard solution for Linux virtualization.” The answer to too much attack surface is a smaller, attested trusted computing base, which is exactly the confidential-computing and paravisor direction these systems are moving toward — and introspection that lives inside the trust boundary stops being a nicety there and becomes the only way to see in at all.

The larger reason is where the engine is headed. The symbol recovery, the page walking, and the integrity checks are written against a tiny interface — read guest-physical memory, give me the page-table root — with no dependency on OpenVMM. That same engine runs host-side today, and it is meant to run inside a paravisor tomorrow, where it is the only way to inspect a confidential VM whose memory the host cannot read. VMCOREINFO works there unchanged, because it lives in the guest’s own encrypted RAM.

One engine behind a two-call interface — read guest-physical memory, and the page-table root. It runs three places: the offline snapshot tool and the live in-process worker today, and a paravisor inside the TEE next, which is the only vantage point that can read a confidential VM’s memory. The symbol bootstrap is identical in all three because VMCOREINFO lives in the guest’s own RAM.

That “read a privileged layer from outside it” pattern is already how people study Windows’ own paravisor-style isolation. A 2025 Radboud University master’s thesis on Windows Secure Kernel security bugs (Jonathan Jagt, Computest / Radboud, June 2025) uses LiveCloudKd as its Secure Kernel debugging setup: it drives Hyper-V, inspects VM memory directly, and recovers the Secure Kernel (VTL1) base address — which the kernel never exposes — to set breakpoints inside it. The thesis notes LiveCloudKd “appears to use an undocumented way of recovering the Secure Kernel virtual base address,” and it is pleasant to see a tool with these roots still being the one researchers reach for to look inside the most isolated layer of Windows. The confidential-VM case is the same shape: a layer the normal OS cannot see, read from a vantage point that can.

Notes and caveats#

This is early and experimental. It is verified end to end on an arm64 Linux guest under macOS and Hypervisor.framework. The x86-64 page walker and the Windows path are exercised by unit tests but have not been run against a real guest of their own yet. The tool ships with 30 unit tests over the protobuf reader, the VMCOREINFO parser, the kallsyms decoder, the ELF layout, and the PE/RSDS parser; the real-snapshot run is a manual check, since nobody wants a gigabyte of memory.bin in a git repository.

The hard limitation is confidential VMs, and it is worth being blunt about: the offline tool does not work on them, and cannot be made to. On SEV-SNP or TDX the guest’s RAM is encrypted with a key the hardware never hands to the host, so a host-side memory.bin is ciphertext, and the register state the tool reads from state.bin lives in the encrypted VMSA the host cannot inspect either. There is no host-side snapshot to analyze — not a format problem, a hardware boundary. The only vantage point that sees plaintext is inside the trust boundary: a paravisor at VTL2 (OpenHCL) or VMPL0 (COCONUT-SVSM). The engine is built to run there precisely because that is the sole place the work is possible, and the VMCOREINFO bootstrap carries over unchanged because it lives in the guest’s own encrypted RAM, which the paravisor can read and the host still cannot.

Building this also turned up a genuine bug in the project’s own page-table code: the AArch64 translator masked the guest-controlled TTBR1 in a way that preserved reserved sub-page bits, which could drive the page-table array index out of bounds and panic the host doing the scan. Fixed by page-aligning the root and bounds-checking the index — a reminder that anything reading guest-controlled registers is parsing hostile input, snapshot or live.

Acknowledgments#

This whole direction came out of conversations with Ryan MacArthur. His Apple Silicon fork of Microsoft’s nvx is what put OpenVMM snapshots in front of me in the first place, and Subvisor exists because of a long run of brainstorming sessions with him about snapshots, OpenVMM internals, and what trustworthy execution environments should actually guarantee. Ryan is one of those rare people who is equally sharp on the systems and the security of them, and he is generous with the hard half-formed ideas most people keep to themselves — the “what if the snapshot itself were the evidence” framing that anchors this post is his as much as mine. If you work on sandboxing, confidential computing, or VMMs, follow him; you will learn things. Thanks also to Arthur “Gerhart” Khudyaev, whose LiveCloudKd work is the standard this tries to live up to.