Blackbird: Defeating PatchGuard, One Layer Deeper.
How do you hook the Windows kernel without actually patching it? Going one layer deeper with VMX, EPT shadow pages & process-specific kernel instrumentation.
Read articleHow do you hook the Windows kernel without actually patching it? Going one layer deeper with VMX, EPT shadow pages & process-specific kernel instrumentation.
How do you hook the Windows kernel without actually patching it? Going one layer deeper with VMX, EPT shadow pages & process-specific kernel instrumentation.
Read articleA long Windows Active Directory chain from LDAP injection and NTLM coercion through delegated ACL abuse, AD CS ESC3, protected-object inheritance repair, and S4U2Self U2U with RBCD.
Read articleWhat happens when an EDR trusts filenames, unauthenticated localhost traffic and world-writable kernel objects? Five vulnerabilities, two SYSTEM LPE's and an RCE, a driver load and a lot of assumptions that should never have crossed a security boundary.
Read articleA complete Windows attack chain through OAuth account-linking abuse, stored browser navigation, SQLite extension loading, DPAPI credential recovery, and a writable SYSTEM service binary.
Read articleSysWhispers, HellsGate, HeavensGate, SidewaysGate, SpoofGate, TFGate, DoomGate, whatever gate your tool is being detected before the initial handle fully opens. How do EDR's detect & deny direct and indirect syscalls?
Read articleHow does Blackbird make Windows lie to malware's faces? Most anti-analysis checks trust the kernel because they have no choice. Blackbird weaponizes this by modifying syscall, timing & registry return data, erasing VM-identifiers and much, much more.
Read articleThere's a reason security products avoid kernel hooking. They are fragile, build-sensitive, and BSOD prone. Advanced malware analysis demands the visibility they provide. This post delves into the hook engine behind Blackbird and the struggle developing it.
Read articleModern defensive tooling doesn’t need to see payloads to stop you, it only needs to see the call path. This post breaks down how Windows system calls are intercepted, how syscall stubs became signatures, and why ActiveBreach takes a fundamentally different approach.
Read articleThis post is pending an update as I have made significant changes since it was made.
My earlier post, Blackbird: Doing What EDRs Won't, covered the first kernel hook engine. It used inline hooksA detour replaces the first instructions of a function with a jump to another handler. The displaced instructions are preserved so the original function can still run. inside ntoskrnl.exeThe Windows kernel executable. It contains core operating-system logic and the kernel implementations of native Nt and Zw routines..
Which, if you've ever touched kernel development, or rootkits, you'd know is a big no-no in PatchGuard's eyes. ntoskrnl.exe is the core Windows kernel executable that contains a huge amount of the functionality that allows Windows to operate, for example the Memory Manager, scheduler, and the kernel implementations of native system services.
Blackbird was placing inline hooks directly inside functions such as ntoskrnl.exe!NtAllocateVirtualMemory, modifying the executable code of the kernel itself. As you could imagine, allowing arbitrary kernel code to modify critical kernel routines would be a pretty horrible security model, especially considering those routines sit directly in paths used to service requests coming from usermode, and these paths are HOT.
So Microsoft implemented something called PatchGuard, AKA KPP (Kernel Patch Protection). I'm not going to get into exactly how KPP works, but the general idea is that KPP periodically validates protected kernel code and critical structures against expected state to detect tampering and, you guessed it, Blackbird tampers with exactly that. If KPP detects a violation, it'll execute a bugcheck (blue screen), usually 0x109 (CRITICAL_STRUCTURE_CORRUPTION).

So this clearly wasn't a long-term solution, especially if I wanted a trustworthy, stable platform to analyze malware on. A bandaid fix would be trying to figure out how KPP works, then patch it out (horrible idea), the normal solution would be to stop trying to intercept actual syscalls, and leverage already supported Windows systems, such as kernel callbacks and Etw-Threat-Intelligence. I'm not normal, and my satisfaction for learning wasn't fulfilled just yet so I started researching other ways to get the amount of observability I wanted.
So I came across a concept called SLAT (Second-Level Address Translation), where a memory page can effectively have two views: the Guest-Page (GP) and the Hypervisor-Page (HP). The GP and HP can differ in their memory protections (eg; the GP can be RWX, while the HP is RX). Under SLAT, the guest's page tables translate guest-virtual addresses to guest-physical addresses, while the hypervisor's second-level page tables translate those guest-physical addresses to host-physical addresses. If something then attempts to write to an address inside that GP while the corresponding second-level mapping is RX, a VM-EXIT can occur due to the second-level page-table violation, triggering a control-flow transfer into the hypervisor, which can then proceed to do further analysis & handling in VMXROOT. The key different is that SLAT never requires a byte inside of ntoskrnl.exe, or any kernel module to be patched, meaning that KPP has nothing to flag.
Now, I'm not going to get into how hypervisors work. The point of this article is how I managed to get the Kernel visibility I wanted without interfering with KPP or causing system instability, so, MOVING ON!
So now I know what SLAT is, but what am I going to target, and how am I going to do this without creating an extremely slow OS? If I put a second-level address translation on each ntoskrnl.exe!Nt* function, a VMEXIT would fire on every... single... API call, which requires the CPU to save-state, transfer control into the hypervisor, then eventually VM-enter the guest again. Imagine that happening across extremely hot kernel paths.
My first idea was naive, "What if I just check the PID on every single access?". Which would still mean a VMEXIT on every call, then checking the PID, deciding and returning, inside of these hotpaths + the full VM-exit/VM-entry overhead.
Then I thought of checking the CR3 before deciding to VMEXIT. Control Register 3 (CR3) identifies the root of the paging structures for the currently active virtual address space. On x86-64 Windows, this generally means CR3 points to the physical base of the process's top-level page table, the PML4, with some additional paging-related bits such as the PCID when enabled. Since each process has its own address space, its CR3 can therefore be used as an identifier for which address space is currently active. That address space also contains the kernel mappings available while that process is executing in kernel mode.
Instead Archangel, the Blackbird hypervisorA privileged layer below the guest operating system that controls virtual processors, memory translation, and selected hardware events.—uses Intel Extended Page Tables (EPT)The Intel implementation of Second Level Address Translation. EPT adds a hypervisor-controlled translation from guest physical memory to the actual machine page. to give selected processes an executable shadow view of that page.

This, is what Archangel ended up doing.
So, the syscall itself is completely normal. ntdll.dll loads the SSN, executes syscall, Windows switches into kernel-mode and eventually reaches the implementation inside ntoskrnl.exe. I did not want to mess with IA32_LSTAR or redirecting the actual syscall entry point. Normally the guest page tables take us from GVA -> GPA, then EPT takes that GPA -> HPA. Windows controls the first translation, Archangel controls the second. With that control I can have NtAllocateVirtualMemory resolve to the exact same virtual address Windows expects, through the exact same PTEs Windows expects, while changing what physical page the hypervisor controls.
Archangel keeps a monitoring policy containing the processes Blackbird cares about, their CR3's, PCID's (see "3" on the diagram), and the relevant EPT configuration. If the currently executing address-space belongs to a Blackbird target, Archangel selects the EPT view containing the shadowed kernel pages, otherwise it'll get the normal EPT and no VMEXIT will occur. Meaning that things like PatchGuard, the memory manager, IO, whatever else will continue executing the original ntoskrnl.exe pages.
Thankfully, almost everything from my old hook engine was still useful. It was never inherently broken, just unusable due to KPP. It already resolved the Nt* routines (weirdly via SSDT walk), decoded instruction boundaries, calculated the overwrite length, built trampolines and knew how to safely enter Blackbird's handler.
Instead of modifying the actual page containing the target routine, Blackbird now allocates another page, aligns it, copies the original kernel page into it and gets the physical address of both;
rawAllocation = ExAllocatePool2(
POOL_FLAG_NON_PAGED,
BK_NTAPI_SLAT_SHADOW_ALLOC_SIZE,
BK_NTAPI_SLAT_SHADOW_POOL_TAG
);
if (rawAllocation == NULL)
return NULL;
shadowPage = BkntkiSlatAlignShadowAllocation(rawAllocation);
RtlCopyMemory(
shadowPage,
(PVOID)kernelPageVa,
PAGE_SIZE
);
livePa = MmGetPhysicalAddress((PVOID)kernelPageVa);
shadowPa = MmGetPhysicalAddress(shadowPage);
The patch is written into shadowPage.
If NtAllocateVirtualMemory and another hooked routine happen to live inside the same 4KB page, creating two separate shadow pages would make absolutely no sense. Blackbird therefore keeps one shadow for the page and applies each required redirect at its respective offset inside that copy.
The hook itself currently uses R11:
mov r11, <Blackbird hook handler>
jmp r11
The hook preparation still validates complete instruction boundaries like before, and it refuses targets where the redirect would walk over the end of the page. Once Blackbird finishes preparing the page it provides Archangel the original GPA and shadow PA through a hypercall. Archangel can now make the same GPA resolve to either the untouched Windows page or Blackbird's modified copy, instead of running every single system call via Blackbird's handler and slowing down the OS.
Reading the EPT entries from kernel-space would reveal Blackbird's hook handler. My first EPT entry for the shadow page was readable + executable (RX), which is standard until something decidesto actually inspect the kernel code. If that occurs, It'd read the Blackbird redirect sitting in the prologue, even if the real kernel page wasn't touched. What I actually wanted was for the guest kernel to not know what the shadow page contained. To do this I could leverage EPT permissions, the shadow mapping is normally execute-only, the CPU can fetch instructions from it without involing a VMEXIT. But, if code attempts to read the same GPA as data, an EPT violation occurs, and control transfers to Archangel. See where this is going?
Archangel sees the read, changes that GPA back to the original physical page and gives it read + execute permission:
if (ReadAccess != FALSE && hookRoot != nullptr) {
const UINT64 clean = EPT_ENTRY_READ | EPT_ENTRY_EXECUTE;
if (EptSetExisting4KbLeafMappingForRoot(
hookRoot,
gpa,
originalPa,
clean) != FALSE &&
EptQueuePendingSlatRestore(
gpa,
shadowPa) != FALSE &&
VmxSetMonitorTrapFlag(TRUE) != FALSE) {
EptInvalidateLocalContext();
return TRUE;
}
}
Though, simply changing the EPT lead to the clean page is insufficient. This is because when the VMEXIT happens the faulting instruction has not completed, so if you change mappings the CPU is mid-execution on the shadow page, then is told to execute the normal page. I needed exactly one instruction to execute against the temporary clean mapping. To do this, I could use RFLAGS.TF (trap-flag for the CPU), but this is visible to the guest architecture, generates #DB and involves me writing guest exception/debug machinery.
Intel has something called the "Monitor Trap Flag", or MTF for short, which exists explicitly at the VMX-root layer. Using the MTF, I could catch EPT violations on data-read via VMEXIT, replace the shadow Host Physical Address (HPA) with the original authentic HPA and record that this GPA needs restoration. Then save the permissions needed by the faulting op, enable the MTF, resume the CPU without advancing the RIP (return-instruction-pointer). At this point the CPU will retry and complete the (yes, retry, the RIP cannot advance at this stage), and the MTF will trigger a VMEXIT, restore the shadow mapping, then disable MTF again, and invalidate the affected EPT context.
case VMX_EXIT_REASON_MONITOR_TRAP_FLAG:
{
EptRestorePendingSlatHooks();
VmxSetMonitorTrapFlag(FALSE);
EptInvalidateLocalContext();
return FALSE;
}
There was still another edge case here though, because the instruction causing the read violation could itself be executing from the same page I'm currently trying to swap out.
EPT permissions apply to the entire 4KB mapping, not to whichever 13 bytes I happened to overwrite for the hook, so if RIP is currently somewhere inside the shadow page and that instruction then reads data from another address inside the same page, Archangel gets the EPT violation while the CPU is still executing code from that mapping. Blindly replacing the GPA with the original HPA at that point means I'm changing the page underneath the instruction which caused the exit, then asking the CPU to resume and retry it against a different physical page, which is obviously something I don't want happening inside random kernel code.
So before doing the normal clean-page + MTF path, Archangel also checks where the guest RIP currently resolves. If RIP belongs to the same GPA page which caused the violation, it has to preserve the executable view long enough for that instruction to finish instead of treating it like a normal read coming from somewhere else.
MTF only really solves the "how do I expose this page for one instruction, then immediately take it away again?" part, it doesn't make swapping the page containing the current instruction magically safe.
Writes were easier, Blackbird has no reason to let anything modify these pages, and letting writes through would also mean deciding how to synchronize the original page with the shadow while keeping Blackbird's redirect intact, so Archangel just records the attempt and refuses to propagate it into either copy.
At this point the basic policy for one of these pages is pretty simple;
Execute -> shadow page
Read -> clean page temporarily + MTF restore
Write -> VMEXIT + reject
Which finally gave me a page that could execute Blackbird's version of the function while still looking completely normal when something tried to inspect it.
There was one fairly massive problem left though, the hook EPT worked for everybody.
If Archangel left the shadow mappings active globally, any process reaching one of those kernel pages would execute Blackbird's redirect, which sort of defeats the whole point of trying to make this selective in the first place. I could've just checked the PID inside Blackbird's handler again like the old hook engine did, but now I'm already sitting underneath the guest paging layer with the ability to decide which physical page gets executed before the handler is ever reached, so doing the filtering afterwards made absolutely no sense.
This brought me back to the CR3 idea from earlier.
Archangel keeps a normal EPT root where kernel GPAs resolve to the real Windows pages, then another EPT root containing the shadow mappings Blackbird needs. When the guest address-space belongs to something Blackbird is monitoring, Archangel switches that logical processor onto the hook EPT, otherwise it stays on the normal one.
The selection path is basically;
GuestCr3 &= AAHV_CR3_ADDRESS_MASK;
selectedEptp = EptGetIdentityMapPointer();
for (LONG i = 0; i < targetCount; ++i) {
if (EptTargetRequiresHookRoot(
&g_ArchangelTargets[i],
GuestCr3)) {
selectedEptp =
EptGetHookMapPointerForCurrentProcessor();
break;
}
}
if (currentEptp != selectedEptp) {
__vmx_vmwrite(
EPT_VMCS_EPT_POINTER,
selectedEptp
);
}
The CPU obviously isn't automatically picking an EPT based on CR3 here, Archangel is watching the guest address-space state & using that as the policy input for which EPTP should be active.
CR3 also isn't literally a process identifier, it's just extremely convenient for what I care about because it points at the paging structures backing the currently active address-space. The base address is the useful part here, and depending on the configuration you've also got things like PCID information living in the low bits, so raw CR3 comparisons without masking or tracking that state correctly are a great way to create bugs that only happen sometimes.
Then Windows reminded me KPTI exists.
My first implementation stored what I thought was the CR3 for a target process, I'd start the target, see the correct CR3, enter usermode, execute a syscall and then the hook would randomly disappear as soon as execution crossed into the kernel.
The process obviously hadn't stopped being monitored in the five instructions between syscall and kernel entry, so I started logging the guest CR3 across the transition and, sure enough, it was changing.
With KPTI, the same process can operate with a user-side paging context and a separate kernel-side one when it transitions into kernel execution. The exact behaviour depends on the Windows build/configuration I'm running, but for Blackbird's environment the important part was that "one process == one CR3" was a bad assumption.
So Archangel's monitoring state ended up tracking both where required, along with the PCID/address-space information I need to identify them properly;
typedef struct _AAHV_TARGET_ADDRESS_SPACE
{
UINT64 UserCr3;
UINT64 KernelCr3;
UINT16 UserPcid;
UINT16 KernelPcid;
} AAHV_TARGET_ADDRESS_SPACE;
Then the actual comparison can operate on the paging structure address instead of pretending the entire raw CR3 value is some globally unique process token;
static BOOLEAN
EptTargetRequiresHookRoot(
const AAHV_TARGET_ADDRESS_SPACE* Target,
UINT64 GuestCr3
)
{
const UINT64 cr3Base =
GuestCr3 & AAHV_CR3_ADDRESS_MASK;
return
cr3Base == Target->UserCr3 ||
cr3Base == Target->KernelCr3;
}
I'm skipping a bunch of boring bookkeeping around process lifetime, stale entries & CR3 reuse here, but that's the important part of the mechanism.
Now if a Blackbird target transitions into kernel-mode, the processor is using the hook EPT and the relevant kernel GPA resolves to the shadow page. If some unrelated process hits the exact same routine, it gets the normal EPT and executes the original Windows page.

That was finally the behaviour I had in my head when I started this whole thing.
Performance was still the thing I was most worried about because a design like this can very easily become technically cool and completely unusable at the same time.
A VMEXIT isn't cheap, the processor has to leave guest execution, save whatever state VMX requires, enter VMXROOT, let Archangel handle the event and then VM-enter back into the guest again. Doing that across every syscall would be horrific, especially considering some of the routines Blackbird monitors are called constantly.
Thankfully, executing the shadow page itself does not cause an EPT violation.
Once the hook EPT is selected, the CPU can just execute that mapping normally. Windows translates the virtual address to the same GPA it always has, then EPT translates that GPA to Blackbird's shadow HPA, the processor fetches the modified instructions and NtAllocateVirtualMemory reaches the redirect without Archangel having to get involved again.
So the common execution path is still basically;
GVA
|
| Windows page tables
v
GPA
|
| Archangel EPT
v
Shadow HPA
|
v
Blackbird hook
No VMEXIT is required just because that last translation points at a shadow page.
The annoying exits are mainly the cases where something violates the policy around those mappings, data reads against an execute-only page, attempted writes, EPT/view changes when the active address-space changes and the MTF exit used to restore a temporary clean mapping.
Still overhead obviously, but nowhere near firing a VMEXIT every time a monitored process calls an API.
I also didn't want one modified 4KB page to destroy the mapping strategy for an entire region. EPT can use larger mappings, and if the only thing I care about is one kernel page sitting inside a 2MB region, splitting absolutely everything for no reason would just increase translation overhead & page-table memory usage.
So Archangel only splits the region containing a page which actually needs separate permissions or a different backing physical address;
if (EptLeafIs2MbLargePage(entry)) {
if (EptSplit2MbMapping(
root,
gpa) == FALSE) {
return FALSE;
}
entry =
EptLookup4KbLeaf(
root,
gpa
);
}
After that split, the one 4KB leaf I care about can point at the shadow HPA while the rest of the region keeps behaving normally.
Changing the EPT entry itself is only half the job though because the processor can cache translations derived from EPT, meaning that writing a new PFN into the leaf does not guarantee the CPU immediately stops using whatever translation it cached previously.
Which matters quite a bit when the entire system depends on me switching;
GPA -> Original HPA
to;
GPA -> Shadow HPA
and expecting that change to actually happen.
So after changing one of those mappings Archangel invalidates the relevant EPT context with INVEPT. I also don't want to blindly invalidate every translation on every change, so the invalidation is scoped around whichever EPT context was actually modified.
Something roughly like;
VOID
EptInvalidateContext(
UINT64 EptPointer
)
{
INVEPT_DESCRIPTOR descriptor = { 0 };
descriptor.EptPointer =
EptPointer;
__invept(
INVEPT_SINGLE_CONTEXT,
&descriptor
);
}
The SMP side of this is where it gets slightly more annoying, because if multiple logical processors are using the same EPT context then a mapping change isn't necessarily local to whichever CPU happened to modify the entry. Any processor which can still hold a cached translation for that context needs to observe the invalidation correctly, otherwise one CPU can be executing against a completely different view of the page than another one.
A lot of the temporary state around this system is per-processor for the same reason.
The pending MTF restore is probably the easiest example, say CPU 0 gets a read violation for GPA A, temporarily exposes the clean mapping and records that A needs to be restored, then CPU 1 gets a completely unrelated violation for GPA B before CPU 0 receives its MTF exit. If that restore state is global, CPU 1 just overwrote CPU 0's entry and now CPU 0 is about to restore the wrong page.
So each logical processor gets its own state for things like the active EPT root and pending SLAT restoration;
typedef struct _AAHV_PROCESSOR_CONTEXT
{
UINT64 CurrentEptp;
struct
{
BOOLEAN Active;
UINT64 Gpa;
UINT64 OriginalPa;
UINT64 ShadowPa;
} PendingSlatRestore;
} AAHV_PROCESSOR_CONTEXT;
I also don't silently overwrite an already-pending restore, if Archangel reaches a state where it wants to queue another one while the previous one still exists, something has gone wrong and continuing execution with corrupt translation state is probably going to make the eventual crash ten times harder to understand.
Hypervisor bugs have a habit of doing that.
The title says defeating PatchGuard, but Archangel never actually attacks PatchGuard.
I don't locate it, I don't patch it, I don't disable any of its checks & I don't need to care when it happens to run.
The original ntoskrnl.exe physical page is still there, Blackbird never writes the redirect into it and all of the normal Windows execution contexts continue using that copy. The only thing Archangel changes is the second-stage translation used for the address-spaces I explicitly decided to monitor.
So Windows still does;
GVA -> GPA
exactly like it normally would, with the same virtual addresses and the same guest PTEs, then Archangel gets the final say over;
GPA -> HPA
for whichever EPT hierarchy is currently active.
Conceptually;
Windows
GVA
|
v
GPA
/ \
/ \
v v
Identity EPT Hook EPT
| |
v v
Original HPA Shadow HPA
I should probably be careful saying KPP has "nothing to flag" because PatchGuard is intentionally undocumented, changes between Windows versions and I don't pretend to know every single internal path it can ever use.
What actually matters for Blackbird is that I'm no longer persistently modifying the kernel text which caused the original 0x109 bugchecks. The canonical Windows page stays untouched, unrelated execution keeps seeing that page, and the instrumentation exists in a second-stage view which only the monitored execution context receives.
This doesn't magically make shitty hooks safe either.
If the trampoline relocates something incorrectly, Windows will still crash, if I screw up VMCS state, Windows will still crash, if I restore the wrong EPT entry, Windows will probably crash somewhere completely unrelated twenty milliseconds later and make me question every decision I've ever made.
Hypervisors gave me another layer to put the instrumentation in, they didn't suddenly make kernel development forgiving.
But I wasn't fighting CRITICAL_STRUCTURE_CORRUPTION anymore.
After getting this working I had the fairly obvious question, if I can do this inside Blackbird, why don't EDRs just do the same thing?
Some security products absolutely use virtualization, this isn't some brand new concept I invented in my bedroom, but putting your own hypervisor underneath arbitrary customer Windows machines creates a completely different set of problems compared to doing it inside an analysis environment you control.
Windows itself may already be using the virtualization extensions for Hyper-V, VBS, HVCI/Memory Integrity, Credential Guard, nested virtualization or whatever else is enabled on that machine. Archangel wants direct control over VMX and EPT, Windows may also want direct control over VMX and EPT, and that becomes significantly less fun when your product has to coexist with whatever configuration the customer happens to be running.
There are supported interfaces for software which needs to operate alongside the Windows hypervisor, but that's a different architecture from Archangel directly owning the VMCS, EPT hierarchy & VMX operation.
There's also the hardware side.
Archangel is currently Intel-specific, so everything in this article is talking about VMX + EPT. AMD has the equivalent concepts through SVM + NPT, but supporting both means maintaining two virtualization backends with different control structures, different exit behaviour and a whole new pile of edge cases.
Then you've got the compatibility matrix.
An EDR vendor has to deal with whatever Windows build, firmware, CPU generation, BIOS configuration, OEM driver collection, power-management implementation, security product & general pile of nonsense happens to exist on the customer's endpoint.
Blackbird doesn't.
Blackbird runs inside an analysis environment I control, I know which Windows image is installed, I know which CPU features are exposed, I know whether nested virtualization is available, I know which drivers are loaded & if Archangel requires VBS/HVCI to be disabled inside that VM, I can just make that part of the image configuration.
Telling millions of customers to reconfigure Windows around your hypervisor would be insane.
Doing it inside a malware analysis appliance is completely reasonable.
That's really the difference for me, virtualization-backed monitoring itself isn't the interesting discovery here, the fact Blackbird runs in a controlled environment means I can actually use the design without needing to support the entire Windows ecosystem at once.
I also didn't want my performance testing to consist of "well the desktop still feels responsive", so Archangel keeps counters around the exits I actually care about;
typedef struct _AAHV_EXIT_STATISTICS
{
volatile UINT64 Total;
volatile UINT64 EptViolations;
volatile UINT64 MonitorTrapFlags;
volatile UINT64 AddressSpaceChanges;
volatile UINT64 ReadReveals;
volatile UINT64 RejectedWrites;
} AAHV_EXIT_STATISTICS;
Then the common exit path records what actually caused Archangel to run;
++Processor->Stats.Total;
switch (exitReason)
{
case VMX_EXIT_REASON_EPT_VIOLATION:
++Processor->Stats.EptViolations;
break;
case VMX_EXIT_REASON_MONITOR_TRAP_FLAG:
++Processor->Stats.MonitorTrapFlags;
break;
}
The numbers I care about aren't just how many syscalls happened, I want to know how many hook executions happened without a VMEXIT, how many reads caused temporary clean-page exposure, how many MTF restorations occurred, how often the EPT view actually changed & whether unrelated processes pay any measurable cost just because Archangel is active.
If the exit count starts scaling directly with every syscall Blackbird observes, then I've basically rebuilt the design I was trying to avoid in the first place.
Obviously none of this is universal.
Archangel currently assumes Intel VMX + EPT, it assumes a Windows environment I control, the CR3/KPTI handling is built around the configurations I actually test, and the hook engine still has all of the normal problems an inline hook engine has once execution reaches the shadow page.
Instruction decoding still has to be correct, RIP-relative operands still need to resolve to the same effective addresses when they're relocated into a trampoline, relative calls/branches still need fixing up, the redirect needs enough complete instructions at the target and I still refuse cases where that overwrite would run across the 4KB page boundary.
Hooks sharing the same physical page also need to share the same shadow, Windows updates can change the implementation/layout of routines I'm targeting, CR3 values aren't permanent process identities, EPT invalidation has to be correct across processors & any bug in VMXROOT has a considerably larger blast radius than a bug in some random usermode telemetry DLL.
So yeah, there are plenty of ways to break this.
Just fewer ways involving PatchGuard.
This ended up being a much larger detour than I expected when the old hook engine started throwing 0x109.
Initially I thought the solution would be learning enough about PatchGuard to survive it, finding whichever checks were catching me, figuring out how they worked & trying to coexist with them somehow, which would've just turned Blackbird into a permanent game of chasing Windows internals every time Microsoft changed something.
The actually useful part of the old hook engine was never the fact it modified ntoskrnl.exe, it was everything around that modification, resolving the routines I cared about, decoding instructions, building trampolines, capturing useful arguments & feeding that information back into the analysis system.
Archangel let me keep all of that while moving the physical modification somewhere Windows doesn't normally depend on.
The monitored process still calls the same Nt* routines through the same syscall path, Windows still uses the same virtual addresses, the guest page tables still produce the same GPAs and the original kernel pages are still sitting there completely untouched.
Blackbird just gets control over the last translation for the execution contexts I care about.
Which is significantly nicer than trying to convince PatchGuard that modifying the Windows kernel is somehow fine.
Special thanks to: