| # Lightweight Fault Isolation (LFI) in LLVM |
| |
| ## Introduction |
| |
| Lightweight Fault Isolation (LFI) is a compiler-based sandboxing technology for |
| native code. Like WebAssembly and Native Client, LFI isolates sandboxed code in-process |
| (i.e., in the same address space as a host application). |
| |
| LFI is designed from the ground up to sandbox existing code, such as C/C++ |
| libraries (including assembly code) and device drivers. |
| |
| LFI aims for the following goals: |
| |
| - Compatibility: LFI can be used to sandbox nearly all existing C/C++/assembly |
| libraries unmodified (they just need to be recompiled). Sandboxed libraries |
| work with existing system call interfaces, and are compatible with existing |
| development tools such as profilers, debuggers, and sanitizers. |
| - Performance: LFI aims for minimal overhead vs. unsandboxed code. |
| - Security: The LFI runtime and compiler elements aim to be simple and |
| verifiable when possible. |
| - Usability: LFI aims to make it as easy as possible to retrofit sandboxing, |
| i.e., to migrate from unsandboxed to sandboxed libraries with minimal effort. |
| |
| When building a program for the LFI target the compiler is designed to ensure |
| that the program will only be able to access memory within a limited region of |
| the virtual address space, starting from where the program is loaded (the |
| current design sets this region to a size of 4GiB of virtual memory). Programs |
| built for the LFI target are restricted to using a subset of the instruction |
| set, designed so that the programs can be soundly confined to their sandbox |
| region. LFI programs must run inside of an "emulator" (usually called the LFI |
| runtime), responsible for initializing the sandbox region, loading the program, |
| and servicing system call requests, or other forms of runtime calls. |
| |
| LFI uses an architecture-specific sandboxing scheme based on the general |
| technique of Software-Based Fault Isolation (SFI). LLVM currently supports LFI |
| for the AArch64 and X86-64 platforms. The AArch64 version is designed to |
| support the Armv8.1 AArch64 architecture. |
| |
| See [https://github.com/lfi-project](https://github.com/lfi-project/) for |
| details about the LFI project and additional software needed to run LFI |
| programs. |
| |
| ## Compiler Requirements |
| |
| When building for an LFI target (`aarch64_lfi` or `x86_64_lfi`), the |
| compiler must restrict use of the instruction set to a subset of instructions, |
| which are known to be safe from a sandboxing perspective. To do this, we apply a |
| set of simple rewrites at the assembly language level to transform standard |
| native assembly into LFI-compatible assembly. |
| |
| These rewrites (also called "expansions") are applied at the very end of the |
| LLVM compilation pipeline (during the assembler step). This allows the rewrites |
| to be applied to hand-written assembly, including inline assembly. |
| |
| ## Context Register |
| |
| Both architectures designate a context register that points to a block of |
| thread-local memory managed by the LFI runtime. The context register is `x25` |
| on AArch64 and `r15` on X86-64. The layout is as follows: |
| |
| | Offset | Size | Description | |
| | --- | --- | --- | |
| | 0 | 8 | Reserved for future use. | |
| | 8 | 8 | Reserved for use by the LFI runtime. | |
| | 16 | 8 | Virtual thread pointer (used for TP access). | |
| |
| ## Linker Support |
| |
| In the initial version, LFI only supports static linking, and only supports |
| creating `static-pie` binaries. There is nothing that fundamentally precludes |
| support for dynamic linking on the LFI target, but such support would require |
| that the code generated by the linker for PLT entries be slightly modified in |
| order to conform to the LFI architecture subset. |
| |
| ## Assembler Directives |
| |
| The following directives are supported for controlling the rewriter. |
| |
| ### `.lfi_rewrite_disable` |
| |
| Disables LFI assembly rewrites for all subsequent instructions, until |
| `.lfi_rewrite_enable` is used. This can be useful for hand-written assembly |
| that is already safe and should not be modified by the rewriter. |
| |
| ### `.lfi_rewrite_enable` |
| |
| Re-enables LFI assembly rewrites after a previous `.lfi_rewrite_disable`. |
| |
| Example: |
| |
| ```gas |
| .lfi_rewrite_disable |
| // No rewrites applied here. |
| ldr x0, [x27, w1, uxtw] |
| .lfi_rewrite_enable |
| ``` |
| |
| ## Compiler Options |
| |
| The LFI target has several configuration options, specified via `-mattr=`: |
| |
| - `+no-lfi-loads`: Disable sandboxing for load instructions (stores-only mode). |
| - `+no-lfi-stores`: Disable sandboxing for store instructions. |
| |
| Use `+no-lfi-loads` to create a "stores-only" sandbox that may read, but not |
| write, outside the sandbox region. |
| |
| Use `+no-lfi-loads,+no-lfi-stores` to create a "jumps-only" sandbox that may |
| read/write outside the sandbox region but may not transfer control outside |
| (e.g., may not execute system calls directly). This is primarily useful in |
| combination with some other form of memory sandboxing, such as Intel MPK. |
| |
| ## AArch64 |
| |
| The AArch64 LFI target is `aarch64_lfi`. This is the first part of a target |
| triple that can be used with `--triple=aarch64_lfi-<rest of triple>`. |
| |
| ### Reserved Registers |
| |
| The AArch64 LFI target uses a custom ABI that reserves additional registers for |
| the platform. The registers are listed below, along with the security invariant |
| that must be maintained. |
| |
| - `x27`: always holds the sandbox base address (must be aligned to the size |
| of the sandbox). |
| - `x28`: always holds an address within the sandbox. |
| - `sp`: always holds an address within the sandbox. |
| - `x30`: always holds an address within the sandbox. |
| - `x26`: scratch register. |
| - `x25`: context register (see [Context Register](#context-register)). |
| |
| The current design only supports 4GiB sandboxes, which requires the sandbox |
| base address to be 4GiB-aligned. This is because LFI's ABI stores pointers as |
| their full 64-bit values, rather than just 32-bit offsets from the base. This |
| enables stores-only mode, where loads are not sandboxed but stores are, and |
| allows the host to directly pass pointers to the sandbox. |
| |
| ### Assembly Rewrites |
| |
| #### Terminology |
| |
| In the following assembly rewrites, some shorthand is used. |
| |
| - `xN` or `wN`: refers to any general-purpose non-reserved register. |
| - `{a,b,c}`: matches any of `a`, `b`, or `c`. |
| - `LDSTr`: a load/store instruction that supports register-register addressing modes, with one source/destination register. |
| - `LDSTx`: a load/store instruction not matched by `LDSTr`. This covers load/store pairs (`ldp`/`stp`), SIMD load/stores (`ld1`, `st1`, ...), atomics, exclusives, load/store-release, and unscaled (`ldur`/`stur`) forms. These instructions have a more limited set of addressing modes than `LDSTr`. |
| |
| #### Control flow |
| |
| Indirect branches get rewritten to branch through register `x28`, which must |
| always contain an address within the sandbox. An `add` is used to safely |
| update `x28` with the destination address. Since `ret` uses `x30` by |
| default, which already must contain an address within the sandbox, it does not |
| require any rewrite. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| {br,blr,ret} xN |
| ``` |
| - ```gas |
| add x28, x27, wN, uxtw |
| {br,blr,ret} x28 |
| ``` |
| * - ```gas |
| ret |
| ``` |
| - ```gas |
| ret |
| ``` |
| ::: |
| |
| #### Memory accesses |
| |
| Memory accesses are rewritten to use the `[x27, wM, uxtw]` addressing mode if |
| it is available, which is automatically safe. Otherwise, rewrites fall back to |
| using `x28` along with an instruction to safely load it with the target |
| address. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| LDSTr xN, [xM] |
| ``` |
| - ```gas |
| LDSTr xN, [x27, wM, uxtw] |
| ``` |
| * - ```gas |
| LDSTr xN, [xM, #I] |
| ``` |
| - ```gas |
| add x28, x27, wM, uxtw |
| LDSTr xN, [x28, #I] |
| ``` |
| * - ```gas |
| LDSTr xN, [xM, #I]! |
| ``` |
| - ```gas |
| add xM, xM, #I |
| LDSTr xN, [x27, wM, uxtw] |
| ``` |
| * - ```gas |
| LDSTr xN, [xM], #I |
| ``` |
| - ```gas |
| LDSTr xN, [x27, wM, uxtw] |
| add xM, xM, #I |
| ``` |
| * - ```gas |
| LDSTr xN, [xM1, xM2] |
| ``` |
| - ```gas |
| add x26, xM1, xM2 |
| LDSTr xN, [x27, w26, uxtw] |
| ``` |
| * - ```gas |
| LDSTr xN, [xM1, xM2, MOD #I] |
| ``` |
| - ```gas |
| add x26, xM1, xM2, MOD #I |
| LDSTr xN, [x27, w26, uxtw] |
| ``` |
| * - ```gas |
| LDSTx ..., [xM] |
| ``` |
| - ```gas |
| add x28, x27, wM, uxtw |
| LDSTx ..., [x28] |
| ``` |
| * - ```gas |
| LDSTx ..., [xM, #I] |
| ``` |
| - ```gas |
| add x28, x27, wM, uxtw |
| LDSTx ..., [x28, #I] |
| ``` |
| * - ```gas |
| LDSTx ..., [xM, #I]! |
| ``` |
| - ```gas |
| add x28, x27, wM, uxtw |
| LDSTx ..., [x28, #I] |
| add xM, xM, #I |
| ``` |
| * - ```gas |
| LDSTx ..., [xM], #I |
| ``` |
| - ```gas |
| add x28, x27, wM, uxtw |
| LDSTx ..., [x28] |
| add xM, xM, #I |
| ``` |
| * - ```gas |
| LDSTx ..., [xM1], xM2 |
| ``` |
| - ```gas |
| add x28, x27, wM1, uxtw |
| LDSTx ..., [x28] |
| add xM1, xM1, xM2 |
| ``` |
| ::: |
| |
| #### Stack pointer modification |
| |
| When the stack pointer is modified, we write the modified value to a temporary, |
| before moving it back into `sp` with a safe `add`. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| mov sp, xN |
| ``` |
| - ```gas |
| add sp, x27, wN, uxtw |
| ``` |
| * - ```gas |
| {add,sub} sp, sp, {#I,xN} |
| ``` |
| - ```gas |
| {add,sub} x26, sp, {#I,xN} |
| add sp, x27, w26, uxtw |
| ``` |
| ::: |
| |
| #### Link register modification |
| |
| When the link register is modified, it is guarded back into the sandbox with a |
| safe `add x30, x27, w30, uxtw`. This guard is deferred until the next |
| control-flow instruction rather than emitted immediately after the |
| modification. Deferral keeps a signed return address intact so that a following |
| authentication instruction (such as `autiasp`) can run before the guard, |
| which would otherwise destroy the pointer authentication signature. See |
| [Pointer Authentication Code (PAC) support](#pointer-authentication-code-pac-support). |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| ldr x30, [...] |
| ret |
| ``` |
| - ```gas |
| ldr x30, [...] |
| add x30, x27, w30, uxtw |
| ret |
| ``` |
| * - ```gas |
| ldp xN, x30, [...] |
| ret |
| ``` |
| - ```gas |
| ldp xN, x30, [...] |
| add x30, x27, w30, uxtw |
| ret |
| ``` |
| ::: |
| |
| #### Pointer Authentication Code (PAC) support |
| |
| LFI is compatible with Arm Pointer Authentication Code (PAC) instructions, |
| which are used to sign and authenticate `x30` to protect against control-flow |
| hijacking. |
| |
| The typical use is `-mbranch-protection=pac-ret`, which signs only the return |
| address in `x30` using the hint-space `paciasp` and `autiasp` |
| instructions. The combined authenticate-and-branch and authenticate-and-return |
| instructions covered below require Armv8.3-a and are not produced by |
| `-mbranch-protection`. They can appear in hand-written assembly or from |
| environments that sign all code pointers, so the rewriter still sandboxes them |
| rather than passing them through unmodified. |
| |
| To gain the security benefit of PAC under LFI, the hardware must implement |
| `FEAT_FPAC`, so that authentication failures fault immediately. Without |
| `FEAT_FPAC`, a failed authentication produces a poisoned pointer, which LFI |
| still keeps confined to the sandbox by masking it, but the mask overwrites the |
| poison caused by the authentication failure. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| paciasp |
| ``` |
| - ```gas |
| paciasp |
| ``` |
| * - ```gas |
| autiasp |
| ret |
| ``` |
| - ```gas |
| autiasp |
| add x30, x27, w30, uxtw |
| ret |
| ``` |
| ::: |
| |
| Authenticated returns (`retaa`/`retab`) combine authentication with return, |
| and must be expanded during rewriting. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| retaa |
| ``` |
| - ```gas |
| autiasp |
| add x30, x27, w30, uxtw |
| ret |
| ``` |
| * - ```gas |
| retab |
| ``` |
| - ```gas |
| autibsp |
| add x30, x27, w30, uxtw |
| ret |
| ``` |
| ::: |
| |
| Authenticated branches (`braa`/`brab`/`braaz`/`brabz`) and calls |
| (`blraa`/`blrab`/`blraaz`/`blrabz`) combine authentication with an |
| indirect branch or call. They are expanded by first authenticating the target |
| register in place, then performing a normal sandboxed branch or call. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| braa xN, xM |
| ``` |
| - ```gas |
| autia xN, xM |
| add x28, x27, wN, uxtw |
| br x28 |
| ``` |
| * - ```gas |
| braaz xN |
| ``` |
| - ```gas |
| autiza xN |
| add x28, x27, wN, uxtw |
| br x28 |
| ``` |
| * - ```gas |
| blraa xN, xM |
| ``` |
| - ```gas |
| autia xN, xM |
| add x28, x27, wN, uxtw |
| blr x28 |
| ``` |
| ::: |
| |
| Authenticated exception returns (`eret`/`eretaa`/`eretab`) are privileged |
| and are not supported: the rewriter reports an error for them. |
| |
| #### System instructions |
| |
| System calls are rewritten into a sequence that loads the address of the first |
| runtime call entrypoint and jumps to it. The runtime call entrypoint table is |
| stored at a negative offset from the sandbox base, so it can be referenced by |
| `x27`. The rewrite also saves and restores the link register, since it is |
| used for branching into the runtime. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| svc #0 |
| ``` |
| - ```gas |
| mov x26, x30 |
| ldur x30, [x27, #-8] |
| blr x30 |
| add x30, x27, w26, uxtw |
| ``` |
| ::: |
| |
| #### Thread pointer (TP) |
| |
| TP accesses are rewritten into loads/stores from the context register |
| (`x25`), which holds the virtual thread pointer at offset 16 (see |
| [Context Register](#context-register)). |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - `mrs xN, tpidr_el0` |
| - `ldr xN, [x25, #16]` |
| * - `msr tpidr_el0, xN` |
| - `str xN, [x25, #16]` |
| ::: |
| |
| ### Optimizations |
| |
| #### Basic guard elimination |
| |
| If a register is guarded multiple times in the same basic block without any |
| modifications to it during the intervening instructions, then subsequent guards |
| can be removed. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| add x28, x27, wN, uxtw |
| ldur xN, [x28] |
| add x28, x27, wN, uxtw |
| ldur xN, [x28, #8] |
| add x28, x27, wN, uxtw |
| ldur xN, [x28, #16] |
| ``` |
| - ```gas |
| add x28, x27, wN, uxtw |
| ldur xN, [x28] |
| ldur xN, [x28, #8] |
| ldur xN, [x28, #16] |
| ``` |
| ::: |
| |
| #### Address generation |
| |
| **Note**: not yet implemented. |
| |
| Addresses to global symbols in position-independent executables are frequently |
| generated via `adrp` followed by `ldr`. Since the address generated by |
| `adrp` can be statically guaranteed to be within the sandbox, it is safe to |
| directly target `x28` for these sequences. This allows the omission of a |
| guard instruction before the `ldr`. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| adrp xN, target |
| ldr xN, [xN, imm] |
| ``` |
| - ```gas |
| adrp x28, target |
| ldr xN, [x28, imm] |
| ``` |
| ::: |
| |
| #### Stack guard elimination |
| |
| **Note**: not yet implemented. |
| |
| If the stack pointer is modified by adding/subtracting a small immediate, and |
| then later used to perform a memory access without any intervening jumps, then |
| the guard on the stack pointer modification can be removed. This is because the |
| load/store is guaranteed to trap if the stack pointer has been moved outside of |
| the sandbox region. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| add x26, sp, #8 |
| add sp, x27, w26, uxtw |
| ... (same basic block) |
| ldr xN, [sp] |
| ``` |
| - ```gas |
| add sp, sp, #8 |
| ... (same basic block) |
| ldr xN, [sp] |
| ``` |
| ::: |
| |
| #### Guard hoisting |
| |
| **Note**: not yet implemented. |
| |
| In certain cases, guards may be hoisted outside of loops. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| mov w8, #10 |
| mov w9, #0 |
| .loop: |
| add w9, w9, #1 |
| ldr xN, [xM] |
| cmp w9, w8 |
| b.lt .loop |
| .end: |
| ``` |
| - ```gas |
| mov w8, #10 |
| mov w9, #0 |
| add x28, x27, wM, uxtw |
| .loop: |
| add w9, w9, #1 |
| ldr xN, [x28] |
| cmp w9, w8 |
| b.lt .loop |
| .end: |
| ``` |
| ::: |
| |
| ## X86-64 |
| |
| The X86-64 LFI target is `x86_64_lfi`. |
| |
| ### Reserved Registers |
| |
| The X86-64 LFI target reserves the following registers: |
| |
| - `r14`: always holds the sandbox base address. Also used as the runtime call |
| table pointer (the runtime call table is stored at the sandbox base). |
| - `gs`: always holds the sandbox base address (used as a segment register for |
| memory access sandboxing). |
| - `rsp`: always holds an address within the sandbox. |
| - `r15`: context register (see [Context Register](#context-register)). |
| - `r11`: scratch register. |
| |
| ### Bundling |
| |
| The X86-64 LFI target confines control flow using 32-byte bundles. Indirect |
| branch targets are masked to a 32-byte boundary, so control can only enter code |
| at a bundle-aligned address. For this to constrain the instruction stream, no |
| instruction may span a bundle boundary. |
| |
| LFI object files are therefore assembled with {doc}`aligned instruction |
| bundling <AlignedBundling>` enabled, which inserts `nop` padding wherever an |
| instruction would otherwise cross a boundary. Bundling is enabled implicitly |
| for the `x86_64_lfi` target. |
| |
| A rewrite that expands one instruction into a sequence which must not be |
| entered in the middle wraps that sequence in `.bundle_lock` / `.bundle_unlock`. |
| This keeps the whole group within a single bundle, so no masked indirect branch |
| can land between its instructions. |
| |
| **Note**: the masking of indirect branch targets is part of the control flow |
| rewrites, which have not been implemented yet. |
| |
| Bundling is specific to X86-64. AArch64 instructions are fixed-width and |
| naturally aligned, and the AArch64 LFI target confines indirect branches by |
| guarding the target register instead. |
| |
| ### Assembly Rewrites |
| |
| #### Terminology |
| |
| In the following assembly rewrites, some shorthand is used. |
| |
| - `%rN` or `%eN`: refers to any general-purpose non-reserved register. |
| - `{a,b,c}`: matches any of `a`, `b`, or `c`. |
| |
| #### Control flow |
| |
| **Note**: these rewrites have not been implemented. |
| |
| #### Memory accesses |
| |
| **Note**: these rewrites have not been implemented. |
| |
| #### String instructions |
| |
| **Note**: these rewrites have not been implemented. |
| |
| #### Stack modification |
| |
| **Note**: these rewrites have not been implemented. |
| |
| #### System instructions |
| |
| System calls are rewritten into a sequence that loads the return address into |
| the scratch register and jumps to the runtime call handler. The runtime call |
| handler table is stored at the address pointed to by `r14`. The `r11` |
| register stores the return address (marked by the label `.Ltmp` in the |
| block below). |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| syscall |
| ``` |
| - ```gas |
| leaq .Ltmp(%rip), %r11 |
| jmpq *-8(%r14) |
| .Ltmp: |
| ``` |
| ::: |
| |
| #### Thread pointer |
| |
| Thread pointer accesses via the `%fs` segment (used for TLS) are rewritten to |
| use the virtual thread pointer from the context register (`r15`) at offset 16 |
| (see [Context Register](#context-register)). The rewrite handles any load or |
| store instruction with an `%fs`-segment memory operand. `Op` represents any |
| such instruction. |
| |
| :::{list-table} |
| :header-rows: 1 |
| |
| * - Original |
| - Rewritten |
| * - ```gas |
| Op %fs:0, %rD |
| ``` |
| - ```gas |
| Op 16(%r15), %rD |
| ``` |
| * - ```gas |
| Op %fs:(%rX), %rD |
| ``` |
| - ```gas |
| movq 16(%r15), %rD |
| Op (%rD, %rX), %rD |
| ``` |
| * - ```gas |
| Op %rS, %fs:(%rX) |
| ``` |
| - ```gas |
| movq 16(%r15), %r11 |
| Op %rS, (%r11, %rX) |
| ``` |
| * - ```gas |
| Op %fs:N(%rX, %rY, S), %rD |
| ``` |
| - ```gas |
| movq 16(%r15), %r11 |
| leaq (%r11, %rX), %r11 |
| Op N(%r11, %rY, S), %rD |
| ``` |
| * - ```gas |
| Op %rS, %fs:N(%rX, %rY, S) |
| ``` |
| - ```gas |
| movq 16(%r15), %r11 |
| leaq (%r11, %rX), %r11 |
| Op %rS, N(%r11, %rY, S) |
| ``` |
| ::: |
| |
| ## References |
| |
| For more information, please see the following resources: |
| |
| - [LFI project page](https://github.com/lfi-project/) |
| - [LFI RFC](https://discourse.llvm.org/t/rfc-lightweight-fault-isolation-lfi-efficient-native-code-sandboxing-upstream-lfi-target-and-compiler-changes/88380) |
| - [LFI paper](https://zyedidia.github.io/papers/lfi_asplos24.pdf) |
| |
| Contact info: |
| |
| - Zachary Yedidia - [zyedidia@cs.stanford.edu](mailto:zyedidia@cs.stanford.edu) |
| - Tal Garfinkel - [tgarfinkel@google.com](mailto:tgarfinkel@google.com) |
| - Sharjeel Khan - [sharjeelkhan@google.com](mailto:sharjeelkhan@google.com) |