blob: 29e2addb03b881f8d4db2f6254335c2cfc74d3f6 [file] [view]
# Formatter Bytecode
## Background
LLDB provides rich customization options to display data types (see {doc}`/use/variable/`). To use custom data formatters, developers need to edit the global `~/.lldbinit` file to make sure they are found and loaded. In addition to this rather manual workflow, developers or library authors can ship ship data formatters with their code in a format that allows LLDB automatically find them and run them securely.
An end-to-end example of such a workflow is the Swift `DebugDescription` macro (see <https://www.swift.org/blog/announcing-swift-6/#debugging> ) that translates Swift string interpolation into LLDB summary strings, and puts them into a `.lldbsummaries` section, where LLDB can find them.
This document describes a minimal bytecode tailored to running LLDB formatters. It defines a human-readable assembler representation for the language, an efficient binary encoding, a virtual machine for evaluating it, and format for embedding formatters into binary containers.
### Goals
Provide an efficient and secure encoding for data formatters that can be used as a compilation target from user-friendly representations (such as DIL, Swift DebugDescription, or NatVis).
### Non-goals
While humans could write the assembler syntax, making it user-friendly is not a goal. It is meant to be used as a compilation target for higher-level, language-specific affordances.
## Design of the virtual machine
The LLDB formatter virtual machine uses a stack-based bytecode, comparable with DWARF expressions, but with higher-level data types and functions.
The virtual machine has two stacks, a data and a control stack. The control stack is kept separate to make it easier to reason about the security aspects of the virtual machine.
### Data types
All objects on the data stack must have one of the following data types. These data types are "host" data types, in LLDB parlance.
- *String* (UTF-8)
- *Integer* (arbitrary precision, always signed)
- *Int* (64 bit) (deprecated: use *Integer*)
- *UInt* (64 bit) (deprecated: use *Integer*)
- *Object* (Basically an `SBValue`)
- *Type* (Basically an `SBType`)
- *Selector* (One of the predefine functions)
- *Dictionary* (A mutable mapping from `String` keys to values of any data type)
*Object* and *Type* are opaque, they can only be used as a parameters of `call`.
## Instruction set
### Stack operations
These instructions manipulate the data stack directly.
```{eval-rst}
======== ========== ===========================
Opcode Mnemonic Stack effect
-------- ---------- ---------------------------
0x00 ``dup`` ``(x -> x x)``
0x01 ``drop`` ``(x y -> x)``
0x02 ``pick`` ``(x ... UInt -> x ... x)``
0x03 ``over`` ``(x y -> x y x)``
0x04 ``swap`` ``(x y -> y x)``
0x05 ``rot`` ``(x y z -> z x y)``
======== ========== ===========================
```
### Control flow
These manipulate the control stack and program counter. Both `if` and `ifelse` expect an `Integer` or `UInt` at the top of the data stack to represent the condition.
```{eval-rst}
======== ============ ============================================================
Opcode Mnemonic Description
-------- ------------ ------------------------------------------------------------
0x10 ``{`` push a code block address onto the control stack
-- ``}`` (technically not an opcode) syntax for end of code block
0x11 ``if`` ``(Integer|UInt -> )`` pop a block from the control stack,
if the top of the data stack is nonzero, execute it
0x12 ``ifelse`` ``(Integer|UInt -> )`` pop two blocks from the control stack, if
the top of the data stack is nonzero, execute the first,
otherwise the second.
0x13 ``return`` pop the entire control stack and return
======== ============ ============================================================
```
### Literals for basic types
```{eval-rst}
======== ============= ============================================================
Opcode Mnemonic Description
-------- ------------- ------------------------------------------------------------
0x20 ``123u`` ``( -> UInt)`` push an unsigned 64-bit host integer (deprecated: use ``lit_integer``)
0x21 ``123`` ``( -> Int)`` push a signed 64-bit host integer (deprecated: use ``lit_integer``)
0x22 ``"abc"`` ``( -> String)`` push a UTF-8 host string
0x23 ``@strlen`` ``( -> Selector)`` push one of the predefined function
selectors. See ``call``.
0x24 ``123`` ``( -> Integer)`` push an arbitrary precision signed integer
======== ============= ============================================================
```
### Conversion operations
```{eval-rst}
======== ============= ================================================================
Opcode Mnemonic Description
-------- ------------- ----------------------------------------------------------------
0x2a ``as_int`` ``( UInt -> Int)`` reinterpret a UInt as an Int (deprecated)
0x2b ``as_uint`` ``( Int -> UInt)`` reinterpret an Int as a UInt (deprecated)
0x2c ``is_null`` ``( Object -> UInt )`` check an object for null ``(object ? 0 : 1)``
======== ============= ================================================================
```
### Arithmetic, logic, and comparison operations
Every `Integer` value on the data stack is signed. `+`, `-`, `*`, `/`, `%`, `=`, `!=`, `<`, `>`, `=<`, `>=` are defined for `Integer` (and deprecated `Int`/`UInt`) and operate on the operands' mathematical values.
`<<`, `>>`, `&`, `|`, `^`, `~` are bitwise operations and operate on an `Integer`'s underlying two's complement bit pattern rather than its mathematical value. Because a bitwise operation never treats its operands as having a sign, `>>` is always a logical (zero-filling) shift, not an arithmetic shift, and a negative operand is not an error.
```{eval-rst}
======== ========== ===========================
Opcode Mnemonic Stack effect
-------- ---------- ---------------------------
0x30 ``+`` ``(x y -> [x+y])``
0x31 ``-`` etc ...
0x32 ``*``
0x33 ``/``
0x34 ``%``
0x35 ``<<``
0x36 ``>>``
0x40 ``~``
0x41 ``|``
0x42 ``^``
0x50 ``=``
0x51 ``!=``
0x52 ``<``
0x53 ``>``
0x54 ``=<``
0x55 ``>=``
======== ========== ===========================
```
### Function calls
For security reasons the list of functions callable with `call` is predefined. The supported functions are either existing methods on `SBValue`, or string formatting operations.
```{eval-rst}
======== ========== ============================================
Opcode Mnemonic Stack effect
-------- ---------- --------------------------------------------
0x60 ``call`` ``(Object argN ... arg0 Selector -> retval)``
======== ========== ============================================
```
Method is one of a predefined set of *Selectors*.
```{eval-rst}
==== =============================== ==================================================================== ======================================
Sel. Mnemonic Stack Effect Description
---- ------------------------------- -------------------------------------------------------------------- --------------------------------------
0x00 ``summary`` ``(Object @summary -> String)`` ``SBValue::GetSummary``
0x01 ``type_summary`` ``(Object @type_summary -> String)`` ``SBValue::GetTypeSummary``
0x10 ``get_num_children`` ``(Object @get_num_children -> Integer)`` ``SBValue::GetNumChildren``
0x11 ``get_child_at_index`` ``(Object Integer @get_child_at_index -> Object)`` ``SBValue::GetChildAtIndex``
0x12 ``get_child_with_name`` ``(Object String @get_child_with_name -> Object)`` ``SBValue::GetChildMemberWithName``
0x13 ``get_child_index`` ``(Object String @get_child_index -> Integer)`` ``SBValue::GetChildIndex``
0x14 ``get_parent`` ``(Object @get_parent -> Object)`` ``SBValue::GetParent``
0x15 ``get_type`` ``(Object @get_type -> Type)`` ``SBValue::GetType``
0x16 ``get_template_argument_type`` ``(Type Integer @get_template_argument_type -> Type)`` ``SBValue::GetTemplateArgumentType``
0x17 ``cast`` ``(Object Type @cast -> Object)`` ``SBValue::Cast``
0x18 ``get_synthetic_value`` ``(Object @get_synthetic_value -> Object)`` ``SBValue::GetSyntheticValue``
0x19 ``get_non_synthetic_value`` ``(Object @get_non_synthetic_value -> Object)`` ``SBValue::GetNonSyntheticValue``
0x20 ``get_value`` ``(Object @get_value -> Object)`` ``SBValue::GetValue``
0x21 ``get_value_as_unsigned`` ``(Object @get_value_as_unsigned -> Integer)`` ``SBValue::GetValueAsUnsigned``
0x22 ``get_value_as_signed`` ``(Object @get_value_as_signed -> Integer)`` ``SBValue::GetValueAsSigned``
0x23 ``get_value_as_address`` ``(Object @get_value_as_address -> Integer)`` ``SBValue::GetValueAsAddress``
0x24 ``clone`` ``(Object String @clone -> Object)`` ``SBValue::Clone``
0x25 ``get_pointee_type`` ``(Type @get_pointee_type -> Type)`` ``SBType::GetPointeeType``
0x26 ``get_byte_size`` ``(Type @get_byte_size -> Integer)`` ``SBType::GetByteSize``
0x27 ``create_child_at_offset`` ``(Object String Integer Type @create_child_at_offset -> Object)`` ``SBValue::CreateChildAtOffset``
0x40 ``read_memory_byte`` ``(UInt @read_memory_byte -> UInt)`` ``Target::ReadMemory``
0x41 ``read_memory_uint32`` ``(UInt @read_memory_uint32 -> UInt)`` ``Target::ReadMemory``
0x42 ``read_memory_int32`` ``(UInt @read_memory_int32 -> Int)`` ``Target::ReadMemory``
0x43 ``read_memory_uint64`` ``(UInt @read_memory_uint64 -> UInt)`` ``Target::ReadMemory``
0x44 ``read_memory_int64`` ``(UInt @read_memory_int64 -> Int)`` ``Target::ReadMemory``
0x45 ``read_memory_address`` ``(UInt @read_memory_uint64 -> UInt)`` ``Target::ReadMemory``
0x46 ``read_memory`` ``(UInt Type @read_memory -> Object)`` ``Target::ReadMemory``
0x50 ``fmt`` ``(String arg0 ... @fmt -> String)`` ``llvm::format``
0x51 ``sprintf`` ``(String arg0 ... sprintf -> String)`` ``sprintf``
0x52 ``strlen`` ``(String strlen -> Integer)`` ``strlen in bytes``
==== =============================== ==================================================================== ======================================
```
### Dictionary objects
`Dictionary` objects are key-value containers, with `String` value keys, and values of any data type. `Dictionary` is a reference type, mutating it through one reference is visible through any other reference to the same dictionary (e.g. one obtained earlier with `dup`). Empty `Dictionary` objects are created with `dict`. `Dictionary` objects are populated with `dict_set`. Values are retrieved with `dict_get`. When `dict_get` is called with a key that is not present in the dictionary, an error is emitted. Use `dict_has` first to check for a key's existence. Dictionary operations consumes the `Dictionary` argument, so `dup` it first if the `Dictionary` is needed afterward. For example, to set multiple keys in a row:
```
dict dup "a" 1 dict_set dup "b" 2 dict_set
```
A `Dictionary` may be stored as a value in another `Dictionary`, but `dict_set` emits an error if doing so would make a dictionary contain itself, directly or through nested dictionaries.
```{eval-rst}
======== ============= ============================================================
Opcode Mnemonic Stack effect
-------- ------------- ------------------------------------------------------------
0x70 ``dict`` ``( -> Dictionary)`` create an empty dictionary
0x71 ``dict_set`` ``(Dictionary String x -> )`` set a key to a value
0x72 ``dict_get`` ``(Dictionary String -> x)`` look up the value for a key
0x73 ``dict_has`` ``(Dictionary String -> Integer)`` check whether a key is present
======== ============= ============================================================
```
### Byte Code
Most instructions are just a single byte opcode. The only exceptions are the literals:
- *String*: Length in bytes encoded as ULEB128, followed length bytes
- *Int*: LEB128
- *UInt*: ULEB128
- *Integer*: LEB128, sign-extended to a signed value of at least 64 bits
- *Selector*: ULEB128
### Embedding
Expression programs are embedded into an `.lldbformatters` section (an evolution of the Swift `.lldbsummaries` section) that is a dictionary of type names/regexes and descriptions. It consists of a list of records. Each record starts with the following header:
- Version number (ULEB128)
- Remaining size of the record (minus the header) (ULEB128)
The version number is increased whenever an incompatible change is made, either to the layout of the record, or to the formatter ABI (see [Calling conventions](#calling-conventions)). Adding new opcodes or selectors is not an incompatible change since consumers can unambiguously detect this and report an error.
Space between two records may be padded with NULL bytes.
A record consists of a dictionary key, which is a type name or regex.
- Length of the key in bytes (ULEB128)
- The key (UTF-8)
A regex has to start with `^`, which is part of the regular expression.
After this comes a flag bitfield, which is a ULEB-encoded `lldb::TypeOptions` bitfield.
- Flags (ULEB128)
This is followed by one or more dictionary values that immediately follow each other and entirely fill out the record size from the header. Each expression program has the following layout:
- Function signature (1 byte)
- Length of the program (ULEB128)
- The program bytecode
### Calling conventions
A record's version number (see [Embedding](#embedding)) also determines the calling convention of its methods. This includes how `self` (aka `this`) is represented. This section describes version 2; see [Version 1](#version-1) for the differences in version 1.
The runtime owns `self`, a `Dictionary` that starts out empty. The runtime passes a reference to `self` as the first argument to every method except `@summary`, followed by that method's arguments (if any). Methods may modify `self` in place, and do not return it. The runtime's ownership of `self` ensures modifications to the `Dictionary` are visible to subsequent method calls.
The possible function signatures are:
```{eval-rst}
========= ========================= ==============================================
Signature Mnemonic Stack Effect
--------- ------------------------- ----------------------------------------------
0x00 ``@summary`` ``(Object -> ... String)``
0x01 ``@init`` ``(Dictionary Object -> ...)``
0x02 ``@get_num_children`` ``(Dictionary -> ... Integer)``
0x03 ``@get_child_index`` ``(Dictionary String -> ... Integer)``
0x04 ``@get_child_at_index`` ``(Dictionary Integer -> ... Object)``
0x05 ``@get_value`` ``(Dictionary -> ... String)``
0x06 ``@update`` ``(Dictionary -> ... Integer)``
========= ========================= ==============================================
```
The `@init` method must only be used for one time setup work. Any computation that needs to be reperformed should happen in `@update`. If not specified, initialization will save the given Object to the `self` dictionary using the idiomatic key `"valobj"`.
The `@update` method performs computation that may change over the course of a value's lifetime, such as interpreting a value's state, and (re)computing children. The return value is an `Integer`, where 1 means the previously computed children can be reused, and 0 means they must be refetched.
While it is more efficient to store multiple programs per type key, this is not a requirement. LLDB will merge all entries. If there are conflicts the result is undefined.
### Execution model
Execution begins at the first byte in the program. The program counter of the virtual machine starts at offset 0 of the bytecode and may never move outside the range of the program as defined in the header. The data stack starts with the method's arguments, as listed in the signature table in [Calling conventions](#calling-conventions).
### Error handling
Errors are unrecoverable, the entire expression will fail if any kind of error is encountered.
## Version 1
Version 1 records use the same record layout as version 2 (see [Embedding](#embedding)), but an older calling convention:
- There is no `self` dictionary. Instead, `self` is whatever is left on the data stack after `@init` and `@update` run. Nothing constrains its shape: it can be zero, one, or many values (written `Object+` below), and it is passed to the other methods in place of the `Dictionary`.
- If `@init` is not specified, `self` is the given Object.
- The `@update` reply is optional. If `@update` does not leave an `Integer` 0 or 1 (or a `UInt` or `Int` 0 or 1) on top of the data stack, the children are refetched.
- Selectors use the deprecated `UInt` and `Int` types in place of `Integer`: `get_num_children`, `get_child_index`, `get_value_as_unsigned`, `get_value_as_address`, and `strlen` return a `UInt`, `get_value_as_signed` returns an `Int`, and `get_child_at_index` and `get_template_argument_type` take a `UInt` index.
The version 1 function signatures are:
```{eval-rst}
========= ========================= ==============================================
Signature Mnemonic Stack Effect
--------- ------------------------- ----------------------------------------------
0x00 ``@summary`` ``(Object -> String)``
0x01 ``@init`` ``(Object -> Object+)``
0x02 ``@get_num_children`` ``(Object+ -> UInt)``
0x03 ``@get_child_index`` ``(Object+ String -> UInt)``
0x04 ``@get_child_at_index`` ``(Object+ UInt -> Object)``
0x05 ``@get_value`` ``(Object+ -> String)``
0x06 ``@update`` ``(Object+ -> Object+ [Integer])``
========= ========================= ==============================================
```