# Anti-Disassembly Sequences of executable code can have multiple disassembly representations, some may be invalid and some may obscure the real functionality of the program. > Anti-disassembly techniques work by taking advantage of the assumptions and limitations of disassemblers. > > For example, disassemblers can only represent each byte of a program as part of one instruction at a time. If the disassembler is tricked into disassembling at the wrong offset, a valid instruction could be hidden from view. ### Linear Disassembly Linear disassembly strategy iterates over a block of code, disassembling one instruction at a time linearly, without deviating. - It uses the size of the disassembled instruction to determine which byte to disassemble next, with no regard for flow control instructions. ### Flow-Oriented Disassembly This method is used by IDA - The key difference between linear and flow-oriented is that the disassembler doesn’t blindly iterate over a buffer, assuming the data is nothing but instructions packed neatly together - Instead it examines each instruction and builds a list of locations to disassemble - Most flow-oriented disassemblers will process the false branch of a conditional jump - Pressing the `C` key turns the cursor location into code - Pressing the `D` key turns the cursor location into data ### Anti-Disassembler Techniques #### Jump Instructions with the same Target The most common anti-disassembly technique seen in the wild is two back-to-back conditional jump instructions that both *point to the same target*. - For example the instruction `jz loc_512` followed by `jnz loc_512` will always be executed - However if the disassembler favours the false branch it could disassemble code that will never be reached #### Jump Instruction with a Constant Condition Another anti-disassembly technique commonly found in the wild is composed of a single conditional jump instruction placed where the condition will always be the same. #### Impossible Disassembly Under some conditions, no traditional assembly listing will accurately represent the instructions that are executed. We use the term *impossible disassembly* for such conditions, but the term isn’t strictly accurate. You could disassemble these techniques, but you would need a vastly different representation of code than what is currently provided by disassemblers. - A *rogue byte* is a byte placed after a conditional jump instruction - This means the real instruction that follows will not be disassembled ![1647963505.png](img/1647963505.png) 1. The first instruction moves data `0xEB05` into the `AX` register 2. The second instruction zeros out this register and sets the zero flag 3. The third is a conditional jump - but is actually an unconditional jump as the zero flag will always be set 4. The disassembler will continue disassembling the fake `CALL` instruction that will never be reached ```assembly 66 B8 EB 05 mov ax, 5EBh 31 C0 xor eax, eax 74 F9 jz short near ptr sub_4011C0+1 loc_4011C8: E8 58 C3 90 90 call near ptr 98A8D525h ``` What it could look like in IDA, note the `+1`. - However after converting it to data and back to code, so that the only instructions visible are the `xor` instruction and the hidden instructions ```assembly 66 byte_4011C0 db 66h B8 db 0B8h EB db 0EBh 05 db 5 ; ------------------------------------------------------- 31 C0 xor eax, eax ; ------------------------------------------------------- 74 db 74h F9 db 0F9h E8 db 0E8h ; ------------------------------------------------------- 58 pop eax C3 retn ``` - This only shows the instructions that are relevant to understanding the program - However this solution may interfere with flow graphs. - Since it’s difficult to tell how the `xor`, `pop` and `retn` instructions are used ### Obscuring Flow Control #### The Function Pointer Problem If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverse-engineer without dynamic analysis. ```assembly 004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p 004011D0 ; sub_401040+8Bp 004011D0 004011D0 var_4 = dword ptr -4 004011D0 arg_0 = dword ptr 8 004011D0 004011D0 push ebp 004011D1 mov ebp, esp 004011D3 push ecx 004011D4 push esi 004011D5 mov [ebp+var_4], offset sub_4011C0 ;1 004011DC push 2Ah 004011DE call [ebp+var_4] ;2 004011E1 add esp, 4 004011E4 mov esi, eax 004011E6 mov eax, [ebp+arg_0] 004011E9 push eax 004011EA call [ebp+var_4] ;3 004011ED add esp, 4 004011F0 lea eax, [esi+eax+1] 004011F4 pop esi 004011F5 mov esp, ebp 004011F7 pop ebp 004011F8 retn 004011F8 sub_4011D0 endp ``` Here `sub_4011C0` is called three times but IDA only recognised it once at `1`. ###### Adding Missing Code Cross-References in IDA We can manually add these in using `AddCodeXref` #### Return Pointer Abuse - `Call` is a combination of `jmp` and `push` - As it jumps to the new function and pushes a return address onto the stack - `retn` instruction pops the value from the top of the stack and jumps to it. - Typically used to return a function call - However no reason why malware authors can’t use it to obscure code ```assembly 004011C0 sub_4011C0 proc near ; CODE XREF: _main+19p 004011C0 ; sub_401040+8Bp 004011C0 004011C0 var_4 = byte ptr -4 004011C0 004011C0 call $+5 004011C5 add [esp+4+var_4], 5 004011C9 retn 004011C9 sub_4011C0 endp ; sp-analysis failed 004011CA ; ---------------------------------------------- 004011CA push ebp 004011CB mov ebp, esp 004011CD mov eax, [ebp+8] 004011D0 imul eax, 2Ah 004011D3 mov esp, ebp 004011D5 pop ebp 004011D6 retn ``` - Here `var_4` is set to the constant `-4` - This means `add [esp+4+var_4], 5` is actually `add [esp+4+(-4)]` - `0x4011C9 + 0x5 = 0x4011CA` - The `retn` instruction jumps to that memory location