Add the rest of university notes

This commit is contained in:
John Gatward committed 2026-10-04 14:02:35 +01:00
1 parent c1b84c7f7d
commit d0f27f276b
366 files changed
+9844 -110

No files matched your search

@@ -0,0 +1,156 @@
# Anti-Disassembly
Sequences of executable code can have multiple disassembly representations, some may be invalid and some may obscure the real functionality of the program.
> Anti-disassembly techniques work by taking advantage of the assumptions and limitations of disassemblers.
>
> For example, disassemblers can only represent each byte of a program as part of one instruction at a time. If the disassembler is tricked into disassembling at the wrong offset, a valid instruction could be hidden from view.
### Linear Disassembly
Linear disassembly strategy iterates over a block of code, disassembling one instruction at a time linearly, without deviating.
- It uses the size of the disassembled instruction to determine which byte to disassemble next, with no regard for flow control instructions.
### Flow-Oriented Disassembly
This method is used by IDA
- The key difference between linear and flow-oriented is that the disassembler doesn’t blindly irate over a buffer, assuming the data is noting but instructions packed neatly together
- Instead it examines each instruction and builds a list of locations to disassemble
- Most flow-oriented disassemblers will process the false branch of a conditional jump
- Pressing the `C` key turns the cursor location into code
- Pressing the `D` key turns the cursor location into data
### Anti-Disassembler Techniques
#### Jump Instructions with the same Target
The most common anti-disassembly technique seen in the wild is two back-to-back conditional jump instructions that both *point to the same target*.
- For example the instruction `jz loc_512` followed by `jnz loc_512` will always be executed
- However if the disassembler favours the false branch it could disassemble code that will never be reached
#### Jump Instruction with a Constant Condition
Another anti-disassembly technique commonly found in the wild is composed of a single conditional jump instruction placed where the condition will always be the same.
#### Impossible Disassembly
Under some conditions, no traditional assembly listing will accurately represent the instructions that are executed. We use the term *impossible disassembly* for such conditions, but the term isn’t strictly accurate. You could disassemble these techniques, but you would need a vastly different representation of code than what is currently provided by disassemblers.
- A *rogue byte* is a byte placed after a conditional jump instruction
- This means the real instruction that follows will not be disassembled
![1647963505.png](img/1647963505.png)
1. The first instruction moves data `0xEB05` into the `AX` register
2. The second instruction zeros out this register and sets the zero flag
3. The third is a conditional jump - but is actually an unconditional jump as the zero flag will always be set
4. The disassembler will continue disassembling the fake `CALL` instruction that will never be reached
```assembly
66 B8 EB 05 mov ax, 5EBh
31 C0 xor eax, eax
74 F9 jz short near ptr sub_4011C0+1
loc_4011C8:
E8 58 C3 90 90 call near ptr 98A8D525h
```
What it could look like in IDA, note the `+1`.
- However after converting it to data and back to code, so that the only instructions visible are the `xor` instruction and the hidden instructions
```assembly
66 byte_4011C0 db 66h
B8 db 0B8h
EB db 0EBh
05 db 5
; -------------------------------------------------------
31 C0 xor eax, eax
; -------------------------------------------------------
74 db 74h
F9 db 0F9h
E8 db 0E8h
; -------------------------------------------------------
58 pop eax
C3 retn
```
- This only shows the instructions that are relevent to understanding the program
- However this solution may interfere with flow graphs.
- Since its difficult to tell how the `xor`, `pop` and `retn` instructions are used
### Obscuring Flow Control
#### The Function Pointer Problem
If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverseengineer without dynamic analysis.
```assembly
004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p
004011D0 ; sub_401040+8Bp
004011D0
004011D0 var_4 = dword ptr -4
004011D0 arg_0 = dword ptr 8
004011D0
004011D0 push ebp
004011D1 mov ebp, esp
004011D3 push ecx
004011D4 push esi
004011D5 mov [ebp+var_4], offset sub_4011C0 ;1
004011DC push 2Ah
004011DE call [ebp+var_4] ;2
004011E1 add esp, 4
004011E4 mov esi, eax
004011E6 mov eax, [ebp+arg_0]
004011E9 push eax
004011EA call [ebp+var_4] ;3
004011ED add esp, 4
004011F0 lea eax, [esi+eax+1]
004011F4 pop esi
004011F5 mov esp, ebp
004011F7 pop ebp
004011F8 retn
004011F8 sub_4011D0 endp
```
Here `sub_4011C0` is called three times but IDA only recognised it once at `1`.
###### Adding Missing Code Cross-References in IDA
We can manually add these in using `AddCodeXref`
#### Return Pointer Abuse
- `Call` is a combination of `jmp` and `push`
- As it jumps to the new function and pushes a return address onto the stack
- `retn` instruction pops the value from the top of the stack and jumps to it.
- Typically used to return a function call
- However no reason why malware authors can’t use it to obscure code
```assembly
004011C0 sub_4011C0 proc near ; CODE XREF: _main+19p
004011C0 ; sub_401040+8Bp
004011C0
004011C0 var_4 = byte ptr -4
004011C0
004011C0 call $+5
004011C5 add [esp+4+var_4], 5
004011C9 retn
004011C9 sub_4011C0 endp ; sp-analysis failed
004011CA ; ----------------------------------------------
004011CA push ebp
004011CB mov ebp, esp
004011CD mov eax, [ebp+8]
004011D0 imul eax, 2Ah
004011D3 mov esp, ebp
004011D5 pop ebp
004011D6 retn
```
- Here `var_4` is set to the constant `-4`
- This means `add [esp+4+var_4], 5` is actually `add [esp+4+(-4)]`
- `0x4011C9 + 0x5 = 0x4011CA`
- The `retn` instruction jumps to that memory location