Low-Level Languages
Master low-level programming. Learn machine code, binary fundamentals, x86 Assembly language, and hands-on assembly development using the NASM assembler.

Low-Level Languages
Content Overview
Ultimate Guide to Low-Level Programming: From Binary to Assembly
Table of Contents
- Introduction: Why Bits and Bytes Rule the World
- Number Systems — The Language of Bits
- Binary Arithmetic and Bitwise Operations
- ASCII — How Text Becomes Binary
- Machine Language — The CPU’s Native Tongue
- How a Computer Actually Works (Hardware Fundamentals)
- Assembly Language with NASM — Complete Guide
- Practical Assembly Programs
- Performance Optimization at the Low Level
- Tools Every Low-Level Programmer Needs
- Conclusion and Next Steps
PART ONE: THE FOUNDATIONS
1. Introduction: Why Bits and Bytes Rule the World
1.1 Binary: The Foundation of All Computation
Binary is a numerical system that uses only two digits: 0 and 1. In electronics, these correspond directly to two physical states of a transistor: off (no current, representing 0) and on (current flowing, representing 1).
Because every modern digital circuit is built from billions of transistors that can only be off or on, binary is the only language the hardware can natively speak. Any information – text, images, video, program code – must ultimately be encoded as sequences of binary digits (bits) before the CPU can process it.
Machine Language is the set of binary instruction codes that a specific CPU model understands and can execute directly, without any translation or interpretation. Each instruction is a binary pattern (e.g., 10110000 01100001) that the CPU’s control unit decodes into signals that activate the correct internal circuits.
Why this matters: Every program you have ever run, from a Python script to a video game to your operating system, is eventually converted into machine language before the CPU touches it. Understanding machine language means understanding the ultimate reality of all software.
1.2 What Is a Low-Level Language?
A low-level language is a programming language that provides little or no abstraction from the computer’s hardware. “Abstraction” means hiding hardware details behind convenient high-level constructs. For example, Python’s print("Hello") hides the system call, the buffer management, and the character encoding. A low-level language exposes those details: you control exactly which registers are used, exactly which memory addresses are accessed, and exactly which CPU instructions are executed.
Important: “Low-level” refers to the level of proximity to the machine, not to difficulty. Writing correct, efficient low-level code requires deep understanding of computer architecture and is one of the most challenging skills in software development.
Two Main Categories
| Category | Also called | Definition |
|---|---|---|
| Machine language | First-generation (1GL) | The native binary instruction set of a specific CPU. Consists entirely of 1s and 0s (or hexadecimal representation). Executed directly by hardware – no translation needed. |
| Assembly language | Second-generation (2GL) | A human-readable notation for machine language. Uses short text mnemonics (e.g., MOV, ADD, JMP) to represent each machine instruction. An assembler program translates these mnemonics into binary machine code. There is a one-to-one correspondence between an assembly instruction and a machine instruction. |
Low-Level vs High-Level – Detailed Comparison
| Feature | Low-level | High-level (e.g., Python, Java) |
|---|---|---|
| Abstraction | None or minimal – you see every CPU cycle | Extensive – memory management, types, I/O are hidden |
| Memory control | Complete, manual – you allocate and free every byte | Limited or automatic (garbage collection) |
| Portability | Non-portable – code written for x86 will not run on ARM | Highly portable – same source runs on many CPUs |
| Performance | Maximum possible – no hidden overhead | Good to moderate – interpreter or JIT adds overhead |
| Development speed | Slow – many instructions needed for simple tasks | Fast – one line of Python does the work of dozens of assembly instructions |
| Error risk | Very high – one wrong memory access crashes the program | Lower – language runtime prevents many errors |
| Typical use | OS kernels, device drivers, firmware, game engines, cryptography | Web apps, data science, automation, business logic |
1.3 Why Learn Low-Level Programming Today?
Understanding low-level programming transforms you as a developer in six concrete ways:
- You understand what your code actually does. When a Python programmer writes
x = [1,2,3], they may not know that this allocates heap memory, creates a pointer, stores a reference count, and sets up a dynamic array structure. A low-level programmer knows exactly what every operation costs in CPU cycles and memory bytes. - You become a better debugger. The most mysterious bugs – segmentation faults, memory corruption, race conditions, security vulnerabilities – all exist at the machine level. Being able to read a stack trace, interpret a core dump, or step through disassembled code makes you capable of solving problems that are completely opaque to those who never leave the high-level world.
- You can write faster code. Performance-critical software – game engines, real-time audio processing, financial trading systems, cryptographic libraries – ultimately depends on what happens at the CPU level. Understanding instruction latency, cache behaviour, and SIMD operations allows you to write code that runs orders of magnitude faster than naive high-level implementations.
- You understand security. Every major class of software vulnerability – buffer overflows, use-after-free, stack smashing, return-oriented programming (ROP), format string attacks – is a low-level concept. Security researchers and exploit developers live in assembly. You cannot fully understand how to prevent attacks without understanding how they work at the machine level.
- You understand how compilers work. Compilers are among the most complex software systems ever built. Understanding assembly lets you read compiler output, understand compiler optimisations, and write high-level code that the compiler can optimise effectively.
- Every high-level language sits on low-level foundations. Python is written in C. JavaScript engines are written in C++. The Linux kernel is written in C with assembly. The Java Virtual Machine is written in C++. Going down to the foundation makes everything above it clearer.
2. Number Systems — The Language of Bits
A number system (or numeral system) is a way to represent numbers using a set of symbols (digits). Every number system has a base (or radix) – the number of unique digits it uses before carrying over to the next position.
Here is how the four systems compare:
| Number System | Base | Digits Used | Example Counting |
|---|---|---|---|
| Decimal | 10 | 0–9 | 0, 1, 2, …, 9, 10, 11… |
| Binary | 2 | 0, 1 | 0, 1, 10, 11, 100… |
| Octal | 8 | 0–7 | 0, 1, 2, …, 7, 10, 11… |
| Hexadecimal | 16 | 0–9, A–F | 0, 1, …, 9, A, B, …F, 10 |
Tip: Hexadecimal digits beyond 9 are represented by letters. A=10, B=11, C=12, D=13, E=14, F=15. So when you see 0xFF, that’s two hex digits each worth 15, giving you the decimal value 255.
2.1 The Positional Value Formula – The Master Key
This is the single most important formula in this entire guide. Everything else builds on it.
Value = Σ (digitᵢ × baseⁱ) for i = 0 to n-1
Positions are numbered from 0 (rightmost) to n-1 (leftmost). Master this one formula and you can work in any number system.
Let’s apply it to all four systems so you see the pattern clearly:
| System | Example | Calculation |
|---|---|---|
| Decimal | 345 | 3×10² + 4×10¹ + 5×10⁰ = 300 + 40 + 5 = 345 |
| Binary | 1011₂ | 1×2³ + 0×2² + 1×2¹ + 1×2⁰ = 8 + 0 + 2 + 1 = 11 |
| Octal | 127₈ | 1×8² + 2×8¹ + 7×8⁰ = 64 + 16 + 7 = 87 |
| Hexadecimal | 3F₁₆ | 3×16¹ + 15×16⁰ = 48 + 15 = 63 |
Notice that the formula is identical every time – only the base changes.
2.2 Binary Number System (Base-2) – Deep Explanation
Binary is the native language of all digital hardware. Every transistor inside your CPU, RAM, and storage device holds exactly one bit – a 0 or a 1. This is because transistors are switches: they are either off (0) or on (1). There is no third state.
Key Properties:
- A single binary digit is called a bit
- A group of 8 bits is called a byte. A byte can hold 2⁸ = 256 different values (0 through 255)
- A group of 4 bits is called a nibble
- All modern data sizes (kilobytes, megabytes, gigabytes) are powers of two multiples of bytes
Important rule: n bits can represent exactly 2ⁿ different values, from 0 to 2ⁿ−1. So 4 bits = 16 values, 8 bits = 256 values, 16 bits = 65,536 values, and 32 bits = over 4 billion values.
Powers of Two – The Building Blocks
Memorise these. They appear constantly in binary work:
| 2ⁿ | Value |
|---|---|
| 2⁰ | 1 |
| 2¹ | 2 |
| 2² | 4 |
| 2³ | 8 |
| 2⁴ | 16 |
| 2⁵ | 32 |
| 2⁶ | 64 |
| 2⁷ | 128 |
| 2⁸ | 256 |
| 2⁹ | 512 |
| 2¹⁰ | 1024 |
| 2¹¹ | 2048 |
| 2¹² | 4096 |
Each value is exactly double the one before it. That doubling pattern is the heartbeat of binary.
Counting in Binary (0–15)
| Decimal | Binary | Hex |
|---|---|---|
| 0 | 0000 | 0x0 |
| 1 | 0001 | 0x1 |
| 2 | 0010 | 0x2 |
| 3 | 0011 | 0x3 |
| 4 | 0100 | 0x4 |
| 5 | 0101 | 0x5 |
| 6 | 0110 | 0x6 |
| 7 | 0111 | 0x7 |
| 8 | 1000 | 0x8 |
| 9 | 1001 | 0x9 |
| 10 | 1010 | 0xA |
| 11 | 1011 | 0xB |
| 12 | 1100 | 0xC |
| 13 | 1101 | 0xD |
| 14 | 1110 | 0xE |
| 15 | 1111 | 0xF |
Binary to Decimal – Conversion with Detailed Steps
Method: Sum the positional values of all bits that are 1.
Example 1: Convert 1101₂ to decimal
Position 3: 1 × 8 = 8
Position 2: 1 × 4 = 4
Position 1: 0 × 2 = 0 (skip)
Position 0: 1 × 1 = 1
Sum: 8 + 4 + 0 + 1 = 13
Example 2: Convert 10110111₂ to decimal
- Write powers from left: 128, 64, 32, 16, 8, 4, 2, 1
- Bits: 1, 0, 1, 1, 0, 1, 1, 1
- Sum only where bit=1: 128 + 0 + 32 + 16 + 0 + 4 + 2 + 1 = 183
Shortcut: Only add the positional values where the bit is 1. Completely skip every position that has a 0. This makes mental calculation much faster.
Decimal to Binary – Repeated Division Method
Method: Repeatedly divide the decimal number by 2, recording the remainder after each division. Read the remainders from bottom to top.
Example: Convert 45 to binary
| Division | Quotient | Remainder |
|---|---|---|
| 45 ÷ 2 | 22 | 1 (least significant bit) |
| 22 ÷ 2 | 11 | 0 |
| 11 ÷ 2 | 5 | 1 |
| 5 ÷ 2 | 2 | 1 |
| 2 ÷ 2 | 1 | 0 |
| 1 ÷ 2 | 0 | 1 (most significant bit) |
Read remainders from bottom to top: 101101₂
Verify: 32 + 8 + 4 + 1 = 45 ✓
Example: Convert 200 to binary
200 → 100 r 0
100 → 50 r 0
50 → 25 r 0
25 → 12 r 1
12 → 6 r 0
6 → 3 r 0
3 → 1 r 1
1 → 0 r 1
Read bottom to top: 11001000₂
Verify: 128 + 64 + 8 = 200 ✓
2.3 Octal Number System (Base-8)
Octal uses digits 0 through 7. It was popular in early computing because 3 binary bits map perfectly to one octal digit, making it a compact way to write binary. Today it is most commonly seen in Unix/Linux file permissions – when you run chmod 755, those three digits are octal numbers describing permission bits.
One octal digit = exactly 3 binary bits.
| Octal | Decimal | Binary |
|---|---|---|
| 0 | 0 | 000 |
| 1 | 1 | 001 |
| 2 | 2 | 010 |
| 3 | 3 | 011 |
| 4 | 4 | 100 |
| 5 | 5 | 101 |
| 6 | 6 | 110 |
| 7 | 7 | 111 |
| 10 | 8 | 1000 |
| 11 | 9 | 1001 |
| 12 | 10 | 1010 |
| 13 | 11 | 1011 |
| 14 | 12 | 1100 |
| 15 | 13 | 1101 |
| 16 | 14 | 1110 |
| 17 | 15 | 1111 |
Example: Convert Octal 127₈ to decimal:
1×8² + 2×8¹ + 7×8⁰ = 64 + 16 + 7 = 87₁₀
Binary to Octal shortcut: Group binary bits in sets of 3 from the right. Convert each group directly.
Example: 110 111 010₂ → 6, 7, 2 → 672₈
Octal to Binary: Simply expand each octal digit into its 3-bit binary equivalent.
Example: 5₈ = 101, 3₈ = 011 → 53₈ = 101011₂
2.4 Hexadecimal Number System (Base-16) – Why It Matters
Hexadecimal is the most practically important non-decimal number system in computing today. It is everywhere – memory addresses, colour codes in CSS (#FF5733), MAC addresses, error codes, and raw machine code dumps all use hex.
The reason hex is so useful is that one hex digit maps perfectly to exactly 4 binary bits (called a nibble). Since a byte is 8 bits, one byte is always exactly 2 hex digits. This makes hex a perfectly compact, human-readable way to write binary data.
Hex to Binary Conversion Table (Memorise the first 16)
| Hex | Decimal | Binary | Hex | Decimal | Binary |
|---|---|---|---|---|---|
| 0 | 0 | 0000 | 8 | 8 | 1000 |
| 1 | 1 | 0001 | 9 | 9 | 1001 |
| 2 | 2 | 0010 | A | 10 | 1010 |
| 3 | 3 | 0011 | B | 11 | 1011 |
| 4 | 4 | 0100 | C | 12 | 1100 |
| 5 | 5 | 0101 | D | 13 | 1101 |
| 6 | 6 | 0110 | E | 14 | 1110 |
| 7 | 7 | 0111 | F | 15 | 1111 |
Binary to Hex shortcut: Group binary bits into sets of 4 from the right, convert each group to a hex digit.
Example: 10110110₂ → group as 1011 0110 → B 6 → 0xB6
Example: 111100001010₂ → 1111 0000 1010 → F 0 A → 0xF0A
Hex to Binary: Expand each hex digit to its 4-bit binary equivalent.
Example: 0x2F → 0010 1111 → 00101111₂
Example: 0xA9 → 1010 1001 → 10101001₂
Hex to Decimal: Use the positional formula with base 16.
Example: 0x3F → 3×16¹ + 15×16⁰ = 48 + 15 = 63
Example: 0xFF → 15×16¹ + 15×16⁰ = 240 + 15 = 255 (the maximum value of one byte)
2.5 Two’s Complement – How Computers Handle Negative Numbers
Here is something that surprises many beginners: computers have no minus sign in hardware. A transistor is either on or off. There is no “negative” transistor. So how do CPUs handle negative numbers?
The answer is two’s complement – a clever encoding that allows both positive and negative integers to be stored in binary, and that lets the same addition circuit handle both without any modification.
Why it is brilliant: The same addition circuit that adds positive numbers also works correctly when one operand is represented in two’s complement. No special “subtraction” hardware is needed.
How to Compute Two’s Complement of a Negative Number
Given an n-bit system (e.g., 8 bits, 16 bits, 32 bits, 64 bits), to represent -x:
- Write the positive value x in binary using exactly n bits (pad with leading zeros if necessary)
- Flip all bits – change every 0 to 1 and every 1 to 0. This is the one’s complement
- Add 1 to the result (binary addition, ignoring any final carry out of the n-th bit)
Example: Represent -7 in 8-bit two’s complement:
Step 1: +7 in 8 bits = 00000111
Step 2: Flip all bits = 11111000 (one's complement)
Step 3: Add 1 = 11111001
Therefore, in an 8-bit system, the binary pattern 11111001 represents the integer -7.
Verification: Add +7 (00000111) and -7 (11111001):
00000111 + 11111001 = 100000000
The result is 9 bits, but we only keep the lower 8 bits (00000000), which is zero. So 7 + (-7) = 0 ✓
Range of n-Bit Two’s Complement
| Bits | Range |
|---|---|
| 8 bits | -128 to +127 |
| 16 bits | -32,768 to +32,767 |
| 32 bits | -2,147,483,648 to +2,147,483,647 |
| 64 bits | -9.22×10¹⁸ to +9.22×10¹⁸ |
Formula:
- Most positive value: 2^(n-1) − 1 (binary 0111…111)
- Most negative value: −2^(n-1) (binary 1000…000)
- Zero: 0000…000 (unique representation)
Interpreting the Sign Bit
In two’s complement, the leftmost (most significant) bit tells you the sign:
- If the MSB is 0, the number is zero or positive
- If the MSB is 1, the number is negative
Common 8-Bit Two’s Complement Values
| Decimal | 8-bit Binary |
|---|---|
| 3 | 00000011 |
| 2 | 00000010 |
| 1 | 00000001 |
| 0 | 00000000 |
| −1 | 11111111 |
| −2 | 11111110 |
| −3 | 11111101 |
| −7 | 11111001 |
| −128 | 10000000 |
| 127 | 01111111 |
Subtraction Using Two’s Complement
To compute A – B, the CPU computes A + (two’s complement of B). This is literally how every CPU on earth performs subtraction.
Example: 13 − 5 in 8 bits:
Two's complement of 5 (00000101):
flip → 11111010
add 1 → 11111011
Add: 00001101 (13) + 11111011 (-5) = 1 00001000
Keep lower 8 bits = 00001000 = 8 ✓
3. Binary Arithmetic and Bitwise Operations
3.1 Binary Addition
Binary addition follows four simple rules:
| A | B | Result | Carry |
|---|---|---|---|
| 0 | 0 | 0 | 0 |
| 0 | 1 | 1 | 0 |
| 1 | 0 | 1 | 0 |
| 1 | 1 | 0 | 1 |
When both bits are 1, the result in that column is 0 and you carry a 1 to the next column left – exactly like carrying in decimal when digits exceed 9.
Step-by-step example: Add 1011₂ (11) + 1101₂ (13)
Carry: 1 1 1 0
1 0 1 1
+ 1 1 0 1
─────────
1 1 0 0 0 (= 24)
Work column by column from right to left:
- Column 0: 1+1 = 0 carry 1
- Column 1: 1+0+carry1 = 0 carry 1
- Column 2: 0+1+carry1 = 0 carry 1
- Column 3: 1+1+carry1 = 1 carry 1
- Final carry 1 gives the 5-bit result
Verify: 11 + 13 = 24 ✓
3.2 Binary Subtraction
Subtraction is performed using two’s complement as described in Section 2.5. Convert the number being subtracted to its two’s complement form, then add. This is not just a trick – this is literally how every CPU on earth performs subtraction.
Example: Subtract 1011 − 0110 (11 − 6):
Two's complement of 0110: flip → 1001, add 1 → 1010
Add: 1011 + 1010 = 1 0101 → drop carry → 0101 = 5
And 11 − 6 = 5 ✓
3.3 Binary Multiplication
Binary multiplication uses the same long multiplication method as decimal, but far simpler because you only ever multiply by 0 or 1.
- Any number multiplied by 0 = 0
- Any number multiplied by 1 = itself (just copy it down, shifted left by position)
- Then add all the partial results together
Example: Multiply 101₂ × 11₂ (5 × 3):
1 0 1
× 1 1
─────────
1 0 1 (101 × 1, shift 0)
+ 1 0 1 0 (101 × 1, shift 1)
───────────
= 1 1 1 1 (= 15)
And 5 × 3 = 15 ✓
3.4 Binary Division
Binary division is performed by repeated subtraction (or shift-and-subtract algorithm). Division by a power of two (2ᵏ) is simply a right shift by k bits, which is why CPUs prefer it.
3.5 Bitwise Operations – Definitions and Practical Uses
Bitwise operation: An operation that treats two binary numbers as sequences of individual bits and applies a logical function (AND, OR, XOR, NOT) to each corresponding pair of bits independently. These operations execute in a single CPU clock cycle and are used for low-level control, graphics, cryptography, and systems programming.
| Operation | Symbol | Example (Binary) | Result |
|---|---|---|---|
| AND | & | 1010 & 1100 | 1000 |
| OR | | | 1010 | 1100 | 1110 |
| XOR | ^ | 1010 ^ 1100 | 0110 |
| NOT | ~ | ~1010 | 0101 |
| Left Shift | << | 0001 << 3 | 1000 |
| Right Shift | >> | 1000 >> 2 | 0010 |
AND (&) – Masking
Definition: AND produces a 1 output only when both input bits are 1.
Truth table: 0&0=0, 0&1=0, 1&0=0, 1&1=1
Practical uses:
- Check if a number is odd:
num & 1– if result is 1, the number is odd (because only the least significant bit determines oddness) - Extract lower 4 bits:
byte & 0x0F– clears the upper 4 bits, passes lower 4 bits unchanged - Clear specific bits: To clear bits 3-5, create a mask with 0s in those positions and 1s elsewhere, then AND
Example: 1010 & 1100 = 1000
OR (|) – Setting Bits
Definition: OR produces a 1 output when at least one input bit is 1.
Truth table: 0|0=0, 0|1=1, 1|0=1, 1|1=1
Practical uses:
- Set bit 3 (value 8) to 1:
value | 0x08– forces bit 3 to 1 regardless of its previous value - Combine flags: Multiple boolean flags can be ORed together into a single byte
Example: 1010 | 1100 = 1110
XOR (^) – Toggling
Definition: XOR (exclusive OR) produces a 1 output when the two input bits are different.
Truth table: 0^0=0, 0^1=1, 1^0=1, 1^1=0
Properties:
a ^ a = 0(any number XOR itself equals zero)a ^ 0 = a(any number XOR zero equals itself)a ^ b ^ b = a(XOR is its own inverse – useful for encryption)
Practical uses:
- Toggle bits:
value ^ 0x0Fflips the lower 4 bits - Swap two variables without a temporary:
a ^= b; b ^= a; a ^= b;
- Simple encryption: XOR a message with a key; XOR again with the same key to decrypt
Example: 1010 ^ 1100 = 0110
NOT (~) – Bit Inversion
Definition: NOT (one’s complement) flips every bit: 0 becomes 1, 1 becomes 0.
Practical uses:
- Compute one’s complement (first step of two’s complement negation)
- Create inverted masks
Example (8 bits): ~00001010 = 11110101
Bit Shifts – Multiply and Divide by Powers of Two
Left shift (<<): Moves all bits to the left by a specified number of positions. Zeros fill the vacated lower bits. Each left shift multiplies the value by 2.
Right shift (>>): Moves all bits to the right. For unsigned numbers, zeros fill the vacated upper bits (logical shift). For signed numbers, the sign bit may be replicated (arithmetic shift). Each right shift divides the value by 2 (integer division).
Formulas:
n << k = n × 2ᵏn >> k = n ÷ 2ᵏ(integer division, truncating toward zero for unsigned)
Example: 5 << 3 = 5 × 8 = 40
Binary: 00000101 → 00101000
Example: 40 >> 3 = 40 ÷ 8 = 5
Binary: 00101000 → 00000101
Why compilers love shifts: Multiplication and division by constants that are powers of two are automatically replaced with shift instructions because shifts are orders of magnitude faster than general multiplication/division circuits.
4. ASCII — How Text Is Stored in Binary
Numbers fit naturally into binary. But what about text? How does a computer store the letter ‘A’?
The answer is ASCII – the American Standard Code for Information Interchange. ASCII assigns a unique number to every printable character and control character. That number is then stored as binary. So text is really just a sequence of numbers, which are sequences of bits.
Selected ASCII values (memorise these common ones):
| Character | Decimal | Binary (8-bit) |
|---|---|---|
| ‘A’ | 65 | 01000001 |
| ‘B’ | 66 | 01000010 |
| ‘C’ | 67 | 01000011 |
| ‘Z’ | 90 | 01011010 |
| ‘a’ | 97 | 01100001 |
| ‘z’ | 122 | 01111010 |
| ‘0’ | 48 | 00110000 |
| ‘9’ | 57 | 00111001 |
| Space | 32 | 00100000 |
| Newline | 10 | 00001010 |
| ‘!’ | 33 | 00100001 |
A clever pattern worth knowing: Uppercase and lowercase letters differ by exactly 32 – or in binary, by a single bit in position 5 (the 32’s place).
A = 01000001, a = 01100001
The only difference is bit 5. This means you can convert between uppercase and lowercase using a single XOR or bit flip operation – something fast string processing routines use constantly.
Example – The word “Hello” in binary:
H = 01001000
e = 01100101
l = 01101100
l = 01101100
o = 01101111
"Hello" = 01001000 01100101 01101100 01101100 01101111
Note on modern encodings: ASCII originally defined 128 characters (7 bits). Modern systems use UTF-8, which extends ASCII to support every character in every human language – Arabic, Chinese, emoji, and everything else. But the first 128 UTF-8 values are identical to ASCII, so all the principles above still apply.
5. Machine Language — The CPU’s Native Tongue
This is where everything comes together. Machine language is the set of binary instructions that the CPU fetches from memory and executes directly. There is no further translation. The processor reads the binary, decodes it using hardware logic, and acts on it in nanoseconds.
Every program you have ever run – operating system, browser, game – was compiled down to this level before it ran.
5.1 What Is Machine Language? (Deep Definition)
Machine language: The set of binary instruction codes that a specific CPU model can execute directly, without any translation, interpretation, or further processing. Each machine instruction is a binary pattern (e.g., B8 05 00 00 00 on x86) that the CPU’s control unit decodes into control signals that activate the ALU, registers, and memory system.
Key characteristics:
- CPU-specific: Machine code written for an x86 CPU will not run on an ARM CPU, and vice versa. The binary format is tied to the Instruction Set Architecture (ISA)
- No portability: A compiled executable for Windows (PE format) will not run on Linux (ELF format) even on the same CPU, because the operating system also defines how programs are loaded
- Direct execution: No interpreter, no just-in-time compiler – the CPU fetches, decodes, and executes each instruction as hardware
5.2 Instruction Structure: Opcode and Operands
Opcode (Operation Code): The part of a machine language instruction that specifies which operation to perform (add, move, jump, compare, etc.). The opcode is a binary number that the CPU’s decoder recognises and maps to the appropriate internal circuit.
Operands: The data that the instruction operates on. An instruction may have zero, one, two, or (rarely) three operands.
| Operand Type | Definition | Example |
|---|---|---|
| Immediate | A constant value encoded directly in the instruction bytes | MOV EAX, 5 – the value 5 is stored in the instruction itself |
| Register | Specifies one of the CPU’s registers | MOV EBX, EAX – both operands refer to registers |
| Memory | Specifies a memory address (direct, indirect, indexed, etc.) | MOV RAX, [RBX] – square brackets mean “the value at this address” |
5.3 Example x86 Machine Code – Broken Down
| Machine Code (hex) | Assembly | Explanation |
|---|---|---|
B8 05 00 00 00 | MOV EAX, 5 | Opcode B8 means “move immediate into EAX”. The next 4 bytes are the little-endian representation of the number 5 |
89 C3 | MOV EBX, EAX | Opcode 89 means “move register to register”. The second byte C3 encodes source (EAX) and destination (EBX) |
01 D8 | ADD EAX, EBX | Opcode 01 is ADD. ModRM byte D8 encodes EBX as source, EAX as destination |
EB 05 | JMP +5 | Opcode EB is short jump (relative, 8-bit offset). 05 means jump forward 5 bytes |
C3 | RET | Single-byte opcode. Pops return address from stack and jumps there |
90 | NOP | No operation – the CPU does nothing for one cycle |
Complete instruction sequence:
B8 05 00 00 00 ; mov eax, 5
B9 03 00 00 00 ; mov ecx, 3
01 C8 ; add eax, ecx (eax becomes 8)
5.4 CPU Architectures – Different Machine Languages
| Architecture | Used in | Characteristics |
|---|---|---|
| x86-64 (AMD64) | Desktops, laptops, servers | CISC (Complex Instruction Set Computer). Variable instruction length (1-15 bytes). Large number of instructions, many can access memory directly |
| ARM (AArch64) | Smartphones, Apple Silicon, AWS Graviton, embedded | RISC (Reduced Instruction Set Computer). Fixed 32-bit instruction length. Load/store architecture – only LDR and STR access memory; all other instructions work on registers. Very energy-efficient |
| RISC-V | Embedded systems, research, emerging products | Open-source, royalty-free ISA. Clean, modern RISC design. Fixed length |
| MIPS | Routers, network equipment, old game consoles | Classic RISC design; popular in computer architecture education |
Important: A binary executable compiled for x86-64 will not run on an ARM processor, and vice versa. The ISA defines the machine language, and they are fundamentally incompatible.
Same operation, different machine code:
| Architecture | Machine Code | Assembly |
|---|---|---|
| x86-64 | B8 05 00 00 00 (5 bytes) | MOV EAX, 5 |
| ARM | E3A00005 (4 bytes, fixed length) | MOV R0, #5 |
5.5 From High-Level Code to Electrical Signals – The Full Journey
High-level source code: x = a + b; (in Python or C)
Compiler: Translates to assembly language – ADD R1, R2, R3
Assembler: Encodes assembly into binary machine code – 00000001 11011000
CPU fetch: The control unit reads those binary bytes from memory
CPU decode: The opcode bits tell the decoder "this is an ADD instruction"
CPU execute: Control signals activate the ALU, which adds the values
Transistors: Billions of transistors switch on and off in response to the
binary values, moving electrical signals through adder circuits
Example – step-by-step for x = 5 + 3 in C:
- C code:
int x = 5 + 3; - Compiler output (x86 assembly):
MOV EAX, 5thenADD EAX, 3thenMOV [x], EAX - Assembler output (hex):
B8 05 00 00 00 83 C0 03 89 45 FC - CPU: fetches
B8, decodes “move immediate to EAX”, fetches the next four bytes as the value 5; then fetches83 C0 03as ADD immediate, etc.
Every abstraction – Python, Java, your operating system – collapses down to this: transistors reading 0s and 1s at billions of operations per second.
6. How a Computer Actually Works (Hardware Fundamentals)
Before writing a single line of assembly, you must understand the hardware. Assembly language is a direct expression of hardware behaviour – you cannot write it well without understanding the machine.
6.1 The CPU and Its Parts
Central Processing Unit (CPU): The primary component of a computer that performs most of the processing. It fetches instructions from memory, decodes them, executes the required operations, and stores the results.
| Component | Deep Definition |
|---|---|
| ALU (Arithmetic Logic Unit) | A digital circuit that performs arithmetic (addition, subtraction, multiplication, division) and logical operations (AND, OR, NOT, XOR, comparisons). It takes two input values (operands), performs the selected operation, and outputs the result. The ALU also sets status flags (zero, carry, overflow, sign) that describe the result. When you write ADD EAX, EBX in assembly, the ALU is the physical circuit that adds the two numbers. |
| Control Unit (CU) | A finite state machine that directs the entire operation of the processor. It fetches the next instruction from memory (using the Program Counter), decodes the opcode to determine what operation to perform, generates control signals that tell the ALU, registers, and memory bus what to do, and then updates the Program Counter. It is the “conductor” of the CPU’s orchestra. |
| Registers | The fastest storage locations in the entire computer – tiny memory cells built directly into the CPU die, operating at the full speed of the processor clock. A modern x86-64 CPU has 16 general-purpose 64-bit registers, plus dozens of special-purpose registers. Accessing a register takes approximately 1 CPU cycle. Accessing main memory (RAM) takes 100–300 cycles because the electrical signals must travel off-chip. This massive speed difference is why assembly programming centres around registers. |
| Program Counter (PC) / Instruction Pointer (RIP) | A special register that always holds the memory address of the next instruction to be fetched and executed. After each instruction is fetched, the control unit automatically increments the Program Counter by the length (in bytes) of that instruction. Jump and call instructions modify the Program Counter directly to implement loops, conditionals, and function calls. |
| Flags Register (RFLAGS) | A special register where each individual bit represents a condition or status of the last arithmetic or comparison operation. For example: the Zero Flag (ZF) is set to 1 if the last result was zero; the Carry Flag (CF) is set if an unsigned addition overflowed; the Sign Flag (SF) reflects the most significant bit of the result (indicating negativity in signed arithmetic). |
6.2 Memory Hierarchy – Deep Definition
Memory hierarchy: The organisation of computer storage from the fastest, smallest, most expensive (closest to the CPU) to the slowest, largest, cheapest (farthest from the CPU).
| Level | Typical size | Access time (cycles) | Technology | Definition |
|---|---|---|---|---|
| Registers | ~128 bytes total | 1 | Flip-flops inside CPU | Storage locations built into the CPU’s execution core. The CPU operates directly on register values without any memory bus delay. |
| L1 cache | 32–64 KB per core | 4–5 | SRAM on-chip | Level-1 cache is split into instruction cache (L1i) and data cache (L1d). It holds the most frequently accessed data and code. |
| L2 cache | 256 KB – 1 MB per core | 12–15 | SRAM on-chip | Larger but slower than L1. Acts as a backup for data that does not fit in L1. |
| L3 cache | 8–64 MB shared | 30–40 | SRAM on-chip | Shared among multiple CPU cores. Larger than L2 but slower. |
| RAM (DRAM) | 8–32 GB | 100–300 | Dynamic RAM off-chip | Main memory. Stores the program code and data currently in use. Volatile – loses all data when power is turned off. |
| SSD/NVMe | 256 GB – 2 TB | 100,000+ | Flash memory | Non-volatile persistent storage. Much slower than RAM but much larger. |
| Hard disk | 1–10 TB | 10,000,000+ | Magnetic platters | Slowest persistent storage. Mechanical parts must physically move to read data. |
Practical implication for assembly programming: If your data fits in registers and your code fits in L1 cache, your program runs at the full speed of the CPU. If you access a memory location not in any cache, the CPU must stall and wait for RAM – this can make your program 100 times slower, no matter how clever your instructions are.
6.3 The Fetch-Decode-Execute Cycle – Step by Step
Fetch-Decode-Execute cycle: The fundamental operational loop of every stored-program computer. The CPU continuously repeats three steps for as long as it is powered on:
- Fetch – The control unit reads the memory address stored in the Program Counter (RIP on x86-64). It sends that address to the memory system, and the bytes of the instruction are loaded into the CPU’s instruction register. The Program Counter is then advanced to the next address.
- Decode – The control unit examines the opcode (operation code) – the first byte or bytes of the instruction – to determine what operation to perform. It also parses the addressing modes: where are the operands? In registers? In memory? As immediate constants encoded in the instruction? This decoding step generates the internal control signals that will be sent to the ALU, registers, and memory bus.
- Execute – The appropriate functional unit carries out the operation. The ALU performs arithmetic; the memory controller loads from or stores to RAM; the branch unit updates the Program Counter for jumps. The result is written to the destination (register or memory). Then the cycle repeats with the next instruction.
Modern complexity: On modern superscalar CPUs, multiple instructions are fetched, decoded, and executed simultaneously through pipelining and out-of-order execution. However, the logical model of fetch-decode-execute remains the correct mental model for understanding assembly programming.
6.4 Von Neumann Architecture – The Blueprint of Every Computer
Von Neumann architecture: A computer architecture proposed by John von Neumann in 1945, where both program instructions and data are stored together in the same memory space and accessed through the same bus. This is the design used by virtually all modern computers.
Consequences of storing code and data together:
- Self-modifying code – A program can write new instructions into memory and then execute them (rarely used today, but possible)
- Just-in-time compilation – A runtime system can generate machine code for a high-level language and then jump to it, as JavaScript engines do
- Security vulnerabilities – Because data can be placed in memory and then accidentally executed as code, attacks like buffer overflows and return-oriented programming (ROP) are possible
The four main components:
- CPU – Contains the ALU and Control Unit
- Memory – Stores both program instructions and data in a single address space
- Input/Output – Devices that communicate with the outside world (keyboard, display, disk)
- Bus – The set of electrical wires that connect all components and carry data, addresses, and control signals
PART TWO: ASSEMBLY LANGUAGE WITH NASM
7. Assembly Language with NASM — Complete Guide
7.1 What Is Assembly Language? (Deep Definition)
Assembly language: A human-readable notation for machine language. Each assembly instruction (like MOV RAX, 5) corresponds directly to exactly one machine instruction (like B8 05 00 00 00). The assembler (e.g., NASM) translates these mnemonics into binary machine code.
Comparison of levels:
| Level | Example | Translator | Execution |
|---|---|---|---|
| High-level (C) | x = a + b; | Compiler → assembly → machine code | Indirect |
| Assembly | ADD RAX, RBX | Assembler → machine code | One-to-one |
| Machine code | 01 D8 | None – direct | CPU executes directly |
7.2 Why Learn Assembly? (Expanded)
- Complete hardware control – You decide exactly which registers to use, exactly how memory is accessed, and exactly what the CPU does at every step
- Maximum performance – Hand-written assembly can be faster than compiled code for critical inner loops
- Understanding compilers – When you read compiler output (disassembly), you understand why your C code generates certain instructions
- Reverse engineering and security – Malware analysis, vulnerability research, and exploit development all require reading and writing assembly
- Embedded and OS development – Bootloaders, operating system kernels, and firmware require assembly for the portions that run before any runtime environment exists
7.3 NASM Setup and Installation
Installing NASM:
On Linux (Ubuntu/Debian):
sudo apt update
sudo apt install nasm
On macOS (with Homebrew):
brew install nasm
On Windows: Download the installer from nasm.us, run it, and add the NASM directory to your system PATH.
Verify the installation with:
nasm --version
Assembling and Linking:
On Linux (64-bit):
nasm -f elf64 program.asm -o program.o
ld program.o -o program
./program
On macOS (64-bit):
nasm -f macho64 program.asm -o program.o
ld program.o -o program -macosx_version_min 10.13 -lSystem
./program
The -f flag specifies the output format – elf64 for Linux, macho64 for macOS.
7.4 Program Structure – Sections Deep Explanation
Section (or segment): A named region of an assembly program with a specific purpose: code, initialised data, or uninitialised data.
section .data
; Initialised data – variables that already have values when the program starts.
; These bytes are stored in the executable file and loaded into memory at runtime.
section .bss
; Uninitialised data – space reserved for variables that will be set at runtime.
; No space is taken in the executable; the loader allocates zeroed memory.
section .text
global _start ; Make the _start label visible to the linker as the entry point.
; Your code – executable instructions.
Detailed example of each section:
section .data
message db "Hello, World!", 10 ; db = define byte. 10 is ASCII newline.
number dw 1234 ; dw = define word (2 bytes)
pi dd 3.14 ; dd = define double word (4 bytes)
section .bss
buffer resb 64 ; reserve 64 bytes. Uninitialised – content is undefined.
counter resd 1 ; reserve 1 double word (4 bytes) for an integer counter.
section .text
global _start
_start:
; Code goes here
Data directives summary:
| Directive | Size | Full name |
|---|---|---|
db | 1 byte | Define Byte |
dw | 2 bytes | Define Word |
dd | 4 bytes | Define Double Word |
dq | 8 bytes | Define Quad Word |
resb n | n bytes | Reserve Bytes |
resw n | n words | Reserve Words |
resd n | n double words | Reserve Double Words |
7.5 Registers – Detailed Reference
General-purpose register: A storage location inside the CPU that can hold data (integers, addresses) and can be used as an operand in most arithmetic and logic instructions. The x86-64 architecture provides 16 general-purpose registers, each 64 bits wide.
Sub-register naming (for partial access):
| 64-bit | 32-bit | 16-bit | 8-bit low | 8-bit high | Conventional use |
|---|---|---|---|---|---|
| RAX | EAX | AX | AL | AH | Accumulator – arithmetic, return value |
| RBX | EBX | BX | BL | BH | Base – general, preserved across calls |
| RCX | ECX | CX | CL | CH | Counter – loop counts, shift amounts |
| RDX | EDX | DX | DL | DH | Data – I/O, extended arithmetic |
| RSI | ESI | SI | SIL | (none) | Source index – string/memory source |
| RDI | EDI | DI | DIL | (none) | Destination index – string/memory dest |
| RSP | ESP | SP | SPL | (none) | Stack pointer – top of stack |
| RBP | EBP | BP | BPL | (none) | Base pointer – stack frame base |
| R8-R15 | R8D-R15D | R8W-R15W | R8B-R15B | (none) | Extra general-purpose registers |
Visual of RAX breakdown:
RAX (64 bits)
┌──────────────────────────┬──────────────────────────┐
│ Upper 32 │ EAX │
└──────────────────────────┴──────────┬────────────────┘
│ AX │
├────────┬────────┤
│ AH │ AL │
└────────┴────────┘
Important rules for partial register access:
- Writing to a 32-bit register (e.g., EAX) zeroes the upper 32 bits of the full 64-bit register (RAX). This is a common source of bugs.
- Writing to a 16-bit (AX) or 8-bit (AL, AH) register does not clear the remaining bits – they retain their previous values.
7.6 Core Instructions – Expanded with Definitions
MOV – The Most Fundamental Instruction
Definition: MOV destination, source copies the value from source to destination. The source remains unchanged. This is a copy, not a move – the original data stays where it was.
Allowed operand combinations:
- Register ← immediate
- Register ← register
- Memory ← register
- Register ← memory
Not allowed: Memory ← memory directly (must go through a register).
Examples:
MOV RAX, 10 ; Immediate → register
MOV RBX, RAX ; Register → register
MOV [var], RAX ; Register → memory (store)
MOV RAX, [var] ; Memory → register (load)
Critical rule: Both operands must be the same size. You cannot MOV a 64-bit register into an 8-bit memory location without using a smaller register first.
ADD and SUB
Definition: ADD dest, src adds src to dest and stores the result in dest. SUB dest, src subtracts src from dest and stores the result in dest. Both update the flags register (ZF, SF, CF, OF) based on the result.
MOV RAX, 15
MOV RBX, 7
ADD RAX, RBX ; RAX = 22
SUB RAX, 5 ; RAX = 17
INC and DEC
Definition: INC reg adds 1 to the register. DEC reg subtracts 1. They are compact (shorter encoding) than ADD reg, 1.
INC RAX ; RAX = RAX + 1
DEC RBX ; RBX = RBX - 1
MUL and IMUL – Multiplication
MUL source (unsigned multiplication): Multiplies RAX by source and stores the result in RDX:RAX (a 128-bit value split across two 64-bit registers). RDX holds the high 64 bits, RAX holds the low 64 bits.
MOV RAX, 6
MOV RBX, 7
MUL RBX ; RAX = 42, RDX = 0 (result fits in 64 bits)
IMUL (signed multiplication): More flexible syntax. Can multiply any two registers and optionally a third immediate.
IMUL RAX, RBX ; RAX = RAX × RBX
IMUL RAX, RBX, 10 ; RAX = RBX × 10
DIV and IDIV – Division
DIV source (unsigned division): Divides the 128-bit value in RDX:RAX (high 64 bits in RDX, low 64 bits in RAX) by source. Quotient goes to RAX, remainder to RDX.
Critical: You must zero out RDX before an unsigned division unless you are intentionally dividing a 128-bit value.
; Divide 20 by 4
MOV RAX, 20
XOR RDX, RDX ; Zero RDX (same as MOV RDX,0)
MOV RBX, 4
DIV RBX ; RAX = 5, RDX = 0
; Divide 17 by 5 → quotient 3, remainder 2
MOV RAX, 17
XOR RDX, RDX
MOV RBX, 5
DIV RBX ; RAX = 3, RDX = 2
7.7 Flags and Conditional Jumps – Deep Explanation
Flags: Individual bits in the RFLAGS register that are set or cleared automatically by the CPU after arithmetic and comparison instructions. They store metadata about the result: was it zero? Was it negative? Did it overflow?
Key flags:
| Flag | Name | Set when |
|---|---|---|
| ZF (Zero Flag) | Zero | Result of last operation was exactly zero |
| SF (Sign Flag) | Sign | Result is negative (most significant bit = 1) |
| CF (Carry Flag) | Carry | Unsigned addition produced a carry out of the MSB, or unsigned subtraction required a borrow |
| OF (Overflow Flag) | Overflow | Signed addition produced a result outside the representable range |
| PF (Parity Flag) | Parity | The result has an even number of 1 bits |
Example flag behaviour:
mov rax, 5
sub rax, 5 ; RAX = 0 → ZF = 1 (set)
mov al, 127 ; Maximum signed 8-bit value
add al, 1 ; 127+1=128 → in 8-bit signed, this is -128 → OF = 1
mov al, 255 ; Maximum unsigned 8-bit value
add al, 1 ; 255+1=0 (wraps) → CF = 1
CMP – The Comparison Instruction
Definition: CMP dest, src subtracts src from dest, updates the flags, and discards the result. It is used solely to set flags for a subsequent conditional jump.
CMP RAX, RBX ; Compute RAX - RBX, set flags, throw away result.
Conditional Jumps – The Decision Makers
| Instruction | Condition | Flags | Meaning |
|---|---|---|---|
JE / JZ | Equal / Zero | ZF=1 | Jump if equal (signed or unsigned) |
JNE / JNZ | Not equal | ZF=0 | Jump if not equal |
JG / JNLE | Greater (signed) | ZF=0 and SF=OF | Signed greater than |
JGE / JNL | Greater or equal (signed) | SF=OF | Signed greater or equal |
JL / JNGE | Less (signed) | SF≠OF | Signed less than |
JLE / JNG | Less or equal (signed) | ZF=1 or SF≠OF | Signed less or equal |
JA / JNBE | Above (unsigned) | CF=0 and ZF=0 | Unsigned greater than |
JB / JC | Below (unsigned) | CF=1 | Unsigned less than |
JMP | Unconditional | – | Always jump |
Practical example:
MOV RAX, 10
CMP RAX, 10
JE equal ; ZF is 1 because 10-10=0, so we jump
JMP not_equal ; This line is skipped
equal:
; code for equality
7.8 Loops – Building Iteration from Jumps
Loop (in assembly): There is no dedicated loop keyword. A loop is constructed manually using a counter, a label, and conditional jumps that jump backward to the label while the counter is not zero.
Basic loop pattern:
MOV RCX, 5 ; counter = 5
loop_start:
; loop body
DEC RCX
CMP RCX, 0
JNZ loop_start ; if RCX != 0, jump back
; loop ends
Using the LOOP instruction: LOOP label automatically decrements RCX and jumps to label if RCX ≠ 0. It combines DEC RCX, CMP RCX, 0, and JNZ into one instruction.
MOV RCX, 10
sum_loop:
ADD RAX, RCX
LOOP sum_loop ; RCX--; if RCX != 0, jump
; RAX now holds 55 (1+2+...+10)
Nested loops: You must save and restore RCX because the inner loop will overwrite it.
MOV RCX, 3 ; outer counter
outer:
PUSH RCX ; save outer counter
MOV RCX, 4 ; inner counter
inner:
; inner work
LOOP inner
POP RCX ; restore outer counter
LOOP outer
7.9 Memory Addressing Modes – Accessing Data in RAM
Addressing mode: The method by which an instruction specifies the location of its operand. Square brackets [] in NASM mean “the value at this memory address” – they are the dereference operator.
| Mode | Example | Definition |
|---|---|---|
| Immediate | MOV RAX, 42 | The value 42 itself is encoded in the instruction. No memory access. |
| Register | MOV RBX, RAX | The operand is the value inside a register. |
| Direct memory | MOV RAX, [number] | The address of number is fixed; load the value stored at that address. |
| Register indirect | MOV RAX, [RBX] | RBX holds a memory address; load the value at that address. |
| Base + offset | MOV RAX, [RBX + 8] | Start at address in RBX, add 8 bytes, then load. Used for struct fields. |
| Base + index × scale | MOV RAX, [RBX + RCX*8] | Used for array indexing. Scale is element size (1,2,4,8). |
7.10 String Operations – Efficient Block Processing
String instructions: Specialised x86 instructions that operate on blocks of memory (strings) using RSI as source pointer and RDI as destination pointer. They automatically increment or decrement the pointers based on the direction flag (DF). The REP prefix repeats the instruction RCX times.
Core string instructions:
| Instruction | Effect |
|---|---|
MOVSB | [RDI] = [RSI]; then RSI++, RDI++ (copy one byte) |
STOSB | [RDI] = AL; then RDI++ (store AL into memory) |
LODSB | AL = [RSI]; then RSI++ (load from memory into AL) |
CMPSB | Compare [RSI] with [RDI], set flags; then RSI++, RDI++ |
SCASB | Compare AL with [RDI], set flags; then RDI++ |
Example – Copy 100 bytes:
cld ; clear direction flag (process left to right)
mov rsi, src_data ; RSI points to source
mov rdi, dst_data ; RDI points to destination
mov rcx, 100 ; count
rep movsb ; repeat MOVSB RCX times
Example – Find length of null-terminated string:
cld
mov rdi, src_data
mov al, 0 ; search for null byte
mov rcx, 256 ; maximum length
repne scasb ; repeat while not equal (scan until AL matches)
; After loop: RCX = (max - length - 1)
; Length = max - RCX - 1
7.11 The Stack – PUSH, POP, and Function Calls
Stack: A region of memory that operates on the Last-In, First-Out (LIFO) principle. It is used to store temporary values, return addresses, function arguments, and saved register states. The stack pointer (RSP) always points to the current top of the stack. On x86-64, the stack grows downward (toward lower memory addresses).
- PUSH:
PUSH srcdecrements RSP by 8 (for 64-bit) and writes src to the memory address now pointed to by RSP - POP:
POP destreads the value from the memory address pointed to by RSP, then increments RSP by 8
PUSH RAX ; RSP = RSP - 8; [RSP] = RAX
POP RBX ; RBX = [RSP]; RSP = RSP + 8
Critical rule: Every PUSH must eventually have a matching POP (or an adjustment to RSP), otherwise the stack becomes unbalanced and your program will crash when it tries to return.
Saving and restoring registers:
_start:
MOV RAX, 100
PUSH RAX ; save RAX
MOV RAX, 999 ; use RAX for something else
; ... work ...
POP RAX ; restore original RAX (now 100)
7.12 Procedures – CALL and RET
Definition – CALL label: Pushes the address of the next instruction (the return address) onto the stack, then jumps to label. This allows the procedure to return later.
Definition – RET: Pops the return address from the stack and jumps to that address, resuming execution after the original CALL.
Simple procedure:
_start:
CALL my_proc ; pushes return address, jumps
; execution resumes here after RET
my_proc:
MOV RAX, 42
RET ; pops return address, jumps back
Function prologue and epilogue (for stack frames and local variables):
my_function:
; Prologue
PUSH RBP ; save caller's base pointer
MOV RBP, RSP ; set our own base pointer
SUB RSP, 32 ; allocate 32 bytes for local variables
; Function body – access locals at [RBP-8], [RBP-16], etc.
MOV QWORD [RBP-8], 10
; Epilogue
MOV RSP, RBP ; deallocate locals
POP RBP ; restore caller's base pointer
RET
Linux 64-bit Calling Convention (System V AMD64 ABI):
When your assembly procedure is called from C code, or when you call C functions from assembly, both sides must agree on how arguments are passed and where return values live.
On Linux 64-bit:
- First 6 integer arguments are passed in registers: RDI, RSI, RDX, RCX, R8, R9
- Additional arguments are pushed onto the stack
- Return value is in RAX (or RDX:RAX for 128-bit values)
- Caller-saved registers (caller must preserve if needed): RAX, RCX, RDX, RSI, RDI, R8, R9, R10, R11
- Callee-saved registers (function must preserve): RBX, RBP, R12, R13, R14, R15
7.13 System Calls – Talking to the Operating System
System call: A controlled way for a user-mode program to request a service from the operating system kernel (e.g., reading from a file, writing to the screen, allocating memory). The CPU transitions from user mode (Ring 3) to kernel mode (Ring 0) to execute the request safely.
Linux x86-64 system call convention:
- System call number → RAX
- Arguments (max 6) → RDI, RSI, RDX, R10, R8, R9 (in that order)
- Execute syscall instruction
- Return value → RAX (negative if error)
Common Linux system call numbers (x86-64):
| System Call | Number (RAX) | RDI | RSI | RDX |
|---|---|---|---|---|
sys_read | 0 | file descriptor | buffer address | byte count |
sys_write | 1 | file descriptor | buffer address | byte count |
sys_open | 2 | filename address | flags | mode |
sys_close | 3 | file descriptor | — | — |
sys_exit | 60 | exit code | — | — |
File descriptors: 0 = stdin, 1 = stdout, 2 = stderr
Complete “Hello, World!” with explanations:
section .data
msg db "Hello, NASM!", 10 ; 10 = ASCII newline
len equ $ - msg ; $ = current address; subtract start → length
section .text
global _start
_start:
; sys_write(1, msg, len)
mov rax, 1 ; syscall number for write
mov rdi, 1 ; file descriptor 1 = stdout
mov rsi, msg ; pointer to string
mov rdx, len ; number of bytes to write
syscall ; call kernel
; sys_exit(0)
mov rax, 60 ; syscall number for exit
xor rdi, rdi ; exit code 0 (xor with itself is a common way to zero a register)
syscall
7.14 Memory Management at the Low Level – Deep Dive
7.14.1 Memory Map of a Running Process
Virtual address space: Each process on a modern operating system has its own private, isolated view of memory, called the virtual address space. The OS and CPU (via the Memory Management Unit, MMU) map virtual addresses to physical RAM addresses.
Typical layout of a 64-bit Linux process (simplified):
High addresses
┌─────────────────────────────────┐
│ Kernel space (inaccessible) │ ← user code cannot read/write this
├─────────────────────────────────┤
│ Environment variables & args │
├─────────────────────────────────┤
│ Stack (grows DOWNWARD) │ ← RSP points here
│ [stack frames, local vars] │
├─────────────────────────────────┤
│ (unmapped gap – segfault if accessed) │
├─────────────────────────────────┤
│ Memory-mapped files / libraries │ ← shared libraries, mmap()
├─────────────────────────────────┤
│ Heap (grows UPWARD) │ ← malloc() / new
├─────────────────────────────────┤
│ BSS (uninitialised globals) │ ← zeroed at startup
├─────────────────────────────────┤
│ Data (initialised globals) │ ← values set in source
├─────────────────────────────────┤
│ Text (code, read+exec) │ ← your instructions
├─────────────────────────────────┤
│ NULL region (unmapped) │ ← accessing 0x0 causes segfault
Low addresses
7.14.2 Stack vs Heap – Detailed Comparison
| Feature | Stack | Heap |
|---|---|---|
| Allocation | Automatic – when you call a function, the stack frame is created; when you return, it is destroyed | Manual – you must call malloc (C) or new (C++) or use system calls like mmap |
| Speed | Extremely fast – just moving RSP (a single subtraction or addition) | Slower – the allocator must find a free block, possibly asking the OS for more memory |
| Lifetime | Tied to function scope – local variables exist only while the function is active | Until explicitly freed with free or delete. Can outlive the function that allocated it |
| Size | Limited – typically 1–8 MB per thread. Exceeding it causes a stack overflow | Large – limited by available RAM and swap |
| Fragmentation | None – stack allocations are strictly LIFO, so no fragmentation | Can become fragmented over time if blocks are allocated and freed in different orders |
| Usage | Local variables, return addresses, saved registers | Data that must outlive a function call, large arrays, dynamically sized structures |
7.14.3 Buffer Overflows – The Classic Low-Level Vulnerability
Buffer overflow: A condition where a program writes more data into a fixed-size buffer than the buffer can hold, causing the excess data to overwrite adjacent memory.
Stack layout during a vulnerable function:
Low addresses
[buffer[0..63]] ← 64 bytes allocated
[saved RBP] ← 8 bytes (if frame pointer is used)
[return address] ← 8 bytes ← overflow reaches here at 72 bytes
High addresses
Why it is dangerous: If an attacker can control the data written into buffer, they can overwrite the return address with the address of malicious code (shellcode) that they also inject, or they can chain together existing code snippets (Return-Oriented Programming).
Modern defenses:
| Defense | Definition |
|---|---|
| Stack canaries | A random value placed between the buffer and the return address. If a buffer overflow occurs, the canary is overwritten. Before returning, the program checks the canary; if it changed, the program aborts. |
| ASLR (Address Space Layout Randomization) | The OS randomises the base addresses of the stack, heap, and shared libraries each time a program runs. This makes it difficult for an attacker to predict where their shellcode or ROP gadgets are located. |
| NX bit (No-Execute) | Memory pages are marked as either writable or executable, never both. This prevents an attacker from placing shellcode in the stack or heap and then executing it. |
| CFI (Control-Flow Integrity) | The compiler inserts checks before every indirect jump/call to ensure the target address is a valid function entry point, breaking ROP attacks. |
8. Complete Example Programs
Program 1 – Print “Hello, World!”
section .data
msg db "Hello, World!", 10
len equ $ - msg
section .text
global _start
_start:
; write to stdout
MOV RAX, 1
MOV RDI, 1
MOV RSI, msg
MOV RDX, len
syscall
; exit cleanly
MOV RAX, 60
MOV RDI, 0
syscall
What each line does:
section .data– defines the data sectionmsg db "Hello, World!", 10– defines a byte sequence with the string and a newline (ASCII 10)len equ $ - msg–$is the current address, so this calculates the length of the stringMOV RAX, 1– system call number forsys_writeMOV RDI, 1– file descriptor 1 = stdoutMOV RSI, msg– pointer to the stringMOV RDX, len– number of bytes to writesyscall– invoke the kernelMOV RAX, 60– system call number forsys_exitMOV RDI, 0– exit code 0 = successsyscall– invoke the kernel again
Program 2 – Add Two Numbers and Exit with the Result
section .text
global _start
_start:
MOV RAX, 14 ; first number
MOV RBX, 28 ; second number
ADD RAX, RBX ; RAX = 14 + 28 = 42
MOV RDI, RAX ; exit code = result
MOV RAX, 60 ; sys_exit
syscall
After running this, check the exit code with echo $? in your terminal. It will print 42.
Program 3 – Loop and Count
section .data
msg db "Counting...", 10
len equ $ - msg
section .text
global _start
_start:
; print message
MOV RAX, 1
MOV RDI, 1
MOV RSI, msg
MOV RDX, len
syscall
; count from 10 down to 1 using LOOP
MOV RCX, 10
count_loop:
PUSH RCX ; save loop counter before syscall clobbers registers
; (in a full program you would print RCX here)
POP RCX
LOOP count_loop
; exit
MOV RAX, 60
MOV RDI, 0
syscall
Program 4 – Simple Procedure Call
section .text
global _start
; Procedure: square
; Input: RDI = number to square
; Output: RAX = RDI * RDI
square:
MOV RAX, RDI
IMUL RAX, RDI
RET
_start:
MOV RDI, 7 ; argument: square the number 7
CALL square ; call procedure – result in RAX
; RAX = 49
MOV RDI, RAX ; exit with result
MOV RAX, 60
syscall
Check with echo $? – it prints 49.
9. Performance Optimization at the Low Level
9.1 Avoid Data Hazards – Let the Pipeline Breathe
Pipeline hazard: A condition where the next instruction cannot execute in the next clock cycle because it depends on the result of a previous instruction that has not yet completed. The CPU must insert “bubble” cycles (stalls), reducing performance.
Bad – dependent chain:
mov rax, [mem] ; load takes 200 cycles if cache miss
add rax, 5 ; must wait for load to complete
add rax, 10 ; must wait for previous add
Better – independent operations:
mov rax, [mem] ; load element 0
mov rbx, [mem+8] ; independent load – can happen in parallel
mov rcx, [mem+16] ; independent
add rax, 5 ; operates on rax, may overlap with other loads
add rbx, 10
add rcx, 15
9.2 Cache Locality – The #1 Optimisation
Locality of reference: The tendency of a program to access the same memory addresses repeatedly (temporal locality) or to access addresses that are close to each other (spatial locality). Caches exploit both forms.
Spatial locality example – row-major vs column-major traversal:
// BAD: column-major traversal – jumps by entire row size each time
for (int col = 0; col < 1000; col++)
for (int row = 0; row < 1000; row++)
sum += matrix[row][col]; // memory stride = 1000*4 = 4000 bytes
// GOOD: row-major traversal – sequential memory access
for (int row = 0; row < 1000; row++)
for (int col = 0; col < 1000; col++)
sum += matrix[row][col]; // accesses consecutive addresses
On large matrices, the second version can be 10x faster or more because it makes optimal use of cache lines.
9.3 Reading Compiler Output – Learn from the Best
Modern compilers (GCC, Clang) generate highly optimised assembly, often better than what a beginner would write by hand. Use this to learn.
# Generate assembly from C with optimisations
gcc -O2 -S myfile.c -o myfile.s
# View disassembly of an executable
objdump -d -M intel myprogram | less
# Use Compiler Explorer online – paste code, see assembly instantly
# https://godbolt.org
10. Tools Every Low-Level Programmer Needs
| Tool | Purpose | Definition / Use |
|---|---|---|
| NASM | Assembler | Translates assembly source into object files (.o). Use -f elf64 for Linux 64-bit. |
| GDB | GNU Debugger | Allows you to step through assembly instructions, inspect registers, examine memory, set breakpoints. Essential for understanding what your code actually does. |
| objdump | Disassembler | Displays machine code from object files or executables in human-readable assembly. objdump -d -M intel program |
| strace | System call tracer | Shows every system call a program makes, along with arguments and return values. strace ./program |
| ltrace | Library call tracer | Shows calls to dynamic libraries (e.g., printf, malloc). |
| readelf | ELF file examiner | Displays detailed information about ELF executable headers, sections, symbols. |
| xxd / hexdump | Hex dump utilities | Show raw binary data as hexadecimal. xxd program or hexdump -C program |
| Radare2 | Reverse engineering framework | Advanced disassembler, debugger, and binary analysis tool. |
| Compiler Explorer | Online tool | godbolt.org – instantly see the assembly generated by any compiler for any language. |
| CPUlator | CPU simulator | Browser-based CPU simulator. Write assembly code, step through it instruction by instruction, and watch registers and memory change in real time. |
| Intel x86 reference | Manual | felixcloutier.com/x86 – complete instruction set reference. |
GDB Quick Reference for Assembly
gdb ./program
(gdb) layout asm # show disassembly view
(gdb) layout regs # show register window
(gdb) break _start # break at symbol
(gdb) break *0x401000 # break at absolute address
(gdb) run
(gdb) stepi # execute one instruction (step into)
(gdb) nexti # execute one instruction (step over calls)
(gdb) info registers # show all registers
(gdb) info registers rax rbx
(gdb) x/10xb $rsp # examine 10 bytes in hex at stack pointer
(gdb) x/4gx $rsp # examine 4 quad-words (8 bytes each)
(gdb) x/s 0x402000 # examine as string
(gdb) disassemble # disassemble current function
(gdb) continue # run until next breakpoint
11. NASM Directives and Useful Features
Constants with EQU
EQU defines a constant name for a value. Unlike variables, constants cannot change – they are substituted at assembly time.
STDOUT equ 1
NEWLINE equ 10
SYS_WRITE equ 1
SYS_EXIT equ 60
; now use the names instead of magic numbers
MOV RAX, SYS_WRITE
MOV RDI, STDOUT
TIMES – Repeat Data
section .data
zeros times 16 db 0 ; 16 zero bytes
buffer times 64 db ' ' ; 64 space characters
Macros
NASM macros let you define reusable code patterns with parameters. They are expanded at assembly time – no function call overhead.
; Define a macro that prints a message
%macro print 2 ; 2 parameters: address and length
MOV RAX, 1
MOV RDI, 1
MOV RSI, %1 ; first argument
MOV RDX, %2 ; second argument
syscall
%endmacro
section .data
msg db "Using a macro!", 10
len equ $ - msg
section .text
global _start
_start:
print msg, len ; single clean line instead of 4 MOV + syscall
MOV RAX, 60
MOV RDI, 0
syscall
External Labels and Linking with C
You can call C standard library functions from NASM and vice versa. This lets you use printf, malloc, and other libc functions from assembly.
extern printf ; declare printf as external
section .data
fmt db "Value: %d", 10, 0 ; printf format string, null-terminated
section .text
global _start
_start:
MOV RDI, fmt ; first arg: format string
MOV RSI, 42 ; second arg: value to print
XOR RAX, RAX ; RAX = 0 (no floating point args)
CALL printf
MOV RAX, 60
MOV RDI, 0
syscall
Link with: nasm -f elf64 prog.asm -o prog.o && gcc prog.o -o prog -no-pie
Quick Reference – Most Used Instructions
| Instruction | Syntax | Effect |
|---|---|---|
MOV | MOV dst, src | dst = src |
ADD | ADD dst, src | dst = dst + src |
SUB | SUB dst, src | dst = dst − src |
INC | INC dst | dst = dst + 1 |
DEC | DEC dst | dst = dst − 1 |
MUL | MUL src | RDX:RAX = RAX × src |
IMUL | IMUL dst, src | dst = dst × src (signed) |
DIV | DIV src | RAX = RDX:RAX ÷ src, RDX = remainder |
AND | AND dst, src | dst = dst AND src (bitwise) |
OR | OR dst, src | dst = dst OR src (bitwise) |
XOR | XOR dst, src | dst = dst XOR src (bitwise) |
NOT | NOT dst | dst = bitwise NOT dst |
SHL | SHL dst, n | dst = dst << n (left shift) |
SHR | SHR dst, n | dst = dst >> n (right shift) |
CMP | CMP a, b | set flags from a − b, discard result |
JMP | JMP label | unconditional jump |
JE | JE label | jump if equal (ZF=1) |
JNE | JNE label | jump if not equal (ZF=0) |
JG | JG label | jump if greater (signed) |
JL | JL label | jump if less (signed) |
PUSH | PUSH src | RSP−=8; [RSP]=src |
POP | POP dst | dst=[RSP]; RSP+=8 |
CALL | CALL label | push return addr; jump to label |
RET | RET | pop return addr; jump there |
syscall | syscall | invoke OS kernel (Linux x86-64) |
NOP | NOP | do nothing for one cycle |
HLT | HLT | halt the CPU |
Conclusion
Assembly language removes every layer of abstraction between you and the CPU. When you write NASM, you are placing exact instructions into memory for the processor to execute. There is no runtime, no garbage collector, no virtual machine – just your binary and the hardware.
The journey you have taken through this guide covers the complete foundation:
- Number systems – binary, octal, hexadecimal, and two’s complement
- Binary arithmetic – addition, subtraction, multiplication, division
- Bitwise operations – AND, OR, XOR, NOT, and shifts
- Machine language – opcodes, operands, and how the CPU executes them
- Hardware fundamentals – CPU, memory hierarchy, fetch-decode-execute cycle
- NASM assembly – program structure, registers, instructions, flags, loops, memory addressing, stack, procedures, and system calls
- Performance optimization – data hazards, cache locality, reading compiler output
- Security – buffer overflows and modern defenses
What makes assembly genuinely rewarding is the clarity it provides. When your program works, you know exactly why. When it breaks, you know exactly where to look. There is no mystery – just instructions, registers, and memory.
“The computer was born to solve problems that did not exist before.” — Bill Gates