Low-Level Languages

Low-Level Languages

Content Overview

Ultimate Guide to Low-Level Programming: From Binary to Assembly

Table of Contents

  1. Introduction: Why Bits and Bytes Rule the World
  2. Number Systems — The Language of Bits
  3. Binary Arithmetic and Bitwise Operations
  4. ASCII — How Text Becomes Binary
  5. Machine Language — The CPU’s Native Tongue
  6. How a Computer Actually Works (Hardware Fundamentals)
  7. Assembly Language with NASM — Complete Guide
  8. Practical Assembly Programs
  9. Performance Optimization at the Low Level
  10. Tools Every Low-Level Programmer Needs
  11. Conclusion and Next Steps

PART ONE: THE FOUNDATIONS

1. Introduction: Why Bits and Bytes Rule the World

1.1 Binary: The Foundation of All Computation

Binary is a numerical system that uses only two digits: 0 and 1. In electronics, these correspond directly to two physical states of a transistor: off (no current, representing 0) and on (current flowing, representing 1).

Because every modern digital circuit is built from billions of transistors that can only be off or on, binary is the only language the hardware can natively speak. Any information – text, images, video, program code – must ultimately be encoded as sequences of binary digits (bits) before the CPU can process it.

Machine Language is the set of binary instruction codes that a specific CPU model understands and can execute directly, without any translation or interpretation. Each instruction is a binary pattern (e.g., 10110000 01100001) that the CPU’s control unit decodes into signals that activate the correct internal circuits.

Why this matters: Every program you have ever run, from a Python script to a video game to your operating system, is eventually converted into machine language before the CPU touches it. Understanding machine language means understanding the ultimate reality of all software.

1.2 What Is a Low-Level Language?

A low-level language is a programming language that provides little or no abstraction from the computer’s hardware. “Abstraction” means hiding hardware details behind convenient high-level constructs. For example, Python’s print("Hello") hides the system call, the buffer management, and the character encoding. A low-level language exposes those details: you control exactly which registers are used, exactly which memory addresses are accessed, and exactly which CPU instructions are executed.

Important: “Low-level” refers to the level of proximity to the machine, not to difficulty. Writing correct, efficient low-level code requires deep understanding of computer architecture and is one of the most challenging skills in software development.

Two Main Categories

CategoryAlso calledDefinition
Machine languageFirst-generation (1GL)The native binary instruction set of a specific CPU. Consists entirely of 1s and 0s (or hexadecimal representation). Executed directly by hardware – no translation needed.
Assembly languageSecond-generation (2GL)A human-readable notation for machine language. Uses short text mnemonics (e.g., MOV, ADD, JMP) to represent each machine instruction. An assembler program translates these mnemonics into binary machine code. There is a one-to-one correspondence between an assembly instruction and a machine instruction.

Low-Level vs High-Level – Detailed Comparison

FeatureLow-levelHigh-level (e.g., Python, Java)
AbstractionNone or minimal – you see every CPU cycleExtensive – memory management, types, I/O are hidden
Memory controlComplete, manual – you allocate and free every byteLimited or automatic (garbage collection)
PortabilityNon-portable – code written for x86 will not run on ARMHighly portable – same source runs on many CPUs
PerformanceMaximum possible – no hidden overheadGood to moderate – interpreter or JIT adds overhead
Development speedSlow – many instructions needed for simple tasksFast – one line of Python does the work of dozens of assembly instructions
Error riskVery high – one wrong memory access crashes the programLower – language runtime prevents many errors
Typical useOS kernels, device drivers, firmware, game engines, cryptographyWeb apps, data science, automation, business logic

1.3 Why Learn Low-Level Programming Today?

Understanding low-level programming transforms you as a developer in six concrete ways:

  1. You understand what your code actually does. When a Python programmer writes x = [1,2,3], they may not know that this allocates heap memory, creates a pointer, stores a reference count, and sets up a dynamic array structure. A low-level programmer knows exactly what every operation costs in CPU cycles and memory bytes.
  2. You become a better debugger. The most mysterious bugs – segmentation faults, memory corruption, race conditions, security vulnerabilities – all exist at the machine level. Being able to read a stack trace, interpret a core dump, or step through disassembled code makes you capable of solving problems that are completely opaque to those who never leave the high-level world.
  3. You can write faster code. Performance-critical software – game engines, real-time audio processing, financial trading systems, cryptographic libraries – ultimately depends on what happens at the CPU level. Understanding instruction latency, cache behaviour, and SIMD operations allows you to write code that runs orders of magnitude faster than naive high-level implementations.
  4. You understand security. Every major class of software vulnerability – buffer overflows, use-after-free, stack smashing, return-oriented programming (ROP), format string attacks – is a low-level concept. Security researchers and exploit developers live in assembly. You cannot fully understand how to prevent attacks without understanding how they work at the machine level.
  5. You understand how compilers work. Compilers are among the most complex software systems ever built. Understanding assembly lets you read compiler output, understand compiler optimisations, and write high-level code that the compiler can optimise effectively.
  6. Every high-level language sits on low-level foundations. Python is written in C. JavaScript engines are written in C++. The Linux kernel is written in C with assembly. The Java Virtual Machine is written in C++. Going down to the foundation makes everything above it clearer.

2. Number Systems — The Language of Bits

A number system (or numeral system) is a way to represent numbers using a set of symbols (digits). Every number system has a base (or radix) – the number of unique digits it uses before carrying over to the next position.

Here is how the four systems compare:

Number SystemBaseDigits UsedExample Counting
Decimal100–90, 1, 2, …, 9, 10, 11…
Binary20, 10, 1, 10, 11, 100…
Octal80–70, 1, 2, …, 7, 10, 11…
Hexadecimal160–9, A–F0, 1, …, 9, A, B, …F, 10

Tip: Hexadecimal digits beyond 9 are represented by letters. A=10, B=11, C=12, D=13, E=14, F=15. So when you see 0xFF, that’s two hex digits each worth 15, giving you the decimal value 255.

2.1 The Positional Value Formula – The Master Key

This is the single most important formula in this entire guide. Everything else builds on it.

Value = Σ (digitᵢ × baseⁱ) for i = 0 to n-1

Positions are numbered from 0 (rightmost) to n-1 (leftmost). Master this one formula and you can work in any number system.

Let’s apply it to all four systems so you see the pattern clearly:

SystemExampleCalculation
Decimal3453×10² + 4×10¹ + 5×10⁰ = 300 + 40 + 5 = 345
Binary1011₂1×2³ + 0×2² + 1×2¹ + 1×2⁰ = 8 + 0 + 2 + 1 = 11
Octal127₈1×8² + 2×8¹ + 7×8⁰ = 64 + 16 + 7 = 87
Hexadecimal3F₁₆3×16¹ + 15×16⁰ = 48 + 15 = 63

Notice that the formula is identical every time – only the base changes.

2.2 Binary Number System (Base-2) – Deep Explanation

Binary is the native language of all digital hardware. Every transistor inside your CPU, RAM, and storage device holds exactly one bit – a 0 or a 1. This is because transistors are switches: they are either off (0) or on (1). There is no third state.

Key Properties:

  • A single binary digit is called a bit
  • A group of 8 bits is called a byte. A byte can hold 2⁸ = 256 different values (0 through 255)
  • A group of 4 bits is called a nibble
  • All modern data sizes (kilobytes, megabytes, gigabytes) are powers of two multiples of bytes

Important rule: n bits can represent exactly 2ⁿ different values, from 0 to 2ⁿ−1. So 4 bits = 16 values, 8 bits = 256 values, 16 bits = 65,536 values, and 32 bits = over 4 billion values.

Powers of Two – The Building Blocks

Memorise these. They appear constantly in binary work:

2ⁿValue
2⁰1
2¹2
2²4
2³8
2⁴16
2⁵32
2⁶64
2⁷128
2⁸256
2⁹512
2¹⁰1024
2¹¹2048
2¹²4096

Each value is exactly double the one before it. That doubling pattern is the heartbeat of binary.

Counting in Binary (0–15)

DecimalBinaryHex
000000x0
100010x1
200100x2
300110x3
401000x4
501010x5
601100x6
701110x7
810000x8
910010x9
1010100xA
1110110xB
1211000xC
1311010xD
1411100xE
1511110xF

Binary to Decimal – Conversion with Detailed Steps

Method: Sum the positional values of all bits that are 1.

Example 1: Convert 1101₂ to decimal

Position 3: 1 × 8 = 8
Position 2: 1 × 4 = 4
Position 1: 0 × 2 = 0 (skip)
Position 0: 1 × 1 = 1
Sum: 8 + 4 + 0 + 1 = 13

Example 2: Convert 10110111₂ to decimal

  • Write powers from left: 128, 64, 32, 16, 8, 4, 2, 1
  • Bits: 1, 0, 1, 1, 0, 1, 1, 1
  • Sum only where bit=1: 128 + 0 + 32 + 16 + 0 + 4 + 2 + 1 = 183

Shortcut: Only add the positional values where the bit is 1. Completely skip every position that has a 0. This makes mental calculation much faster.

Decimal to Binary – Repeated Division Method

Method: Repeatedly divide the decimal number by 2, recording the remainder after each division. Read the remainders from bottom to top.

Example: Convert 45 to binary

DivisionQuotientRemainder
45 ÷ 2221 (least significant bit)
22 ÷ 2110
11 ÷ 251
5 ÷ 221
2 ÷ 210
1 ÷ 201 (most significant bit)

Read remainders from bottom to top: 101101₂

Verify: 32 + 8 + 4 + 1 = 45 ✓

Example: Convert 200 to binary

200 → 100 r 0
100 → 50  r 0
50  → 25  r 0
25  → 12  r 1
12  → 6   r 0
6   → 3   r 0
3   → 1   r 1
1   → 0   r 1

Read bottom to top: 11001000₂
Verify: 128 + 64 + 8 = 200 ✓

2.3 Octal Number System (Base-8)

Octal uses digits 0 through 7. It was popular in early computing because 3 binary bits map perfectly to one octal digit, making it a compact way to write binary. Today it is most commonly seen in Unix/Linux file permissions – when you run chmod 755, those three digits are octal numbers describing permission bits.

One octal digit = exactly 3 binary bits.

OctalDecimalBinary
00000
11001
22010
33011
44100
55101
66110
77111
1081000
1191001
12101010
13111011
14121100
15131101
16141110
17151111

Example: Convert Octal 127₈ to decimal:
1×8² + 2×8¹ + 7×8⁰ = 64 + 16 + 7 = 87₁₀

Binary to Octal shortcut: Group binary bits in sets of 3 from the right. Convert each group directly.

Example: 110 111 010₂ → 6, 7, 2 → 672₈

Octal to Binary: Simply expand each octal digit into its 3-bit binary equivalent.

Example: 5₈ = 101, 3₈ = 011 → 53₈ = 101011₂

2.4 Hexadecimal Number System (Base-16) – Why It Matters

Hexadecimal is the most practically important non-decimal number system in computing today. It is everywhere – memory addresses, colour codes in CSS (#FF5733), MAC addresses, error codes, and raw machine code dumps all use hex.

The reason hex is so useful is that one hex digit maps perfectly to exactly 4 binary bits (called a nibble). Since a byte is 8 bits, one byte is always exactly 2 hex digits. This makes hex a perfectly compact, human-readable way to write binary data.

Hex to Binary Conversion Table (Memorise the first 16)

HexDecimalBinaryHexDecimalBinary
000000881000
110001991001
220010A101010
330011B111011
440100C121100
550101D131101
660110E141110
770111F151111

Binary to Hex shortcut: Group binary bits into sets of 4 from the right, convert each group to a hex digit.

Example: 10110110₂ → group as 1011 0110 → B 6 → 0xB6

Example: 111100001010₂ → 1111 0000 1010 → F 0 A → 0xF0A

Hex to Binary: Expand each hex digit to its 4-bit binary equivalent.

Example: 0x2F → 0010 1111 → 00101111₂

Example: 0xA9 → 1010 1001 → 10101001₂

Hex to Decimal: Use the positional formula with base 16.

Example: 0x3F → 3×16¹ + 15×16⁰ = 48 + 15 = 63

Example: 0xFF → 15×16¹ + 15×16⁰ = 240 + 15 = 255 (the maximum value of one byte)

2.5 Two’s Complement – How Computers Handle Negative Numbers

Here is something that surprises many beginners: computers have no minus sign in hardware. A transistor is either on or off. There is no “negative” transistor. So how do CPUs handle negative numbers?

The answer is two’s complement – a clever encoding that allows both positive and negative integers to be stored in binary, and that lets the same addition circuit handle both without any modification.

Why it is brilliant: The same addition circuit that adds positive numbers also works correctly when one operand is represented in two’s complement. No special “subtraction” hardware is needed.

How to Compute Two’s Complement of a Negative Number

Given an n-bit system (e.g., 8 bits, 16 bits, 32 bits, 64 bits), to represent -x:

  1. Write the positive value x in binary using exactly n bits (pad with leading zeros if necessary)
  2. Flip all bits – change every 0 to 1 and every 1 to 0. This is the one’s complement
  3. Add 1 to the result (binary addition, ignoring any final carry out of the n-th bit)

Example: Represent -7 in 8-bit two’s complement:

Step 1: +7 in 8 bits = 00000111
Step 2: Flip all bits = 11111000 (one's complement)
Step 3: Add 1 = 11111001

Therefore, in an 8-bit system, the binary pattern 11111001 represents the integer -7.

Verification: Add +7 (00000111) and -7 (11111001):

00000111 + 11111001 = 100000000

The result is 9 bits, but we only keep the lower 8 bits (00000000), which is zero. So 7 + (-7) = 0 ✓

Range of n-Bit Two’s Complement

BitsRange
8 bits-128 to +127
16 bits-32,768 to +32,767
32 bits-2,147,483,648 to +2,147,483,647
64 bits-9.22×10¹⁸ to +9.22×10¹⁸

Formula:

  • Most positive value: 2^(n-1) − 1 (binary 0111…111)
  • Most negative value: −2^(n-1) (binary 1000…000)
  • Zero: 0000…000 (unique representation)

Interpreting the Sign Bit

In two’s complement, the leftmost (most significant) bit tells you the sign:

  • If the MSB is 0, the number is zero or positive
  • If the MSB is 1, the number is negative

Common 8-Bit Two’s Complement Values

Decimal8-bit Binary
300000011
200000010
100000001
000000000
−111111111
−211111110
−311111101
−711111001
−12810000000
12701111111

Subtraction Using Two’s Complement

To compute A – B, the CPU computes A + (two’s complement of B). This is literally how every CPU on earth performs subtraction.

Example: 13 − 5 in 8 bits:

Two's complement of 5 (00000101):
  flip → 11111010
  add 1 → 11111011

Add: 00001101 (13) + 11111011 (-5) = 1 00001000
Keep lower 8 bits = 00001000 = 8 ✓

3. Binary Arithmetic and Bitwise Operations

3.1 Binary Addition

Binary addition follows four simple rules:

ABResultCarry
0000
0110
1010
1101

When both bits are 1, the result in that column is 0 and you carry a 1 to the next column left – exactly like carrying in decimal when digits exceed 9.

Step-by-step example: Add 1011₂ (11) + 1101₂ (13)

    Carry:   1 1 1 0
             1 0 1 1
           + 1 1 0 1
           ─────────
           1 1 0 0 0   (= 24)

Work column by column from right to left:

  • Column 0: 1+1 = 0 carry 1
  • Column 1: 1+0+carry1 = 0 carry 1
  • Column 2: 0+1+carry1 = 0 carry 1
  • Column 3: 1+1+carry1 = 1 carry 1
  • Final carry 1 gives the 5-bit result

Verify: 11 + 13 = 24 ✓

3.2 Binary Subtraction

Subtraction is performed using two’s complement as described in Section 2.5. Convert the number being subtracted to its two’s complement form, then add. This is not just a trick – this is literally how every CPU on earth performs subtraction.

Example: Subtract 1011 − 0110 (11 − 6):

Two's complement of 0110: flip → 1001, add 1 → 1010
Add: 1011 + 1010 = 1 0101 → drop carry → 0101 = 5
And 11 − 6 = 5 ✓

3.3 Binary Multiplication

Binary multiplication uses the same long multiplication method as decimal, but far simpler because you only ever multiply by 0 or 1.

  • Any number multiplied by 0 = 0
  • Any number multiplied by 1 = itself (just copy it down, shifted left by position)
  • Then add all the partial results together

Example: Multiply 101₂ × 11₂ (5 × 3):

         1 0 1
       ×   1 1
       ─────────
         1 0 1   (101 × 1, shift 0)
     + 1 0 1 0   (101 × 1, shift 1)
     ───────────
     = 1 1 1 1   (= 15)

And 5 × 3 = 15 ✓

3.4 Binary Division

Binary division is performed by repeated subtraction (or shift-and-subtract algorithm). Division by a power of two (2ᵏ) is simply a right shift by k bits, which is why CPUs prefer it.

3.5 Bitwise Operations – Definitions and Practical Uses

Bitwise operation: An operation that treats two binary numbers as sequences of individual bits and applies a logical function (AND, OR, XOR, NOT) to each corresponding pair of bits independently. These operations execute in a single CPU clock cycle and are used for low-level control, graphics, cryptography, and systems programming.

OperationSymbolExample (Binary)Result
AND&1010 & 11001000
OR|1010 | 11001110
XOR^1010 ^ 11000110
NOT~~10100101
Left Shift<<0001 << 31000
Right Shift>>1000 >> 20010

AND (&) – Masking

Definition: AND produces a 1 output only when both input bits are 1.

Truth table: 0&0=0, 0&1=0, 1&0=0, 1&1=1

Practical uses:

  • Check if a number is odd: num & 1 – if result is 1, the number is odd (because only the least significant bit determines oddness)
  • Extract lower 4 bits: byte & 0x0F – clears the upper 4 bits, passes lower 4 bits unchanged
  • Clear specific bits: To clear bits 3-5, create a mask with 0s in those positions and 1s elsewhere, then AND

Example: 1010 & 1100 = 1000

OR (|) – Setting Bits

Definition: OR produces a 1 output when at least one input bit is 1.

Truth table: 0|0=0, 0|1=1, 1|0=1, 1|1=1

Practical uses:

  • Set bit 3 (value 8) to 1: value | 0x08 – forces bit 3 to 1 regardless of its previous value
  • Combine flags: Multiple boolean flags can be ORed together into a single byte

Example: 1010 | 1100 = 1110

XOR (^) – Toggling

Definition: XOR (exclusive OR) produces a 1 output when the two input bits are different.

Truth table: 0^0=0, 0^1=1, 1^0=1, 1^1=0

Properties:

  • a ^ a = 0 (any number XOR itself equals zero)
  • a ^ 0 = a (any number XOR zero equals itself)
  • a ^ b ^ b = a (XOR is its own inverse – useful for encryption)

Practical uses:

  • Toggle bits: value ^ 0x0F flips the lower 4 bits
  • Swap two variables without a temporary:
  a ^= b; b ^= a; a ^= b;
  • Simple encryption: XOR a message with a key; XOR again with the same key to decrypt

Example: 1010 ^ 1100 = 0110

NOT (~) – Bit Inversion

Definition: NOT (one’s complement) flips every bit: 0 becomes 1, 1 becomes 0.

Practical uses:

  • Compute one’s complement (first step of two’s complement negation)
  • Create inverted masks

Example (8 bits): ~00001010 = 11110101

Bit Shifts – Multiply and Divide by Powers of Two

Left shift (<<): Moves all bits to the left by a specified number of positions. Zeros fill the vacated lower bits. Each left shift multiplies the value by 2.

Right shift (>>): Moves all bits to the right. For unsigned numbers, zeros fill the vacated upper bits (logical shift). For signed numbers, the sign bit may be replicated (arithmetic shift). Each right shift divides the value by 2 (integer division).

Formulas:

  • n << k = n × 2ᵏ
  • n >> k = n ÷ 2ᵏ (integer division, truncating toward zero for unsigned)

Example: 5 << 3 = 5 × 8 = 40
Binary: 00000101 → 00101000

Example: 40 >> 3 = 40 ÷ 8 = 5
Binary: 00101000 → 00000101

Why compilers love shifts: Multiplication and division by constants that are powers of two are automatically replaced with shift instructions because shifts are orders of magnitude faster than general multiplication/division circuits.

4. ASCII — How Text Is Stored in Binary

Numbers fit naturally into binary. But what about text? How does a computer store the letter ‘A’?

The answer is ASCII – the American Standard Code for Information Interchange. ASCII assigns a unique number to every printable character and control character. That number is then stored as binary. So text is really just a sequence of numbers, which are sequences of bits.

Selected ASCII values (memorise these common ones):

CharacterDecimalBinary (8-bit)
‘A’6501000001
‘B’6601000010
‘C’6701000011
‘Z’9001011010
‘a’9701100001
‘z’12201111010
‘0’4800110000
‘9’5700111001
Space3200100000
Newline1000001010
‘!’3300100001

A clever pattern worth knowing: Uppercase and lowercase letters differ by exactly 32 – or in binary, by a single bit in position 5 (the 32’s place).

A = 01000001, a = 01100001

The only difference is bit 5. This means you can convert between uppercase and lowercase using a single XOR or bit flip operation – something fast string processing routines use constantly.

Example – The word “Hello” in binary:

H = 01001000
e = 01100101
l = 01101100
l = 01101100
o = 01101111

"Hello" = 01001000 01100101 01101100 01101100 01101111

Note on modern encodings: ASCII originally defined 128 characters (7 bits). Modern systems use UTF-8, which extends ASCII to support every character in every human language – Arabic, Chinese, emoji, and everything else. But the first 128 UTF-8 values are identical to ASCII, so all the principles above still apply.

5. Machine Language — The CPU’s Native Tongue

This is where everything comes together. Machine language is the set of binary instructions that the CPU fetches from memory and executes directly. There is no further translation. The processor reads the binary, decodes it using hardware logic, and acts on it in nanoseconds.

Every program you have ever run – operating system, browser, game – was compiled down to this level before it ran.

5.1 What Is Machine Language? (Deep Definition)

Machine language: The set of binary instruction codes that a specific CPU model can execute directly, without any translation, interpretation, or further processing. Each machine instruction is a binary pattern (e.g., B8 05 00 00 00 on x86) that the CPU’s control unit decodes into control signals that activate the ALU, registers, and memory system.

Key characteristics:

  • CPU-specific: Machine code written for an x86 CPU will not run on an ARM CPU, and vice versa. The binary format is tied to the Instruction Set Architecture (ISA)
  • No portability: A compiled executable for Windows (PE format) will not run on Linux (ELF format) even on the same CPU, because the operating system also defines how programs are loaded
  • Direct execution: No interpreter, no just-in-time compiler – the CPU fetches, decodes, and executes each instruction as hardware

5.2 Instruction Structure: Opcode and Operands

Opcode (Operation Code): The part of a machine language instruction that specifies which operation to perform (add, move, jump, compare, etc.). The opcode is a binary number that the CPU’s decoder recognises and maps to the appropriate internal circuit.

Operands: The data that the instruction operates on. An instruction may have zero, one, two, or (rarely) three operands.

Operand TypeDefinitionExample
ImmediateA constant value encoded directly in the instruction bytesMOV EAX, 5 – the value 5 is stored in the instruction itself
RegisterSpecifies one of the CPU’s registersMOV EBX, EAX – both operands refer to registers
MemorySpecifies a memory address (direct, indirect, indexed, etc.)MOV RAX, [RBX] – square brackets mean “the value at this address”

5.3 Example x86 Machine Code – Broken Down

Machine Code (hex)AssemblyExplanation
B8 05 00 00 00MOV EAX, 5Opcode B8 means “move immediate into EAX”. The next 4 bytes are the little-endian representation of the number 5
89 C3MOV EBX, EAXOpcode 89 means “move register to register”. The second byte C3 encodes source (EAX) and destination (EBX)
01 D8ADD EAX, EBXOpcode 01 is ADD. ModRM byte D8 encodes EBX as source, EAX as destination
EB 05JMP +5Opcode EB is short jump (relative, 8-bit offset). 05 means jump forward 5 bytes
C3RETSingle-byte opcode. Pops return address from stack and jumps there
90NOPNo operation – the CPU does nothing for one cycle

Complete instruction sequence:

B8 05 00 00 00    ; mov eax, 5
B9 03 00 00 00    ; mov ecx, 3
01 C8             ; add eax, ecx   (eax becomes 8)

5.4 CPU Architectures – Different Machine Languages

ArchitectureUsed inCharacteristics
x86-64 (AMD64)Desktops, laptops, serversCISC (Complex Instruction Set Computer). Variable instruction length (1-15 bytes). Large number of instructions, many can access memory directly
ARM (AArch64)Smartphones, Apple Silicon, AWS Graviton, embeddedRISC (Reduced Instruction Set Computer). Fixed 32-bit instruction length. Load/store architecture – only LDR and STR access memory; all other instructions work on registers. Very energy-efficient
RISC-VEmbedded systems, research, emerging productsOpen-source, royalty-free ISA. Clean, modern RISC design. Fixed length
MIPSRouters, network equipment, old game consolesClassic RISC design; popular in computer architecture education

Important: A binary executable compiled for x86-64 will not run on an ARM processor, and vice versa. The ISA defines the machine language, and they are fundamentally incompatible.

Same operation, different machine code:

ArchitectureMachine CodeAssembly
x86-64B8 05 00 00 00 (5 bytes)MOV EAX, 5
ARME3A00005 (4 bytes, fixed length)MOV R0, #5

5.5 From High-Level Code to Electrical Signals – The Full Journey

High-level source code:     x = a + b; (in Python or C)

Compiler:                   Translates to assembly language – ADD R1, R2, R3

Assembler:                  Encodes assembly into binary machine code – 00000001 11011000

CPU fetch:                  The control unit reads those binary bytes from memory

CPU decode:                 The opcode bits tell the decoder "this is an ADD instruction"

CPU execute:                Control signals activate the ALU, which adds the values

Transistors:                Billions of transistors switch on and off in response to the 
                            binary values, moving electrical signals through adder circuits

Example – step-by-step for x = 5 + 3 in C:

  1. C code: int x = 5 + 3;
  2. Compiler output (x86 assembly): MOV EAX, 5 then ADD EAX, 3 then MOV [x], EAX
  3. Assembler output (hex): B8 05 00 00 00 83 C0 03 89 45 FC
  4. CPU: fetches B8, decodes “move immediate to EAX”, fetches the next four bytes as the value 5; then fetches 83 C0 03 as ADD immediate, etc.

Every abstraction – Python, Java, your operating system – collapses down to this: transistors reading 0s and 1s at billions of operations per second.

6. How a Computer Actually Works (Hardware Fundamentals)

Before writing a single line of assembly, you must understand the hardware. Assembly language is a direct expression of hardware behaviour – you cannot write it well without understanding the machine.

6.1 The CPU and Its Parts

Central Processing Unit (CPU): The primary component of a computer that performs most of the processing. It fetches instructions from memory, decodes them, executes the required operations, and stores the results.

ComponentDeep Definition
ALU (Arithmetic Logic Unit)A digital circuit that performs arithmetic (addition, subtraction, multiplication, division) and logical operations (AND, OR, NOT, XOR, comparisons). It takes two input values (operands), performs the selected operation, and outputs the result. The ALU also sets status flags (zero, carry, overflow, sign) that describe the result. When you write ADD EAX, EBX in assembly, the ALU is the physical circuit that adds the two numbers.
Control Unit (CU)A finite state machine that directs the entire operation of the processor. It fetches the next instruction from memory (using the Program Counter), decodes the opcode to determine what operation to perform, generates control signals that tell the ALU, registers, and memory bus what to do, and then updates the Program Counter. It is the “conductor” of the CPU’s orchestra.
RegistersThe fastest storage locations in the entire computer – tiny memory cells built directly into the CPU die, operating at the full speed of the processor clock. A modern x86-64 CPU has 16 general-purpose 64-bit registers, plus dozens of special-purpose registers. Accessing a register takes approximately 1 CPU cycle. Accessing main memory (RAM) takes 100–300 cycles because the electrical signals must travel off-chip. This massive speed difference is why assembly programming centres around registers.
Program Counter (PC) / Instruction Pointer (RIP)A special register that always holds the memory address of the next instruction to be fetched and executed. After each instruction is fetched, the control unit automatically increments the Program Counter by the length (in bytes) of that instruction. Jump and call instructions modify the Program Counter directly to implement loops, conditionals, and function calls.
Flags Register (RFLAGS)A special register where each individual bit represents a condition or status of the last arithmetic or comparison operation. For example: the Zero Flag (ZF) is set to 1 if the last result was zero; the Carry Flag (CF) is set if an unsigned addition overflowed; the Sign Flag (SF) reflects the most significant bit of the result (indicating negativity in signed arithmetic).

6.2 Memory Hierarchy – Deep Definition

Memory hierarchy: The organisation of computer storage from the fastest, smallest, most expensive (closest to the CPU) to the slowest, largest, cheapest (farthest from the CPU).

LevelTypical sizeAccess time (cycles)TechnologyDefinition
Registers~128 bytes total1Flip-flops inside CPUStorage locations built into the CPU’s execution core. The CPU operates directly on register values without any memory bus delay.
L1 cache32–64 KB per core4–5SRAM on-chipLevel-1 cache is split into instruction cache (L1i) and data cache (L1d). It holds the most frequently accessed data and code.
L2 cache256 KB – 1 MB per core12–15SRAM on-chipLarger but slower than L1. Acts as a backup for data that does not fit in L1.
L3 cache8–64 MB shared30–40SRAM on-chipShared among multiple CPU cores. Larger than L2 but slower.
RAM (DRAM)8–32 GB100–300Dynamic RAM off-chipMain memory. Stores the program code and data currently in use. Volatile – loses all data when power is turned off.
SSD/NVMe256 GB – 2 TB100,000+Flash memoryNon-volatile persistent storage. Much slower than RAM but much larger.
Hard disk1–10 TB10,000,000+Magnetic plattersSlowest persistent storage. Mechanical parts must physically move to read data.

Practical implication for assembly programming: If your data fits in registers and your code fits in L1 cache, your program runs at the full speed of the CPU. If you access a memory location not in any cache, the CPU must stall and wait for RAM – this can make your program 100 times slower, no matter how clever your instructions are.

6.3 The Fetch-Decode-Execute Cycle – Step by Step

Fetch-Decode-Execute cycle: The fundamental operational loop of every stored-program computer. The CPU continuously repeats three steps for as long as it is powered on:

  1. Fetch – The control unit reads the memory address stored in the Program Counter (RIP on x86-64). It sends that address to the memory system, and the bytes of the instruction are loaded into the CPU’s instruction register. The Program Counter is then advanced to the next address.
  2. Decode – The control unit examines the opcode (operation code) – the first byte or bytes of the instruction – to determine what operation to perform. It also parses the addressing modes: where are the operands? In registers? In memory? As immediate constants encoded in the instruction? This decoding step generates the internal control signals that will be sent to the ALU, registers, and memory bus.
  3. Execute – The appropriate functional unit carries out the operation. The ALU performs arithmetic; the memory controller loads from or stores to RAM; the branch unit updates the Program Counter for jumps. The result is written to the destination (register or memory). Then the cycle repeats with the next instruction.

Modern complexity: On modern superscalar CPUs, multiple instructions are fetched, decoded, and executed simultaneously through pipelining and out-of-order execution. However, the logical model of fetch-decode-execute remains the correct mental model for understanding assembly programming.

6.4 Von Neumann Architecture – The Blueprint of Every Computer

Von Neumann architecture: A computer architecture proposed by John von Neumann in 1945, where both program instructions and data are stored together in the same memory space and accessed through the same bus. This is the design used by virtually all modern computers.

Consequences of storing code and data together:

  • Self-modifying code – A program can write new instructions into memory and then execute them (rarely used today, but possible)
  • Just-in-time compilation – A runtime system can generate machine code for a high-level language and then jump to it, as JavaScript engines do
  • Security vulnerabilities – Because data can be placed in memory and then accidentally executed as code, attacks like buffer overflows and return-oriented programming (ROP) are possible

The four main components:

  1. CPU – Contains the ALU and Control Unit
  2. Memory – Stores both program instructions and data in a single address space
  3. Input/Output – Devices that communicate with the outside world (keyboard, display, disk)
  4. Bus – The set of electrical wires that connect all components and carry data, addresses, and control signals

PART TWO: ASSEMBLY LANGUAGE WITH NASM

7. Assembly Language with NASM — Complete Guide

7.1 What Is Assembly Language? (Deep Definition)

Assembly language: A human-readable notation for machine language. Each assembly instruction (like MOV RAX, 5) corresponds directly to exactly one machine instruction (like B8 05 00 00 00). The assembler (e.g., NASM) translates these mnemonics into binary machine code.

Comparison of levels:

LevelExampleTranslatorExecution
High-level (C)x = a + b;Compiler → assembly → machine codeIndirect
AssemblyADD RAX, RBXAssembler → machine codeOne-to-one
Machine code01 D8None – directCPU executes directly

7.2 Why Learn Assembly? (Expanded)

  • Complete hardware control – You decide exactly which registers to use, exactly how memory is accessed, and exactly what the CPU does at every step
  • Maximum performance – Hand-written assembly can be faster than compiled code for critical inner loops
  • Understanding compilers – When you read compiler output (disassembly), you understand why your C code generates certain instructions
  • Reverse engineering and security – Malware analysis, vulnerability research, and exploit development all require reading and writing assembly
  • Embedded and OS development – Bootloaders, operating system kernels, and firmware require assembly for the portions that run before any runtime environment exists

7.3 NASM Setup and Installation

Installing NASM:

On Linux (Ubuntu/Debian):

sudo apt update
sudo apt install nasm

On macOS (with Homebrew):

brew install nasm

On Windows: Download the installer from nasm.us, run it, and add the NASM directory to your system PATH.

Verify the installation with:

nasm --version

Assembling and Linking:

On Linux (64-bit):

nasm -f elf64 program.asm -o program.o
ld program.o -o program
./program

On macOS (64-bit):

nasm -f macho64 program.asm -o program.o
ld program.o -o program -macosx_version_min 10.13 -lSystem
./program

The -f flag specifies the output format – elf64 for Linux, macho64 for macOS.

7.4 Program Structure – Sections Deep Explanation

Section (or segment): A named region of an assembly program with a specific purpose: code, initialised data, or uninitialised data.

section .data
    ; Initialised data – variables that already have values when the program starts.
    ; These bytes are stored in the executable file and loaded into memory at runtime.

section .bss
    ; Uninitialised data – space reserved for variables that will be set at runtime.
    ; No space is taken in the executable; the loader allocates zeroed memory.

section .text
    global _start    ; Make the _start label visible to the linker as the entry point.
    ; Your code – executable instructions.

Detailed example of each section:

section .data
    message db "Hello, World!", 10   ; db = define byte. 10 is ASCII newline.
    number  dw 1234                  ; dw = define word (2 bytes)
    pi      dd 3.14                  ; dd = define double word (4 bytes)

section .bss
    buffer  resb 64      ; reserve 64 bytes. Uninitialised – content is undefined.
    counter resd 1       ; reserve 1 double word (4 bytes) for an integer counter.

section .text
    global _start
_start:
    ; Code goes here

Data directives summary:

DirectiveSizeFull name
db1 byteDefine Byte
dw2 bytesDefine Word
dd4 bytesDefine Double Word
dq8 bytesDefine Quad Word
resb nn bytesReserve Bytes
resw nn wordsReserve Words
resd nn double wordsReserve Double Words

7.5 Registers – Detailed Reference

General-purpose register: A storage location inside the CPU that can hold data (integers, addresses) and can be used as an operand in most arithmetic and logic instructions. The x86-64 architecture provides 16 general-purpose registers, each 64 bits wide.

Sub-register naming (for partial access):

64-bit32-bit16-bit8-bit low8-bit highConventional use
RAXEAXAXALAHAccumulator – arithmetic, return value
RBXEBXBXBLBHBase – general, preserved across calls
RCXECXCXCLCHCounter – loop counts, shift amounts
RDXEDXDXDLDHData – I/O, extended arithmetic
RSIESISISIL(none)Source index – string/memory source
RDIEDIDIDIL(none)Destination index – string/memory dest
RSPESPSPSPL(none)Stack pointer – top of stack
RBPEBPBPBPL(none)Base pointer – stack frame base
R8-R15R8D-R15DR8W-R15WR8B-R15B(none)Extra general-purpose registers

Visual of RAX breakdown:

RAX (64 bits)
┌──────────────────────────┬──────────────────────────┐
│         Upper 32          │           EAX             │
└──────────────────────────┴──────────┬────────────────┘
                                      │       AX        │
                                      ├────────┬────────┤
                                      │   AH   │   AL   │
                                      └────────┴────────┘

Important rules for partial register access:

  • Writing to a 32-bit register (e.g., EAX) zeroes the upper 32 bits of the full 64-bit register (RAX). This is a common source of bugs.
  • Writing to a 16-bit (AX) or 8-bit (AL, AH) register does not clear the remaining bits – they retain their previous values.

7.6 Core Instructions – Expanded with Definitions

MOV – The Most Fundamental Instruction

Definition: MOV destination, source copies the value from source to destination. The source remains unchanged. This is a copy, not a move – the original data stays where it was.

Allowed operand combinations:

  • Register ← immediate
  • Register ← register
  • Memory ← register
  • Register ← memory

Not allowed: Memory ← memory directly (must go through a register).

Examples:

MOV RAX, 10          ; Immediate → register
MOV RBX, RAX         ; Register → register
MOV [var], RAX       ; Register → memory (store)
MOV RAX, [var]       ; Memory → register (load)

Critical rule: Both operands must be the same size. You cannot MOV a 64-bit register into an 8-bit memory location without using a smaller register first.

ADD and SUB

Definition: ADD dest, src adds src to dest and stores the result in dest. SUB dest, src subtracts src from dest and stores the result in dest. Both update the flags register (ZF, SF, CF, OF) based on the result.

MOV RAX, 15
MOV RBX, 7
ADD RAX, RBX       ; RAX = 22
SUB RAX, 5         ; RAX = 17

INC and DEC

Definition: INC reg adds 1 to the register. DEC reg subtracts 1. They are compact (shorter encoding) than ADD reg, 1.

INC RAX    ; RAX = RAX + 1
DEC RBX    ; RBX = RBX - 1

MUL and IMUL – Multiplication

MUL source (unsigned multiplication): Multiplies RAX by source and stores the result in RDX:RAX (a 128-bit value split across two 64-bit registers). RDX holds the high 64 bits, RAX holds the low 64 bits.

MOV RAX, 6
MOV RBX, 7
MUL RBX        ; RAX = 42, RDX = 0 (result fits in 64 bits)

IMUL (signed multiplication): More flexible syntax. Can multiply any two registers and optionally a third immediate.

IMUL RAX, RBX         ; RAX = RAX × RBX
IMUL RAX, RBX, 10     ; RAX = RBX × 10

DIV and IDIV – Division

DIV source (unsigned division): Divides the 128-bit value in RDX:RAX (high 64 bits in RDX, low 64 bits in RAX) by source. Quotient goes to RAX, remainder to RDX.

Critical: You must zero out RDX before an unsigned division unless you are intentionally dividing a 128-bit value.

; Divide 20 by 4
MOV RAX, 20
XOR RDX, RDX     ; Zero RDX (same as MOV RDX,0)
MOV RBX, 4
DIV RBX          ; RAX = 5, RDX = 0

; Divide 17 by 5 → quotient 3, remainder 2
MOV RAX, 17
XOR RDX, RDX
MOV RBX, 5
DIV RBX          ; RAX = 3, RDX = 2

7.7 Flags and Conditional Jumps – Deep Explanation

Flags: Individual bits in the RFLAGS register that are set or cleared automatically by the CPU after arithmetic and comparison instructions. They store metadata about the result: was it zero? Was it negative? Did it overflow?

Key flags:

FlagNameSet when
ZF (Zero Flag)ZeroResult of last operation was exactly zero
SF (Sign Flag)SignResult is negative (most significant bit = 1)
CF (Carry Flag)CarryUnsigned addition produced a carry out of the MSB, or unsigned subtraction required a borrow
OF (Overflow Flag)OverflowSigned addition produced a result outside the representable range
PF (Parity Flag)ParityThe result has an even number of 1 bits

Example flag behaviour:

mov rax, 5
sub rax, 5      ; RAX = 0 → ZF = 1 (set)

mov al, 127     ; Maximum signed 8-bit value
add al, 1       ; 127+1=128 → in 8-bit signed, this is -128 → OF = 1

mov al, 255     ; Maximum unsigned 8-bit value
add al, 1       ; 255+1=0 (wraps) → CF = 1

CMP – The Comparison Instruction

Definition: CMP dest, src subtracts src from dest, updates the flags, and discards the result. It is used solely to set flags for a subsequent conditional jump.

CMP RAX, RBX    ; Compute RAX - RBX, set flags, throw away result.

Conditional Jumps – The Decision Makers

InstructionConditionFlagsMeaning
JE / JZEqual / ZeroZF=1Jump if equal (signed or unsigned)
JNE / JNZNot equalZF=0Jump if not equal
JG / JNLEGreater (signed)ZF=0 and SF=OFSigned greater than
JGE / JNLGreater or equal (signed)SF=OFSigned greater or equal
JL / JNGELess (signed)SF≠OFSigned less than
JLE / JNGLess or equal (signed)ZF=1 or SF≠OFSigned less or equal
JA / JNBEAbove (unsigned)CF=0 and ZF=0Unsigned greater than
JB / JCBelow (unsigned)CF=1Unsigned less than
JMPUnconditional–Always jump

Practical example:

MOV RAX, 10
CMP RAX, 10
JE  equal        ; ZF is 1 because 10-10=0, so we jump
JMP not_equal    ; This line is skipped

equal:
    ; code for equality

7.8 Loops – Building Iteration from Jumps

Loop (in assembly): There is no dedicated loop keyword. A loop is constructed manually using a counter, a label, and conditional jumps that jump backward to the label while the counter is not zero.

Basic loop pattern:

MOV RCX, 5        ; counter = 5

loop_start:
    ; loop body
    DEC RCX
    CMP RCX, 0
    JNZ loop_start    ; if RCX != 0, jump back
; loop ends

Using the LOOP instruction: LOOP label automatically decrements RCX and jumps to label if RCX ≠ 0. It combines DEC RCX, CMP RCX, 0, and JNZ into one instruction.

MOV RCX, 10
sum_loop:
    ADD RAX, RCX
    LOOP sum_loop     ; RCX--; if RCX != 0, jump
; RAX now holds 55 (1+2+...+10)

Nested loops: You must save and restore RCX because the inner loop will overwrite it.

MOV RCX, 3          ; outer counter
outer:
    PUSH RCX         ; save outer counter
    MOV RCX, 4       ; inner counter
inner:
    ; inner work
    LOOP inner
    POP RCX          ; restore outer counter
    LOOP outer

7.9 Memory Addressing Modes – Accessing Data in RAM

Addressing mode: The method by which an instruction specifies the location of its operand. Square brackets [] in NASM mean “the value at this memory address” – they are the dereference operator.

ModeExampleDefinition
ImmediateMOV RAX, 42The value 42 itself is encoded in the instruction. No memory access.
RegisterMOV RBX, RAXThe operand is the value inside a register.
Direct memoryMOV RAX, [number]The address of number is fixed; load the value stored at that address.
Register indirectMOV RAX, [RBX]RBX holds a memory address; load the value at that address.
Base + offsetMOV RAX, [RBX + 8]Start at address in RBX, add 8 bytes, then load. Used for struct fields.
Base + index × scaleMOV RAX, [RBX + RCX*8]Used for array indexing. Scale is element size (1,2,4,8).

7.10 String Operations – Efficient Block Processing

String instructions: Specialised x86 instructions that operate on blocks of memory (strings) using RSI as source pointer and RDI as destination pointer. They automatically increment or decrement the pointers based on the direction flag (DF). The REP prefix repeats the instruction RCX times.

Core string instructions:

InstructionEffect
MOVSB[RDI] = [RSI]; then RSI++, RDI++ (copy one byte)
STOSB[RDI] = AL; then RDI++ (store AL into memory)
LODSBAL = [RSI]; then RSI++ (load from memory into AL)
CMPSBCompare [RSI] with [RDI], set flags; then RSI++, RDI++
SCASBCompare AL with [RDI], set flags; then RDI++

Example – Copy 100 bytes:

cld                     ; clear direction flag (process left to right)
mov rsi, src_data       ; RSI points to source
mov rdi, dst_data       ; RDI points to destination
mov rcx, 100            ; count
rep movsb               ; repeat MOVSB RCX times

Example – Find length of null-terminated string:

cld
mov rdi, src_data
mov al, 0               ; search for null byte
mov rcx, 256            ; maximum length
repne scasb             ; repeat while not equal (scan until AL matches)
; After loop: RCX = (max - length - 1)
; Length = max - RCX - 1

7.11 The Stack – PUSH, POP, and Function Calls

Stack: A region of memory that operates on the Last-In, First-Out (LIFO) principle. It is used to store temporary values, return addresses, function arguments, and saved register states. The stack pointer (RSP) always points to the current top of the stack. On x86-64, the stack grows downward (toward lower memory addresses).

  • PUSH: PUSH src decrements RSP by 8 (for 64-bit) and writes src to the memory address now pointed to by RSP
  • POP: POP dest reads the value from the memory address pointed to by RSP, then increments RSP by 8
PUSH RAX       ; RSP = RSP - 8; [RSP] = RAX
POP  RBX       ; RBX = [RSP]; RSP = RSP + 8

Critical rule: Every PUSH must eventually have a matching POP (or an adjustment to RSP), otherwise the stack becomes unbalanced and your program will crash when it tries to return.

Saving and restoring registers:

_start:
    MOV RAX, 100
    PUSH RAX          ; save RAX
    MOV RAX, 999      ; use RAX for something else
    ; ... work ...
    POP RAX           ; restore original RAX (now 100)

7.12 Procedures – CALL and RET

Definition – CALL label: Pushes the address of the next instruction (the return address) onto the stack, then jumps to label. This allows the procedure to return later.

Definition – RET: Pops the return address from the stack and jumps to that address, resuming execution after the original CALL.

Simple procedure:

_start:
    CALL my_proc      ; pushes return address, jumps
    ; execution resumes here after RET

my_proc:
    MOV RAX, 42
    RET               ; pops return address, jumps back

Function prologue and epilogue (for stack frames and local variables):

my_function:
    ; Prologue
    PUSH RBP          ; save caller's base pointer
    MOV  RBP, RSP     ; set our own base pointer
    SUB  RSP, 32      ; allocate 32 bytes for local variables

    ; Function body – access locals at [RBP-8], [RBP-16], etc.
    MOV  QWORD [RBP-8], 10

    ; Epilogue
    MOV  RSP, RBP     ; deallocate locals
    POP  RBP          ; restore caller's base pointer
    RET

Linux 64-bit Calling Convention (System V AMD64 ABI):

When your assembly procedure is called from C code, or when you call C functions from assembly, both sides must agree on how arguments are passed and where return values live.

On Linux 64-bit:

  • First 6 integer arguments are passed in registers: RDI, RSI, RDX, RCX, R8, R9
  • Additional arguments are pushed onto the stack
  • Return value is in RAX (or RDX:RAX for 128-bit values)
  • Caller-saved registers (caller must preserve if needed): RAX, RCX, RDX, RSI, RDI, R8, R9, R10, R11
  • Callee-saved registers (function must preserve): RBX, RBP, R12, R13, R14, R15

7.13 System Calls – Talking to the Operating System

System call: A controlled way for a user-mode program to request a service from the operating system kernel (e.g., reading from a file, writing to the screen, allocating memory). The CPU transitions from user mode (Ring 3) to kernel mode (Ring 0) to execute the request safely.

Linux x86-64 system call convention:

  1. System call number → RAX
  2. Arguments (max 6) → RDI, RSI, RDX, R10, R8, R9 (in that order)
  3. Execute syscall instruction
  4. Return value → RAX (negative if error)

Common Linux system call numbers (x86-64):

System CallNumber (RAX)RDIRSIRDX
sys_read0file descriptorbuffer addressbyte count
sys_write1file descriptorbuffer addressbyte count
sys_open2filename addressflagsmode
sys_close3file descriptor——
sys_exit60exit code——

File descriptors: 0 = stdin, 1 = stdout, 2 = stderr

Complete “Hello, World!” with explanations:

section .data
    msg db "Hello, NASM!", 10    ; 10 = ASCII newline
    len equ $ - msg              ; $ = current address; subtract start → length

section .text
    global _start

_start:
    ; sys_write(1, msg, len)
    mov rax, 1          ; syscall number for write
    mov rdi, 1          ; file descriptor 1 = stdout
    mov rsi, msg        ; pointer to string
    mov rdx, len        ; number of bytes to write
    syscall             ; call kernel

    ; sys_exit(0)
    mov rax, 60         ; syscall number for exit
    xor rdi, rdi        ; exit code 0 (xor with itself is a common way to zero a register)
    syscall

7.14 Memory Management at the Low Level – Deep Dive

7.14.1 Memory Map of a Running Process

Virtual address space: Each process on a modern operating system has its own private, isolated view of memory, called the virtual address space. The OS and CPU (via the Memory Management Unit, MMU) map virtual addresses to physical RAM addresses.

Typical layout of a 64-bit Linux process (simplified):

High addresses
┌─────────────────────────────────┐
│ Kernel space (inaccessible)     │ ← user code cannot read/write this
├─────────────────────────────────┤
│ Environment variables & args    │
├─────────────────────────────────┤
│ Stack (grows DOWNWARD)          │ ← RSP points here
│   [stack frames, local vars]    │
├─────────────────────────────────┤
│ (unmapped gap – segfault if accessed) │
├─────────────────────────────────┤
│ Memory-mapped files / libraries │ ← shared libraries, mmap()
├─────────────────────────────────┤
│ Heap (grows UPWARD)             │ ← malloc() / new
├─────────────────────────────────┤
│ BSS (uninitialised globals)     │ ← zeroed at startup
├─────────────────────────────────┤
│ Data (initialised globals)      │ ← values set in source
├─────────────────────────────────┤
│ Text (code, read+exec)          │ ← your instructions
├─────────────────────────────────┤
│ NULL region (unmapped)          │ ← accessing 0x0 causes segfault
Low addresses

7.14.2 Stack vs Heap – Detailed Comparison

FeatureStackHeap
AllocationAutomatic – when you call a function, the stack frame is created; when you return, it is destroyedManual – you must call malloc (C) or new (C++) or use system calls like mmap
SpeedExtremely fast – just moving RSP (a single subtraction or addition)Slower – the allocator must find a free block, possibly asking the OS for more memory
LifetimeTied to function scope – local variables exist only while the function is activeUntil explicitly freed with free or delete. Can outlive the function that allocated it
SizeLimited – typically 1–8 MB per thread. Exceeding it causes a stack overflowLarge – limited by available RAM and swap
FragmentationNone – stack allocations are strictly LIFO, so no fragmentationCan become fragmented over time if blocks are allocated and freed in different orders
UsageLocal variables, return addresses, saved registersData that must outlive a function call, large arrays, dynamically sized structures

7.14.3 Buffer Overflows – The Classic Low-Level Vulnerability

Buffer overflow: A condition where a program writes more data into a fixed-size buffer than the buffer can hold, causing the excess data to overwrite adjacent memory.

Stack layout during a vulnerable function:

Low addresses
   [buffer[0..63]]     ← 64 bytes allocated
   [saved RBP]         ← 8 bytes (if frame pointer is used)
   [return address]    ← 8 bytes ← overflow reaches here at 72 bytes
High addresses

Why it is dangerous: If an attacker can control the data written into buffer, they can overwrite the return address with the address of malicious code (shellcode) that they also inject, or they can chain together existing code snippets (Return-Oriented Programming).

Modern defenses:

DefenseDefinition
Stack canariesA random value placed between the buffer and the return address. If a buffer overflow occurs, the canary is overwritten. Before returning, the program checks the canary; if it changed, the program aborts.
ASLR (Address Space Layout Randomization)The OS randomises the base addresses of the stack, heap, and shared libraries each time a program runs. This makes it difficult for an attacker to predict where their shellcode or ROP gadgets are located.
NX bit (No-Execute)Memory pages are marked as either writable or executable, never both. This prevents an attacker from placing shellcode in the stack or heap and then executing it.
CFI (Control-Flow Integrity)The compiler inserts checks before every indirect jump/call to ensure the target address is a valid function entry point, breaking ROP attacks.

8. Complete Example Programs

Program 1 – Print “Hello, World!”

section .data
    msg db "Hello, World!", 10
    len equ $ - msg

section .text
    global _start

_start:
    ; write to stdout
    MOV RAX, 1
    MOV RDI, 1
    MOV RSI, msg
    MOV RDX, len
    syscall

    ; exit cleanly
    MOV RAX, 60
    MOV RDI, 0
    syscall

What each line does:

  • section .data – defines the data section
  • msg db "Hello, World!", 10 – defines a byte sequence with the string and a newline (ASCII 10)
  • len equ $ - msg – $ is the current address, so this calculates the length of the string
  • MOV RAX, 1 – system call number for sys_write
  • MOV RDI, 1 – file descriptor 1 = stdout
  • MOV RSI, msg – pointer to the string
  • MOV RDX, len – number of bytes to write
  • syscall – invoke the kernel
  • MOV RAX, 60 – system call number for sys_exit
  • MOV RDI, 0 – exit code 0 = success
  • syscall – invoke the kernel again

Program 2 – Add Two Numbers and Exit with the Result

section .text
    global _start

_start:
    MOV RAX, 14        ; first number
    MOV RBX, 28        ; second number
    ADD RAX, RBX       ; RAX = 14 + 28 = 42

    MOV RDI, RAX       ; exit code = result
    MOV RAX, 60        ; sys_exit
    syscall

After running this, check the exit code with echo $? in your terminal. It will print 42.

Program 3 – Loop and Count

section .data
    msg db "Counting...", 10
    len equ $ - msg

section .text
    global _start

_start:
    ; print message
    MOV RAX, 1
    MOV RDI, 1
    MOV RSI, msg
    MOV RDX, len
    syscall

    ; count from 10 down to 1 using LOOP
    MOV RCX, 10

count_loop:
    PUSH RCX           ; save loop counter before syscall clobbers registers
    ; (in a full program you would print RCX here)
    POP  RCX
    LOOP count_loop

    ; exit
    MOV RAX, 60
    MOV RDI, 0
    syscall

Program 4 – Simple Procedure Call

section .text
    global _start

; Procedure: square
; Input:  RDI = number to square
; Output: RAX = RDI * RDI
square:
    MOV RAX, RDI
    IMUL RAX, RDI
    RET

_start:
    MOV RDI, 7         ; argument: square the number 7
    CALL square        ; call procedure – result in RAX
    ; RAX = 49

    MOV RDI, RAX       ; exit with result
    MOV RAX, 60
    syscall

Check with echo $? – it prints 49.

9. Performance Optimization at the Low Level

9.1 Avoid Data Hazards – Let the Pipeline Breathe

Pipeline hazard: A condition where the next instruction cannot execute in the next clock cycle because it depends on the result of a previous instruction that has not yet completed. The CPU must insert “bubble” cycles (stalls), reducing performance.

Bad – dependent chain:

mov rax, [mem]   ; load takes 200 cycles if cache miss
add rax, 5       ; must wait for load to complete
add rax, 10      ; must wait for previous add

Better – independent operations:

mov rax, [mem]     ; load element 0
mov rbx, [mem+8]   ; independent load – can happen in parallel
mov rcx, [mem+16]  ; independent
add rax, 5         ; operates on rax, may overlap with other loads
add rbx, 10
add rcx, 15

9.2 Cache Locality – The #1 Optimisation

Locality of reference: The tendency of a program to access the same memory addresses repeatedly (temporal locality) or to access addresses that are close to each other (spatial locality). Caches exploit both forms.

Spatial locality example – row-major vs column-major traversal:

// BAD: column-major traversal – jumps by entire row size each time
for (int col = 0; col < 1000; col++)
    for (int row = 0; row < 1000; row++)
        sum += matrix[row][col];   // memory stride = 1000*4 = 4000 bytes

// GOOD: row-major traversal – sequential memory access
for (int row = 0; row < 1000; row++)
    for (int col = 0; col < 1000; col++)
        sum += matrix[row][col];   // accesses consecutive addresses

On large matrices, the second version can be 10x faster or more because it makes optimal use of cache lines.

9.3 Reading Compiler Output – Learn from the Best

Modern compilers (GCC, Clang) generate highly optimised assembly, often better than what a beginner would write by hand. Use this to learn.

# Generate assembly from C with optimisations
gcc -O2 -S myfile.c -o myfile.s

# View disassembly of an executable
objdump -d -M intel myprogram | less

# Use Compiler Explorer online – paste code, see assembly instantly
# https://godbolt.org

10. Tools Every Low-Level Programmer Needs

ToolPurposeDefinition / Use
NASMAssemblerTranslates assembly source into object files (.o). Use -f elf64 for Linux 64-bit.
GDBGNU DebuggerAllows you to step through assembly instructions, inspect registers, examine memory, set breakpoints. Essential for understanding what your code actually does.
objdumpDisassemblerDisplays machine code from object files or executables in human-readable assembly. objdump -d -M intel program
straceSystem call tracerShows every system call a program makes, along with arguments and return values. strace ./program
ltraceLibrary call tracerShows calls to dynamic libraries (e.g., printf, malloc).
readelfELF file examinerDisplays detailed information about ELF executable headers, sections, symbols.
xxd / hexdumpHex dump utilitiesShow raw binary data as hexadecimal. xxd program or hexdump -C program
Radare2Reverse engineering frameworkAdvanced disassembler, debugger, and binary analysis tool.
Compiler ExplorerOnline toolgodbolt.org – instantly see the assembly generated by any compiler for any language.
CPUlatorCPU simulatorBrowser-based CPU simulator. Write assembly code, step through it instruction by instruction, and watch registers and memory change in real time.
Intel x86 referenceManualfelixcloutier.com/x86 – complete instruction set reference.

GDB Quick Reference for Assembly

gdb ./program
(gdb) layout asm           # show disassembly view
(gdb) layout regs          # show register window
(gdb) break _start         # break at symbol
(gdb) break *0x401000      # break at absolute address
(gdb) run
(gdb) stepi                # execute one instruction (step into)
(gdb) nexti                # execute one instruction (step over calls)
(gdb) info registers       # show all registers
(gdb) info registers rax rbx
(gdb) x/10xb $rsp          # examine 10 bytes in hex at stack pointer
(gdb) x/4gx $rsp           # examine 4 quad-words (8 bytes each)
(gdb) x/s 0x402000         # examine as string
(gdb) disassemble          # disassemble current function
(gdb) continue             # run until next breakpoint

11. NASM Directives and Useful Features

Constants with EQU

EQU defines a constant name for a value. Unlike variables, constants cannot change – they are substituted at assembly time.

STDOUT  equ 1
NEWLINE equ 10
SYS_WRITE equ 1
SYS_EXIT  equ 60

; now use the names instead of magic numbers
MOV RAX, SYS_WRITE
MOV RDI, STDOUT

TIMES – Repeat Data

section .data
    zeros  times 16 db 0     ; 16 zero bytes
    buffer times 64 db ' '   ; 64 space characters

Macros

NASM macros let you define reusable code patterns with parameters. They are expanded at assembly time – no function call overhead.

; Define a macro that prints a message
%macro print 2              ; 2 parameters: address and length
    MOV RAX, 1
    MOV RDI, 1
    MOV RSI, %1             ; first argument
    MOV RDX, %2             ; second argument
    syscall
%endmacro

section .data
    msg db "Using a macro!", 10
    len equ $ - msg

section .text
    global _start

_start:
    print msg, len          ; single clean line instead of 4 MOV + syscall

    MOV RAX, 60
    MOV RDI, 0
    syscall

External Labels and Linking with C

You can call C standard library functions from NASM and vice versa. This lets you use printf, malloc, and other libc functions from assembly.

extern printf               ; declare printf as external

section .data
    fmt db "Value: %d", 10, 0    ; printf format string, null-terminated

section .text
    global _start

_start:
    MOV RDI, fmt            ; first arg: format string
    MOV RSI, 42             ; second arg: value to print
    XOR RAX, RAX            ; RAX = 0 (no floating point args)
    CALL printf

    MOV RAX, 60
    MOV RDI, 0
    syscall

Link with: nasm -f elf64 prog.asm -o prog.o && gcc prog.o -o prog -no-pie

Quick Reference – Most Used Instructions

InstructionSyntaxEffect
MOVMOV dst, srcdst = src
ADDADD dst, srcdst = dst + src
SUBSUB dst, srcdst = dst − src
INCINC dstdst = dst + 1
DECDEC dstdst = dst − 1
MULMUL srcRDX:RAX = RAX × src
IMULIMUL dst, srcdst = dst × src (signed)
DIVDIV srcRAX = RDX:RAX ÷ src, RDX = remainder
ANDAND dst, srcdst = dst AND src (bitwise)
OROR dst, srcdst = dst OR src (bitwise)
XORXOR dst, srcdst = dst XOR src (bitwise)
NOTNOT dstdst = bitwise NOT dst
SHLSHL dst, ndst = dst << n (left shift)
SHRSHR dst, ndst = dst >> n (right shift)
CMPCMP a, bset flags from a − b, discard result
JMPJMP labelunconditional jump
JEJE labeljump if equal (ZF=1)
JNEJNE labeljump if not equal (ZF=0)
JGJG labeljump if greater (signed)
JLJL labeljump if less (signed)
PUSHPUSH srcRSP−=8; [RSP]=src
POPPOP dstdst=[RSP]; RSP+=8
CALLCALL labelpush return addr; jump to label
RETRETpop return addr; jump there
syscallsyscallinvoke OS kernel (Linux x86-64)
NOPNOPdo nothing for one cycle
HLTHLThalt the CPU

Conclusion

Assembly language removes every layer of abstraction between you and the CPU. When you write NASM, you are placing exact instructions into memory for the processor to execute. There is no runtime, no garbage collector, no virtual machine – just your binary and the hardware.

The journey you have taken through this guide covers the complete foundation:

  • Number systems – binary, octal, hexadecimal, and two’s complement
  • Binary arithmetic – addition, subtraction, multiplication, division
  • Bitwise operations – AND, OR, XOR, NOT, and shifts
  • Machine language – opcodes, operands, and how the CPU executes them
  • Hardware fundamentals – CPU, memory hierarchy, fetch-decode-execute cycle
  • NASM assembly – program structure, registers, instructions, flags, loops, memory addressing, stack, procedures, and system calls
  • Performance optimization – data hazards, cache locality, reading compiler output
  • Security – buffer overflows and modern defenses

What makes assembly genuinely rewarding is the clarity it provides. When your program works, you know exactly why. When it breaks, you know exactly where to look. There is no mystery – just instructions, registers, and memory.

“The computer was born to solve problems that did not exist before.” — Bill Gates

Scroll to Top