R language
R is a high-level programming language and environment primarily designed for statistical computing, data analysis, and graphical visualization. It provides a rich collection of built-in capabilities and specialized packages that enable users to process data, perform statistical tests, build analytical models, and create detailed visualizations. R is widely used in fields such as research, finance, bioinformatics, data science, and academic statistics.
This section explores the fundamental concepts of R, including variables, data types, vectors, matrices, data frames, functions, packages, statistical analysis, and data visualization. It also examines R’s strengths, limitations, and practical applications, providing a foundation for understanding how R can be used to analyze complex datasets and communicate meaningful insights through statistical and visual methods.

Introduction To R language
wwww
Introduction to R
R has emerged as one of the most powerful programming languages for statistical computing, data exploration, and creating stunning visualizations. From Fortune 500 companies to academic research labs, professionals across every industry turn to R when they need to make sense of complex information and uncover hidden patterns in their data.
What Exactly Is R?
At its core, R is both a programming language and an interactive computing environment. Unlike general-purpose languages, R was purpose-built for statistical analysis and creating graphical representations of data. When you work with R, you are not just writing code — you are exploring, questioning, and discovering insights that might otherwise remain hidden.
The arrow symbol <- assigns values to variables, similar to how you might use = in other programming languages. The print() function displays your results. Think of R as a sophisticated calculator that can also create beautiful charts and perform complex statistical tests with just a few lines of code.
# Basic R operations
first_number <- 25
second_number <- 30
total <- first_number + second_number
print(total) # Output: [1] 55
Understanding the output is important — the [1] indicates that this is the first element of the output vector. R displays results with index numbers in brackets, which becomes particularly useful when working with vectors containing multiple values.
How R Came to Be
The story of R begins in 1993 at the University of Auckland in New Zealand, where statisticians Ross Ihaka and Robert Gentleman created a language that made statistical computing more intuitive and accessible. The name “R” pays homage to its predecessor — the S programming language developed at Bell Laboratories.
What made R special from the beginning was its open-source philosophy. The creators decided to make R completely free, meaning anyone could use it, study it, modify it, and share it freely. This decision transformed R from a niche academic tool into a global phenomenon. Today, thousands of developers contribute to R’s development, constantly adding new features and fixing issues.
Why R Matters in Today’s World
R matters because data matters. In almost every field — business, healthcare, education, finance, science — organizations collect massive amounts of information.Without analysis, data is nothing more than a collection of numbers waiting to be understood. R provides the tools to transform raw data into actionable insights.
In business and finance, companies use R to analyze customer behavior, forecast sales, detect fraud, and optimize pricing strategies. Financial institutions rely on R for risk modeling and portfolio management. In healthcare and medicine, researchers use R to analyze clinical trial data, identify disease patterns, and develop treatment recommendations. During public health crises, R helps epidemiologists track infection rates.
In scientific research, biologists analyze DNA sequences, physicists model complex systems, and chemists simulate molecular interactions using R. It has become the standard tool for data analysis in virtually every scientific discipline. Technology companies use R for quality assurance, user behavior analysis, and recommendation systems. Engineers use R for signal processing and simulation.
How R Compares to Other Languages
R excels in statistical analysis and visualization. If your primary goal is to analyze data and create compelling graphics, R is often the best choice. Python offers more flexibility for general programming and building applications. SAS dominates in corporate environments where compliance and auditing are critical. MATLAB serves engineering and simulation needs. The good news is that these languages complement each other, and many professionals learn both R and Python, using each for what it does best.
| Aspect | R | Python | SAS | MATLAB |
|---|---|---|---|---|
| Primary Focus | Statistics and visualization | General purpose and data science | Business analytics | Mathematics and engineering |
| Learning Curve | Gentle for statistics | Moderate | Steep for beginners | Steep for non-engineers |
| Visualizations | Outstanding quality | Good | Limited | Good |
| Community Support | Strong among statisticians | Massive and diverse | Corporate focused | Engineering focused |
| Cost | Free and open-source | Free and open-source | Very expensive | Expensive |
Key Features That Make R Powerful
R comes packed with features that make it ideal for data work. It is completely open-source, meaning you can download, use, and modify it freely. R handles extensive datasets efficiently and includes thousands of community-contributed packages. Its visualization capabilities are unmatched, producing professional graphics and dashboards. The language includes virtually every statistical method and runs on Windows, macOS, and Linux seamlessly.
# Installing and loading a package
install.packages("ggplot2")
library(ggplot2)
Understanding package management is essential. The install.packages() function downloads and installs a package from CRAN. You only need to install a package once. The library() function loads the package into your current R session, making its functions available for use. You must load packages in every new session.
Setting Up R on Your Computer
Before you can start using R, you need to install the software on your computer. Installing R involves visiting the CRAN (Comprehensive R Archive Network) website, choosing your operating system, downloading the installer, and running it. The installation wizard will guide you through the process. Accept the default settings unless you have specific reasons to change them.
While R works perfectly on its own, RStudio enhances the experience significantly. RStudio is what programmers call an Integrated Development Environment (IDE). It brings together the R console, a text editor for writing code, a file manager, and a plot viewer all in one window. Visit the RStudio official website, download the free version, and run the installer.
For beginners, RStudio reduces confusion and makes learning easier. You can write a script, run it, see the results, and view your plots all in one place. The interface helps you understand what is happening at each step of your analysis.
Your First R Program
Once you have R and RStudio installed, you are ready to write and run your first program. The simplest approach is using the console. In RStudio, the console is the panel where you see the > prompt. You can type commands directly and see results immediately.
print("Welcome to R programming!")
When you type this and press Enter, R responds with [1] "Welcome to R programming!".
For longer work, use scripts. In RStudio, click File → New File → R Script. This opens a blank editor panel where you can type multiple lines of code.
# My first R script
message <- "Hello, this is my first program!"
print(message)
# Performing a calculation
number <- 42
answer <- number * 2
print(answer)
To run this script, click the “Run” button at the top of the editor panel, or press Ctrl+Enter (Command+Enter on Mac). The code executes line by line, and you will see the outputs in the console.
Tips for Success
Learning R takes time, but the journey is rewarding. Start with simple operations. Master variables, data types, and basic arithmetic before moving to complex analyses. Always comment your code to remember what you were thinking and help others understand your work. Practice consistently — even fifteen minutes daily builds muscle memory and reinforces concepts. Do not be afraid to use packages; R’s power comes from its ecosystem.
R Fundamentals
This section covers the essential building blocks of R programming — the concepts you will use in every program you write.
Understanding R Syntax and Comments
Syntax refers to the rules that govern how R code must be written. Every instruction in R is called a statement, and statements can be written one per line. R is case-sensitive, so variable and Variable are considered completely different names.
Comments are lines of text that R ignores, existing only for human readers. In R, the hash symbol # starts a comment. Any text written after # on the same line is ignored when the code is executed. Good comments explain the purpose of your code, making it easier to understand, modify, and debug.
# This script demonstrates R syntax and comments
# All lines starting with # are comments ignored by R
x <- 10 # Assign value to x
y <- 5 # Assign value to y
result <- x * y # Perform multiplication
print(result) # Display result: [1] 50
Each line of the code represents a complete statement that R can execute independently. The # symbol designates comments, spaces between elements improve readability, and R executes statements sequentially from top to bottom.
Variables, Identifiers, and Data Types
Variables are containers that store data. Each variable has a name, and you use that name to refer to the data later. Variable names in R must start with a letter or a dot, can contain letters, digits, dots, and underscores, and cannot contain spaces or special characters.
R supports several core data types. Numeric types store numbers with decimals. Integer types store whole numbers (use L after the number). Character types store text in quotes. Logical types store Boolean values (TRUE or FALSE). Complex types store numbers with imaginary parts.
# Variable naming - valid and invalid examples
student_name <- "John" # Valid - uses underscore
.age <- 25 # Valid - starts with dot
studentAge <- 30 # Valid - camelCase
# 1st_student <- "Alice" # Invalid - cannot start with number
# student-name <- "Bob" # Invalid - cannot contain hyphen
# R is case-sensitive
name <- "Alice"
Name <- "Bob" # Different variable
# Data types in R
numeric_var <- 25.5 # Numeric (decimals)
integer_var <- 10L # Integer (whole numbers, note the L)
character_var <- "Hello" # Character (text)
logical_var <- TRUE # Logical (TRUE/FALSE)
complex_var <- 3 + 4i # Complex (imaginary numbers)
# Checking data types
class(numeric_var) # "numeric"
class(integer_var) # "integer"
class(character_var) # "character"
class(logical_var) # "logical"
# Constants (by convention, use UPPERCASE)
PI <- 3.14159
MAX_SCORE <- 100
# Reserved words cannot be used as variable names
# if <- 10 # This would cause an error
Each data type serves a specific purpose. Numeric values store decimal numbers like 25.5, 3.14, or -10.2. Integer values store whole numbers like 10, -5, or 100 with the L suffix. Character values store text like "Hello", 'R', or "123" as text. Logical values represent Boolean conditions and can store either TRUE or FALSE in R. Complex values store numbers with imaginary parts like 3+4i.
Data Structures
Understanding R’s data structures is essential because they organize how data is stored and accessed.
Vectors
Vectors are one-dimensional collections of elements that generally contain values of the same data type. In R, they are commonly created with the c() function, which combines multiple values into a single vector. Vectors are one of R’s fundamental data structures and are widely used for storing and processing data. Every variable in R is actually a vector — even a single number is a vector of length 1.
# Vectors - same type elements
numeric_vector <- c(10, 20, 30, 40, 50)
character_vector <- c("A", "B", "C", "D")
logical_vector <- c(TRUE, FALSE, TRUE, TRUE)
# Accessing vector elements (indices start at 1)
numeric_vector[1] # First element: 10
numeric_vector[2:4] # Elements 2-4: 20, 30, 40
# Vector operations (applied to all elements)
numeric_vector * 2 # [1] 20 40 60 80 100
mean(numeric_vector) # [1] 30
sum(numeric_vector) # [1] 150
# Creating sequences
1:10 # [1] 1 2 3 4 5 6 7 8 9 10
seq(1, 10, by=2) # [1] 1 3 5 7 9
rep(5, times=3) # [1] 5 5 5
The c() function combines multiple elements into a vector. Vector indices start at 1, not 0 (unlike many other languages). The colon operator : creates sequences with step 1. The seq() function creates sequences with custom step sizes. The rep() function repeats values a specified number of times. Operations applied to vectors affect every element simultaneously.
Lists
Lists are more flexible than vectors — each element can hold different types of data. Lists can contain numbers, text, logical values, other lists, or any combination. This flexibility makes lists ideal for storing heterogeneous data.
# Lists - can hold mixed types
mixed_list <- list(25, "R Programming", TRUE, c(85, 90, 92))
mixed_list[[2]] # Access second element: "R Programming"
# Named lists
named_list <- list(name="Alice", age=25, scores=c(88, 92, 85))
named_list$name # "Alice"
named_list$scores # [1] 88 92 85
Use double brackets [[ ]] to access list elements. Single brackets [ ] return a sublist, not the element itself. Named lists allow accessing elements with $ notation. Lists can contain other lists, creating nested structures.
Matrices
Matrices are two-dimensional data structures in R in which all elements have the same data type. They are especially useful for mathematical operations, including linear algebra. The matrix() function creates a matrix from a vector of values by specifying its number of rows and columns.
# Matrices - 2D structures with same type
matrix_example <- matrix(1:6, nrow=2, ncol=3)
# [,1] [,2] [,3]
# [1,] 1 3 5
# [2,] 2 4 6
# Matrix by rows
matrix_byrow <- matrix(1:6, nrow=2, ncol=3, byrow=TRUE)
# [,1] [,2] [,3]
# [1,] 1 2 3
# [2,] 4 5 6
# Accessing elements [row, column]
matrix_example[1, 2] # Row 1, Column 2: 3
matrix_example[2, ] # Entire second row: 2, 4, 6
matrix_example[, 3] # Entire third column: 5, 6
# Matrix operations
matrix_example * 2 # Multiplies all elements
# Matrix multiplication
A <- matrix(1:4, nrow=2)
B <- matrix(5:8, nrow=2)
A %*% B # Matrix product
Elements are filled column-wise by default (first column, then second). Use byrow=TRUE to fill row-wise. Access elements using [row, column] notation. Leaving a dimension blank (e.g., [2, ]) returns the entire row or column. The %*% operator performs matrix multiplication, not element-wise multiplication.
Arrays
Arrays extend matrices by allowing data to be organized across more than two dimensions. Like matrices, all elements in an array must have the same data type. They are useful for multi-dimensional data like time series across multiple locations or image data with color channels.
# Arrays - multi-dimensional
array_example <- array(1:12, dim=c(2, 3, 2))
# Access [row, column, layer]
array_example[1, 2, 2] # Layer 2, Row 1, Column 2
The dim parameter specifies dimensions: rows, columns, layers, and so on. Arrays can have 3, 4, or more dimensions. Access uses comma-separated indices for each dimension.
Factors
Factors are R’s way of handling categorical data — data that falls into distinct categories like gender, education level, or product category. They store both the values and the possible categories (levels). Factors are essential for statistical modeling because many analyses treat categorical variables differently from numeric ones.
# Factors - for categorical data
gender <- factor(c("Male", "Female", "Female", "Male"))
levels(gender) # "Female" "Male"
table(gender) # Female: 2, Male: 2
# Ordered factors
size <- factor(c("Small", "Large", "Medium"),
levels=c("Small", "Medium", "Large"), ordered=TRUE)
The levels() function shows the unique categories. The table() function counts occurrences in each category. Ordered factors have a meaningful sequence. Factors are stored as integers internally for efficiency.
Data Frames
Data frames are one of the most commonly used data structures in R for working with structured data. They organize information in rows and columns, similar to a spreadsheet or database table, while allowing each column to contain a different data type.Data frames are used for virtually all data analysis tasks.
# Data frames - tabular data
employees <- data.frame(
Name = c("Alice", "Bob", "Carol", "David"),
Age = c(25, 30, 28, 35),
Department = c("Sales", "IT", "Marketing", "IT"),
Salary = c(50000, 60000, 52000, 65000)
)
# Accessing columns
employees$Name # [1] "Alice" "Bob" "Carol" "David"
employees$Salary # [1] 50000 60000 52000 65000
# Accessing rows and cells
employees[1, ] # First row
employees[, 3] # Third column (Department)
employees[2, 4] # Row 2, Column 4: 60000
# Adding and modifying columns
employees$Bonus <- employees$Salary * 0.10
employees$Age <- employees$Age + 1
# Adding a new row
new_employee <- data.frame(Name="Eve", Age=29,
Department="HR", Salary=55000, Bonus=5500)
employees <- rbind(employees, new_employee)
# Viewing structure
str(employees) # Shows column types
summary(employees) # Statistical summary
The $ operator accesses columns by name. The [row, column] syntax accesses specific cells. The str() function displays the structure (types of each column). The summary() function provides statistical summaries for numeric columns. The rbind() function adds rows, while cbind() adds columns.
Tibbles
Tibbles are a modern alternative to data frames introduced by the tidyverse package. They offer better printing and stricter adherence to data types. Tibbles do not convert strings to factors automatically, and they only print the first few rows and columns to avoid overwhelming output.
# Tibbles - modern data frames
library(tibble)
employees_tbl <- tibble(
Name = c("Alice", "Bob", "Carol"),
Age = c(25, 30, 28),
Department = c("Sales", "IT", "Marketing")
)
# Tibbles don't convert strings to factors automatically
Tibbles are part of the tidyverse ecosystem, offer more consistent behavior than data frames, provide better printing with colored output, and only show the first 10 rows and columns by default.
Operators and Expressions
Operators perform computations, comparisons, and logical evaluations. Expressions combine operators with values to produce results.
Arithmetic Operators
Arithmetic operators handle mathematical calculations. They work element-wise on vectors and return results of the same length.
# Arithmetic operators
x <- 20
y <- 6
x + y # Addition: 26
x - y # Subtraction: 14
x * y # Multiplication: 120
x / y # Division: 3.333333
x ^ 2 # Exponentiation: 400
x %% y # Modulus (remainder): 2
x %/% y # Integer division: 3
The + operator adds two numbers or vectors. The - operator subtracts the second from the first. The * operator multiplies values element-wise. The / operator divides the first by the second. The ^ operator raises the first to the power of the second. The %% operator returns the remainder after division. The %/% operator returns the integer quotient, discarding the remainder.
Relational Operators
Relational operators compare values and return TRUE or FALSE. They are fundamental for conditional statements and filtering data.
# Relational operators
a <- 15
b <- 20
a == b # Equal to: FALSE
a != b # Not equal to: TRUE
a < b # Less than: TRUE
a > b # Greater than: FALSE
a <= 15 # Less than or equal: TRUE
b >= 20 # Greater than or equal: TRUE
The == operator checks if values are exactly equal (use two equals signs). The != operator checks if values are not equal. The < and > operators check for less than or greater than. The <= and >= operators check for less/greater than or equal. Results can be used directly in if conditions and work element-wise on vectors.
Logical Operators
Logical operators combine or negate conditions. They are essential for complex conditional logic.
# Logical operators
x <- 12
y <- 8
(x > 10) & (y < 10) # AND: TRUE & TRUE → TRUE
(x > 15) & (y < 10) # AND: FALSE & TRUE → FALSE
(x > 15) | (y < 10) # OR: FALSE | TRUE → TRUE
!(x > 15) # NOT: NOT FALSE → TRUE
The & operator (AND) returns TRUE only when both conditions evaluate to TRUE; otherwise, it returns FALSE. The | operator (OR) returns TRUE if at least one condition is TRUE. The ! operator (NOT) reverses the logical value. Parentheses help group conditions clearly. The && and || operators are short-circuit versions of AND and OR in R. They evaluate only the first element of each logical expression and are mainly used when working with single logical values.
Assignment Operators
Assignment operators store values in variables. The <- operator is the most common and preferred in R.
# Assignment operators
score <- 95 # Left assignment (most common)
95 -> score # Right assignment
score = 95 # Equal assignment
The <- operator is the standard R assignment operator. The -> operator assigns in the opposite direction (less common). The = operator also works but can be confused with argument passing. RStudio has a keyboard shortcut for <- (Alt + – on Windows/Linux, Option + – on Mac).
Special Operators
Special operators provide unique functionality in R. The most common are %in% for membership checking and %*% for matrix multiplication.
# Special operators
colors <- c("red", "blue", "green")
"red" %in% colors # Membership check: TRUE
"yellow" %in% colors # Membership check: FALSE
# Matrix multiplication
A <- matrix(1:4, nrow=2)
B <- matrix(5:8, nrow=2)
A %*% B # Matrix product
The %in% operator checks if a value exists in a vector. The %*% operator performs matrix multiplication. The %/% and %% operators are also special operators.
Operator Precedence
Operator precedence determines the order of evaluation in expressions. Parentheses can override the default order, making your code clearer and ensuring correct results.
# Operator precedence
2 + 3 * 4 # 14 (multiplication first)
(2 + 3) * 4 # 20 (parentheses override)
The precedence order from highest to lowest is: exponentiation (^), then multiplication, division, and modulus (*, /, %%, %/%), then addition and subtraction (+, -), then relational operators (<, >, <=, >=), then equality operators (==, !=), then logical NOT (!), then logical AND (&), and finally logical OR (|).
Making Decisions and Repeating Actions
Control structures let your programs make decisions and repeat operations. These are the tools that make programs dynamic rather than linear.
Conditional Statements
Conditional statements execute code blocks based on specific conditions. The if statement executes code when a condition is TRUE. The if-else structure provides alternative paths for TRUE and FALSE conditions. Nested if-else handles multiple conditions. The switch function selects from multiple options.
# if statement - executes when condition is TRUE
score <- 82
if (score >= 70) {
print("Passing score achieved")
}
# Output: [1] "Passing score achieved"
# if-else structure - chooses between two paths
score <- 65
if (score >= 70) {
print("Passing score achieved")
} else {
print("Score is below passing threshold")
}
# Output: [1] "Score is below passing threshold"
# Nested conditions - handles multiple possibilities
score <- 78
if (score >= 90) {
grade <- "A"
} else if (score >= 80) {
grade <- "B"
} else if (score >= 70) {
grade <- "C"
} else if (score >= 60) {
grade <- "D"
} else {
grade <- "F"
}
print(paste("Grade:", grade)) # [1] "Grade: C"
# switch function - selects from multiple options
day_number <- 3
day_name <- switch(day_number,
"Monday", "Tuesday", "Wednesday",
"Thursday", "Friday", "Saturday", "Sunday")
print(day_name) # [1] "Wednesday"
# switch with named cases
weather_code <- "rainy"
weather_description <- switch(weather_code,
sunny = "It's a beautiful day",
rainy = "Remember to take an umbrella",
cloudy = "It might rain later",
"Weather forecast not available")
print(weather_description) # [1] "Remember to take an umbrella"
Conditions must evaluate to TRUE or FALSE. The if statement executes only the TRUE block. The else statement executes when the condition is FALSE. The else if statement allows multiple conditions to be checked sequentially. The switch function is cleaner than multiple else if statements for multiple discrete cases.
Loop Structures
Loop structures repeat code execution. The for loop executes a fixed number of iterations, useful when you know exactly how many times to repeat. The while loop continues while a condition remains TRUE. The repeat loop executes indefinitely until a break condition is met.
# For loops
for (i in 1:5) {
print(i)
}
# Output: 1, 2, 3, 4, 5
# For loop over a vector
fruits <- c("Apple", "Banana", "Cherry")
for (fruit in fruits) {
print(paste("I like", fruit))
}
# While loop
counter <- 1
while (counter <= 5) {
print(counter)
counter <- counter + 1
}
# Output: 1, 2, 3, 4, 5
# While loop for summation
total <- 0
i <- 1
while (i <= 10) {
total <- total + i
i <- i + 1
}
# total is 55
# Repeat loop
counter <- 1
repeat {
print(counter)
counter <- counter + 1
if (counter > 5) {
break
}
}
# Output: 1, 2, 3, 4, 5
# Loop control commands
for (i in 1:5) {
if (i == 3) next # Skip 3
if (i == 5) break # Exit at 5
print(i)
}
# Output: 1, 2, 4
# Nested loops
for (i in 1:3) {
for (j in 1:3) {
cat(i, "*", j, "=", i*j, "\n")
}
}
The for loop iterates over sequences or vector elements. The while loop checks the condition before each iteration. The repeat loop runs until explicitly stopped with break. The next statement skips the current iteration and goes to the next. The break statement exits the loop completely.
Vectorized Operations
Vectorized operations are a defining feature of R. Unlike loops that process one element at a time, vectorized operations apply to entire vectors simultaneously. They are faster, more readable, and preferred in R programming.
# Vectorized operations (preferred over loops)
numbers <- c(5, 10, 15, 20)
doubled <- numbers * 2 # [1] 10 20 30 40
squared <- numbers ^ 2 # [1] 25 100 225 400
# Compare loop vs vectorized
# Loop approach (slower)
result <- c()
for (i in 1:4) {
result[i] <- numbers[i] * 2
}
# Vectorized approach (faster and cleaner)
result <- numbers * 2
Operations apply to every element simultaneously. Vectorized operations are much faster than loops in R, more concise and readable, and work with all arithmetic, relational, and logical operators.
Functions in R
Functions are reusable blocks of code that perform specific tasks. Instead of writing the same code repeatedly, you define a function once and call it whenever you need it.
Creating and Using Functions
A function definition includes a name, a list of arguments (parameters), and the code to execute. Functions can have default argument values, accept named arguments, and return values either explicitly with return() or implicitly as the last evaluated expression.
Variable scope determines where variables are accessible. Local variables exist only inside functions. Global variables are accessible everywhere. Understanding scope helps you write predictable, bug-free code.
# Basic function definition
greet <- function(name) {
paste("Hello,", name, "!")
}
greet("Alice") # "Hello, Alice !"
# Function with default arguments
greet <- function(name = "Friend") {
paste("Hello,", name, "!")
}
greet() # "Hello, Friend !"
greet("Bob") # "Hello, Bob !"
# Function with multiple arguments
calculate_discount <- function(price, discount_rate = 0.10) {
return(price * (1 - discount_rate))
}
calculate_discount(100) # 90
calculate_discount(100, discount_rate = 0.20) # 80
# Named arguments (order doesn't matter)
calculate_discount(discount_rate = 0.20, price = 100)
# Implicit return (last expression is returned)
square <- function(x) {
x ^ 2
}
square(4) # 16
# Anonymous functions (no name)
sapply(1:5, function(x) x ^ 2) # [1] 1 4 9 16 25
# Variable scope - local variables
my_function <- function() {
local_var <- 10
print(local_var)
}
my_function() # Prints 10
# print(local_var) # Error: local_var doesn't exist outside
# Global variables
global_var <- 20
my_function <- function() {
print(global_var)
}
my_function() # Prints 20
# Recursive functions (function calls itself)
factorial <- function(n) {
if (n <= 1) {
return(1)
} else {
return(n * factorial(n - 1))
}
}
factorial(5) # 120
Functions are defined using the function() syntax. Arguments are passed inside parentheses. Default values make functions more flexible. Named arguments allow calling with any argument order. The return() function explicitly returns a value. The last expression is returned if return() is omitted.
The Apply Family
The apply family of functions in R provides convenient ways to perform operations across vectors, lists, matrices, and data structures, often reducing the need for explicit loops. These functions apply operations to data structures efficiently.
# apply() - works on matrices
mat <- matrix(1:9, nrow=3)
apply(mat, 1, mean) # Row means
apply(mat, 2, mean) # Column means
# lapply() - returns a list
list_data <- list(a = 1:3, b = 4:6)
lapply(list_data, sum)
# sapply() - returns a simplified structure
sapply(list_data, sum)
The apply(mat, 1, function) applies to rows (1 = rows). The apply(mat, 2, function) applies to columns (2 = columns). The lapply() function always returns a list. The sapply() function simplifies the output if possible.
Data Manipulation
Data manipulation is the process of getting your data ready for analysis. This includes importing, cleaning, and transforming data.
Importing Data
R can read data from many file formats including CSV, Excel, text, RDS, and JSON. Each format has its own import function with specific parameters.
# CSV Files
data <- read.csv("data.csv")
# Excel Files
library(readxl)
data <- read_excel("data.xlsx")
# Text Files
data <- read.table("data.txt", header=TRUE, sep="\t")
# RDS Files (R's native format)
data <- readRDS("data.rds")
# JSON Files
library(jsonlite)
data <- fromJSON("data.json")
The read.csv() function imports CSV files with header detection. The read_excel() function requires the readxl package. The read.table() function is flexible for various text formats. The sep="\t" parameter indicates tab-separated values. The header=TRUE parameter indicates the first row contains column names.
Exporting Data
Saving results is equally straightforward using corresponding write functions.
# CSV
write.csv(data, "output.csv", row.names=FALSE)
# Excel
library(writexl)
write_xlsx(data, "output.xlsx")
# RDS
saveRDS(data, "output.rds")
# JSON
toJSON(data, pretty=TRUE)
The row.names=FALSE parameter prevents R from saving row numbers. The pretty=TRUE parameter formats JSON for readability.
Data Cleaning
Real-world data often has problems that need fixing. Missing values can be replaced with means or other values. Duplicates should be removed to avoid skewed results. Outliers can be identified and handled using statistical methods like the IQR approach.
# Create data with issues
data <- data.frame(
id = 1:5,
age = c(25, NA, 30, 22, NA),
salary = c(50000, 60000, NA, 55000, 70000)
)
# Handle missing values
data$age[is.na(data$age)] <- mean(data$age, na.rm=TRUE)
data$salary[is.na(data$salary)] <- mean(data$salary, na.rm=TRUE)
# Remove duplicates
data <- data[!duplicated(data), ]
# Handle outliers using IQR
Q1 <- quantile(data$salary, 0.25)
Q3 <- quantile(data$salary, 0.75)
IQR <- Q3 - Q1
data <- data[data$salary >= (Q1 - 1.5*IQR) &
data$salary <= (Q3 + 1.5*IQR), ]
The is.na() function identifies missing values. The na.rm=TRUE parameter ignores missing values in calculations. The !duplicated() function creates a logical vector for unique rows. The IQR method removes values beyond 1.5×IQR from quartiles.
Data Transformation
Data transformation includes subsetting rows and columns, sorting data, merging datasets, and aggregating information.
# Subsetting
older_employees <- data[data$age > 25, ]
selected <- data[, c("id", "salary")]
# Sorting
sorted <- data[order(data$salary, decreasing=TRUE), ]
# Merging datasets
data1 <- data.frame(id=1:3, name=c("Alice", "Bob", "Carol"))
data2 <- data.frame(id=1:3, salary=c(50000, 60000, 55000))
merged <- merge(data1, data2, by="id")
# Aggregation
aggregate(salary ~ department, data=data, FUN=mean)
The [rows, columns] syntax is used for subsetting. The order() function sorts by columns, with decreasing=TRUE for descending order. The merge() function combines datasets by a common key column. The aggregate() function computes summary statistics by group.
The dplyr Package
dplyr provides intuitive verbs for data manipulation, making code more readable and maintainable.
library(dplyr)
data <- data.frame(
name = c("Alice", "Bob", "Carol", "Dave", "Eve"),
age = c(25, 32, 28, 35, 29),
department = c("Sales", "IT", "Sales", "IT", "HR"),
salary = c(50000, 60000, 48000, 65000, 52000)
)
# Filter rows
filter(data, age > 28, salary > 50000)
# Select columns
select(data, name, salary)
# Create new columns
mutate(data, salary_k = salary / 1000)
# Sort rows
arrange(data, desc(salary))
# Group and summarize
data %>%
group_by(department) %>%
summarise(avg_salary = mean(salary), count = n())
The filter() function selects rows based on conditions. The select() function chooses columns to keep. The mutate() function creates or modifies columns. The arrange() function sorts rows (use desc() for descending). The group_by() function prepares data for grouped operations. The summarise() function collapses groups into summary statistics. The pipe %>% chains operations sequentially.
The tidyr Package
tidyr helps reshape data between wide and long formats, a common task in data analysis.
library(tidyr)
# Wide format (months as columns)
wide_data <- data.frame(year=c(2020,2021), q1=c(100,120), q2=c(110,130))
# Wide to long
long_data <- pivot_longer(wide_data, cols=starts_with("q"),
names_to="quarter", values_to="sales")
# Long to wide
wide_again <- pivot_wider(long_data, names_from="quarter", values_from="sales")
The pivot_longer() function converts multiple columns into two columns (key-value pairs). The pivot_wider() function spreads key-value pairs into multiple columns. The starts_with() function selects columns by prefix.
Data Visualization
Data visualization transforms numbers into pictures. Good visualizations reveal patterns that numbers alone cannot show.
Base R Graphics
R includes built-in plotting functions for creating various charts. These can be customized with colors, labels, titles, and legends.
The plot() function is the most versatile base R plotting tool. It creates scatter plots and line plots.
x <- 1:10
y <- x * 2 + rnorm(10)
# Basic scatter plot
plot(x, y, main="Scatter Plot", xlab="X Axis", ylab="Y Axis")
# Line plot with customization
plot(x, y, type="l", col="blue", lwd=2,
main="Line Plot", xlab="X Values", ylab="Y Values")
The main parameter sets the plot title. The xlab and ylab parameters set axis labels. The type parameter controls the plot style — "p" for points (default), "l" for lines, "b" for both points and lines, "o" for overplotted points and lines, "h" for histogram-like vertical lines, "s" for stair steps, and "n" for no plotting (useful for custom annotations). The col parameter sets color (use names like "red", "blue" or hex codes). The lwd parameter sets line width (1 is default, 2 is thicker). The pch parameter sets point character (1 for circle, 2 for triangle, and so on). The lty parameter sets line type (1 for solid, 2 for dashed, 3 for dotted, and so on).
Bar charts display categorical data with rectangular bars using the barplot() function.
sales <- c(45, 67, 52, 89, 73)
barplot(sales, names.arg=c("A","B","C","D","E"),
col="skyblue", main="Sales by Category",
xlab="Category", ylab="Sales Amount")
The names.arg parameter labels the bars. The col parameter sets bar colors. The main parameter sets the title. The xlab and ylab parameters set axis labels. The horiz=TRUE parameter creates a horizontal bar chart. The beside=TRUE parameter creates grouped bars for matrices.
Histograms show the distribution of numeric data using the hist() function.
ages <- c(22,25,23,28,30,24,29,27,26,25,31,28)
hist(ages, breaks=5, col="lightgreen",
main="Age Distribution", xlab="Age",
ylab="Frequency", border="black")
The breaks parameter controls the number of bins. The col parameter sets the fill color of bars. The border parameter sets the border color of bars.
Boxplots display the five-number summary: minimum, Q1, median, Q3, and maximum, along with outliers.
scores <- c(65,70,72,68,85,90,45,78,82,69)
boxplot(scores, main="Score Distribution", ylab="Scores",
col="lightblue", border="darkblue")
The box shows Q1 to Q3. The line inside is the median. The whiskers extend to min and max (excluding outliers). Points beyond whiskers are outliers. The col parameter sets the fill color. The border parameter sets the border color. Multiple groups can be plotted using boxplot(scores ~ group, data=df).
Pie charts display proportional data using the pie() function.
slices <- c(30, 25, 20, 15, 10)
labels <- c("A", "B", "C", "D", "E")
pie(slices, labels=labels, main="Distribution",
col=c("red","blue","green","yellow","purple"))
The labels parameter sets slice labels. The col parameter sets slice colors. The clockwise=TRUE parameter plots in clockwise direction.
The ggplot2 Package
ggplot2 implements the Grammar of Graphics, creating visualizations by layering components. This systematic approach makes complex plots easier to build and understand.
The basic structure includes ggplot() to initialize the plot with data, aes() to map variables to visual properties (aesthetics), geom_*() to add geometric objects (points, lines, bars), labs() to add labels and titles, and theme_*() to apply predefined themes.
library(ggplot2)
# Basic structure: ggplot(data, aes()) + geom_*()
ggplot(data, aes(x = variable_x, y = variable_y)) +
geom_point() +
labs(title = "My Plot", x = "X Label", y = "Y Label") +
theme_minimal()
Common geometric objects include geom_point() for scatter plots, geom_line() for line plots, geom_bar() for bar charts, geom_histogram() for histograms, geom_boxplot() for boxplots, and geom_density() for density plots.
# Scatter plot
ggplot(data, aes(x=x, y=y)) +
geom_point(color="blue", size=3, alpha=0.7)
# Line plot
ggplot(data, aes(x=x, y=y)) +
geom_line(color="red", size=1.2)
# Points and lines together
ggplot(data, aes(x=x, y=y)) +
geom_point(color="blue", size=3) +
geom_line(color="red", linetype="dashed")
# Bar chart
ggplot(df, aes(x=category, y=value)) +
geom_bar(stat="identity", fill="steelblue")
# Histogram
ggplot(data, aes(x=score)) +
geom_histogram(bins=20, fill="lightgreen", color="black")
# Boxplot
ggplot(df, aes(x=group, y=value, fill=group)) +
geom_boxplot()
# Density plot
ggplot(data, aes(x=value)) +
geom_density(fill="skyblue", alpha=0.5)
The color parameter sets the outline color. The fill parameter sets the fill color. The size parameter controls point or line thickness. The alpha parameter controls transparency (0-1). The linetype parameter sets line style. The stat="identity" parameter is used for bar charts with pre-calculated values.
Comprehensive customizations make plots publication-ready.
# Comprehensive customization
ggplot(df, aes(x=category, y=value, fill=category)) +
geom_bar(stat="identity") +
labs(title="Category Analysis",
x="Product Category",
y="Total Sales") +
theme_minimal() +
theme(legend.position="bottom",
plot.title=element_text(hjust=0.5, face="bold"),
axis.text=element_text(size=12)) +
scale_fill_manual(values=c("blue", "red", "green", "orange"))
# Faceting - split plots by category
ggplot(data, aes(x=x, y=y)) +
geom_point() +
facet_wrap(~category, ncol=2)
The labs() function adds title and axis labels. The theme_minimal() function applies a clean theme. The theme() function allows fine-grained control of all plot elements. The legend.position parameter controls legend placement. The scale_fill_manual() function customizes fill colors. The facet_wrap() function creates multiple panels by a grouping variable.
Interactive Visualizations
The plotly package creates interactive plots that can be explored with hover, zoom, and pan.
library(plotly)
# Interactive scatter plot
data <- data.frame(x=1:20, y=cumsum(rnorm(20)))
plot_ly(data, x=~x, y=~y, type="scatter", mode="lines+markers")
# Interactive bar chart
plot_ly(df, x=~category, y=~value, type="bar")
Statistical Analysis
R was built for statistics. It includes functions for virtually every statistical method you might need.
Descriptive Statistics
Descriptive statistics summarize data’s central tendency and spread. The mean is the average value. The median is the middle value. Standard deviation measures spread from the mean. Variance is the squared standard deviation. Quantiles divide data into parts.
data <- c(12, 15, 18, 20, 22, 25, 28, 30, 32, 35)
# Mean - average value
mean(data) # 23.7
# Median - middle value
median(data) # 23.5
# Standard deviation - spread from mean
sd(data) # 7.33
# Variance - squared standard deviation
var(data) # 53.79
# Quantiles - divide data into parts
quantile(data) # 0%, 25%, 50%, 75%, 100%
# Comprehensive summary
summary(data) # All statistics at once
The mean is sensitive to outliers. The median is robust to outliers. Standard deviation measures how spread out values are — larger values mean more spread. Quantiles divide data into equal parts.
Probability Distributions
Probability distributions describe how data values are spread. R can generate random numbers from various distributions including normal, binomial, and Poisson.
# Normal distribution
rnorm(5, mean=100, sd=15) # Generate 5 random numbers
dnorm(100, mean=100, sd=15) # Probability density at point 100
pnorm(115, mean=100, sd=15) # Cumulative probability ≤ 115
qnorm(0.95, mean=100, sd=15) # Quantile at probability 0.95
# Binomial distribution
rbinom(10, size=5, prob=0.5) # 10 random values from 5 trials
dbinom(3, size=5, prob=0.5) # Probability of exactly 3 successes
pbinom(3, size=5, prob=0.5) # Cumulative probability ≤ 3
# Poisson distribution
rpois(10, lambda=3) # Random values with average 3
The r* functions generate random numbers. The d* functions compute probability density or mass. The p* functions compute cumulative probability. The q* functions compute quantiles.
Hypothesis Testing
Hypothesis testing helps make decisions about data. The t-test compares means of two groups. The chi-square test checks relationships between categorical variables. ANOVA compares means of more than two groups.
# t-test - compare means of two groups
group1 <- c(5, 6, 7, 8, 9)
group2 <- c(8, 9, 10, 11, 12)
t.test(group1, group2)
# Chi-square test - categorical independence
observed <- matrix(c(10, 20, 30, 40), nrow=2)
chisq.test(observed)
# ANOVA - compare means of more than two groups
group <- factor(c("A","A","A","B","B","B","C","C","C"))
values <- c(5,6,7,8,9,10,11,12,13)
result <- aov(values ~ group)
summary(result)
# Correlation
x <- c(1, 2, 3, 4, 5)
y <- c(3, 5, 7, 9, 11)
cor(x, y) # 1.0 (perfect positive)
The t-test p-value less than 0.05 suggests a significant difference. The chi-square test tests association between categories. The ANOVA tests if means across groups differ. Correlation ranges from -1 to +1.
Regression Analysis
Regression models predict outcomes using input variables. Linear regression predicts continuous outcomes. Logistic regression predicts binary outcomes.
# Linear regression
x <- 1:10
y <- 2 * x + 3 + rnorm(10)
model <- lm(y ~ x)
summary(model) # Shows coefficients, R-squared, p-values
# Multiple linear regression
data <- data.frame(x1=c(1,2,3,4,5), x2=c(5,4,3,2,1), y=c(7,8,9,10,11))
model <- lm(y ~ x1 + x2, data=data)
summary(model)
# Logistic regression
data <- data.frame(age=c(25,30,35,40,45,50,55,60),
bought=c(0,0,1,0,1,1,1,1))
model <- glm(bought ~ age, data=data, family=binomial)
summary(model)
The lm() function fits linear models. R-squared measures goodness of fit. P-values test if coefficients are significant. The glm(family=binomial) function performs logistic regression.
Time Series Analysis
Time series analysis focuses on studying data points recorded at different moments to understand patterns, trends, and changes over time. R can decompose series into components and forecast future values.
library(forecast)
# Create time series
ts_data <- ts(c(10,12,14,16,18,20,22,24,26,28,30,32),
frequency=12, start=c(2020,1))
# Decompose into components
decomposed <- decompose(ts_data)
# Auto-ARIMA forecasting
model <- auto.arima(ts_data)
forecast(model, h=6) # Forecast 6 periods ahead
The ts() function creates time series objects. The frequency parameter indicates seasonal pattern (12 for monthly data). The decompose() function separates trend, seasonal, and residual components. The auto.arima() function automatically selects the best ARIMA model.
Machine Learning
Machine learning allows R to learn patterns from data and make predictions. R provides numerous packages for machine learning.
Supervised Learning
Supervised learning uses labeled data to train models. Linear regression predicts continuous values like house prices or sales amounts. Logistic regression predicts binary outcomes like customer churn or loan default. Decision trees split data based on conditions, making them interpretable. Random forest combines many decision trees for better accuracy.Support Vector Machines (SVMs) classify data by finding the best boundary that separates different classes, while maximizing the margin between them.
# Load required libraries
library(caret)
library(randomForest)
library(e1071)
library(rpart)
# --- HOUSE PRICE PREDICTION (Linear Regression) ---
# Real-world scenario: Predicting house prices based on features
set.seed(123)
housing <- data.frame(
sqft = c(1200, 1500, 1800, 2000, 2200, 2500, 2800, 3000, 3200, 3500),
bedrooms = c(2, 3, 3, 3, 4, 4, 4, 5, 5, 5),
bathrooms = c(1, 2, 2, 2, 3, 3, 3, 4, 4, 4),
age = c(10, 8, 5, 3, 2, 1, 15, 12, 8, 5),
price = c(180000, 220000, 260000, 300000, 340000,
380000, 290000, 320000, 360000, 410000)
)
# View first few rows
head(housing, 3)
# sqft bedrooms bathrooms age price
# 1 1200 2 1 10 180000
# 2 1500 3 2 8 220000
# 3 1800 3 2 5 260000
# Train linear regression model
price_model <- lm(price ~ sqft + bedrooms + bathrooms + age, data=housing)
summary(price_model)
# Predict price for a new house
new_house <- data.frame(sqft=2600, bedrooms=4, bathrooms=3, age=6)
predicted_price <- predict(price_model, newdata=new_house)
print(paste("Predicted house price: $", round(predicted_price, 0)))
# Output: "Predicted house price: $ 372500"
# --- CUSTOMER CHURN PREDICTION (Logistic Regression) ---
# Real-world scenario: Predicting which customers will cancel service
churn_data <- data.frame(
tenure_months = c(12, 24, 36, 6, 48, 18, 3, 60, 30, 42),
monthly_charges = c(50, 60, 55, 70, 45, 80, 65, 40, 75, 58),
support_calls = c(1, 2, 0, 3, 0, 4, 5, 0, 2, 1),
churned = c(0, 0, 0, 1, 0, 1, 1, 0, 0, 0) # 1 = churned, 0 = stayed
)
head(churn_data, 3)
# tenure_months monthly_charges support_calls churned
# 1 12 50 1 0
# 2 24 60 2 0
# 3 36 55 0 0
# Train logistic regression model
churn_model <- glm(churned ~ tenure_months + monthly_charges + support_calls,
data=churn_data, family=binomial)
summary(churn_model)
# Predict churn probability for a new customer
new_customer <- data.frame(tenure_months=8, monthly_charges=72, support_calls=4)
churn_probability <- predict(churn_model, newdata=new_customer, type="response")
print(paste("Churn probability:", round(churn_probability * 100, 1), "%"))
# Output: "Churn probability: 82.3 %"
# --- CUSTOMER SEGMENTATION (K-Means Clustering) ---
# Real-world scenario: Grouping customers by purchasing behavior
set.seed(123)
customers <- data.frame(
customer_id = 1:100,
annual_income = sample(30000:120000, 100, replace=TRUE),
spending_score = sample(1:100, 100, replace=TRUE),
purchase_frequency = sample(1:50, 100, replace=TRUE)
)
head(customers, 3)
# customer_id annual_income spending_score purchase_frequency
# 1 1 66112 27 39
# 2 2 95688 74 44
# 3 3 31206 82 29
# Perform K-means clustering (3 segments)
segments <- kmeans(customers[, c("annual_income", "spending_score", "purchase_frequency")],
centers=3, nstart=25)
customers$segment <- segments$cluster
# Analyze each segment
segment_summary <- customers %>%
group_by(segment) %>%
summarise(
avg_income = mean(annual_income),
avg_spending = mean(spending_score),
avg_purchase_freq = mean(purchase_frequency),
count = n()
)
print(segment_summary)
# # A tibble: 3 × 5
# segment avg_income avg_spending avg_purchase_freq count
# <int> <dbl> <dbl> <dbl> <int>
# 1 1 76850 55.2 27.1 33
# 2 2 53326 35.7 16.9 34
# 3 3 92750 76.8 37.2 33
Linear regression uses lm() to predict continuous values. Logistic regression uses glm(family=binomial) to predict binary outcomes. K-means uses kmeans() to group similar data points. The set.seed() function ensures reproducible results. The head() function displays the first few rows of data. The summary() function shows model statistics and coefficients.
Model Evaluation
Model evaluation ensures models perform well on new data. Common metrics include accuracy, precision, recall, and F1-score.
# Model evaluation with caret
library(caret)
# Split data into training and test sets
set.seed(123)
train_index <- createDataPartition(iris$Species, p=0.7, list=FALSE)
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]
# Train model on training data
model <- train(Species ~ ., data=train_data, method="rf")
# Predict on test data
predictions <- predict(model, test_data)
# Evaluate performance
confusionMatrix(predictions, test_data$Species)
The createDataPartition() function splits data (70% training, 30% testing). The train() function fits the model using caret’s unified interface. The predict() function generates predictions on new data. The confusionMatrix() function produces a confusion matrix with accuracy metrics.
Advanced R Programming
This section covers techniques for writing efficient, robust, and professional R code.
Object-Oriented Programming
R supports multiple OOP systems. The S3 system is the simplest, where classes are defined by a class attribute and methods are functions that act on objects. The S4 system is more formal with strict rules. The R6 system provides mutable objects similar to Python or Java classes.
# S3 OOP system
person <- list(name="Alice", age=25)
class(person) <- "Person"
print.Person <- function(obj) {
cat("Name:", obj$name, "\nAge:", obj$age, "\n")
}
print(person)
# S4 OOP system
setClass("Person", slots=list(name="character", age="numeric"))
alice <- new("Person", name="Alice", age=25)
alice@name
# R6 OOP system
library(R6)
Person <- R6Class("Person",
public=list(
name=NULL,
age=NULL,
initialize=function(name, age) {
self$name <- name
self$age <- age
},
greet=function() {
paste("Hello, my name is", self$name)
}
))
alice <- Person$new("Alice", 25)
alice$greet()
Error Handling
Error handling prevents programs from crashing. The try() function catches errors and continues execution. The tryCatch() function allows handling errors with specific actions.
# try() - catches errors and continues
result <- try(log("a"), silent=TRUE)
if (class(result) == "try-error") {
print("There was an error")
}
# tryCatch() - handles errors with specific actions
result <- tryCatch({
log("a")
}, warning=function(w) {
print("Warning occurred")
}, error=function(e) {
print("Error occurred")
return(NA)
})
Functional Programming
Functional programming treats functions as data. Closures remember their environment. Higher-order functions take or return other functions.
# Closures - functions that remember their environment
make_adder <- function(x) {
function(y) x + y
}
add_five <- make_adder(5)
add_five(10) # 15
# Higher-order functions
apply_function <- function(f, x) {
f(x)
}
apply_function(sqrt, 16) # 4
Parallel Computing
Parallel computing runs code on multiple cores for speed. The parallel package provides functions like parLapply() and parSapply().
# Parallel computing
library(parallel)
cl <- makeCluster(2)
result <- parLapply(cl, 1:10, function(x) x^2)
stopCluster(cl)
Memory Management
Memory management optimizes performance for large datasets. The gc() function frees unused memory. The data.table package provides efficient structures.
# Memory management
gc() # Garbage collection
# data.table for large datasets
library(data.table)
dt <- data.table(x=1:1000000, y=rnorm(1000000))
dt[, mean(y)] # Fast aggregation
Packages and Libraries
R’s power comes from its ecosystem of packages — collections of functions, data, and documentation.
Installing and Loading Packages
You only need to install a package once, but you must load it in every new session.
# Installing packages
install.packages("ggplot2")
install.packages("dplyr")
install.packages("tidyr")
# Loading packages
library(ggplot2)
library(dplyr)
Essential Packages for Data Science
Essential packages include dplyr for data manipulation, tidyr for data reshaping, ggplot2 for visualization, data.table for high-performance operations, lubridate for date handling, stringr for text manipulation, and purrr for functional programming.
# dplyr - data manipulation
library(dplyr)
data %>% filter(age > 30) %>% select(name, salary)
# tidyr - data reshaping
library(tidyr)
pivot_longer(data, cols=starts_with("Q"), names_to="Quarter", values_to="Sales")
# ggplot2 - visualization
library(ggplot2)
ggplot(data, aes(x, y)) + geom_point()
# data.table - high performance
library(data.table)
dt <- data.table(x=1:5, y=6:10)
dt[y > 7]
# lubridate - date handling
library(lubridate)
date <- ymd("2024-12-25")
year(date); month(date); day(date)
# stringr - string manipulation
library(stringr)
str_length("Hello R")
str_extract("Hello R", "[A-Z][a-z]+")
# purrr - functional programming
library(purrr)
map(1:5, ~ .x^2)
Creating Your Own Package
Creating your own package helps organize and share your work. The devtools package provides tools for creating package structures.
# Creating your own package
library(devtools)
create("MyPackage")
# Add functions to R/ directory
# Add documentation using roxygen2
install("MyPackage")
Real-World Applications
This section demonstrates how R solves practical problems.
Sales Analysis Dashboard
Sales analysis helps businesses track performance. R can aggregate sales data by product, month, or region, calculate totals and averages, and create visualizations that reveal trends.
library(dplyr)
library(ggplot2)
# Sample sales data
sales <- data.frame(
month = c("Jan","Jan","Feb","Feb","Mar","Mar"),
product = c("A","B","A","B","A","B"),
revenue = c(1000,1500,1200,1600,1300,1800)
)
# Analyze by product
sales_summary <- sales %>%
group_by(product) %>%
summarise(total_revenue = sum(revenue),
avg_revenue = mean(revenue))
print(sales_summary)
# # A tibble: 2 × 3
# product total_revenue avg_revenue
# <chr> <dbl> <dbl>
# 1 A 3500 1167.
# 2 B 4900 1633.
# Visualize
ggplot(sales, aes(x=month, y=revenue, fill=product)) +
geom_bar(stat="identity", position="dodge") +
labs(title="Monthly Sales by Product", x="Month", y="Revenue")
Customer Segmentation
Customer segmentation groups customers based on behavior or demographics. K-means clustering identifies natural groupings. The resulting segments can be analyzed to understand each group’s characteristics.
# Sample customer data
set.seed(123)
customers <- data.frame(
id = 1:100,
age = sample(18:65, 100, replace=TRUE),
income = sample(20000:100000, 100, replace=TRUE),
spending_score = sample(1:100, 100, replace=TRUE)
)
# K-means clustering
segments <- kmeans(customers[, c("age", "income", "spending_score")],
centers=4)
customers$segment <- segments$cluster
# Analyze segments
segment_summary <- customers %>%
group_by(segment) %>%
summarise(avg_age = mean(age),
avg_income = mean(income),
avg_spending = mean(spending_score),
count = n())
print(segment_summary)
Time Series Forecasting
Time series forecasting helps businesses plan for the future. ARIMA models automatically identify patterns in historical data and project them forward.
library(forecast)
# Sample monthly sales data
monthly_sales <- ts(c(100,110,105,120,115,130,125,140,135,150,145,160),
frequency=12, start=c(2023,1))
# Fit ARIMA model
model <- auto.arima(monthly_sales)
# Forecast next 6 months
forecast_result <- forecast(model, h=6)
print(forecast_result)
# Visualize forecast
autoplot(forecast_result)
Interactive Shiny Application
Shiny applications create interactive web interfaces from R code. The UI defines layout and controls, the server handles reactive outputs, and the result is a fully functional dashboard.
library(shiny)
ui <- fluidPage(
titlePanel("Data Explorer"),
sidebarLayout(
sidebarPanel(
sliderInput("size", "Sample Size:", 10, 100, 50),
numericInput("mean", "Mean:", 50, 1, 100),
numericInput("sd", "Std Dev:", 10, 1, 20)
),
mainPanel(
plotOutput("histogram"),
verbatimTextOutput("summary")
)
)
)
server <- function(input, output) {
output$histogram <- renderPlot({
data <- rnorm(input$size, input$mean, input$sd)
hist(data, col="skyblue", border="white",
main="Random Distribution", xlab="Value")
})
output$summary <- renderPrint({
data <- rnorm(input$size, input$mean, input$sd)
summary(data)
})
}
shinyApp(ui=ui, server=server)
The UI defines the user interface layout. The fluidPage() function creates a responsive layout. The sidebarLayout() function creates a sidebar/main panel layout. The sliderInput() function creates interactive sliders. The plotOutput() function reserves space for plots. The server contains reactive logic. The renderPlot() function creates reactive plots. The renderPrint() function creates reactive text output.
Your Learning Roadmap
Learning R is a journey. Here is a recommended path.
Foundations — Start with the basics. Learn R syntax, data types, and how to assign variables. Get comfortable with the RStudio environment. Write and run simple scripts. This foundation supports everything that follows.
Data Structures — Master vectors, lists, matrices, arrays, factors, and data frames. Understanding these structures is essential because R uses them everywhere. Practice creating and manipulating each type.
Control Flow — Learn decision-making with if, else, and switch. Master loops with for, while, and repeat. Understand loop control with break and next. These structures make your programs dynamic and flexible.
Functions — Learn to create your own functions. Understand arguments, default values, and return values. Explore the apply family of functions. Functions make your code reusable and organized.
Data Manipulation — Learn to import data from CSV, Excel, text, and other formats. Master data cleaning to handle missing values, duplicates, and outliers. Use dplyr for intuitive data manipulation. Use tidyr for reshaping data.
Data Visualization — Create plots with base R graphics. Master ggplot2 for professional visualizations. Learn to customize colors, labels, titles, and themes. Explore interactive visualizations with plotly.
Statistics — Understand descriptive statistics. Learn probability distributions. Master hypothesis testing including t-tests, chi-square, and ANOVA. Build regression models. Conduct time series analysis.
Machine Learning — Explore supervised learning including linear regression, logistic regression, decision trees, random forest, and SVM. Study unsupervised learning including clustering and PCA. Learn model evaluation with caret.
Advanced Topics — Learn object-oriented programming in R. Master error handling and functional programming. Explore parallel computing and memory management. Understand package development.
Real-World Projects — Apply your skills to real problems. Build dashboards with Shiny. Create reports with R Markdown. Contribute to open-source packages. Participate in competitions on Kaggle.
Final Thoughts
R is more than just a programming language; it is a powerful environment for data analysis, statistical computing, and visualization. Every dataset tells a story, and R helps you uncover it. Whether you are analyzing business metrics, conducting scientific research, or building machine learning models, R provides the tools you need.
The journey from beginner to advanced takes time. Be patient with yourself. Practice regularly. Each new concept you master opens the door to new possibilities. The R community is large and supportive, and help is always available when you need it.
Remember that even the most experienced R programmers started where you are now. Every line of code you write builds your skills. Every dataset you explore increases your understanding. Keep going, keep learning, and keep discovering.


