{ "nbformat": 4, "nbformat_minor": 0, "metadata": { "kernelspec": { "name": "ir", "display_name": "R", "language": "R" }, "language_info": { "name": "R" }, "colab": { "provenance": [], "toc_visible": true } }, "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Session 1 · Introduction to R for soil science\n", "**Module 1 · Introduction to R, spatial data and soil data preparation** · SoilFER Training\n", "\n", "[Course page](https://training.yigini.net/modules/01-r-soil-data/) · [Manual chapter](https://training.yigini.net/manual/introduction-to-soil-data-preparation-spatial-data.html) · [Original script](https://github.com/SoilFER/SoilFER-Training-Resources/blob/main/02_scripts/module1/Session1_Introduction_to_R.R)\n", "\n", "**Before you start**\n", "1. Check that the runtime is **R**: *Runtime → Change runtime type → R* (this notebook should open in R automatically).\n", "2. Run the **Setup** cell below once per session (≈1–3 minutes). It downloads the training project, installs the R packages and downloads the course rasters and MIR data (≈1.2 GB) and sets the working folder.\n", "3. Then run the cells in order with **Shift + Enter**.\n", "\n", "> Colab resets when you close it or after ~90 minutes without activity. Save results you want to keep with *Files → Download* (left sidebar), or re-run the Setup cell after a reset.\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "## Setup" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# ==== SoilFER · Colab setup (run first, once per session) ==================\n", "options(repos = c(CRAN = \"https://cloud.r-project.org\"), timeout = 3600)\n", "# Works in Google Colab and in any other Jupyter (JupyterHub, JupyterLab on your computer)\n", "on_colab <- nzchar(Sys.getenv(\"COLAB_RELEASE_TAG\")) || dir.exists(\"/content/sample_data\")\n", "root <- if (on_colab) \"/content/SoilFER-Training-Resources\" else path.expand(\"~/SoilFER-Training-Resources\")\n", "Sys.setenv(SOILFER_ROOT = root)\n", "\n", "# 1. Training project (scripts, small data, outputs, assignments)\n", "if (!dir.exists(root))\n", " system(paste(\"git clone --depth 1 https://github.com/SoilFER/SoilFER-Training-Resources\", root))\n", "\n", "# 2. R packages (in Colab the R runtime installs ready-made binaries, so this is fast)\n", "pkgs <- c(\"aqp\", \"dplyr\", \"remotes\", \"sf\", \"source\", \"terra\", \"tidyr\", \"tidyverse\", \"win.binary\")\n", "need <- setdiff(pkgs, rownames(installed.packages()))\n", "if (length(need)) install.packages(need)\n", "\n", "# 3. Course rasters + MIR spectra (Google Drive folder of the SoilFER training, ≈1.2 GB)\n", "td <- file.path(root, \"01_data/module1/training_data\")\n", "if (!file.exists(file.path(td, \"MIR_KANSAS_data.xlsx\"))) {\n", " system(\"python3 -m pip -q install gdown\")\n", " drv <- file.path(dirname(root), \"soilfer_drive\")\n", " system(paste(\"python3 -m gdown --folder --quiet https://drive.google.com/drive/folders/1K7tq9zX5HsqbqWcNoT27WtfPtehcKBCu -O\", shQuote(drv)))\n", " f <- list.files(drv, recursive = TRUE, full.names = TRUE)\n", " file.copy(f, td, overwrite = FALSE)\n", "}\n", "\n", "setwd(root)\n", "cat(\"Ready. Working folder:\", getwd(), \"\\n\")\n", "missing <- setdiff(pkgs, rownames(installed.packages()))\n", "if (length(missing)) message(\"Not installed: \", paste(missing, collapse = \", \"))\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### SoilFER Online Training Programme — Module 1" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "###############################################################################\n", "# SoilFER Online Training Programme — Module 1\n", "# SESSION 1: Introduction to R for Soil Science (1.5 hours)\n", "###############################################################################\n", "#\n", "# LEARNING OBJECTIVES\n", "# -------------------\n", "# By the end of this session, participants will be able to:\n", "# 1. Navigate the RStudio interface\n", "# 2. Install and load R packages\n", "# 3. Create and work with R objects and data types\n", "# 4. Use vectors, factors, matrices, data frames, and lists\n", "# 5. Apply arithmetic, comparison, and logical operators\n", "# 6. Write conditional statements and loops\n", "# 7. Create and call custom functions\n", "# 8. Manipulate data using base R and the tidyverse\n", "# 9. Handle missing values\n", "# 10. Reshape tabular data between wide and long formats\n", "#\n", "# DATASET\n", "# -------\n", "# All examples in this session use small toy datasets constructed inline.\n", "# The KSSL Kansas dataset is introduced in Session 2.\n", "#\n", "# TIMING GUIDE (approximate)\n", "# ---------------------------\n", "# 0:00 – 0:15 R Packages: install, load, find help\n", "# 0:15 – 0:35 Objects, data types, and data structures\n", "# 0:35 – 0:55 Operators and control structures\n", "# 0:55 – 1:15 Functions and data manipulation\n", "# 1:15 – 1:30 Missing values, reshaping, and BDL data\n", "###############################################################################\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 1 — PACKAGES" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 1 — PACKAGES\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 1.1 Installing Packages\n", "# -----------------------------------------------------------------------------\n", "# Packages need to be installed only ONCE. Use install.packages():\n", "\n", "# Install a single package\n", "install.packages(\"tidyverse\")\n", "\n", "# Install multiple packages at once\n", "install.packages(c(\"terra\", \"sf\", \"aqp\"))\n", "\n", "# Check which packages are already installed\n", "installed.packages()\n", "\n", "# -----------------------------------------------------------------------------\n", "# 1.2 Installing from GitHub or Other Sources\n", "# -----------------------------------------------------------------------------\n", "\n", "# First, install remotes package (if not already installed)\n", "install.packages(\"remotes\")\n", "\n", "# Install an R package from GitHub\n", "remotes::install_github(\"rspatial/terra\")\n", "\n", "# -----------------------------------------------------------------------------\n", "# 1.3 Manual Installation (offline environments)\n", "# -----------------------------------------------------------------------------\n", "\n", "# Installing from a compressed source package file (e.g., mypackage_1.0.0.tar.gz)\n", "install.packages(\"path/to/mypackage_1.0.0.tar.gz\", repos = NULL, type = \"source\")\n", "\n", "# Installing from a Local .zip file (Windows Binary)\n", "install.packages(\"path/to/mypackage.zip\", repos = NULL, type = \"win.binary\")\n", "\n", "# -----------------------------------------------------------------------------\n", "# 1.4 Loading Packages\n", "# -----------------------------------------------------------------------------\n", "# After installation, LOAD a package each time you start a new R session:\n", "\n", "# Load tidyverse package\n", "library(tidyverse)\n", "\n", "# Load multiple packages\n", "library(terra)\n", "library(sf)\n", "library(aqp)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 1.5 Finding Help on Packages\n", "# -----------------------------------------------------------------------------\n", "\n", "# Get help on a package\n", "help(package = \"tidyverse\")\n", "\n", "# Or use\n", "?tidyverse\n", "\n", "# View vignettes (tutorials) for a package\n", "vignette(package = \"ggplot2\")\n", "vignette(\"ggplot2\", package = \"ggplot2\")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 2 — OBJECTS AND DATA TYPES" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 2 — OBJECTS AND DATA TYPES\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 2.1 Creating Objects with Assignment\n", "# -----------------------------------------------------------------------------\n", "\n", "# Assign a value to a variable\n", "soil_depth <- 30\n", "soil_depth\n", "\n", "# Assignment `<-` to objects and `=` to function arguments\n", "soil_depth <- c(10,20,30)\n", "mean_soil_depth <- mean(x = soil_depth, na.rm = TRUE)\n", "mean_soil_depth\n", "\n", "# Good naming practices — descriptive, consistent\n", "# - Object names must start with a letter\n", "# - Names can contain letters, numbers, underscores `_`, and periods `.`\n", "# - Names are case-sensitive: `SoilDepth` is different from `soildepth`\n", "# - Avoid using reserved words like `TRUE`, `FALSE`, `NA`, `function`, etc.\n", "\n", "soil_ph <- 6.5\n", "organic_carbon_percent <- 2.1\n", "clay_content_gkg <- 350\n", "\n", "# Use consistent style\n", "plot_id <- \"P001\" # snake_case (recommended)\n", "plotID <- \"P001\" # camelCase (alternative)\n", "\n", "# Bad naming practices — too short, too long, unclear\n", "# Avoid — too short, unclear\n", "x <- 6.5\n", "a <- 2.1\n", "\n", "# Avoid — too long\n", "the_ph_value_of_the_topsoil_at_site_one <- 6.5\n", "\n", "# -----------------------------------------------------------------------------\n", "# 2.2 Data Types in R \n", "# -----------------------------------------------------------------------------\n", "\n", "# Numeric values\n", "ph_value <- 6.8\n", "clay_percent <- 25.5\n", "temperature <- 15.2\n", "\n", "# Integer values (use L suffix)\n", "sample_count <- 100L\n", "plot_number <- 5L\n", "\n", "# Character values (use quotes)\n", "soil_type <- \"Acrisol\"\n", "location <- \"Kansas\"\n", "notes <- \"Sample collected from topsoil\"\n", "\n", "# Logical values\n", "is_valid <- TRUE\n", "has_missing_data <- FALSE\n", "\n", "# Dates: The ISO 8601 date format, YYYY-MM-DD\n", "sampling_date <- \"2024-03-12\"\n", "# Convert to date (%Y = 4-digit year; %y = 2-digit year)\n", "sampling_date <- as.Date(sampling_date, format = \"%Y-%m-%d\") \n", "class(sampling_date)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 2.3 Checking Data Types\n", "# -----------------------------------------------------------------------------\n", "\n", "# Check the class (high-level type)\n", "class(ph_value) # \"numeric\"\n", "class(soil_type) # \"character\"\n", "class(is_valid) # \"logical\"\n", "\n", "# Check the internal storage type\n", "typeof(sample_count)\n", "typeof(ph_value)\n", "\n", "# Helpful checks\n", "is.numeric(ph_value)\n", "is.character(soil_type)\n", "is.logical(is_valid)\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 3 — DATA STRUCTURES" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 3 — DATA STRUCTURES\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 3.1 Vectors\n", "# -----------------------------------------------------------------------------\n", "\n", "#### Creating Vectors ----\n", "\n", "# Create a numeric vector using c() (combine)\n", "ph_values <- c(5.2, 6.5, 7.1, 5.8, 6.9)\n", "print(ph_values)\n", "ph_values\n", "\n", "# Character vector\n", "soil_types <- c(\"Acrisol\", \"Ferralsol\", \"Vertisol\", \"Andosol\", \"Cambisol\")\n", "soil_types\n", "\n", "# Logical vector\n", "valid_samples <- c(TRUE, TRUE, FALSE, TRUE, TRUE)\n", "valid_samples\n", "\n", "# Create sequences\n", "depths <- 0:100 # Integers from 0 to 100\n", "depths_seq <- seq(0, 100, by=10) # 0, 10, 20, ..., 100\n", "\n", "#### Vector Operations ----\n", "\n", "# Arithmetic on vectors (element-wise)\n", "ph_values * 10\n", "ph_values + 1\n", "\n", "# Summary statistics\n", "mean(ph_values)\n", "median(ph_values)\n", "sd(ph_values) # standard deviation\n", "min(ph_values)\n", "max(ph_values)\n", "# NAs affect the statistic results\n", "mean(c(ph_values,NA))\n", "mean(c(ph_values,NA), na.rm=TRUE) # Remove NAs from calculations\n", "\n", "#### Accessing Vector Elements ----\n", "\n", "# Access by index (position)\n", "ph_values # Complete data\n", "ph_values[1] # First element\n", "ph_values[3] # Third element\n", "\n", "# Access multiple elements\n", "ph_values[c(1,3,5)] # Elements 1, 3, and 5\n", "\n", "# Access by logical condition\n", "ph_values[ph_values > 6] # All pH values greater than 6\n", "\n", "# Negative indices exclude elements\n", "ph_values[-1] # All except first\n", "ph_values[-c(1,2)] # All except first two\n", "\n", "# -----------------------------------------------------------------------------\n", "# 3.2 Factors\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create a factor from character vector\n", "soil_class <- factor(c(\"Clay\", \"Loam\", \"Sand\", \"Clay\", \"Loam\"))\n", "soil_class\n", "\n", "# Check levels\n", "levels(soil_class)\n", "\n", "# Count observations per level\n", "table(soil_class)\n", "\n", "# Ordered factors (when order matters)\n", "texture_class <- factor(\n", " c(\"Coarse\", \"Fine\", \"Medium\", \"Fine\", \"Coarse\"),\n", " levels = c(\"Coarse\", \"Medium\", \"Fine\"),\n", " ordered = TRUE\n", ")\n", "texture_class\n", "\n", "# -----------------------------------------------------------------------------\n", "# 3.3 Matrices\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create a matrix\n", "soil_matrix <- matrix(\n", " c(5.2, 25, 30,\n", " 6.5, 30, 28,\n", " 7.1, 18, 35),\n", " nrow = 3,\n", " ncol = 3,\n", " byrow = TRUE\n", ")\n", "soil_matrix\n", "\n", "# Add column names\n", "colnames(soil_matrix) <- c(\"pH\", \"Clay\", \"Sand\")\n", "rownames(soil_matrix) <- c(\"Sample1\", \"Sample2\", \"Sample3\")\n", "soil_matrix\n", "\n", "# Access elements\n", "soil_matrix[1, 2] # Row 1, Column 2\n", "soil_matrix[1, ] # All of row 1\n", "soil_matrix[, 2] # All of column 2\n", "\n", "# -----------------------------------------------------------------------------\n", "# 3.4 Data Frames\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create a data frame\n", "soil_data <- data.frame(\n", " plot_id = c(\"P001\", \"P002\", \"P003\", \"P004\", \"P005\"),\n", " latitude = c(-1.25, -1.27, -1.23, -1.29, -1.26),\n", " longitude = c(36.85, 36.83, 36.87, 36.81, 36.84),\n", " ph = c(5.2, 6.5, 7.1, 5.8, 6.9),\n", " organic_carbon = c(2.1, 3.2, 1.8, 2.7, 2.9),\n", " clay_content = c(25, 30, 18, 42, 35),\n", " soil_type = c(\"Acrisol\", \"Ferralsol\", \"Vertisol\", \"Andosol\", \"Cambisol\")\n", ")\n", "\n", "# View the data frame\n", "soil_data\n", "\n", "# View structure\n", "str(soil_data)\n", "\n", "# View first rows\n", "head(soil_data)\n", "\n", "# View last rows\n", "tail(soil_data)\n", "\n", "summary(soil_data)\n", "\n", "\n", "# Get dimensions\n", "dim(soil_data) # rows, columns\n", "nrow(soil_data) # number of rows\n", "ncol(soil_data) # number of columns\n", "\n", "# Access columns and rows\n", "soil_data$ph\n", "soil_data[[\"ph\"]]\n", "\n", "# Access rows\n", "soil_data[1, ] # First row\n", "soil_data[c(1,3,5), ] # Rows 1, 3, and 5\n", "\n", "# Access specific cells\n", "soil_data[2, 4] # Row 2, Column 4 (pH of second plot)\n", "\n", "# Subset based on conditions\n", "soil_data[soil_data$ph > 6, ] # Plots with pH > 6\n", "soil_data[soil_data$soil_type == \"Acrisol\", ] # Only Acrisols\n", "\n", "# Add a new column\n", "soil_data$silt_content <- c(40, 35, 52, 23, 30)\n", "\n", "# Calculate new columns from existing ones\n", "soil_data$clay_plus_silt <- soil_data$clay_content + soil_data$silt_content\n", "\n", "# View updated data frame\n", "head(soil_data)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 3.5 Lists\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create a list\n", "soil_analysis <- list(\n", " site_name = \"Kansas Field\",\n", " coordinates = c(lat = -1.25, lon = 36.85),\n", " measurements = data.frame(\n", " depth = c(0, 10, 20, 30),\n", " ph = c(6.5, 6.2, 5.8, 5.5)\n", " ),\n", " notes = \"Collected during dry season\"\n", ")\n", "\n", "# View list structure\n", "str(soil_analysis)\n", "\n", "# Access list elements\n", "soil_analysis$site_name\n", "soil_analysis[[1]] # First element\n", "soil_analysis[[\"measurements\"]] # View measurements data frame\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 4 — OPERATORS" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 4 — OPERATORS\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 4.1 Arithmetic Operators\n", "# -----------------------------------------------------------------------------\n", "\n", "# Basic arithmetic\n", "10 + 5 # Addition\n", "10 - 5 # Subtraction\n", "10 * 5 # Multiplication\n", "10 / 5 # Division\n", "10 ^ 2 # Exponentiation (10 squared)\n", "10 %% 3 # Modulus (remainder: 10 mod 3 = 1)\n", "10 %/% 3 # Integer division (10 divided by 3 = 3)\n", "\n", "# Order of operations (PEMDAS)\n", "result <- (10 + 5) * 2 / 4 - 1\n", "result\n", "\n", "# -----------------------------------------------------------------------------\n", "# 4.2 Comparison Operators\n", "# -----------------------------------------------------------------------------\n", "\n", "# Comparison operators return TRUE or FALSE\n", "5 == 5 # Equal to\n", "5 != 3 # Not equal to\n", "5 > 3 # Greater than\n", "5 < 3 # Less than\n", "5 >= 5 # Greater than or equal to\n", "5 <= 6 # Less than or equal to\n", "\n", "# Use in subsetting\n", "ph_values <- c(5.2, 6.5, 7.1, 5.8, 6.9)\n", "ph_values > 6 # Logical vector\n", "ph_values[ph_values > 6] # Values greater than 6\n", "\n", "# -----------------------------------------------------------------------------\n", "# 4.3 Logical Operators\n", "# -----------------------------------------------------------------------------\n", "\n", "# AND operator: & (element-wise) or && (single values)\n", "TRUE & TRUE # TRUE\n", "TRUE & FALSE # FALSE\n", "\n", "# OR operator: | (element-wise) or || (single values)\n", "TRUE | FALSE # TRUE\n", "FALSE | FALSE # FALSE\n", "\n", "# NOT operator: !\n", "!TRUE # FALSE\n", "!FALSE # TRUE\n", "\n", "# Combining conditions\n", "ph_values <- c(5.2, 6.5, 7.1, 5.8, 6.9)\n", "clay_content <- c(25, 30, 18, 42, 35)\n", "\n", "# Find samples with pH > 6 AND clay > 25\n", "ph_values > 6 & clay_content > 25\n", "\n", "# Find samples with pH > 6 OR clay > 40\n", "ph_values > 6 | clay_content > 40\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 5 — CONTROL STRUCTURES" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 5 — CONTROL STRUCTURES\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 5.1 If-Else Statements\n", "# -----------------------------------------------------------------------------\n", "\n", "# Basic if statement\n", "ph_value <- 7.5\n", "\n", "# Basic if statement\n", "if (ph_value > 7) {\n", " print(\"Alkaline soil\")\n", "}\n", "\n", "# If-else\n", "if (ph_value > 7) {\n", " print(\"Alkaline soil\")\n", "} else {\n", " print(\"Neutral or acidic soil\")\n", "}\n", "\n", "# Multiple conditions (only the first TRUE branch runs)\n", "if (ph_value > 7.5) {\n", " print(\"Strongly alkaline\")\n", "} else if (ph_value > 7) {\n", " print(\"Slightly alkaline\")\n", "} else if (ph_value == 7) {\n", " print(\"Neutral\")\n", "} else {\n", " print(\"Acidic\")\n", "}\n", "\n", "# -----------------------------------------------------------------------------\n", "# 5.2 Vectorized If-Else: ifelse()\n", "# -----------------------------------------------------------------------------\n", "\n", "# Vectorized conditional assignment\n", "ph_values <- c(5.2, 6.5, 7.1, 5.8, 7.0)\n", "\n", "# Simple two-class example\n", "soil_reaction <- ifelse(ph_values > 7, \"Alkaline\", \"Not alkaline\")\n", "soil_reaction\n", "\n", "# -----------------------------------------------------------------------------\n", "# 5.3 Vectorized Cut: cut()\n", "# -----------------------------------------------------------------------------\n", "\n", "soil_class <- cut(\n", " ph_values,\n", " breaks = c(-Inf, 5.5, 7.0, Inf),\n", " labels = c(\"Acidic\", \"Neutral\", \"Alkaline\"),\n", " right = TRUE, include.lowest = TRUE\n", ")\n", "soil_class\n", "\n", "# Notes:\n", "# - `right = TRUE` means intervals are **right-closed** (e.g., `(5.5, 7.0]`).\n", "# - `include.lowest = TRUE` ensures the smallest value is included in the first interval.\n", "\n", "# -----------------------------------------------------------------------------\n", "# 5.4 For Loops\n", "# -----------------------------------------------------------------------------\n", "\n", "# Loop through a sequence\n", "for (i in 1:5) {\n", " print(paste(\"Iteration:\", i))\n", "}\n", "\n", "# Loop through a vector\n", "soil_types <- c(\"Acrisol\", \"Ferralsol\", \"Vertisol\")\n", "\n", "for (soil in soil_types) {\n", " print(paste(\"Soil type:\", soil))\n", "}\n", "\n", "# Loop with conditional logic\n", "ph_values <- c(5.2, 6.5, 7.1, 5.8, 6.9)\n", "\n", "for (i in 1:length(ph_values)) {\n", " if (ph_values[i] > 6) {\n", " print(paste(\"Sample\", i, \"has pH =\", ph_values[i], \"(Acceptable)\"))\n", " } else {\n", " print(paste(\"Sample\", i, \"has pH =\", ph_values[i], \"(Too acidic)\"))\n", " }\n", "}\n", "\n", "# -----------------------------------------------------------------------------\n", "# 5.5 While Loops\n", "# -----------------------------------------------------------------------------\n", "\n", "# While loop continues until condition is FALSE\n", "counter <- 1\n", "\n", "while (counter <= 5) {\n", " print(paste(\"Counter value:\", counter))\n", " counter <- counter + 1 # Increment counter\n", "}\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 6 — FUNCTIONS" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 6 — FUNCTIONS\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 6.1 Using Built-in Functions\n", "# -----------------------------------------------------------------------------\n", "\n", "# Statistical functions\n", "mean(c(5, 10, 15, 20))\n", "median(c(5, 10, 15, 20))\n", "sd(c(5, 10, 15, 20))\n", "sum(c(5, 10, 15, 20))\n", "\n", "# String functions\n", "toupper(\"acrisol\")\n", "tolower(\"FERRALSOL\")\n", "nchar(\"soil science\") # Count characters\n", "\n", "# Math functions\n", "sqrt(16)\n", "log(10)\n", "exp(2)\n", "abs(-5)\n", "round(pi, 2)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 6.2 Creating Custom Functions\n", "# -----------------------------------------------------------------------------\n", "\n", "# Example: Bulk density (mass / volume)\n", "calculate_bulk_density <- function(mass, volume) {\n", " if (any(volume <= 0)) stop(\"volume must be > 0\")\n", " mass / volume\n", "}\n", "\n", "calculate_bulk_density(mass = 150, volume = 100)\n", "\n", "# Function with a default argument\n", "classify_soil_ph <- function(ph, threshold = 7) {\n", " if (ph > threshold) {\n", " \"Alkaline\"\n", " } else if (ph == threshold) {\n", " \"Neutral\"\n", " } else {\n", " \"Acidic\"\n", " }\n", "}\n", "\n", "classify_soil_ph(6.5)\n", "classify_soil_ph(6.5, threshold = 6)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 6.3 Function Arguments and Defaults\n", "# -----------------------------------------------------------------------------\n", "\n", "# SOC stock calculation (example)\n", "# Assumptions:\n", "# - soc_percent is in %\n", "# - bulk_density is in g/cm^3\n", "# - depth is in cm\n", "# Output: SOC stock in Mg/ha\n", "\n", "calculate_soc_stock <- function(soc_percent, bulk_density, depth, coarse_fragment = 0) {\n", " if (any(coarse_fragment < 0 | coarse_fragment > 100)) stop(\"coarse_fragment must be between 0 and 100\")\n", " soc_percent * bulk_density * depth * (1 - coarse_fragment / 100)\n", "}\n", "\n", "calculate_soc_stock(soc_percent = 2.5, bulk_density = 1.3, depth = 30)\n", "calculate_soc_stock(soc_percent = 2.5, bulk_density = 1.3, depth = 30, coarse_fragment = 15)\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 7 — DATA MANIPULATION WITH BASE R" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 7 — DATA MANIPULATION WITH BASE R\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 7.1 Subsetting and Filtering\n", "# -----------------------------------------------------------------------------\n", "\n", "# Example data frame\n", "soil_data <- data.frame(\n", " plot_id = c(\"P001\", \"P002\", \"P003\", \"P004\", \"P005\"),\n", " latitude = c(-1.25, -1.27, -1.23, -1.29, -1.26),\n", " longitude = c(36.85, 36.83, 36.87, 36.81, 36.84),\n", " ph = c(5.2, 6.5, 7.1, 5.8, 6.9),\n", " clay_content = c(25, 30, 18, 42, 35),\n", " soil_type = c(\"Acrisol\", \"Ferralsol\", \"Vertisol\", \"Andosol\", \"Cambisol\")\n", ")\n", "\n", "# Select column by index: Show second column\n", "soil_data[, 2]\n", "# Select column by index: Show first and third columns\n", "soil_data[, c(1,3)]\n", "# Select column by name: Show texture columns\n", "soil_data[, c(\"ph\",\"clay_content\",\"soil_type\")]\n", "# Filter rows by index: keep first 2 records in the data frame\n", "soil_data[c(1:2),]\n", "\n", "# Filter rows by index: keep texture properties for the first 100 records in the data frame\n", "soil_data[c(1:2), c(\"ph\",\"clay_content\",\"soil_type\")]\n", "\n", "# Filter rows based on conditions\n", "high_ph <- soil_data[soil_data$ph > 6, ]\n", "high_ph\n", "\n", "# Multiple conditions with & and |\n", "high_ph_clay <- soil_data[soil_data$ph > 6 & soil_data$clay_content > 25, ]\n", "high_ph_clay\n", "\n", "# Select columns by index: Show second column\n", "subset_data <- soil_data[, 2]\n", "# Select several columns by index\n", "subset_data <- soil_data[, c(1:3)]\n", "\n", "# Select by name\n", "subset_data <- soil_data[, c(\"plot_id\", \"ph\", \"soil_type\")]\n", "subset_data\n", "\n", "# -----------------------------------------------------------------------------\n", "# 7.2 Deleting Columns and Rows\n", "# -----------------------------------------------------------------------------\n", "\n", "# Example data frame\n", "data <- data.frame(\n", " Plot = c(\"P001\", \"P002\", \"P003\"),\n", " Clay = c(25, 30, 18),\n", " Silt = c(40, 35, 52),\n", " pH = c(5.2, 6.5, 7.1)\n", ")\n", "\n", "# --- Delete columns ---\n", "\n", "# Delete one column by name\n", "data$Clay <- NULL\n", "data$Clay <- c() # Empty vector\n", "data\n", "\n", "# Delete multiple columns by name\n", "data[, c(\"Silt\", \"pH\")] <- NULL\n", "data\n", "\n", "# --- Delete rows ---\n", "\n", "# Delete a row by index (e.g., remove the 2nd row)\n", "data <- data[-2, ]\n", "data\n", "\n", "# Delete rows based on a condition (e.g., remove rows with pH < 6)\n", "# Rebuild the dataframe\n", "data <- data.frame(\n", " Plot = c(\"P001\", \"P002\", \"P003\"),\n", " Clay = c(25, 30, 18),\n", " Silt = c(40, 35, 52),\n", " pH = c(5.2, 6.5, 7.1)\n", ")\n", "data <- data[data$pH >= 6, ]\n", "data\n", "\n", "# -----------------------------------------------------------------------------\n", "# 7.3 Sorting Data\n", "# -----------------------------------------------------------------------------\n", "\n", "# Sort by pH (ascending)\n", "soil_data[order(soil_data$ph), ]\n", "\n", "# Sort by pH (descending)\n", "soil_data[order(-soil_data$ph), ]\n", "\n", "# Sort by multiple columns\n", "soil_data[order(soil_data$soil_type, soil_data$ph), ]\n", "\n", "# -----------------------------------------------------------------------------\n", "# 7.4 Aggregating Data\n", "# -----------------------------------------------------------------------------\n", "\n", "# Example data frame with repeated soil_type groups\n", "soil_data <- data.frame(\n", " plot_id = paste0(\"P\", sprintf(\"%03d\", 1:12)),\n", " latitude = c(-1.25, -1.26, -1.27, -1.28, -1.23, -1.24, -1.22, -1.21, -1.29, -1.30, -1.31, -1.32),\n", " longitude = c(36.85, 36.86, 36.83, 36.84, 36.87, 36.88, 36.82, 36.81, 36.84, 36.83, 36.86, 36.85),\n", " ph = c(5.2, 5.6, 6.5, 6.2, 7.1, 6.9, 5.8, 6.0, 6.9, 6.7, 5.4, 5.1),\n", " clay_content = c(25, 28, 30, 33, 18, 20, 42, 40, 35, 34, 27, 26),\n", " soil_type = c(\n", " \"Acrisol\",\"Acrisol\",\n", " \"Ferralsol\",\"Ferralsol\",\n", " \"Vertisol\",\"Vertisol\",\n", " \"Andosol\",\"Andosol\",\n", " \"Cambisol\",\"Cambisol\",\n", " \"Acrisol\",\"Acrisol\"\n", " )\n", ")\n", "\n", "# Multiple summary values by group (mean and standard deviation of pH by soil_type)\n", "aggregate(\n", " ph ~ soil_type,\n", " data = soil_data,\n", " FUN = function(x) c(mean = mean(x), sd = sd(x))\n", ")\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 8 — DATA MANIPULATION WITH THE TIDYVERSE" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 8 — DATA MANIPULATION WITH THE TIDYVERSE\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.1 Loading the Tidyverse\n", "# -----------------------------------------------------------------------------\n", "\n", "# Install if not already installed\n", "# install.packages(\"tidyverse\")\n", "\n", "# Load tidyverse\n", "library(tidyverse)\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.2 Tibbles: Modern Data Frames\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create tibble directly\n", "soil_tbl <- tibble(\n", " plot_id = c(\"P001\", \"P002\", \"P003\"),\n", " ph = c(5.2, 6.5, 7.1),\n", " clay = c(25, 30, 18)\n", ")\n", "soil_tbl\n", "\n", "soil_data <- data.frame(\n", " plot_id = c(\"P001\", \"P002\", \"P003\"),\n", " ph = c(5.2, 6.5, 7.1),\n", " clay = c(25, 30, 18)\n", ")\n", "\n", "# Convert data frame to tibble\n", "soil_tbl <- as_tibble(soil_data)\n", "soil_tbl\n", "\n", "\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.3 The Pipe Operator: %>%\n", "# -----------------------------------------------------------------------------\n", "soil_data <- tibble(\n", " plot_id = c(\"P001\", \"P002\", \"P003\"),\n", " ph = c(5.2, 6.5, 7.1),\n", " clay = c(25, 30, 18)\n", ")\n", "\n", "# Without pipe (nested functions)\n", "round(mean(soil_data$ph), 2)\n", "\n", "# With pipe (sequential operations)\n", "soil_data$ph %>%\n", " mean() %>%\n", " round(2)\n", "\n", "# More complex example\n", "soil_data %>%\n", " filter(ph > 6) %>%\n", " select(plot_id, ph, clay) %>%\n", " arrange(desc(ph))\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.4 Data Manipulation with {dplyr}\n", "# -----------------------------------------------------------------------------\n", "\n", "# Select specific columns\n", "soil_data %>%\n", " select(plot_id, ph, clay)\n", "\n", "# Select range of columns\n", "soil_data %>%\n", " select(plot_id:ph)\n", "\n", "# Remove columns\n", "soil_data %>%\n", " select(-plot_id, -ph)\n", "\n", "# Select columns matching pattern\n", "soil_data %>%\n", " select(contains(\"p\"))\n", "\n", "# Filter rows based on condition\n", "soil_data %>%\n", " filter(ph > 6)\n", "\n", "# Multiple conditions\n", "soil_data %>%\n", " filter(ph > 6 & clay > 25)\n", "\n", "# Filter with OR\n", "soil_data %>%\n", " filter(plot_id == \"P001\" | plot_id == \"P002\")\n", "\n", "# Use %in% for multiple values\n", "soil_data %>%\n", " filter(plot_id %in% c(\"P001\", \"P002\"))\n", "\n", "# Rename columns for clarity\n", "rename(soil_data,\n", " site_id = plot_id)\n", "\n", "# You can rename multiple columns at once\n", "soil_data %>%\n", " rename(\n", " location = plot_id,\n", " acidity = ph,\n", " clay_percent = clay\n", " ) \n", "\n", "# Create new columns\n", "soil_data %>%\n", " mutate(\n", " ph_class = ifelse(ph > 7, \"Alkaline\", \"Acidic\"),\n", " clay_gr_kg = clay * 10 \n", " )\n", "\n", "# Modify existing columns\n", "soil_data %>%\n", " mutate(\n", " ph = round(ph, 1),\n", " plot_id = tolower(plot_id)\n", " )\n", "\n", "# Sort ascending\n", "soil_data %>%\n", " arrange(ph)\n", "\n", "# Sort descending\n", "soil_data %>%\n", " arrange(desc(ph))\n", "\n", "# Multiple sort keys\n", "soil_data %>%\n", " arrange(ph, desc(clay))\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.5 Grouping and Summarizing Data\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create example dataset\n", "soil_data <- data.frame(\n", "site = c(\"Forest_A\", \"Forest_B\", \"Grassland_A\", \"Grassland_B\", \"Urban_A\", \"Urban_B\"),\n", "ecosystem = c(\"Forest\", \"Forest\", \"Grassland\", \"Grassland\", \"Urban\", \"Urban\"),\n", "pH = c(6.2, 6.8, 7.1, 6.9, 5.8, 6.0),\n", "organic_carbon = c(3.2, 2.8, 2.1, 2.4, 1.5, 1.8)\n", ")\n", "\n", "# Summarize soil properties\n", "soil_data %>%\n", " summarize(\n", " mean_ph = mean(pH),\n", " sd_ph = sd(pH),\n", " min_soc = min(organic_carbon),\n", " max_soc = max(organic_carbon),\n", " n_samples = n()\n", " )\n", "\n", "# Group by and summarize\n", "soil_data %>%\n", " group_by(ecosystem) %>%\n", " summarize(\n", " mean_ph = mean(pH),\n", " mean_soc = mean(organic_carbon), # Average organic carbon\n", " count = n(), \n", " .groups = \"drop\" # Remove grouping\n", ")\n", "\n", "# Alternative: count()\n", "soil_data <- data.frame(\n", "site = c(\"Forest_A\", \"Forest_B\", \"Grassland_A\", \"Grassland_B\", \"Urban_A\", \"Urban_B\"),\n", "ecosystem = c(\"Forest\", \"Forest\", \"Grassland\", \"Grassland\", \"Urban\", \"Urban\"),\n", "pH = c(6.2, 6.8, 7.1, 6.9, 5.8, 6.0),\n", "organic_carbon = c(3.2, 2.8, 2.1, 2.4, 1.5, 1.8)\n", ")\n", "\n", "# Count observations by group\n", "count(soil_data, ecosystem)\n", "\n", "# Count by multiple groups\n", "count(soil_data, ecosystem, site)\n", "\n", "# Count with weights of a column (sum of carbon instead of count)\n", "soil_data %>%\n", " count(ecosystem, wt = organic_carbon, name = \"total SOC\")\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.6 Combining Data Frames (Joins)\n", "# -----------------------------------------------------------------------------\n", "\n", "# Soil data example\n", "soil_basic <- data.frame(\n", " site_id = c(\"A\", \"B\", \"C\", \"D\"),\n", " pH = c(6.2, 6.8, 7.1, 6.9),\n", " organic_carbon = c(3.2, 2.8, 2.1, 2.4)\n", ")\n", "# Additional measurements (note: includes data on site E, but not on site A)\n", "nutrients <- data.frame(\n", " site_id = c(\"B\", \"C\", \"D\", \"E\"),\n", " phosphorus = c(0.15, 0.12, 0.18, 0.14),\n", " potassium = c(0.8, 0.9, 0.7, 0.6)\n", ")\n", "# Site information\n", "site_info <- data.frame(\n", " site_id = c(\"A\", \"B\", \"C\", \"D\"),\n", " ecosystem = c(\"Forest\", \"Forest\", \"Grassland\", \"Grassland\"),\n", " elevation = c(450, 520, 380, 420)\n", ")\n", "\n", "# Join examples (same keys, different rules for which rows are kept)\n", "\n", "# 1) `left_join()`: Keep all rows from soil_basic\n", "left_join(soil_basic, nutrients, by = \"site_id\")\n", "\n", "# 2) `right_join()`: Keep all rows from nutrients\n", "right_join(soil_basic, nutrients, by = \"site_id\")\n", "\n", "# 3) `inner_join()`: Keep only rows that exist in both tables\n", "inner_join(soil_basic, nutrients, by = \"site_id\")\n", "\n", "# 4) `full_join()`: Keep all rows from both tables (missing values become NA)\n", "full_join(soil_basic, nutrients, by = \"site_id\")\n", "\n", "## Joining more than two tables sequentially\n", "soil_basic %>%\n", " left_join(nutrients, by = \"site_id\") %>%\n", " left_join(site_info, by = \"site_id\")\n", "\n", "# -----------------------------------------------------------------------------\n", "# 8.7 Stacking Data with bind_rows()\n", "# -----------------------------------------------------------------------------\n", "\n", "# Data from different time periods\n", "spring_data <- data.frame(\n", " site = c(\"A\", \"B\"),\n", " season = \"Spring\",\n", " pH = c(6.1, 6.7),\n", " temperature = c(12.5, 11.8)\n", ")\n", "\n", "summer_data <- data.frame(\n", " site = c(\"A\", \"B\"),\n", " season = \"Summer\", \n", " pH = c(6.3, 6.9),\n", " temperature = c(18.2, 17.5)\n", ")\n", "\n", "# Combine the datasets\n", "bind_rows(spring_data, summer_data)\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 9 — MISSING DATA" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 9 — MISSING DATA\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 9.1 Identifying Missing Data\n", "# -----------------------------------------------------------------------------\n", "\n", "# Create data with missing values\n", "soil_data_na <- data.frame(\n", " plot = c(\"P001\", \"P002\", \"P003\", \"P004\"),\n", " ph = c(5.2, NA, 7.1, 6.5),\n", " clay = c(25, 30, NA, 35)\n", ")\n", "\n", "# Check for missing values\n", "is.na(soil_data_na)\n", "\n", "# Count missing values per column\n", "colSums(is.na(soil_data_na))\n", "\n", "# Identify complete cases (rows with no missing values)\n", "complete.cases(soil_data_na)\n", "\n", "# Extract complete cases\n", "soil_complete <- soil_data_na[complete.cases(soil_data_na), ]\n", "soil_complete\n", "\n", "# -----------------------------------------------------------------------------\n", "# 9.2 Handling Missing Data (Base R)\n", "# -----------------------------------------------------------------------------\n", "\n", "# Remove rows with any missing values (base R)\n", "na.omit(soil_data_na)\n", "\n", "# Remove rows where a specific column is missing (keep rows with non-missing pH)\n", "soil_data_na[!is.na(soil_data_na$ph), ]\n", "\n", "# Replace missing pH values with the mean pH (na.rm = TRUE ignores NA in the mean)\n", "soil_data_na$ph[is.na(soil_data_na$ph)] <- mean(soil_data_na$ph, na.rm = TRUE)\n", "soil_data_na\n", "\n", "# -----------------------------------------------------------------------------\n", "# 9.3 Handling Missing Data (Tidyverse)\n", "# -----------------------------------------------------------------------------\n", "\n", "# Remove all rows with any NA\n", "soil_data_na %>%\n", " drop_na()\n", "\n", "# Remove rows where ph is NA\n", "soil_data_na %>%\n", " drop_na(ph)\n", "\n", "# Replace NA values with specified values\n", "soil_data_na %>%\n", " replace_na(list(ph = 6.0, clay = 30))\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 10 — DATA RESHAPING WITH {tidyr}" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 10 — DATA RESHAPING WITH {tidyr}\n", "# =============================================================================\n", "\n", "# -----------------------------------------------------------------------------\n", "# 10.1 pivot_longer(): Wide to Long\n", "# -----------------------------------------------------------------------------\n", "\n", "library(dplyr) # for piping %>%\n", "library(tidyr) # for pivoting\n", "# Wide format data\n", "wide_soil <- data.frame(\n", " site = c(\"A\", \"B\", \"C\"),\n", " ecosystem = c(\"Forest\", \"Grassland\", \"Urban\"),\n", " pH = c(6.2, 7.1, 5.8),\n", " carbon = c(3.2, 2.1, 1.5),\n", " nitrogen = c(0.25, 0.18, 0.12)\n", ")\n", "wide_soil\n", "\n", "# Convert to long format\n", "long_soil <- wide_soil %>%\n", " pivot_longer(\n", " cols = c(pH, carbon, nitrogen), # Columns to pivot\n", " names_to = \"measurement_type\", # Name for the variable column\n", " values_to = \"value\" # Name for the values column\n", " )\n", "long_soil\n", "\n", "# -----------------------------------------------------------------------------\n", "# 10.2 pivot_wider(): Long to Wide\n", "# -----------------------------------------------------------------------------\n", "\n", "# Convert back to wide format\n", "long_soil %>%\n", " pivot_wider(\n", " names_from = measurement_type, # Column containing variable names\n", " values_from = value # Column containing values\n", " )\n", "\n", "# -----------------------------------------------------------------------------\n", "# 10.3 Advanced Pivoting: Multiple Measurements per Site (Replicates)\n", "# -----------------------------------------------------------------------------\n", "\n", "# More complex examples with multiple measurements per site\n", "field_data <- data.frame(\n", " site = rep(c(\"Forest\", \"Grassland\"), each = 6),\n", " measurement = rep(c(\"pH\", \"carbon\", \"nitrogen\"), 4),\n", " replicate = rep(c(\"R1\", \"R2\"), 6),\n", " value = c(6.2, 6.1, 3.2, 3.0, 0.25, 0.23, 7.1, 7.0, 2.1, 2.3, 0.18, 0.19)\n", ")\n", "field_data\n", "\n", "# Pivot to have measurements as columns\n", "print(\"Pivoted data:\")\n", "field_data %>%\n", " pivot_wider(\n", " names_from = measurement,\n", " values_from = value\n", " )\n" ] }, { "cell_type": "markdown", "metadata": {}, "source": [ "### PART 11 — HANDLING BELOW DETECTION LIMIT (BDL) DATA" ] }, { "cell_type": "code", "metadata": {}, "execution_count": null, "outputs": [], "source": [ "# =============================================================================\n", "# PART 11 — HANDLING BELOW DETECTION LIMIT (BDL) DATA\n", "# =============================================================================\n", "\n", "soil_lab <- data.frame(\n", " plot_id = c(\"P001\", \"P002\", \"P003\", \"P004\"),\n", " no3_mgkg_raw = c(\"0.12\", \"<0.05\", \"0.31\", \"<0.05\"),\n", " stringsAsFactors = FALSE\n", ")\n", "soil_lab\n", "\n", "soil_lab <- soil_lab %>%\n", " mutate(\n", " censored = str_detect(no3_mgkg_raw, \"^\\\\s*<\"),\n", " dl = if_else(censored,\n", " as.numeric(str_remove(no3_mgkg_raw, \"^\\\\s*<\\\\s*\")),\n", " NA_real_),\n", " no3_mgkg = if_else(censored, dl / 2, as.numeric(no3_mgkg_raw))\n", " )\n", "\n", "soil_lab\n", "\n", "\n", "###############################################################################\n", "# END OF SESSION 1\n", "# Next session: Session 2 — Soil Data Preparation with KSSL Dataset (Part 1)\n", "###############################################################################\n" ] } ] }