Presentation and basics

Background and Objective

Healthy and productive soils are fundamental to resilient agrifood systems, sustainable land management, and climate adaptation. Yet, in many regions, soil information remains fragmented, outdated, or inaccessible. The Soil Mapping for Resilient Agrifood Systems (SoilFER) programme responds to this challenge by building comprehensive soil information systems that integrate sampling design, laboratory analysis, soil spectroscopy, digital modeling, and decision-support tools. These systems aim to empower governments, researchers, and farmers with actionable knowledge for crop selection, fertilizer recommendations, and soil health management.

This manual provides a step-by-step guide along the entire soil data value chain. It is designed as both a technical reference and a practical training resource, bridging the gap between raw soil data and its functional use in agricultural and environmental decision-making.

The objectives of this manual are to:

  • Present harmonized approaches for soil sampling design used under the SoilFER programme

  • Introduce best practices in soil data preparation and management, aligned with the Global Soil Information System (GloSIS)

  • Demonstrate methods for digital soil modeling and mapping, covering classical statistics, machine learning, and hybrid inference for both continuous and categorical soil properties.

  • Explain how to generate functional soil information to support evidence-based decision-making.

  • Provide guidance on the integration soil spectroscopy-estimated soil parameters into the digital soil mapping process

  • Facilitate data sharing and dissemination, promoting open standards, metadata documentation, and web-based services.

By integrating these components, the manual equips users to move from raw samples to reliable soil information products that inform policy, guide sustainable soil management, and strengthen food and nutrition security.

The SoilFER project is co-funded by the U.S. Department of State and the Ministry of Foreign Affairs of Japan.

The Training Dataset

This manual uses the KSSL (Kellogg Soil Survey Laboratory) soil dataset from Kansas, USA as a comprehensive example for implementing the complete Digital Soil Mapping (DSM) workflow. The dataset provides a realistic foundation for learning DSM techniques, from inspection and cleaning of raw soil laboratory data to advanced spatial modeling and uncertainty assessment. It also demonstrates the integration of soil spectroscopy data and its application in operational DSM production.

The KSSL Kansas dataset represents a collection of laboratory-analyzed soil samples with corresponding mid-infrared spectroscopy measurements. It contains 10,352 soil horizon measurements from 2,584 distinct soil profiles distributed across Kansas, spanning from 37.02°N to 39.99°N latitude and -102.04°W to -94.71°W longitude. The dataset exemplifies common data quality issues encountered in soil databases, providing hands-on experience with:

  • Coordinate validation: Geographic coordinate quality assessment and correction
  • Depth sequence validation: Nested horizon depths requiring logical consistency checks
  • Outlier detection: Property values spanning common feasible ranges
  • Replicate measurements: Resolving data replicates
  • Missing values: Strategic approach for dealing with missing values

These characteristics make the KSSL Kansas dataset an adequate example for acquiring skills in systematic data cleaning, validation, and quality assurance procedures before proceeding to spatial modeling.

Data Structure

The original dataset has been split into two data files:

1) Site and wet chemistry data (KSSL_data.xlsx & KSSL_data.csv): This file includes soil location, sample identifiers, and analytical soil properties (wet chemistry).

The wet chemistry analyses include 14 soil properties commonly measured in soil survey operations. Table 2.2 presents the complete list of available properties, their measurement units, and data completeness.

Table 0.1: Soil properties available in the KSSL Kansas dataset
Category Property Variable Name Units Complete (%) Notes
Carbon & Organic Matter Soil Organic Carbon SOC % (w/w) 97.4 Organic carbon only; excludes carbonates
Total Carbon Carbon_Total % (w/w) 99.4 Includes both organic and inorganic C
Physical Properties Bulk Density (1/3 bar) Bulk.Density_1_3.BAR g/cm^3 60.8 Measured at field capacity (1/3 bar tension)
Bulk Density (oven-dry) Bulk.Density_ovendry g/cm^3 60.8 Measured after oven drying at 105 degrees C
pH pH unitless 99.3 Standard soil pH (1:1 soil:water ratio)
Texture Clay Clay % 99.5 Particles < 0.002 mm
Silt Silt % 99.5 Particles 0.002-0.05 mm
Sand Sand % 99.5 Particles 0.05-2.0 mm (note: sum should = 100%)
Chemical Properties Cation Exchange Capacity CEC cmol(+)/kg 96.7 Nutrient retention capacity; also as meq/100g
Nutrients Calcium Extractable Calcium cmol(+)/kg 96.7 Exchangeable Calcium in CEC; high in Mollisols
Potassium Extractable Potassium cmol(+)/kg 84.7 Exchangeable Potassium in CEC;
Total Nitrogen Nitrogen_Total % (w/w) 93.7 Used to calculate C:N ratio
Phosphorus (Mehlich-3) Phosphorus_Mehlich3 mg/kg 38.6 Multi-element extraction; common in US
Phosphorus (Olsen) Phosphorus_Olsen mg/kg 13.7 Bicarbonate extraction for alkaline soils

Some properties show strategic missing data patterns that reflect standard laboratory protocols in which not all tests are performed on every sample. The varying completeness percentages reflect typical soil laboratory operations rather than random missing data:

  • Fully measured properties (>95% complete) represent standard analyses performed on nearly all samples: organic carbon, texture, pH, and CEC
  • Selectively measured properties (40-94% complete) reflect analyses performed when specific conditions warrant: bulk density, nitrogen, potassium, and calcium carbonate
  • Conditionally measured properties (<40% complete) are performed only when needed

2) Dry chemistry data (MIR_KANSAS_data.xlsx): This file includes replicates of mid-infrared reflectance data (identified by scan_path_name) for each soil sample (linked via the join_key variable). It is available to download from the shared folder in this manual. Multiple spectral scans per horizon enable:

  • Quality control through spectral replicates
  • Training spectroscopy-based prediction models
  • Hybrid DSM approaches combining laboratory measurements with spectral inference
  • Gap-filling strategies for properties with lower completeness using spectral predictions

How to Use This Manual

This tutorial manual introduces the essential concepts of soil data management, covering the complete workflow from soil sampling design through Digital Soil Mapping (DSM) to data sharing. It is organized into 6 modules with multiple sessions each, progressing from fundamental R programming skills through advanced spatial modeling and international data standardization. The manual is designed for hands-on learning, combining conceptual explanations with practical coding exercises, real soil datasets, and spatial data workflows.

Manual Structure

Each module builds upon the previous one, creating a comprehensive learning pathway for operational DSM implementation:

Module 1: Introduction to R and Preparation of Soil Data for Digital Soil Mapping

This module introduces R programming fundamentals for soil science applications. It is divided into two parts covering essential skills for soil data management and analysis. The first part introduces the RStudio interface and R concepts. The second focuses on the use of R for Digital Soil Mapping, including data preparation, quality assessment, and cleaning procedures.

Module 2: Soil Sampling Design

This module introduces sampling methodologies for soil surveys and monitoring, with particular emphasis on the SoilFER project’s three-stage hierarchical hybrid sampling approach. The methodology combines statistical robustness with operational feasibility for designing soil sampling campaigns. The sampling design is founded on the SCORPAN model and developed through successive steps, including sample size optimisation and the identification of representative locations within a three-scale sampling framework. The design is implemented through an R workflow that includes calculation of minimum sampling size, application of the Covariate Space Coverage (CSC) algorithm for selection of large sampling units (PSUs), generation of Secondary and Tertiary sampling units through stratified random sampling, and creation of systematic site identification codes to ensure the traceability of samples across the full sampling strategy.

Module 3: Digital Soil Mapping

This module provides comprehensive training in digital soil mapping (DSM) theory and practice, progressing from fundamental concepts through practical implementation using the standardised dataset produced in Module 1. Feature selection techniques are introduced for the identification of elevant environmental information to explain spatial patterns in soil properties. Moreover, several machine learning algorithms and predictive models to produce raster predictions of soil properties are presented, addressing model training, cross-validation and quantification of the prediction uncertainties.

Module 4: Soil Spectroscopy

This module introduces diffuse reflectance spectroscopy (DRS) as a complementary approach to conventional laboratory methods for estimating soil properties, with the aim of improving the time and cost efficiency of routine analysis. The focus is on the integration of DRS into DSM workflows. The essentials of DRS are presented, such as spectral data preprocessing techniques, quality control procedures and advanced chemometric modelling approaches for building models that translate DRS signals into values of relevant soil properties. A practical exercise then shows how DRS calibration models can be built from relatively small reference datasets to estimate soil organic carbon for samples with only DRS measurements available. Raw spectra from the KSSL dataset are used, and the resulting soil organic carbon predictions are then integrated into a complete DSM workflow.

Module 5: Soil Data Standardisation, Sharing and Dissemination

This module focuses on the organisation, standardisation, sharing and dissemination of soil information through a Soil Information System (SIS), so-called Open-source National Soil Information System (OpenNSIS). It introduces relevant international standards for structuring soil data and metadata and shows how the information produced during the previous modules can be organised and disseminated through a national soil information infrastructure. Attention is given to ISO 28258, the international standard for the digital exchange of soil-related data. The principles of ISO 19115/19139 for describing geospatial datasets through standardised metadata, supporting the documentation, discovery and exchange of soil information, are also introduced. The demonstration shows how data from previous SoilFER modules can be integrated, managed, and disseminated through the system.

Required downloads

All training materials are distributed across two complementary sources.

  1. Scripts and lightweight tabular data (GitHub)

The GitHub repository contains:

  • All scripts used throughout the manual
  • Lightweight tabular datasets
  • The full project folder structure, organised by module

Download the repository as a ZIP file from:
https://github.com/SoilFER/SoilFER-Training-Resources

  1. Large input datasets (Google Drive)

The Google Drive folder contains files that are too large to be hosted on GitHub, including:

  • Raster covariates and maps
  • KSSL spectral dataset (MIR_KANSAS_data.xlsx)

Download the data from:
https://drive.google.com/drive/folders/1K7tq9zX5HsqbqWcNoT27WtfPtehcKBCu

Step-by-step setup instructions

Follow these steps in order before starting any module exercises.

Step 1: Download and extract the GitHub repository

  1. Download the ZIP file from the GitHub link above.
  2. Extract (unzip) the contents on your local machine.
  3. Rename the extracted folder if needed (e.g. SoilFER-Training-Resources).
  4. Use this folder as your main project directory.

Step 2: Download the Google Drive data

  1. Download the full contents of the Google Drive folder.
  2. Keep all file names and folder structures unchanged.
  3. Do not work directly from the downloaded Google Drive folder.

Step 3: Inspect the project folder structure

  1. Open the extracted project directory.
  2. Confirm that the top-level structure matches the following:
SoilFER-Training-Resources/
├── 01_data/
│   ├── module1/
│   ├── module2/
│   ├── module3/
│   ├── module4/
│   └── module5/
├── 02_scripts/
├── 03_outputs/
├── 04_assignments/
├── README.md
└── LICENSE

Verify that each subfolder under 01_data/ corresponds to a module in this manual.

Step 4: Place Google Drive files into the correct module folders

  1. Locate the downloaded Google Drive data on your computer.

  2. Copy the files in the rasters folder and the file MIR_KANSAS_data.xlsx into:

    01_data/module1/training_data

  3. Confirm that the folders now resemble:

   01_data/  
   ├── module1/  
   │   ├── kssl/
   │   │          ├── KSSL_data.csv
   │   │          ├── KSSL_data.xlsx
   │   │          ├── property_thresholds.csv
   │   ├── shapes/
   │   │          ├── Tiger_2020_Counties.shp
   │   │          ├── .... 
   │   ├── training_data/  
   │   │                ├── TerraClimate_Precip_1981_2023_KANSAS.tif  
   │   │                ├── TerraClimate_PET_1981_2023_KANSAS.tif
   │   │                ├── TerraClimate_AvgTemp_1981_2023_KANSAS.tif  
   │   │                ├── ....tif  
   │   │                ├── MIR_KANSAS_data.xlsx  
   │   └── README.md  
  • All files downloaded from the Google Drive folder must be placed inside the 01_data/module1/training_data folder.
  • All code in the manual that uses these files reference that folder.
  • Scripts rely on relative paths and will not run correctly if files are stored elsewhere.