Presentation and basics
Background and Objective
Healthy and productive soils are fundamental to resilient agrifood systems, sustainable land management, and climate adaptation. Yet, in many regions, soil information remains fragmented, outdated, or inaccessible. The Soil Mapping for Resilient Agrifood Systems (SoilFER) programme responds to this challenge by building comprehensive soil information systems that integrate sampling design, laboratory analysis, soil spectroscopy, digital modeling, and decision-support tools. These systems aim to empower governments, researchers, and farmers with actionable knowledge for crop selection, fertilizer recommendations, and soil health management.
This manual provides a step-by-step guide along the entire soil data value chain. It is designed as both a technical reference and a practical training resource, bridging the gap between raw soil data and its functional use in agricultural and environmental decision-making.
The objectives of this manual are to:
Present harmonized approaches for soil sampling design used under the SoilFER programme
Introduce best practices in soil data preparation and management, aligned with the Global Soil Information System (GloSIS)
Demonstrate methods for digital soil modeling and mapping, covering classical statistics, machine learning, and hybrid inference for both continuous and categorical soil properties.
Explain how to generate functional soil information to support evidence-based decision-making.
Provide guidance on the integration soil spectroscopy-estimated soil parameters into the digital soil mapping process
Facilitate data sharing and dissemination, promoting open standards, metadata documentation, and web-based services.
By integrating these components, the manual equips users to move from raw samples to reliable soil information products that inform policy, guide sustainable soil management, and strengthen food and nutrition security.
The SoilFER project is co-funded by the U.S. Department of State and the Ministry of Foreign Affairs of Japan.
The Training Dataset
This manual uses the KSSL (Kellogg Soil Survey Laboratory) soil dataset from Kansas, USA as a comprehensive example for implementing the complete Digital Soil Mapping (DSM) workflow. The dataset provides a realistic foundation for learning DSM techniques, from inspection and cleaning of raw soil laboratory data to advanced spatial modeling and uncertainty assessment. It also demonstrates the integration of soil spectroscopy data and its application in operational DSM production.
The KSSL Kansas dataset represents a collection of laboratory-analyzed soil samples with corresponding mid-infrared spectroscopy measurements. It contains 10,352 soil horizon measurements from 2,584 distinct soil profiles distributed across Kansas, spanning from 37.02°N to 39.99°N latitude and -102.04°W to -94.71°W longitude. The dataset exemplifies common data quality issues encountered in soil databases, providing hands-on experience with:
- Coordinate validation: Geographic coordinate quality assessment and correction
- Depth sequence validation: Nested horizon depths requiring logical consistency checks
- Outlier detection: Property values spanning common feasible ranges
- Replicate measurements: Resolving data replicates
- Missing values: Strategic approach for dealing with missing values
These characteristics make the KSSL Kansas dataset an adequate example for acquiring skills in systematic data cleaning, validation, and quality assurance procedures before proceeding to spatial modeling.
Data Structure
The original dataset has been split into two data files:
1) Site and wet chemistry data (KSSL_data.xlsx & KSSL_data.csv): This file includes soil location, sample identifiers, and analytical soil properties (wet chemistry).
The wet chemistry analyses include 14 soil properties commonly measured in soil survey operations. Table 2.2 presents the complete list of available properties, their measurement units, and data completeness.
| Category | Property | Variable Name | Units | Complete (%) | Notes |
|---|---|---|---|---|---|
| Carbon & Organic Matter | Soil Organic Carbon | SOC | % (w/w) | 97.4 | Organic carbon only; excludes carbonates |
| Total Carbon | Carbon_Total | % (w/w) | 99.4 | Includes both organic and inorganic C | |
| Physical Properties | Bulk Density (1/3 bar) | Bulk.Density_1_3.BAR | g/cm^3 | 60.8 | Measured at field capacity (1/3 bar tension) |
| Bulk Density (oven-dry) | Bulk.Density_ovendry | g/cm^3 | 60.8 | Measured after oven drying at 105 degrees C | |
| pH | pH | unitless | 99.3 | Standard soil pH (1:1 soil:water ratio) | |
| Texture | Clay | Clay | % | 99.5 | Particles < 0.002 mm |
| Silt | Silt | % | 99.5 | Particles 0.002-0.05 mm | |
| Sand | Sand | % | 99.5 | Particles 0.05-2.0 mm (note: sum should = 100%) | |
| Chemical Properties | Cation Exchange Capacity | CEC | cmol(+)/kg | 96.7 | Nutrient retention capacity; also as meq/100g |
| Nutrients | Calcium Extractable | Calcium | cmol(+)/kg | 96.7 | Exchangeable Calcium in CEC; high in Mollisols |
| Potassium Extractable | Potassium | cmol(+)/kg | 84.7 | Exchangeable Potassium in CEC; | |
| Total Nitrogen | Nitrogen_Total | % (w/w) | 93.7 | Used to calculate C:N ratio | |
| Phosphorus (Mehlich-3) | Phosphorus_Mehlich3 | mg/kg | 38.6 | Multi-element extraction; common in US | |
| Phosphorus (Olsen) | Phosphorus_Olsen | mg/kg | 13.7 | Bicarbonate extraction for alkaline soils | |
Some properties show strategic missing data patterns that reflect standard laboratory protocols in which not all tests are performed on every sample. The varying completeness percentages reflect typical soil laboratory operations rather than random missing data:
- Fully measured properties (>95% complete) represent standard analyses performed on nearly all samples: organic carbon, texture, pH, and CEC
- Selectively measured properties (40-94% complete) reflect analyses performed when specific conditions warrant: bulk density, nitrogen, potassium, and calcium carbonate
- Conditionally measured properties (<40% complete) are performed only when needed
2) Dry chemistry data (MIR_KANSAS_data.xlsx): This file includes replicates of mid-infrared reflectance data (identified by scan_path_name) for each soil sample (linked via the join_key variable). It is available to download from the shared folder in this manual. Multiple spectral scans per horizon enable:
- Quality control through spectral replicates
- Training spectroscopy-based prediction models
- Hybrid DSM approaches combining laboratory measurements with spectral inference
- Gap-filling strategies for properties with lower completeness using spectral predictions
How to Use This Manual
This tutorial manual introduces the essential concepts of soil data management, covering the complete workflow from soil sampling design through Digital Soil Mapping (DSM) to data sharing. It is organized into 6 modules with multiple sessions each, progressing from fundamental R programming skills through advanced spatial modeling and international data standardization. The manual is designed for hands-on learning, combining conceptual explanations with practical coding exercises, real soil datasets, and spatial data workflows.
Manual Structure
Each module builds upon the previous one, creating a comprehensive learning pathway for operational DSM implementation:
Module 1: Introduction to R and Preparation of Soil Data for Digital Soil Mapping
This module introduces R programming fundamentals for soil science applications. It is divided into two parts covering essential skills for soil data management and analysis. The first part introduces the RStudio interface and R concepts. The second focuses on the use of R for Digital Soil Mapping, including data preparation, quality assessment, and cleaning procedures.
Module 2: Soil Sampling Design
This module introduces sampling methodologies for soil surveys and monitoring, with particular emphasis on the SoilFER project’s three-stage hierarchical hybrid sampling approach. The methodology combines statistical robustness with operational feasibility for designing soil sampling campaigns. The sampling design is founded on the SCORPAN model and developed through successive steps, including sample size optimisation and the identification of representative locations within a three-scale sampling framework. The design is implemented through an R workflow that includes calculation of minimum sampling size, application of the Covariate Space Coverage (CSC) algorithm for selection of large sampling units (PSUs), generation of Secondary and Tertiary sampling units through stratified random sampling, and creation of systematic site identification codes to ensure the traceability of samples across the full sampling strategy.
Module 3: Digital Soil Mapping
This module provides comprehensive training in digital soil mapping (DSM) theory and practice, progressing from fundamental concepts through practical implementation using the standardised dataset produced in Module 1. Feature selection techniques are introduced for the identification of elevant environmental information to explain spatial patterns in soil properties. Moreover, several machine learning algorithms and predictive models to produce raster predictions of soil properties are presented, addressing model training, cross-validation and quantification of the prediction uncertainties.
Module 4: Soil Spectroscopy
This module introduces diffuse reflectance spectroscopy (DRS) as a complementary approach to conventional laboratory methods for estimating soil properties, with the aim of improving the time and cost efficiency of routine analysis. The focus is on the integration of DRS into DSM workflows. The essentials of DRS are presented, such as spectral data preprocessing techniques, quality control procedures and advanced chemometric modelling approaches for building models that translate DRS signals into values of relevant soil properties. A practical exercise then shows how DRS calibration models can be built from relatively small reference datasets to estimate soil organic carbon for samples with only DRS measurements available. Raw spectra from the KSSL dataset are used, and the resulting soil organic carbon predictions are then integrated into a complete DSM workflow.
Module 5: Soil Data Standardisation, Sharing and Dissemination
This module focuses on the organisation, standardisation, sharing and dissemination of soil information through a Soil Information System (SIS), so-called Open-source National Soil Information System (OpenNSIS). It introduces relevant international standards for structuring soil data and metadata and shows how the information produced during the previous modules can be organised and disseminated through a national soil information infrastructure. Attention is given to ISO 28258, the international standard for the digital exchange of soil-related data. The principles of ISO 19115/19139 for describing geospatial datasets through standardised metadata, supporting the documentation, discovery and exchange of soil information, are also introduced. The demonstration shows how data from previous SoilFER modules can be integrated, managed, and disseminated through the system.
Required downloads
All training materials are distributed across two complementary sources.
- Scripts and lightweight tabular data (GitHub)
The GitHub repository contains:
- All scripts used throughout the manual
- Lightweight tabular datasets
- The full project folder structure, organised by module
Download the repository as a ZIP file from:
https://github.com/SoilFER/SoilFER-Training-Resources
- Large input datasets (Google Drive)
The Google Drive folder contains files that are too large to be hosted on GitHub, including:
- Raster covariates and maps
- KSSL spectral dataset (
MIR_KANSAS_data.xlsx)
Download the data from:
https://drive.google.com/drive/folders/1K7tq9zX5HsqbqWcNoT27WtfPtehcKBCu
Step-by-step setup instructions
Follow these steps in order before starting any module exercises.
Step 1: Download and extract the GitHub repository
- Download the ZIP file from the GitHub link above.
- Extract (unzip) the contents on your local machine.
- Rename the extracted folder if needed (e.g.
SoilFER-Training-Resources).
- Use this folder as your main project directory.
Step 2: Download the Google Drive data
- Download the full contents of the Google Drive folder.
- Keep all file names and folder structures unchanged.
- Do not work directly from the downloaded Google Drive folder.
Step 3: Inspect the project folder structure
- Open the extracted project directory.
- Confirm that the top-level structure matches the following:
SoilFER-Training-Resources/
├── 01_data/
│ ├── module1/
│ ├── module2/
│ ├── module3/
│ ├── module4/
│ └── module5/
├── 02_scripts/
├── 03_outputs/
├── 04_assignments/
├── README.md
└── LICENSE
Verify that each subfolder under 01_data/ corresponds to a module in this manual.
Step 4: Place Google Drive files into the correct module folders
Locate the downloaded Google Drive data on your computer.
Copy the files in the
rastersfolder and the fileMIR_KANSAS_data.xlsxinto:01_data/module1/training_data
Confirm that the folders now resemble:
01_data/
├── module1/
│ ├── kssl/
│ │ ├── KSSL_data.csv
│ │ ├── KSSL_data.xlsx
│ │ ├── property_thresholds.csv
│ ├── shapes/
│ │ ├── Tiger_2020_Counties.shp
│ │ ├── ....
│ ├── training_data/
│ │ ├── TerraClimate_Precip_1981_2023_KANSAS.tif
│ │ ├── TerraClimate_PET_1981_2023_KANSAS.tif
│ │ ├── TerraClimate_AvgTemp_1981_2023_KANSAS.tif
│ │ ├── ....tif
│ │ ├── MIR_KANSAS_data.xlsx
│ └── README.md
- All files downloaded from the Google Drive folder must be placed inside the
01_data/module1/training_datafolder. - All code in the manual that uses these files reference that folder.
- Scripts rely on relative paths and will not run correctly if files are stored elsewhere.
