How to Build a PLS Model in Practice: From Raw Spectra to Validated Calibration

Step-by-step PLS calibration workflow: spectral import, preprocessing selection, outlier removal, factor optimization, cross-validation, and validation set test

Introduction: Why PLS Modeling Matters in NIR Analysis

If you've spent any time working with near-infrared spectroscopy, you already know the fundamental challenge: the spectra you collect are packed with information, but that information isn't obvious to the naked eye. Overlapping absorption bands, light scattering effects, and baseline shifts all conspire to hide the chemical values you actually care about—moisture content, protein levels, fat percentages, or other quality parameters.

This is where Partial Least Squares (PLS) regression comes in. PLS is the workhorse algorithm for quantitative NIR analysis in food and agriculture. It's not the newest method on the block, and it's certainly not the most exotic. But it works, it's robust, and it's well-understood by regulatory bodies and industry standards alike.

How to build a PLS model in practice: from raw spectra to validated calibration is a question every spectroscopist faces when setting up a new application. Whether you're measuring protein in wheat, moisture in cheese, or sugar content in fruit, the workflow follows a similar path. Let's walk through that path step by step, with a practical example you can adapt to your own work.

Step 1: Collecting a Representative Calibration Set

Before you write a single line of code or touch a PLS algorithm, you need samples. This sounds obvious, but the quality of your calibration set determines the ceiling for your model's performance. No amount of chemometric sophistication can fix a poorly chosen sample set.

Think in Terms of Variability

Your calibration samples need to span the full range of the property you're predicting. If you're building a model for protein content in wheat, don't just grab samples from one harvest season or one geographic region. You need:

A common mistake is building a model on a narrow range and then expecting it to predict outside that range. PLS is an interpolation tool, not an extrapolation tool. If your calibration set covers protein from 10% to 14%, don't expect accurate predictions at 8% or 16%.

Reference Method Accuracy Matters

Your PLS model is only as good as your reference data. If your lab reference method has a standard error of 0.5% for moisture, you cannot expect your NIR model to achieve a standard error of prediction better than that. The reference method is your ground truth—make sure it's accurate, precise, and reproducible.

How Many Samples Do You Need?

There's no magic number, but a practical guideline is:

More samples are better, but only if they add genuine variability. Fifty diverse samples will outperform two hundred near-duplicates every time.

Step 2: Spectral Preprocessing—Cleaning Up Your Data

Raw NIR spectra are rarely ready for modeling as-is. Physical effects like particle size differences, path length variations, and temperature fluctuations create baseline shifts and slope changes that have nothing to do with chemistry. Preprocessing helps separate chemical information from physical artifacts.

Common Preprocessing Techniques

Multiplicative Scatter Correction (MSC) corrects for additive and multiplicative scattering effects. It's particularly useful for powdered or granular materials like flour, feed, or ground grain.

Standard Normal Variate (SNV) centers and scales each individual spectrum. It's similar to MSC in purpose but works on each spectrum independently, making it robust to baseline variations.

Savitzky-Golay Derivatives (first or second derivative) remove baseline offsets and resolve overlapping peaks. The first derivative removes constant baseline shifts; the second derivative also removes linear baseline slopes. Derivatives amplify noise, so they're typically combined with smoothing.

Detrending removes a polynomial baseline from each spectrum, useful for spectra with broad, slowly varying baselines.

What Should You Use?

There's no universal answer. The best preprocessing depends on your sample type and the physical effects present. A practical approach:

  1. Start with raw spectra and build a preliminary model
  2. Try each preprocessing method individually
  3. Compare performance metrics (more on these below)
  4. Select the approach that gives the lowest prediction error

Don't over-preprocess. Every preprocessing step removes information along with noise. The goal is to remove the irrelevant variation while keeping the chemical signal intact.

Step 3: Building the PLS Model

Now we get to the heart of how to build a PLS model in practice: from raw spectra to validated calibration. The PLS algorithm works by finding latent variables (also called factors or components) that maximize the covariance between the spectral data (X) and the reference values (Y). Each latent variable captures a direction of variation that is both relevant to the spectra and predictive of the property you're modeling.

Choosing the Number of Latent Variables

This is arguably the most critical decision in PLS modeling. Too few latent variables and your model is underfitted—it misses important variation and has high bias. Too many and it's overfitted—it models noise and random variation, giving great results on calibration data but poor performance on new samples.

The standard approach is cross-validation. In cross-validation, you split your calibration set into training and validation subsets multiple times, build models with different numbers of latent variables, and evaluate their performance on the held-out samples.

Cross-Validation Methods

For food and agriculture applications, random subsets with 5-10 segments is a solid default. The key is to monitor the Root Mean Square Error of Cross-Validation (RMSECV) and choose the number of latent variables where adding another factor doesn't meaningfully reduce the error.

A Practical Rule of Thumb

Choose the smallest number of latent variables that gives an RMSECV within 1-2% of the minimum. This "parsimonious" approach reduces overfitting risk and usually gives more robust models in practice.

Step 4: Validating Your Model

Cross-validation is an internal validation—it tells you how well your model generalizes within your calibration set. But for a truly validated calibration, you need external validation with independent samples that were never used in model development.

Setting Aside a Test Set

If you have enough samples, set aside 15-20% of them before you start modeling. Use these only at the very end to test your final model. This gives you an unbiased estimate of how the model will perform on new, unseen samples.

Key Performance Metrics

Root Mean Square Error of Prediction (RMSEP): The average difference between predicted and reference values for your test set. This is your most important metric—it tells you what error to expect in practice.

R² (Coefficient of Determination): The proportion of variance in the reference values explained by your model. Values above 0.90 are generally considered good for most food and agriculture applications, but this depends heavily on your sample variability.

RPD (Ratio of Performance to Deviation): The ratio of the standard deviation of your reference values to the RMSEP. An RPD above 3 is generally considered good for screening, above 5 is excellent for quality control.

Bias: The average difference between predicted and reference values. Ideally close to zero. A consistent bias suggests a systematic error that might be correctable.

Practical Example: Predicting Moisture in Wheat Flour

Let's bring this together with a concrete example. Suppose you're a flour miller who wants to use NIR spectroscopy for rapid moisture measurement.

Sample Collection

You collect 150 flour samples over several months, covering a moisture range of 10-15%. For each sample, you record the NIR spectrum (1100-2500 nm) and measure moisture using the reference oven-drying method (AACC Method 44-15A).

Preprocessing

You compare raw spectra, SNV, and first derivative (Savitzky-Golay, 15-point window, second-order polynomial). The SNV-treated spectra give the lowest cross-validation error, so you proceed with SNV.

Model Building

Using random subsets cross-validation with 5 segments, you find that RMSECV decreases steadily up to 6 latent variables, then levels off. You choose 6 latent variables as your final model.

Validation Results

Your model metrics:

These results indicate a model suitable for quality control purposes. The RMSEP of 0.21% is well within the accuracy needed for process monitoring and product release.

Common Pitfalls and How to Avoid Them

Sample Selection Bias

If your calibration set doesn't represent the full range of variability your model will encounter, your model will fail in production. Always validate with samples from different batches, seasons, and conditions than your calibration set.

Overfitting

Using too many latent variables is the most common PLS mistake. Always use cross-validation to guide your choice, and be conservative. A model with slightly higher error but better generalization is almost always preferable.

Ignoring Sample Temperature

NIR spectra are temperature-sensitive. If your calibration samples were measured at 20°C but your production samples are at 30°C, you'll see prediction errors. Either measure everything at a consistent temperature or include temperature as a variable in your model.

Not Monitoring Model Performance

A PLS model isn't a "set it and forget it" tool. Instruments drift, sample matrices change, and reference methods evolve. Regularly check your model's performance with known samples and update your calibration when needed.

Closing Takeaway

How to build a PLS model in practice: from raw spectra to validated calibration isn't just about running an algorithm—it's about thoughtful experimental design, careful preprocessing, and rigorous validation. The steps we've covered—representative sampling, appropriate preprocessing, cross-validation, and external validation—form the backbone of any successful NIR calibration.

The good news is that the workflow is well-established and the tools are accessible. Whether you're using a dedicated chemometrics package or a general-purpose programming environment, the principles remain the same. Start with a solid sample set, keep your preprocessing simple, validate honestly, and monitor your model over time. Do that, and you'll have a calibration you can trust for years to come.

Remember: the goal isn't just to build a model that works on your calibration data. It's to build one that works on the next sample you measure, and the one after that, and the one after that. That's what validated calibration really means.

Continue learning: NIR Spectroscopy Training Online | NIR Fundamentals Course — 32 Lessons

← Back to NIR Spectroscopy Blog