How to Run PCA on NIR Spectra: Step-by-Step with Real Data and Score Plots

A complete walkthrough of Principal Component Analysis applied to NIR spectra — from importing raw data to interpreting score plots, loading plots, and variance

How to Run PCA on NIR Spectra: Step-by-Step with Real Data and Score Plots

If you've been following along with our NIR spectroscopy lessons, you know that near-infrared spectra are packed with overlapping, broad bands. Raw spectra are often hard to interpret visually — two samples might look nearly identical to the eye, yet be completely different in moisture or protein content. That's where principal component analysis, or PCA, comes in.

In this guide, I'll show you how to run PCA on NIR spectra: step-by-step with real data and score plots, using a practical example from the grain industry. By the end, you'll understand how PCA reduces hundreds of wavelength variables into a few meaningful components, and how to read the score plots that make sample patterns visible.

Why PCA Is Essential for NIR Data

NIR spectra typically contain 1,000 to 2,000 data points per sample. If you tried to compare samples across all those wavelengths directly, you'd be drowning in numbers. Worse, many of those wavelengths are highly correlated — the absorbance at 1,200 nm tells you something similar to the absorbance at 1,210 nm.

PCA solves this by finding the directions of maximum variance in your data. It creates new variables, called principal components, that are linear combinations of the original wavelengths. The first principal component captures the most variation, the second captures the next most, and so on.

This does two things for you. First, it compresses your data dramatically. Second, it reveals hidden structure — clusters, outliers, and trends — that you simply cannot see in raw spectra.

What You Need Before Running PCA

Before you touch any software, make sure your data is properly prepared. PCA is powerful, but it's not magic. Garbage in, garbage out.

Clean Your Spectra First

Remove any obvious outliers caused by sample handling issues, such as clumping, bubbles, or stray light. If you have spectra with extreme baseline shifts, consider applying a preprocessing method like Standard Normal Variate (SNV) or a first derivative. These help remove physical scattering effects that would otherwise dominate your PCA results.

Organize Your Data Matrix

Your data should be arranged as a matrix where rows are samples and columns are wavelengths. For example, if you have 50 wheat flour samples scanned at 1,100 to 2,500 nm in 2 nm steps, you'll have a 50 × 700 matrix. Each cell contains the absorbance value at that wavelength for that sample.

Mean Center Your Data

Almost every PCA routine will mean-center your data automatically. This means subtracting the average spectrum from every sample spectrum. It ensures that PCA focuses on differences between samples, not on the overall average shape of the spectra.

Step-by-Step: Running PCA on NIR Spectra

Let's walk through the process using a real example: we have 60 corn samples measured by NIR, and we want to see if they naturally group by hybrid variety or moisture level.

Step 1: Load and Inspect Your Spectra

Start by plotting all your raw spectra on one graph. This gives you a quick visual check. Look for obvious outliers — spectra that shoot off in weird directions or have unusual baseline slopes. If you see any, investigate the source before proceeding.

In this example, our corn spectra look fairly consistent, with a few baseline shifts. We'll apply SNV preprocessing to handle the scattering.

Step 2: Preprocess the Data

Apply your chosen preprocessing method. For our corn data, SNV works well because it normalizes each spectrum individually, removing multiplicative scatter effects.

After preprocessing, plot the spectra again. You should see tighter grouping of the curves. If you still see strong baseline drift, try a first derivative with Savitzky-Golay smoothing.

Step 3: Run the PCA Algorithm

In most chemometrics software — whether it's Unscrambler, SIMCA, R, or Python with scikit-learn — running PCA is a one-liner command. But here's what happens under the hood:

For our corn data, let's say the first principal component (PC1) explains 68% of the variance, and PC2 explains 21%. Together, that's 89% of the total variation captured in just two dimensions.

Step 4: Examine the Explained Variance

Always check the explained variance plot, sometimes called a scree plot. It shows you how much variance each PC accounts for. You want to keep enough components to explain at least 95% of the variance, but no more.

In our example, PC1 and PC2 give us 89%. Adding PC3 might push us to 96%. For visualization purposes, we'll stick with PC1 and PC2, but for modeling, we might include PC3.

Step 5: Generate and Interpret the Score Plot

The score plot is where the magic happens. Each point on the plot represents one sample, positioned by its scores on PC1 and PC2. Samples that are similar in composition will cluster together.

For our corn data, here's what we see:

This is the power of PCA. You instantly see relationships that would take hours to uncover by comparing raw spectra.

Step 6: Look at the Loadings to Understand What PCA Found

The loadings tell you which wavelengths contribute most to each principal component. Plot the loadings for PC1 against wavelength. Peaks in the loadings correspond to spectral regions where samples differ the most.

In our corn example, PC1 loadings show strong peaks around 1,450 nm and 1,940 nm — both associated with water absorption. That confirms PC1 is capturing moisture variation. PC2 loadings might show peaks around 1,720 nm and 2,300 nm, which relate to starch and protein content.

This step is crucial because it links your statistical findings back to actual chemistry.

Real-World Example: Sorting Wheat by Protein Content

Let me share a more detailed example from a flour milling operation. A miller wanted to know if NIR could sort incoming wheat loads by protein content before they entered the mill.

They collected NIR spectra from 120 wheat samples, each with a known protein value from reference lab analysis. After running PCA:

The miller was able to use the score plot to visually verify that their protein segregation was working. More importantly, they identified two samples that were misclassified by the supplier — the PCA flagged them before they entered the mill, saving potential quality issues downstream.

This is a perfect example of why how to run PCA on NIR spectra: step-by-step with real data and score plots matters in practice. It's not just an academic exercise — it's a quality control tool.

Common Mistakes to Avoid

Overfitting the Number of Components

Just because you can extract 20 principal components doesn't mean you should. Stick with the minimum number that explains the meaningful variation. If your score plot looks like a cloud with no structure, adding more PCs won't fix it.

Ignoring Preprocessing

PCA on raw, unprocessed NIR data often produces score plots dominated by baseline shifts, not chemistry. Always test a couple of preprocessing methods and see which gives you the cleanest, most interpretable score plots.

Misreading Distance in Score Plots

Samples that are far apart in the score plot are genuinely different in the dimensions shown. But samples that are close together might still differ along PC3 or PC4. Always check the explained variance — if you're only looking at 50% of the total variation, you might be missing important differences.

Software Options for Running PCA

You don't need expensive commercial software to run PCA on NIR spectra. Here are some options:

Whichever you choose, the steps are the same: preprocess, run PCA, examine explained variance, and plot scores and loadings.

Closing Takeaway

PCA is the single most important exploratory tool in NIR spectroscopy. It turns messy, high-dimensional spectra into clean, interpretable visuals. Once you master how to run PCA on NIR spectra: step-by-step with real data and score plots, you'll find yourself using it constantly — for quality control, for spotting outliers, for checking batch consistency, and for understanding what your spectra are really telling you.

Start with a small dataset you know well. Run PCA, look at the score plot, and ask yourself: does the grouping make sense? If it does, you've just unlocked a powerful way to see inside your samples. If it doesn't, dig into the loadings — they'll tell you what's driving the variation.

Either way, you're no longer just staring at wavy lines. You're seeing chemistry.

Continue learning: NIR Spectroscopy Training Online | NIR Fundamentals Course — 32 Lessons

← Back to NIR Spectroscopy Blog