11 jul 2022

Diference between subsamples

Once we have sorted and grouped all the samples and subsamples, we can create two matrices one that contains the odd samples (first subsample) and other with the even subsamples (second subsample), this way we can calculate the spectra difference between subsamples and check the spectra.

This difference spectra matrix is important to calculate the RMS between subsamples, that gave us an idea how similar are the subsamples between them. Just sum the squared values for every spectrum difference, divide by the number of wavelengths and calculate the square root.






6 jul 2022

Soil clay regressions: Looking for the better accuracy (part 3)

 Now we have four calibrations for clay in soil developed with a selection of samples from the LUCAS database (Spanish crop soils). Now we want to see if those calibrations predict with certain accuracy a new set of soil samples (155) from a Spanish region, acquired in a different instrument and at a different laboratory.

In these cases, it is normal to expect a bias or a slope in the predictions, so we can use the model with those adjustments applied, until the database is updated and a new expanded method with new variability (instrument, laboratory, region,) developed.

Well, these are the results of the XY plots  "Lab vs Predictions" for this independent data set:

PLS predictions:


Random Forest Predictions

Cubist Predictions


MBL Predictions


As we see in all the cases some adjustment or calibration update is needed. We can try to reprocess everything trying to find a better configuration which improves these values, but in that case this new set will never be independent again like it is now. 



5 jul 2022

Soil clay regressions: Looking for the better accuracy (part 2)

Let´s select a seed (to fix the training and test set) and develop the regression with four algorithms (PLS, Random Forest, Cubist and Memory Based Learning.

These are the Test Validation XY plots and statistics:

PLS Regression:


Random Forest Regression:


Cubist Regression:


Memory Based Learning Regression


We can see PLS give the better RMSEP, but some samples are outside the Action Limits Warning, while that in the MBL the residual distribution is more stable and there are no samples outside the action limits threshold.

Questions about NIR Modelling (001)

Sometimes I receive mails from the readers, that are very interesting, so I create this post to answer the reader and to keep the post to create comments or add what the readers consider about their own experience.

 

The choice of the wavelength corresponding to the studied parameter (is it better to keep all the scan or to choose a part which represents the targeted parameter? If any, how to do so?)

Normally all the scan is used, and the PLS algorithm, latent variables (PLS terms) which represents the spectral variance and their covariance with the studied parameters. Looking to the regression coefficients, and knowing the wavelengths at which those correlation absorb, you can try to interpret the regression and decide if certain wavelength zones could be excludes (as flat zero zones, …..). Regression coefficients are very difficult to interpret due to the math treatments applied (specially derivatives).

Other option is to choose a few specific wavelengths, when normally the first one is the one at which the parameter absorbs (example: 1940 nm for the water) and continue adding wavelengths of other constituents that interfere with the water, or zones that do not absorb, but scatter is observed. Normally the software helps you with these selections, and you have always the statistics to see if the wavelengths added improve the regression. Normally this type of algorithm is called MLR (Multiple Linear Regression).

 

How to split the samples between calibration and validation (is there a test to do?)

Split randomly 80% of the samples to the Training Set, and the remain 20% to the Test Set. In the case you have a lot of samples from different years you can uses other approaches (the older samples for calibration and the new ones for validation, …..). Anyway, if the calibration is robust, you should get similar results.

 

The criteria for choosing the tests to be performed for pre-processing (2nd derivative, SNV, MSC, etc.).

Normally the criteria is to choose the simplest math treatment. For the scatter, if you have a lot od samples and all the possible variability represented you can try MSC, if not one of the best options is to combine SNV and Detrend.

First derivative is difficult to interpret (the maximum for the raw spectra, becomes a zero crossing), second derivative is better interpretable. The important option is the gap you use (not to long because you can loose information, and not too short because you add noise).