Mostrando entradas con la etiqueta Learning Win ISI. Mostrar todas las entradas
Mostrando entradas con la etiqueta Learning Win ISI. Mostrar todas las entradas

12 ago 2021

Time based validation improvement for pH in cocoa paste

 The best way to check the performance of a calibration is with a new time based validation set. In this case one calibration has been developed to predict pH in cocoa paste. The calibration has been developed choosing the best performance with a cross validation using groups. 

After this the model has been installed in routine and after 2 months, new samples with reference values attached have been collected so we can run the validation to see the statistics. This is the performance:


One of the sample with a high GH is classified as outlier, so we can think that it was a lab error value and we can excluded without a good reason and that must not be done.

Why not to develop again the model with the old training set trying to find a better math-treatment or configuration and validate with this new time based set. The results are surprising with a "None-0-0-1-1" math and 15 PLS terms:

RSQ increase and the sample with the high pH is predicted fine. Take into account that we have a very sort range for pH so we can not expect a high RSQ value.


3 feb 2021

Checking overfitting in validation

 Validation is an important tool to improve the calibration. When developing the calibration we use the cross validation to select the number of terms, and we do not do any further action to avoid overfitting. Later we have the surprise that our predictions for new samples have a bias for certain tip e of samples and the error is much more than expected compared to the SECV (standard error of cross validation) that we use as reference.

This is the case, for example,  for wheat bran where the CV for moisture (Humedad) suggested 9 terms, and we keep that decision. Some time later we have 61 new samples and the validation gives this results:


In the case we have chosen 3 terms when we develop the calibration the results would be:



We can see that we have a high improvement. 
I suggest you use some more techniques apart from the cross validation to select the number of terms. Lately I am trying with some bootstrap techniques.




12 may 2020

Creating a Single Sample Standardization



We create two single sample standardization files, one with the NIR5000 as Master and the DS2500F as Host , and other with the NIR5000 as Host and the DS2500F as Master. 

Depending of the scenario we can use one or the other.

Choosing one sample for the standardization



See first the other three videos:
 Trim Spectra
 RMS of subsamples
 RMS between same samples in different instruments 

In this video we want to select one of the samples to create a standarization file to transfer the data base of soy meal from the NIR5000 to the DS2500.

 One rule of thumb is to select the sample with the lowest GH value looking that the sample is near the average of the spectral population of the soy datababe from which I had developed the calibration in the NIR5000.


11 may 2020

RMS between same samples in different instruments



See first the previous two videos:
Trim Spectra
RMS of subsamples
 
Now we average the repacks and compare the RMS between the samples scanned on different instrument not standardized using the Contrast spectra function of Win ISI. 

Obviously the RMS are higher than the repacks of the samples in one instrument, because we are adding the difference between instruments, but the idea is that after the standardization the RMS of the same sample between two instruments be similar or if possible lower than the sampling error  (RMS between repacks on the same instrument).

RMS of Subsamples



 

See first the previous video:

It is important to know the sampling error, and for this reason the calculation of the RMS of subsamples is very important. 

The RMS is a way to obtain a value for the spectral differences between all the repacks.
For the calculation of the RMS simply scan diiferent repacks of the same sample and homogenize the sample betwee subsamples. Try to do it as best as you can in order to get a low RMS.

Same products are quite heterogeneus and for that reason the RMS can increase. If we grind the sample the RMS wil decrease much more.

The RMS value wil be usefull for some comparisons during the standardization or database transfer.

TRIM SPECTRA



We have a certain number of samples scanned in two instruments (a NIR5000 and a DS2500F). Several repacks of the same sample have been scanned on both instruments, due that they have different sample presentation and different cups.

A sample with a certain ID was well homogenized and the contain was splitted into the two cuvettes, and we repeat the process several times in order to get a higher probability that the same sample has been scanned on both instrument and that will help to see the differences between the instruments.

When we want to compare spectra files from different instruments they must have the same range and the same number of data. In the video we trim the spectra from a DS2500F (850-2500, 0.5) to the range and data points of a NIR5000 (1100-2500, 2nm). 

After that we can overplot or subtract the spectra.

The idea of all this coming videos is to show the process of database transfer from a NIR5000 to a DS2500 or DS2500F instrument.

25 ene 2020

Comparing "R-PLS", "CARET" and "WIN ISI"

I made this exercise just for fun. When we develop a regression we don´t have to look the the XY plot where a sample is predicted from a model where this sample is already included, we have to compare it against a model where she is excluded. So we have to look to see how the samples fits to the regression line to the XY plot of Cross Validation for the number of terms we have selected for the final model.
 
In this case I develop the comparison for a regression of protein in fish meal with the 40 samples I am using in the series of tutorials about "Tidyverse and Chemometrics" using the PLS  and Caret packages and also Win ISI (FOSS Analytical Chemometric software).
 
PLS, Win ISI and Caret recommend 8 terms and as you can see they gave the same LOO Cross Validation XY plots or very similar. Remember that these are the XY plots of predictions where every sample is predicted with the 39 remaining so it´s more realistic of how it will performs in routine.
 
 

4 ene 2020

EQA "Estimated Minimum" and "Estimated Maximum" in Win ISI

One of the statistics we can see after an ".eqa" file is generated are these : "Estimated Minimum "  and the "Estimated Maximum".
 
These values are calculated from the ".cal" file (file with the laboratory values distribution) used to develop the calibration for a certain constituent.
 
Take into account that the  ".cal" used to develop the equation can be different than the original ".cal" due that some samples can be taken out as outliers (the ones specified in the ".oll" file).
 
So from the ".cal" file used for the equation file we calculate the Mean Value and the Standard Deviation Value and the Estimated Minimum is calculates as Mean - 3.SD and the "Estimated Maximum is calculated as Mean + 3.SD.
 
If we get a negative value for the Estimated Minimum we replaced it by cero.
 
 

22 oct 2019

Expand product library with new spectra

One way to expand a calibration is with the option "Expand a product library with new spectra" included in the "Make and use scores" option of Win ISI.
We have a library file (cal file) with the samples we have use to develop a calibration. We have also the ".pca" files and ".lib" files which are giving us the GH and NH values.
 
Now we have new samples analyzed after installing this equation in routine, and we want to select the samples which can expand the calibration in order to make a more robust model. This selection is based on the NH distances, so the  calibration will select empty spaces in the LIB file and will expand the PCs in other directions in the case we get new variability.
 
To expand the library we need the P and T matrices of the current library and the ".nir" file of the new spectra with the candidates to expand and fill the current database. We can stablish a certain cutoff for the NH limits, being 0.6 the default.
 
In this case of the total 46 samples 21 have a NH value higher than 0.600, so they are selected to expand the calibration. 5 samples are similar to the samples we have in the library we have use to develop the current calibration with their ".pca" and ".lib" files and these 5 samples have a "L" at the end. Finally 20 samples are similar to the selected ones, so it does not make sense to choose them, because they are redundant with the selected ones, this samples have an S at the end,
 
So we can send this samples to the lab to create a new cal file that we will merge with the current cal file to expand the calibration.

24 jul 2019

GOOD PRODUCT EVALUATION

 When we create a Good Product Model we want to test it with new samples knowing if they are good or bad. There is of course an uncertain area that we can calculate from experience we get from several evaluation.

In this Excel plot we see the values for "Max Peak T", for the training set (samples before Mars 2019) and new batches we consider are fine from Mars to June. There is a set of bad samples prepared with mixtures out of tolerance for a certain component or components of the mixture.

As we can see the model works in some cases but there are other that are misclassified, so we have to try other treatments or models to check if we can classify them better. Anyway there is always an uncertain zone and we have to check for confidence levels of the prediction.

24 jun 2019

Validation problem (extrapolation)

Sometime when validating a product for a certain constituent (in this case dry matter) we can see this type of X-Y plot:


This a not nice at all validation, but we have to see first that we have like to clusters of lab values for lower and higher dry matter. So the first question is:
Which is the range of the calibration samples in the model which I am validating?.

I check and I see that the range for dry matter  in the model is from 78,700 to 86,800, so I am validating with samples more dried than the ones in the calibration.

I see that it seems like bias effect for those samples. Let´s remove the samples in range and check the statistics for the samples out of range:

We see that we have a bias effect, and some slope caused but one of the samples. So this is a new source of variation to expand the calibration. Merge the validation samples to the database and recalibrate. Try to make robust the new model for extrapolation.

18 may 2019

set.seed function in R and also in Win ISI

It is common to see how at the beginning of some code the "set.feed" function is fixed to a number. The idea of this is to get reproducible results when working with functions which require random sample generation. This is the case for example in Artificial Neural Networks models where the weights are selected randomly at the beginning and after that are changing during the learning process.

Let´s see what happens if set.seed() is not used:
library(nnet)
data(airquality)

model=nnet( Ozone~Wind, airquality  size=4, linout=TRUE )

The results for the weights are:

# weights:  13
initial  value 340386.755571
iter  10 value 125143.482617
iter  20 value 114677.827890
iter  30 value 64060.355881
iter  40 value 61662.633170
final  value 61662.630819
converged

 
If we repeat again the same process:

model=nnet( Ozone~Wind, airquality  size=4, linout=TRUE )

The results for the weights are different:

# weights:  13
initial  value 326114.338213
iter  10 value 125356.496387
iter  20 value 68060.365524
iter  30 value 61671.200838
final  value 61662.628120
converged
 

 
But if we fit the seed to a certain value (whichever you like) .

set.seed(1)
model=nnet( Ozone~Wind, airquality  size=4, linout=TRUE )

 # weights:  13
initial  value 336050.392093
iter  10 value 67199.164471
iter  20 value 61402.103611
iter  30 value 61357.192666
iter  40 value 61356.342240
final  value 61356.324337
converged

 
and repeat the code with the same seed:

set.seed(1)
model=nnet( Ozone~Wind, airquality  size=4, linout=TRUE )

we obtain the same results:

# weights:  13
initial  value 336050.392093
iter  10 value 67199.164471
iter  20 value 61402.103611
iter  30 value 61357.192666
iter  40 value 61356.342240
final  value 61356.324337
converged


SET.SEED es used in Chemometric Programs as Win ISI to select samples randomly:

18 ene 2019

Using RMS statistic in discriminant analysis (.dc4)

In the case we want to check if a certain spectrum belongs to a certain product we can create an algorithm with PCA in such a way that this algorithm try to reconstruct the unknown spectrum with the scores of this unknown spectrum on the PCA space of the product, and the loadings of the product. So we have the reconstructed spectrum of the unknown and the original spectrum of the unknown.
 
If we subtract one from the another we get the Residual spectrum which is really informative. We can calculate the RMS value of this spectrum to see if the unknown spectrum is really well reconstructed so the RMS values is small (RMS is used as statistic to check the noise in the diagnostics of the instrument).
 
Find the right cutoff to check if the sample is well reconstructed depends of the type of sample and sample presentation.
 
Win ISI multiply the RMS by 1000, so the default value for this cutoff which is 100 in reality is 0.1, anyway a smaller or higher value can be used depending of the application.
 
This type of discrimination is known as RMS-X residual in Win ISI 4 and create ".dc4" models.
 
We see in next posts other ways to use this RMS residual.

31 may 2018

Comparing Residuals an Lab Error

One important point before  develop a calibration is to know the laboratory error. This error, known as SEL (Standard Error Laboratory).
This error change depending of the product, because it can be very homogenous or heterogeneous, so in the first case the lab error is lower than in the second case.
In the case of meat meal the error are higher than for other products and this is a case where I am working these days and I want to share in this posts.
A certain number of samples (n) well homogenized had been divided into two or four subsamples and had been send to a laboratory for the Dumas Protein analysis. After receive the results, a formula used for this case, based on the standard deviation of the subsamples and in the number of samples, gives the laboratory error for protein in meat meal.
The result is 1,3 . Probably some of you can think that this is a high value, but really, the meat meal product is quite complex and I think is normal.
 
Every subsample that went to the lab has been analyze in a NIR, and the spectra of the subsamples studied apart with the statistic RMS that shows that the samples were well homogenized.
Now I have the option to average the predictions of the NIR predictions (case 1), or to average the spectra of the subsamples, predict it with the model and get the predicted result (case 2). I use in this case the option 1 and plot a residual plot with the residuals of the average predictions subtracted from the lab average value:

Blue line is +/- “ 1.SEL”, yellow +/- “ 2.SEL” and red +/- “ 3.SEL”.
 

 

20 may 2018

Monitoring LOCAL calibrations (part 2)

Following with the previous post when monitoring LOCAL equations we have to consider several points. Some of them could be.
  1. Which is the reference method and what is the laboratory error for the reference method. Ex:Reference method for protein can be from Dumas or from Kjeldahl .
  2. Do you have data from several laboratories or from just one?. Which reference method use every lab and what is their laboratory error. 
  3. If we have data in the Monitor from both methods, or from different laboratories, split them apart and see which error  SEP do we have. Check also the SEP corrected by the Bias.
  4. Check the performance of the SEP and other statistics for every specie. In meat meal for example you have pork, mix, poultry, feather,....
Sure you can find more knowledge about the data:
  1. Do we have more error with one method than another, with one lab than another?.
  2. Does perform a reference method better for one particular specie than another?.
  3. ................................
Make conclusions from these and other thoughts.
 
For all these, is important to organize well the data and use it fine. Don´t forget that we are comparing reference versus predicted, and that the prediction depends of the model itself apart from the reference values, so we have to be sure that our model is not overfitted or wrong developed.

19 may 2018

Monitoring LOCAL calibrations

Databases to develop Local calibrations has normally a high quantity of spectra with lab values, but we have to take care of  them adding new sources of variance. This way we make them more robust and the standard prediction errors (SEP) decrease when we validate with future independent validation sets.
 
This was the case with a meat meal local database updated with 100 samples with protein values, and with new source of variation as: laboratories,  reference methods (Kj, Dumas), providers, new instruments,...
 
After the LOCAL database update, a set of new samples was received with reference values and I have predicted this values with the Monitor function in Win ISI with the LOCAL database before (blue dots) and after the LOCAL update (red dots).
 
The errors decrease , specially for some types of samples, in an important way when validating with the new set of samples (new samples acquired after the updated Local calibration was installed in the routine software), so even if we have spectra from this instrument, labs, ...., this set has to be considered as an independent set.
 
I don´t give details of the statistics but this picture show the same samples predicted with the LOCAL calibration without update (in blue), and predicted with the LOCAL calibration update (in red), the improvement for the prediction for some samples is important, so the idea is to add this new samples and continuing monitoring the LOCAL database with future validation sets.




30 abr 2018

Validation with LOCAL calibrations

When developing a LOCAL calibration in Win ISI, we use an input file and a Library file with the idea the idea to select the best possible model, so the selection of a well representative input (let´s call validation set) is very important to have success in the development of the model.
 
So the model is conditioned to the input file, if we have choose another input file we could have get another model which performs different, so the need of a Test Set is obvious to check how the model performs with new data.
 
It is important to have this in mind, so one proposal would be to divide (randomly) the data in three sets: 60% for training, 20% for Input or Validation, and another 20% for testing.
 
There are other ways to sort the data in order to select these three Sets (time, seasons, species,...). One thing is clear, some of the models developed will perform better than others, so you can keep several of them and you can check this when you have new data and use an statistic an MSE (Mean Square Error) to compare them.