2 feb 2018

Abstracts ready for Eleventh Winter Symposium on Chemometrics (WSC)


Abstracts for the presentations of the  Eleventh Winter Symposium on Chemometrics (WSC) are ready to be consult on the WSC web page. A lot of readers of this blog are from Russia, and I would like to say thanks to all from here.

Std Dev vs. Mean Centered Spectrum

This is a plot which shows the mean centered spectra treated with Multiple Scatter Correction, and over plotted is the estándar deviation values of the mean centered values at each wavelength.
The idea is to create a plot with different scales and laves in the axes, and the result is quite nice with this code:

## Plotting the Standard Deviation Spectrum
X_msc_mc_sd<-as.matrix(apply(X_msc_mc,2,sd))
par(mar=c(5, 4, 4, 6))
matplot(wavelengths,t(X_msc_mc),type="l",col="grey",
        xlab="",ylab="",axes=FALSE)
mtext("Absorbancia",side=4,col="grey",line=3)
axis(4,col="grey",col.axis="grey",las=1)
par(new=TRUE)
matplot(wavelengths,X_msc_mc_sd,lty=1,pch=NULL,axes=FALSE,
        type="l",col="red",lwd=2,xlab="Wavelength",ylab="",
        main= "Std Dev vs Mean centered spectra")
axis(2,col="red",col.axis="red",las=1)
mtext("Standard Deviation",col="red",side=2,line=3)
axis(1,col="black",col.axis="black",las=1)
box()


Practice with this example.

RStudio Tips and Tricks

If you use R there are videos and tutorials which can help you to improve your work, so I will add some of them  in the blog, under the label VIDEO: R-Tutorials 

1 feb 2018

Correlation between scores and PC / PLS terms

First PC search to explain the maximum variability in the X matrix. Once extracted the second PC explain the maximum variance remaining, and the process is repeated until almost all the important variance is explained and the remaining variance is the noise and we don´t want it to incorporate this variance into the model.

Once we have the terms, samples are projected over the several PC terms and every sample has a score for every term. Therefore, we have a score matrix with “N” samples (rows) and “A” components (columns).

This variance can be due to different sources or mixture of sources.

In the case of PLS we are looking for a compromise explaining the maximum possible variance in X, at the same time that we explain a maximum variance in Y. We have also a score matrix when developing the PLS algorithm and this scores have more correlation with the constituent that the scores calculated with PC.

In the case of the soy meal in the conveyor, we can calculate the correlation between the scores for every  of the four PC and the protein:

> cor(scores_4t_pc[,1:4],soy_ift_prot1r1$Prot)

               [,1]
PC term 1  -0.2105997
PC term 2   0.3445256
PC term 3   0.1647146
PC term 4  -0.6888083

We can do the same, but with the scores of the PLS regression:

> cor(Prot_plsr_r1$scores[,1:4],soy_ift_prot1r1$Prot)

            [,1]
Comp 1 0.2129742
Comp 2 0.4193727
Comp 3 0.5348858
Comp 4 0.4425912

As I can see the correlations are higher for the PLS, but there are some curiosities about the PC scores that we can try to check yin future posts.

31 ene 2018

Trying to understand the Loadings

The PLS regression in R, give the values of the loadings. In previous posts we decide to use 4 terms, so we have 4 loadings, one for every term.
See that the first term explain almost all the variability in X, while the rest explain quite few. This does not mean that we don´t need them and they are important.
The first loading looks like an spectrum of a sample, so can be related and affected by the scatter. The others look more similar to the regression coefficients we have seen in the previous post.
So we have lo look to the different loadings shapes carefully.
plot(Prot_plsr,"loadings",comps=1:4,legendpos="top",
     lty=c(1,2,4,5),col=c(1,2,4,5),ylim=c(-0.3,0.7))
We can compare individually every loading with the regression coefficients, and we can see how the third and fourth loadings seem to contribute to the regression coefficients.
par(mfrow=c(2,2))
plot(Prot_plsr,"loadings",comps=2,legendpos="top",
     lty=c(2),col=c(2),ylim=c(-0.3,0.7))
plot(Prot_plsr,"loadings",comps=3,legendpos="top",
     lty=c(4),col=c(4),ylim=c(-0.3,0.7))
plot(Prot_plsr,"loadings",comps=4,legendpos="top",
     lty=c(5),col=c(5),ylim=c(-0.3,0.7))
plot(coefficients[,,4],type="l")