2 feb 2018
Abstracts ready for Eleventh Winter Symposium on Chemometrics (WSC)
Abstracts for the presentations of the Eleventh Winter Symposium on Chemometrics (WSC) are ready to be consult on the WSC web page. A lot of readers of this blog are from Russia, and I would like to say thanks to all from here.
Std Dev vs. Mean Centered Spectrum
This is a plot which shows the mean centered spectra treated with Multiple Scatter Correction, and over plotted is the estándar deviation values of the mean centered values at each wavelength.
The idea is to create a plot with different scales and laves in the axes, and the result is quite nice with this code:
## Plotting the Standard Deviation Spectrum
X_msc_mc_sd<-as.matrix(apply(X_msc_mc,2,sd))
par(mar=c(5, 4, 4, 6))
matplot(wavelengths,t(X_msc_mc),type="l",col="grey",
xlab="",ylab="",axes=FALSE)
mtext("Absorbancia",side=4,col="grey",line=3)
axis(4,col="grey",col.axis="grey",las=1)
par(new=TRUE)
matplot(wavelengths,X_msc_mc_sd,lty=1,pch=NULL,axes=FALSE,
type="l",col="red",lwd=2,xlab="Wavelength",ylab="",
main= "Std Dev vs Mean centered spectra")
axis(2,col="red",col.axis="red",las=1)
mtext("Standard Deviation",col="red",side=2,line=3)
axis(1,col="black",col.axis="black",las=1)
box()
Practice with this example.
The idea is to create a plot with different scales and laves in the axes, and the result is quite nice with this code:
## Plotting the Standard Deviation Spectrum
X_msc_mc_sd<-as.matrix(apply(X_msc_mc,2,sd))
par(mar=c(5, 4, 4, 6))
matplot(wavelengths,t(X_msc_mc),type="l",col="grey",
xlab="",ylab="",axes=FALSE)
mtext("Absorbancia",side=4,col="grey",line=3)
axis(4,col="grey",col.axis="grey",las=1)
par(new=TRUE)
matplot(wavelengths,X_msc_mc_sd,lty=1,pch=NULL,axes=FALSE,
type="l",col="red",lwd=2,xlab="Wavelength",ylab="",
main= "Std Dev vs Mean centered spectra")
axis(2,col="red",col.axis="red",las=1)
mtext("Standard Deviation",col="red",side=2,line=3)
axis(1,col="black",col.axis="black",las=1)
box()
Practice with this example.
RStudio Tips and Tricks
If you use R there are videos and tutorials which can help you to improve your work, so I will add some of them in the blog, under the label VIDEO: R-Tutorials
1 feb 2018
Correlation between scores and PC / PLS terms
First PC
search to explain the maximum variability in the X matrix. Once extracted the
second PC explain the maximum variance remaining, and the process is repeated until
almost all the important variance is explained and the remaining variance is
the noise and we don´t want it to incorporate this variance into the model.
PC term 2 0.3445256
PC term 3 0.1647146
PC term 4 -0.6888083
Comp 2 0.4193727
Comp 3 0.5348858
Comp 4 0.4425912
Once we
have the terms, samples are projected over the several PC terms and every
sample has a score for every term. Therefore, we have a score matrix with “N”
samples (rows) and “A” components (columns).
This variance
can be due to different sources or mixture of sources.
In the
case of PLS we are looking for a compromise explaining the maximum possible
variance in X, at the same time that we explain a maximum variance in Y. We
have also a score matrix when developing the PLS algorithm and this scores have
more correlation with the constituent that the scores calculated with PC.
In the
case of the soy meal in the conveyor, we can calculate the correlation between
the scores for every of the four PC and
the protein:
> cor(scores_4t_pc[,1:4],soy_ift_prot1r1$Prot)
[,1]
PC term 1 -0.2105997PC term 2 0.3445256
PC term 3 0.1647146
PC term 4 -0.6888083
We can
do the same, but with the scores of the PLS regression:
> cor(Prot_plsr_r1$scores[,1:4],soy_ift_prot1r1$Prot)
[,1]
Comp 1 0.2129742Comp 2 0.4193727
Comp 3 0.5348858
Comp 4 0.4425912
As I can
see the correlations are higher for the PLS, but there are some curiosities
about the PC scores that we can try to check yin future posts.
31 ene 2018
Trying to understand the Loadings
The PLS regression in R, give the values of the loadings. In previous posts we decide to use 4 terms, so we have 4 loadings, one for every term.
See that the first term explain almost all the variability in X, while the rest explain quite few. This does not mean that we don´t need them and they are important.
The first loading looks like an spectrum of a sample, so can be related and affected by the scatter. The others look more similar to the regression coefficients we have seen in the previous post.
So we have lo look to the different loadings shapes carefully.
plot(Prot_plsr,"loadings",comps=1:4,legendpos="top",
lty=c(1,2,4,5),col=c(1,2,4,5),ylim=c(-0.3,0.7))
We can compare individually every loading with the regression coefficients, and we can see how the third and fourth loadings seem to contribute to the regression coefficients.
par(mfrow=c(2,2))
plot(Prot_plsr,"loadings",comps=2,legendpos="top",
plot(Prot_plsr,"loadings",comps=2,legendpos="top",
lty=c(2),col=c(2),ylim=c(-0.3,0.7))
plot(Prot_plsr,"loadings",comps=3,legendpos="top",
plot(Prot_plsr,"loadings",comps=3,legendpos="top",
lty=c(4),col=c(4),ylim=c(-0.3,0.7))
plot(Prot_plsr,"loadings",comps=4,legendpos="top",
plot(Prot_plsr,"loadings",comps=4,legendpos="top",
lty=c(5),col=c(5),ylim=c(-0.3,0.7))
plot(coefficients[,,4],type="l")
plot(coefficients[,,4],type="l")
Suscribirse a:
Entradas (Atom)


