Mostrando entradas con la etiqueta Unscrambler. Mostrar todas las entradas
Mostrando entradas con la etiqueta Unscrambler. Mostrar todas las entradas

30 jul 2016

Looking to the scatter effects with Unscrambler


In the left side of the previous plots, you can see NIT spectra of wheat kernels, that I have download from the database available at:
This web page is very interesting and also the YouTube Chanel where Rasmus Bro (Professor, Dept. of Food Science, University of Copenhagen) explain PCA and PLS concepts with Unscrambler apart from other Chemometric Lessons.
The data is in Matlab and I have to play with it to import it to Unscrambler. Once in Unscrambler I check the option to look to the scatter effects that I saw in an Unscrambler Camo video in YouTube. The video use the new Unscrambler X, but I have the 9.1, and this function to check the scatter effects is also in this old version.
So in the left side we can see the scatter effect for every sample, and it is clear that we have an add effect that we have to remove.
We want to see the chemical effects and not the physical effects like the scatter. So I apply the S. Golay math treatment and look to the effects again and I see this plot:

and something curious happen, because we continue seeing scatter effects in a multiplicative way from the center to the extremes, so SG could help to improve the correlation with the constituent  of interest, but not the scatter removal, so I add to the SG transformation the MSC transformation and we can see how the scatter is almost removed.

29 jul 2016

Importing WinISI data into Unscrambler

I use the Win ISI "demo.cal" data to export it as an ASCII file, and I tried to import it into the 9.1 Unscrambler version which has the option to import from ASCII, and a window appears to configure how the ASCII file exported from Win ISI is designed.
This way, I have all data in place so I can start to configure the sample and variable sets.
This is a well known sample set so it is quite interesting for a tutorial of Unscrambler.
 


24 jun 2015

Unscrambler Video: Chemometrics applied to NIR data

It´s nice to see always this kind of videos (software chemometric tools), explaining how to treat and proceed NIR data in this case with Unscrambler software.

22 nov 2013

Interview Brad Swarbrick (CAMO): Introduction to Multivariate Data Analysis


From YouTube: "Brad Swarbrick, Vice President of Business Development at CAMO Software, gives a short introduction to multivariate data analysis, discusses some of its applications and how these powerful analytical tools are being used to improve products and manufacturing processes in a wide range of industries".

14 sept 2012

Unscrambler (Jam Exercise) - 004

In the posts:
I ´ve been practicing Unscramber with some of the Demo files (Jam), used in the book “Multivariate Data Analysis - in practice” and following the tutorials.
I continue in this post with an important part: Compare the models in order to be sure which one is better, PCR or PLS1, to predict the Y parameter “preference”. For this is clear that we have to look to the residual variance left by the models, taking into account of course the number of terms, over-fitting,…
If we have a look to the plot for the Y residual variance for the PCR, we see an increase in the residual variance for the first PC. That is not good….but think about it.
The PCA does not take into account the Y matrix, so the first PC can be related to some important X structure which cannot be related to the Y parameter. Once extracted, the second PC correlates better with the Y matrix,but still not as good as the first PLS1 term . So this type of plots helps us to understand what is happening.
 Let´s see now the PLS1 residual variance plot for Y, we have a much better prediction with the first term, because the Y matrix was a part of the calculation process in the PLS1.
We have to decide for the model, the best number of terms, and software’s as Unscrambler can decide by you the best option, but you can change the number up or down. You have the control, but we have to check more plots and statistics, before to decide the best option.

4 sept 2012

Unscrambler (Jam Exercise) - 002

PCR is a MLR regression, but indeed using X matrix, we use the T (scores) matrix. We know that the X explained variance is normally quite high for this first PC and decrease with the others, but in the PCR there is no guarantee that the explained variance for the Y follow that order and in the case of the post “Unscrambler (Jam Exercise) – 001 is just 1% of explained variance for the first PC, 57% for the second and 34% for the third. This does not happen for the PLS.
Let´s develop a PLS1 regression for the same X (sensory) and Y (preference) than in “Unscrambler (Jam Exercise) – 001.
We see how the first PLS term explain 91% of the variance, 3% the second and 2% the third.
The first term is very influence by parameters as Thickness which is inverse correlated with others which are preferenced by the consumers (Redness, Colour, Sweetness and Juiciness) . Other parameters as Chewiness, Bitterness, Rasp smell and flavor, do not have influence in the preferences of the consumers.
PLS terms 1 and 2, model quite well the groups for harvesting time H1, H2 and H3, which explain the important parameters for the customers. So,...., 2 terms seem enough for the model.
See in this link details from CAMO about this Jam data set: http://www.camo.com/products/unscrambler/trial.swf
 


3 sept 2012

Unscrambler (Jam Exercise) - 001

I will write during the next days some posts about a famous exercise of Unscrambler describe in the book "Multivariate Data Analysis - in practice", in order to help myself improving my knowledge about this software.
This exercise has raspberry samples from 4 different locations and harvested at 3 different times.
The names of the files are C”a”H”b” where C is the indication for the location and "a" has a value of 1 for location 1, 2 for location 2, 3 for location 3, and 4 for location 4.
H is the indication for Harvest time and "b" has a value of 1 for the early harvest, 2 for the middle harvest and 3 for the late harvest.
When developing a PCR (X variables= a serial of sensory parameters, Y =average value of the preference of 114 consumers for each sample), the scores and loadings are calculated as PCA.
We visualize a group for samples harvested early (H1), clearly in the plot PC1 vs PC2.
We see the variance explained by the taste variations along PC1 (48%), 28% along PC2 and 21% along PC3.
“Y” variable is not well represented by PC1 (only 1%), but the variance explained for “Y” in PC2 is 57% and 34% for PC3.
We see how sweetness has a small loading in PC1 vs PC2 (consider as not important), but  it becomes an important variable along PC3.
We can see correlations between the “X” variables:
Which variable/s, is/are inverse correlated with “thickness”? Redness and color are inverse correlated, which is a characteristic of maturity (late harvest).
Which samples are more thickness (harvested early, middle or late), why? It is clear that sample harvested early, and the samples harvested late are less value for this parameter.
We can get a lot of conclusions from these plots if we study them carefully, as which samples and from which places are preferred by the consumers. We see how samples from places 1 and 3 harvested late are preferred by their red intensity color.
See in this link details from CAMO about this Jam data set: http://www.camo.com/products/unscrambler/trial.swf

2 dic 2011

Cross Validation Groups

La selección de número de grupos, así como el número de muestras que pertenecen a cada grupo, tiene cierta importancia a la hora de desarrollar nuestro modelo de calibración. Pongamos por ejemplo el caso de una calibración con 1000 muestras y creamos 2 grupos, por lo que 500 irán a un grupo y las otras 500 al otro. ¿Como lo hace?.
Se pueden seleccionar de una manera aleatoria.
Las pares a un grupo y las impares al otro.
Cuando se  crean mas grupos (3 por ejemplo) existen mas combinaciones para poder seleccionarlas:
La 1ª, 4ª, 7ª,......al primer grupo. La 2ª, 5ª, 8ª ,.....al segundo y la 3ª, 6ª, 9ª,..... al tercero
De manera aleatoria.
Cada tercio a un grupo.
....................................
Y no olvidemos la Full Cross Validation (leave one out) en la que hay tantos grupos como muestras.
Debemos de observar las opciones que tiene nuestro software para el desarrollo de la Cross Validation y seleccionar el que nos parezca más adecuado.
También debemos de tener en cuenta en función del que elijamos si ordenamos nuestro conjunto de calibración (en el caso de "leave one out" no tendría importancia) nuestras muestras por valor de constituyente, o por algún otro criterio como podría ser su distancia de Mahalanobis al centro de la población.
Lo recomendable es seguir algún tipo de orden y después seleccionar una selección de tipo sistemática, o seleccionar para los grupos las muestras de una manera aleatoria.
Debemos de tener en cuenta de que en el caso de que una "única muestra" sea muy especial en cierto momento formará parte del conjunto de validación y no habrá ninguna como ella en el de calibración, por lo que será detectada como anómala. Es por tanto conveniente que cuando una muestra esté en el grupo de validación existan muestras similares (en variedad, concentración de analito, procedencia,...,etc) en el conjunto de calibración.
Los estadísticos que obtengamos serán diferentes en función de los grupos, pero si la base de datos es lo suficiente robusta y la validación cruzada está bien estructurada, no debería de haber grandes diferencias.

Manera de selección de las muestras para validación cruzada en Unscrambler:

30 nov 2011

Unscrambler: PLS Regression (Part 3)

En Unscrambler podemos marcar las muestras anómalas y recalcular sin ellas. Hacemos esto sin las muestras que comentamos anteriormente.
Las cosas cambian, ahora solo dos términos son necesarios para explicar la mayoría de la varianza. se observan también agrupaciones en función del valor de número de octano.
 Se observan también agrupaciones en función del valor de número de octano.
Los estadísticos de validación "leverage" para la regresión son:

29 nov 2011

Unscrambler: PLS Regression (part 2)

Después de desarrollar la regresión PLS, debemos fijarnos en el gráfico de Varianza Explicada o Varianza residual (depende del que más nos guste). En el caso de la varianza explicada, lo normal es que aumente en función de los términos que vayamos añadiendo. No olvidemos que este cálculo lo realiza con el método de validación que hayamos elegido (leverage o cross validation). En este caso estamos usando leverage.

26 nov 2011

Unscrambler: PLS Regression (Part 1)

Ya os he comentado la imprescindible ayuda que proporciona el libro "Multivariate Data Analysis - in practice" del Kim H. Esbensen para iniciarte en la práctica de Unscrambler. Lo hemos estado haciendo en la serie "Repasando Unscrambler" con datos espectroscópicos de espectros de gasolina.
Esta serie de espectros los estudia también A.M.C. Davies en la Tony Davies Column en tres artículos:
The value of pictures 10/4 (1998).
More pictures from PLS regression analysis 10/6 (1998).
Uncertainty testing in PLS regression 13/2 (2001).
Podéis encontrarlos y descargarlos en Spectroscopy Europe.
En ellos comenta la importancia de mirar a los gráficos que se generan y no simplemente a los ya conocidos X-Y plot, así como a estadísticos como el RSQ y SEP.
También comenta en estos artículos, un comentario de Ian Cowe, que realmente debemos de tener en cuenta "What we (chemometricians) do is mainly to look at the pictures". 
Es importante por tanto, para nosotros que tratamos de introducirnos y profundizar en lo posible en este complejo mundo de la quimiometría de interpretar y entender los gráficos que se generan con el proceso de calibración y no solo quedarnos con el resultado final.
Hemos estado trabajando con los espectros con los análisis PCA, y ahora lo haremos con la regresión  PLS1.
En la regresión se recomiendan usar 3 factores, por lo que podemos ver mapas de dos dimensiones:
Term 1 vs Term 2
Term 1 vs Term 3
Term 2 vs Term 3
Observemos el 1 vs 2:

Volvemos a encontrar las muestras M52 y H59, que son las dos muestras aditivadas.
Estas muestras en el gráfico de influencia (usando los tres términos)  y en el eje de varianza residual Y, muestran un valor practicamente de cero:
Teniendo la H59 una gran influencia sobre el modelo.
Sin quitar ninguna muestra y con los tres factores PLS, el gráfico X-Y  es:

25 nov 2011

Repasando Unscrambler - 006 (Sample Residuals)

En Unscrambler podemos ver los residuales espectrales para cada muestra a medida que se añaden PCs. No olvidemos que partimos de las muestra centrada.
Observemos (de izquierda a derecha y de arriba a abajo) los residuales espectrales a medida que se añaden PCs. El primer gráfico es de la muestra centrada, el 2º el residual con 1 PC, el 3º con 2 PCs, el 4º con 3 PCs (el recomendado por el modelo) y el 5º con 4 PCs (uno más del recomendado con lo que podiamos entrar en problemas de overfitting).

23 nov 2011

Repasando Unscrambler - 005 (X- Loadings)

Hemos visto en 001, los espectros y en que zonas había una mayor varianza, con los márgenes de desviación estándar y los cuartiles. Como sabemos al calcular los componentes principales,, calculamos los "loadings" (la ya varias veces comentada Matriz "P"). Estos Loadings son espectros que nos servirán para reconstruir otros espectros. Disponemos de tantos loadings como PCs, y en el ejemplo que estamos desarrollando en este repaso, la recomendación es de 3 PCs.
El primer PC, esta relacionado con la mayor fuente de varianza espectral, una vez extraida esta varianza, se calculan los demás extrayendo las demas fuentes de varianza,...
Observar estos Loadings requiere un cierto conocimiento de espectroscopía (donde están las bandas de los diferentes analitos, interferencias,.....).

21 nov 2011

Repasando Unscrambler - 004 (Residual X - Leverage)

Este gráfico es de gran importancia, a la hora de determinar si descartamos muestras como anómalas o las mantenemos. Algunas muestras tienen un alto residual y separan del resto de muestras, pero pueden hacerlo de distinta forma. Pongamos un simple ejemplo.
Una serie de muestras se describen perfectamente con dos componentes principales, que como sabemos describen un plano. Sus proyecciones sobre dicho plano serán mas o menos pequeñas en función de su residual X, pudiendo caer sobre el mismo plano, en cuyo caso su residual sería cero. En el caso de que alguna muestra, tenga un alto residual y su proyección sobre el plano sea muy grande, dicha muestra es un anómalo por alto residual y es muy probable que tengamos que descartar dicha muestra.
Por otra parte puede haber muestras con un residual, pequeño, pero que se aparta del resto de muestras considerablemente, esta muestra se considera de gran influencia en el modelo y tiene un gran peso al describir los componentes principales. En este caso debemos de considerar el mantenerla previo estudio de la muestra (¿pertenece a la misma población?, ¿es una muestra con una concentración de analito alta?,....).
Puede ocurrir el caso en que la muestra tenga un alto residual, así como una alta influencia.
Para el ejemplo que estamos viendo en la serie "repasando unscrambler", el gráfico "Residual X vs Leverage" es:

El gráfico de barras nos muestra los residuales de validación en rojo para cada una de las muestras. Claramente destacan los de las muestras M52 (penúltimo) y H59 (último).Vemos que la muestra H59, tiene un menor residual en el modelo final (azul), pero un alto residual en el de validación, pero esto tiene su lógica, porque hemos utilizado la validación cruzada (leave one out), y cuando esta muestra esta en el grupo de validación, no hay ninguna como ella en el de calibración y el residual de validación es muy alto.


19 nov 2011

Repasando Unscrambler - 003

Realizamos el análisis de Componentes Principales para una mejor comprensión de nuestra base de datos. Los gráficos de "varianza explicada" , nos indican que tres PCs son suficiente para explicar la variabilidad de nuestros espectros.

Repasando Unscrambler - 002

Una vez observados los datos de la matriz X (espectros), observaremos los de la matriz Y (valores de referencia), y para ello como ya hemos hecho en otras ocasiones la mejor opción es el histograma.

18 nov 2011

Repasando Unscrambler - 001

A lo largo de una serie de entradas, iré repasando conceptos de Unscrambler para datos espectroscópicos. Existe un maravilloso libro "Multivariate Data Analysis - in practice" del profesor Kim H. Esbensen que es una guía perfecta para repasar algunos conceptos y meterse posteriormente en mas profundidad con otra serie de datos.