18 nov 2011

Repasando Unscrambler - 001

A lo largo de una serie de entradas, iré repasando conceptos de Unscrambler para datos espectroscópicos. Existe un maravilloso libro "Multivariate Data Analysis - in practice" del profesor Kim H. Esbensen que es una guía perfecta para repasar algunos conceptos y meterse posteriormente en mas profundidad con otra serie de datos.

14 nov 2011

Histogramas (Skewness - Kurtosis)

We have seen some of the statistics linked to the histograms, but these two (in this case given by Unscrambler) can be very usefull for a better understanding of our constituent database.

Skewness & Kurtosis

Wikipedia:
Skewness.....Asimetría
Kurtosis......Curtosis

11 nov 2011

Shoot-out 2008_parte 006

Al igual que en las entradas previas, desarrollo la calibración con el software VISION, los tratamientos usados anteriormente no dan tan buenos estadísticos y en esta ocasión funciona mejor el tratamiento de 2ª derivada, para los mismos segmentos espectrales.
Same as previous posts, I have developed the calibration with another software (VISION), here the same math treatment an wavelengths regions as in the others, does not give the same statistics (a little bit worse), but the the 2º derivative gives a SEP similar.

Estadísticos de calibración / Calibration Statistics:


Para no tener overfitting VISION utiliza el estadístico PRESS, del que hablaremos proximamente:
To avoid overfitting VISION use the PRESS Statistic (We will talk about it soon)

Validamos con el conjunto de validación de la Campaña del 99, obteniendo:
Validation with the validation set.


El error de predicción SEP es de 0,1799.
The Standard Error of Prediction is: SEP = 0,1799
Visión nos dá los de Bias Pendiente e Intercepto y nos dice que en caso de ajuste el SEP bajaría a 0,1626. No obstante esto se debe ignorar.
Proximamente probaremos con Matlab y Unscrambrer para dar una evaluación general de los estadísticos de este parámetro.

5 nov 2011

Shoot-out 2008_parte 005

En Shoot-out 2008_parte 004 usamos una ecuación PLS, vamos a probar que pasa con las LOCAL. ¿Mejorará la predicción?.
In Shoot-out 2008_parte 004 we used PLS to develop the equation, now we are to deveop the equation with LOCAL. Will it improve the predictions?.
Minimun number of samples: We can use 75.
Maximun number of samples: We will use Batch Mode (100, 150, 200, 250, 300 and 350).
SEP values are quite similar with 200, 250, 300 and 350.
We can select 200 for our LOCAL model.
Now we have to find the best configuration for "minimun & Máximun numder of factors". Wé will check all along the allowed values (minimun 1, maximun 50):
The best combination found was: Min = 5, Max = 27
We used  SNV-Detrend 1-4-4-1 as Math treatment which it seems to work quite well for this product.
For some reasons (can be explained in a near future if comment are added) when put it into routine statistics change a little bit, and this are the values that we will compare with the other models.



4 nov 2011

Shoot-out 2008_parte 004

Desarrollando la ecuación (para proteína).
Developing the equation (for protein).

Modified PLS Regression Statistics    
Input File……………………………………   FRED2008.CAL 
Validation File…………………………….  fred99a.cal           
Math Treatment ……………………….  1, 4, 4, 1            
Number of variables…………………..  768
Scatter Corr. …………………………….   SNV and Detrend       
Downweight outliers………………….  No
Constituent ……………………………....  WHTPRO                 
Number of samples…………………….  774
Mean    ………………………………………..13.670              
 Range………………………………………….10.00 - 17.00     
Std Dev……………………………………….. 1.367

 
CALIBRACIÓNVALIDACIÓN
Terms SECRSQSECV1-VRSEV BIASSEV(C)
150.1600.9860.1800.9830.1790.0380.176
 
El RPD para la validación es:
RPD = 1,536 : 0,179 = 8.58