User Tools

Site Tools


rvkde:usage

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
rvkde:usage [2008/08/19 01:50] – dirtyrvkde:usage [2008/08/19 04:21] (current) – dirty
Line 28: Line 28:
 </code> </code>
  
-====== Model selection ====== +====== Parameter selection ====== 
-Most machine learning tools provide some parameters for users.  For example, the __k__ in knn classification algorithm and the __k__ in k-means clustering algorithm.  From the optimistic view, these parameters provide flexibility and make machine learning tools more powerful.  However, from another point of view, these machine learning techniques cannot determine (or learn) some parameters automatically so that users must specify by themselves.  To find a good parameter combination is so-called model selection.  See [[wp>model_selection|Model Selection]] for more information.+Most machine learning tools provide some parameters for users.  For example, the __k__ in knn classification algorithm and the __k__ in k-means clustering algorithm.  From the optimistic view, these parameters provide flexibility and make machine learning tools more powerful.  However, from another point of view, these machine learning techniques cannot determine (or learn) some parameters automatically so that users must specify by themselves.
  
-RVKDE provides two alternative ways to do its model selection as described in the following two sections.+RVKDE provides two alternative ways to do its parameter selection as described in the two following sections.
  
 ====== Cross-validation ====== ====== Cross-validation ======
   * Change the current path to __~/tmp__.  All following steps are supposed to execute in this path.   * Change the current path to __~/tmp__.  All following steps are supposed to execute in this path.
   * Cross-validation on __satimage.scale__.<code>   * Cross-validation on __satimage.scale__.<code>
-rvkde-0.2.3-final/rvkde --cv --classify --acc -n 5 -v rvkde-0.2.3-final/satimage.scale+rvkde-0.2.3-final/rvkde --cv --classify --acc -n 5 -v rvkde-0.2.3-final/satimage.scale -a 1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30
 </code> </code>
 Let's take a look at the command. Let's take a look at the command.
-^ --cv | Switch rvkde into cross-validation mode. | +^  --cv | Switch rvkde into cross-validation mode. | 
-^ --classify | Tell rvkde we want to do classification rather than regression now. | +^  --classify | Tell rvkde we want to do classification rather than regression now. | 
-^ --acc | Use [[wp>Accuracy|Accuracy]] as the evaluation index. | +^  --acc | Use [[wp>accuracy|accuracy]] as the evaluation index. | 
-^ -n | Do //n//-fold cross-validation. | +^  -n | Do __n__-fold cross-validation. | 
-^ -v | Followed by the dataset for cross-validation. |+^  -v | Followed by the dataset for cross-validation. | 
 +^  -a | Set the range (begin, end and step) of __alpha__ values of RVKDE.  In this example, 1 is the only possible __alpha__ value. | 
 +^  -b | Set the range (begin, end and step) of __beta__ values of RVKDE.  In this example, the possible __beta__ values are 1, 1.5 and 2. | 
 +^  --ks | Set the range (begin, end and step) of __ks__ values of RVKDE.  In this example, the possible __ks__ values are 1, 2, ... 30. | 
 +^  --kt | Set the range (begin, end and step) of __kt__ values of RVKDE.  In this example, the possible __kt__ values are also 1, 2, ... 30 since the default step is 1. |
  
   * The result looks like<code>   * The result looks like<code>
-[0.914994] a=1 b=1 s=8 t=10... +[0.918602] a=1 b=1 s=8 t=21... 
-</code>Which tell us the best parameter combination is __alpha__ = 1, __beta__ = 1, __ks__ = 9 and __kt__ = 10.  In addition, the Accuracy under the best model is 0.41195.+</code>Which tell us the best parameter combination is __alpha__ = 1, __beta__ = 1, __ks__ = 8 and __kt__ = 21.  In addition, the accuracy under the best parameters is 0.918602.
  
-===== Predict with the selected model =====+===== Predict with the selected parameters =====
 Now we have a parameter combination derived from cross-validation. Now we have a parameter combination derived from cross-validation.
-  * Use this model to predict //te.x5//.<code> +  * Use these parameters to predict __satimage.scale.t__.<code> 
-rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1 -b 1 --ks 9 --kt 10+rvkde-0.2.3-final/rvkde --predict --classify --acc -v rvkde-0.2.3-final/satimage.scale -V rvkde-0.2.3-final/satimage.scale.t -a 1 -b 1 --ks 8 --kt 21
 </code> </code>
 Let's take a look at the command. Let's take a look at the command.
-^ --predict | Switch rvkde into prediction mode (rather than cross-validation). |+^ --predict | Switch RVKDE into prediction mode (rather than cross-validation). |
 ^ -v | Followed by the training dataset. | ^ -v | Followed by the training dataset. |
 ^ -V | Followed by the testing dataset. | ^ -V | Followed by the testing dataset. |
  
   * The result looks like<code>   * The result looks like<code>
-[0.333333] a=1 b=1 s=9 t=10... +[0.9175] a=1 b=1 s=8 t=21... 
-</code>It seems that this model is not good for //te.x5//. +</code>It indicates that RVKDE can yield a accuracy of 0.9175 under this parameter combination when using __satimage.scale__ to predict __satimage.scale.t__.
- +
-===== How good cross-validation is ===== +
-We can enumerate many parameter combinations (as in cross-validation mode) to see the prediction performance. +
-  * Predict with a range of parameter combinations.<code> +
-rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1,5,1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30,1 +
-</code> +
-  * The result looks like<code> +
-[0.428986] a=1 b=1 s=10 t=30... +
-</code>It seems that the model selected with cross-validation is not the best one.+
  
 ====== Train, validate, and then test ====== ====== Train, validate, and then test ======
-Another common procedure for model selection is to create an independent validation set.  For example, you can use //tr.x5// and //va.x5// to do model selection and see how good the model is when applying on //te.x5//. +Another common procedure for parameter selection is to create an independent validation set.  For example, you can use __satimage.scale.tr__ and __satimage.scale.val__ to do parameter selection and see how good the parameters are when applying on __satimage.scale.t__. 
-  * Use //tr.x5// (as training set) and //va.x5// (as validation set) to select the model.<code> +  * Use __satimage.scale.tr__ (as training set) and __satimage.scale.val__ (as validation set) to select the parameters.<code> 
-rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/va.x5 -a 1,5,1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30,1+rvkde-0.2.3-final/rvkde --predict --classify --acc -v rvkde-0.2.3-final/satimage.scale.tr -V rvkde-0.2.3-final/satimage.scale.val -a 1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30,1
 </code> </code>
   * The result looks like<code>   * The result looks like<code>
-[0.381232] a=1 b=1.5 s=17 t=10...+[0.913599] a=1 b=1 s=8 t=23...
 </code> </code>
-  * Predict set2 with the selected model.<code> +  * Predict __satimage.scale.t__ with the selected parameters.<code> 
-rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1 -b 1.5 --ks 17 --kt 10+rvkde-0.2.3-final/rvkde --predict --classify --acc -v rvkde-0.2.3-final/satimage.scale.tr -V rvkde-0.2.3-final/satimage.scale.t -a 1 -b 1 --ks 8 --kt 23
 </code> </code>
   * The result looks like<code>   * The result looks like<code>
-[0.365651] a=1 b=1.5 s=17 t=10... +[0.917] a=1 b=1 s=8 t=23... 
-</code>It seems that the model selected is better than the previous one. +</code>It reveals that RVKDE yields very close accuracies with these two parameter selection schemes.
- +
-====== Summary of RVKDE ====== +
-Until now, you know that RVKDE has four parameters (//alpha//, //beta//, //ks//, and //kt//).  rvkde is a sophisticated machine learning package which has built-in functionalities for cross-validation and parameter-enumeration.  We will see some other parameters of rvkde in future exercises.  Of course, you could check [[http://zoro.ee.ncku.edu.tw/mbincku/rvkde/|the homepage of rvkde]] if you want to learn these facilities now. +
- +
-In addition, you learn two common model selection procedures.  In this exercise, you must try to find the best model for prediction the whole dataset, that is, F-measures of both //te.x5// and //te.x10// are good using only //tr.x5//, //va.x5//, //tr.x10//, and //va.x10//.+
rvkde/usage.1219110650.txt.gz · Last modified: by dirty