User Tools

Site Tools


rvkde:usage

This is an old revision of the document!


About

If you are new to RVKDE, we suggest to read the README first, and then this document. The README introduces using two wrapper scripts (kde-train.pl and kde-predict.pl) to execute RVKDE, which is very similar to the procedure of using the well-know LIBSVM. An alternative way to use RVKDE is directly executing the rvkde (or rvkde.exe on Windows system) binary executable file. There are two benefits by doing so:

  1. The usage of rvkde is much powerful and flexible.
  2. You don't need to install the Perl environment, which might be a little annoying on Windows system.

The following sections are written mainly for Linux users. The mapping of most operations are trivial on Windows system.

Preparation

  • Create a directory for this document. Here we use ~/tmp as an example.
    cd ~
    mkdir tmp
  • Download and install RVKDE.
    cd ~/tmp
    wget http://mbi.ee.ncku.edu.tw/rvkde/res/rvkde-current-linux32.tgz
    tar zxvf rvkde-current-linux32.tgz

Classify the satimage dataset

  • Change the current path ~/tmp. All following steps are supposed to execute in this directory.
    cd ~/tmp
  • Classify the satimage dataset. Notice that your RVKDE version might be different.
    rvkde-0.2.3-final/rvkde --classify --train -v rvkde-0.2.3-final/satimage.scale -m rvkde-0.2.3-final/satimage.scale.model --ks 10 # training
    rvkde-0.2.3-final/rvkde --classify --predict -m rvkde-0.2.3-final/satimage.scale.model -V rvkde-0.2.3-final/satimage.scale.t -a 1 -b 1 --ks 10 --kt 10 # testing
  • Classify the satimage dataset in one-step.
    rvkde-0.2.3-final/rvkde --classify --predict -v rvkde-0.2.3-final/satimage.scale -V rvkde-0.2.3-final/satimage.scale.t -a 1 -b 1 --ks 10 --kt 10

Model selection

Most machine learning tools provide some parameters for users. For example, the k in knn classification algorithm and the k in k-means clustering algorithm. From the optimistic view, these parameters provide flexibility and make machine learning tools more powerful. However, from another point of view, these machine learning techniques cannot determine (or learn) some parameters automatically so that users must specify by themselves. To find a good parameter combination is so-called model selection. See Model Selection for more information.

RVKDE provides two alternative ways to do its model selection as described in the following two sections.

Cross-validation

  • Change the current path to ~/tmp. All following steps are supposed to execute in this path.
  • Cross-validation on tr.x5.
    rvkde/rvkde --cv --classify --f-measure -n 5 -v res/ly221/tr.x5

Let's take a look at the command.

–cv Switch rvkde into cross-validation mode.
–classify Tell rvkde we want to do classification rather than regression now.
–f-measure Use F-measure as the evaluation index.
-n Do n-fold cross-validation.
-v Followed by the dataset for cross-validation.
  • The result looks like
    [0.41195] a=1 b=1 s=9 t=10...

    Which tell us the best parameter combination is alpha = 1, beta = 1, ks = 9 and kt = 10. In addition, the F-measure under the best model is 0.41195.

Predict with the selected model

Now we have a parameter combination derived from cross-validation.

  • Use this model to predict te.x5.
    rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1 -b 1 --ks 9 --kt 10

Let's take a look at the command.

–predict Switch rvkde into prediction mode (rather than cross-validation).
-v Followed by the training dataset.
-V Followed by the testing dataset.
  • The result looks like
    [0.333333] a=1 b=1 s=9 t=10...

    It seems that this model is not good for te.x5.

How good cross-validation is

We can enumerate many parameter combinations (as in cross-validation mode) to see the prediction performance.

  • Predict with a range of parameter combinations.
    rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1,5,1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30,1
  • The result looks like
    [0.428986] a=1 b=1 s=10 t=30...

    It seems that the model selected with cross-validation is not the best one.

Train, validate, and then test

Another common procedure for model selection is to create an independent validation set. For example, you can use tr.x5 and va.x5 to do model selection and see how good the model is when applying on te.x5.

  • Use tr.x5 (as training set) and va.x5 (as validation set) to select the model.
    rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/va.x5 -a 1,5,1 -b 1,2,0.5 --ks 1,30,1 --kt 1,30,1
  • The result looks like
    [0.381232] a=1 b=1.5 s=17 t=10...
  • Predict set2 with the selected model.
    rvkde/rvkde --predict --classify --f-measure -v res/tr.x5 -V res/te.x5 -a 1 -b 1.5 --ks 17 --kt 10
  • The result looks like
    [0.365651] a=1 b=1.5 s=17 t=10...

    It seems that the model selected is better than the previous one.

Summary of RVKDE

Until now, you know that RVKDE has four parameters (alpha, beta, ks, and kt). rvkde is a sophisticated machine learning package which has built-in functionalities for cross-validation and parameter-enumeration. We will see some other parameters of rvkde in future exercises. Of course, you could check the homepage of rvkde if you want to learn these facilities now.

In addition, you learn two common model selection procedures. In this exercise, you must try to find the best model for prediction the whole dataset, that is, F-measures of both te.x5 and te.x10 are good using only tr.x5, va.x5, tr.x10, and va.x10.

rvkde/usage.1219110197.txt.gz · Last modified: by dirty