显示标签为“mining”的博文。显示所有博文
显示标签为“mining”的博文。显示所有博文

2012年2月12日星期日

Basket analysis & Association Mining

I'm looking for suggestions on the right design approach in relation to a problem that resembles Basket analysis. The data to be analyzed is a dimension Attribute_DIM and contains an ID, Attribute and Attribute_Value. Some examples of the data are :

ID Attribute Attribute_Value

1 Color Black

1 Movie Men in Black

1 Book Of Human Bondage

2 Color White

2 Movie Men in Black

2 Book Grapes of Wrath

We need to be able to analyze multiple selections of the dimension. For example,

Men In Black

Grapes Of Wrath Of Human Bondage

Men In Black Black 1 1

White 1 0

I have had some success using the Association Algorithm Mining Model. I think It is an overkill since I only need descriptive and no predictive analysis.

I'm looking for some ideas on the right approach to this problem. Ideally, we need to present the data in a cube and have the possibility to perform member analysis of the dimension.

I have looked at several articles (including http://msdn2.microsoft.com/en-us/library/aa902637(sql.80).aspx and http://www.aspnetpro.net/newsletterarticle/2004/10/asp200410ri_l/asp200410ri_l.asp). I'm not convinced those are the solutions and would appreciate any insight into this problem.

Thank you,

Anna.

You might try an OLAP cube with a many to many dimension - described here: http://msdn2.microsoft.com/en-us/library/ms170463.aspx. There's also a short book dedicated to the feature by Marco Russo (http://www.lulu.com/content/812235)

|||

So, if I apply your solution, I would have two dimensions from the same source and establish a many-to-many relationship between the two, right? Wouldn't this limit the analysis to only two dimensions; i.e., if I needed to analyze 3 attributes and how the basket looks in that case; I need to be able to analyze an open-ended number of attributes on separate axes.

This article comes closest to describing my problem: http://msdn2.microsoft.com/en-us/library/aa902637(sql.80).aspx (Analysis Services: DISTINCT COUNT, Basket Analysis, and Solving the Multiple Selection of Members Problem). The article is based on SQL Server 2000. I would like to know if there is a simpler and different approach with SQL Server 2005.

|||Actually, in response to your original question, Association Rules is generally used for descriptive analysis, not predictive analysis. It is a rather recent innovation that has allowed AR to be used for predictive purposes. I think your best bet is to use AR.|||Thank you for the suggestion. I will go ahead with the idea of using a Mining model with Association rules.

2012年2月11日星期六

Basic steps as how to web usage mine in sql server 2005

Hi there,

I am doing a project on web usage mining of my universities server logs and im just wondering how i go about mining them in sql server 2005?

Do i mine them in one table? do i normalise the web log data? what algorithms will i use on them as im trying to get usage patterns from the users and also where most of the users come from.

Thanks in advance

Gary

Hi

Here’re some thoughts that might help you design your Data Mining for weblogs:

1. You can put the data in one or more tables as per the semantics of the data. Data which is an entity by itself should be put in a flat table with one key per entry (user sessions on the web server for example), whereas some data naturally will have a many to one relation with the primary data (page visits per session for example) and can be put in a separate table with primary/foreign key relationship. SQL Server 2005 will model them as case table/nested table for the purpose of mining.

2. Normalize: Depends on what your data looks like. If you want to run clustering algorithm and your data has two attributes, A and B and A is 10 times more important than B, you should normalize accordingly. If the actual values of A are in order of thousand and actual values ob B are in order of tens and they are equally important, you should again normalize. However, if you want to use the as-is value without a weight, you do not have to normalize the data.

3. Usage Patterns, like a sequence of page visits can be modeled using sequence clustering algorithm. User categorization based on attributes might be a clustering algorithm. If you have more information on what you want to find out, I might be able to suggest more specific choices.

Hope this helps

Shuvro