数据分析学习常用的数据源

“巧妇难为无米之炊”,很多人在学习数据挖掘时,总是感觉没有数据,然后就不了了之了。事实上,在互联网迅猛发展的今天,我们并不缺少数据,而是被数据所淹没,从而迷失了方向。在此总结一下获取数据的途径,我将其分为屠龙式和倚天式。所谓屠龙式也就是动态获取,可以写一些爬虫程序,从互联网上抓取数据,或者是通过程序自动生成数据;而倚天式就是静态获取,现在有很多公开测试数据挖掘算法的数据,还有一些数据挖掘相关竞赛会提供数据,并且有些竞赛有商业性质,所以数据真实性也很可靠。

在KDNuggets上有Datasets栏目,提供一些数据集,网址为:http://www.kdnuggets.com/datasets/

还有另外一个很好的资源网址为:http://kdd.ics.uci.edu/,里面包含的数据资源如下(按应用领域划分):

Direct Marketing 
  KDD CUP 1998 Data

GIS 
  Forest CoverType

Indexing 
  Corel Image Features

  Pseudo Periodic Synthetic Time Series

Intrusion Detection 
  KDD CUP 1999 Data

Process Control 
  Synthetic Control Chart Time Series

Recommendation Systems 
  Entree Chicago Recommendation Data

Robots 
  Pioneer-1 Mobile Robot Data

  Robot Execution Failures

Sign Language Recognition 
  Australian Sign Language Data

  High-quality Australian Sign Language Data

Text Categorization 
  20 Newsgroups Data

  Reuters-21578 Text Categorization Collection

  NSF Research Awards Abstracts 199 0-2003

World Wide Web 
  Microsoft Anonymous Web Data

  MSNBC Anonymous Web Data

  Syskill Webert Web Data

http://www.cs.toronto.edu/~roweis/data.html
http://www.cs.toronto.edu/~roweis/data.html
http://kdd.ics.uci.edu/summary.task.type.html
http://www-2.cs.cmu.edu/afs/cs.cmu.edu/project/theo-20/www/data/
http://www-2.cs.cmu.edu/afs/cs.cmu.edu/project/theo-11/www/wwkb/
http://www.phys.uni.torun.pl/~duch/software.html
在下面的网址可以找到reuters数据集http://www.research.att.com/~lewis/reuters21578.html

以下网址上有各种数据集:
http://kdd.ics.uci.edu/summary.data.type.html

进行文本分类,还有一个数据集是可以用的,即rainbow的数据集
http://www-2.cs.cmu.edu/afs/cs/project/theo-11/www/naive-bayes.html

UCI收集的机器学习数据集
ftp://pami.sjtu.edu.cn/
http://www.ics.uci.edu/~mlearn//MLRepository.htm

statlib

http://liama.ia.ac.cn/SCILAB/scilabindexgb.htm
http://lib.stat.cmu.edu/

样本数据库
http://kdd.ics.uci.edu/
http://www.ics.uci.edu/~mlearn/MLRepository.html

关于基金的数据挖掘的网站
http://www.gotofund.com/index.asp

http://lans.ece.utexas.edu/~strehl/

各种数据集:
http://kdd.ics.uci.edu/summary.data.type.html
http://www.mlnet.org/cgi-bin/mlnetois.pl/?File=datasets.html
http://lib.stat.cmu.edu/datasets/
http://dctc.sjtu.edu.cn/adaptive/datasets/

http://fimi.cs.helsinki.fi/data/
http://www.almaden.ibm.com/software/quest/Resources/index.shtml
http://miles.cnuce.cnr.it/~palmeri/datam/DCI/

进行文本分类&WEB
http://www-2.cs.cmu.edu/afs/cs/project/theo-11/www/naive-bayes.html

http://www.w3.org/TR/WD-logfile-960221.html
http://www.w3.org/Daemon/User/Config/Logging.html#AccessLog
http://www.w3.org/1998/11/05/WC-workshop/Papers/bala2.html
http://www-2.cs.cmu.edu/afs/cs.cmu.edu/project/theo-11/www/wwkb/
http://www.web-caching.com/traces-logs.html
http://www-2.cs.cmu.edu/webkb
http://www.cs.auc.dk/research/DP/tdb/TimeCenter/TimeCenterPublications/TR-75.pdf
http://www.cs.cornell.edu/projects/kddcup/index.html
时间序列数据的网址
http://www.stat.wisc.edu/~reinsel/bjr-data/

apriori算法的测试数据
http://www.almaden.ibm.com/cs/quest/syndata.html

数据生成器的链接
http://www.cse.cuhk.edu.hk/~kdd/data_collection.html
http://www.almaden.ibm.com/cs/quest/syndata.html
关联:
http://flow.dl.sourceforge.net/sourceforge/weka/regression-datasets.jar
http://www.almaden.ibm.com/software/quest/Resources/datasets/syndata.html#assocSynData

WEKA:
http://flow.dl.sourceforge.net/sourceforge/weka/regression-datasets.jar
1。A jarfile containing 37 classification problems, originally obtained from the UCI repository
http://prdownloads.sourceforge.net/weka/datasets-UCI.jar
2。A jarfile containing 37 regression problems, obtained from various sources
http://prdownloads.sourceforge.net/weka/datasets-numeric.jar
3。A jarfile containing 30 regression datasets collected by Luis Torgo
http://prdownloads.sourceforge.net/weka/regression-datasets.jar

癌症基因:
http://www.broad.mit.edu/cgi-bin/cancer/datasets.cgi

金融数据:
http://lisp.vse.cz/pkdd99/Challenge/chall.htm

以下网址上有各种数据集:
http://kdd.ics.uci.edu/summary.data.type.html

数据挖掘相关比赛以及数据集

2005 University of California data mining contest, predicting bad accounts and their churn date using real-world CRM data, deadline June 30, 2005.

·
ILP 2005 Challenge, on the prediction of functional classes of genes.

·
KDD Cup 2005, on classifying internet user search queries, deadline July 8.

·
Data Mining Cup 2005 (Chemnitz, Germany), for students; topic: How data mining can ascertain the risk of loss of payments andreduce this risk.

·
KDD Cup 2004, focuses on data-mining for a several performance criteria using datasets from bioinformatics and quantum physics.

·
InfoVis 2004 Contest, The History of InfoVis.

·
DATA MINING CUP 2004 (Chemnitz, Germany), for students.

·
InfoVis 2003 Contest: Visualization and Pair Wise Comparison of Trees, results announced Sep 5, 2003.

·
KDD Cup 2003, focuses on problems motivated by network mining and the analysis of usage logs.

·
DATA MINING CUP 2003 (Chemnitz, Germany). The task is to identify spam emails before they reach the user′s mailbox.

·
KDD Cup 2002, focus on data mining in molecular biology.

·
Student Data Mining Cup (2002), Chemnitz University and Prudential Systems.

posted on 2011-08-22 16:52  xuq  阅读(491)  评论(0编辑  收藏  举报

导航