词频统计

请编写程序，对一段英文文本，统计其中所有不同单词的个数，以及词频最大的前10%的单词。

所谓“单词”，是指由不超过80个单词字符组成的连续字符串，但长度超过15的单词将只截取保留前15个单词字符。而合法的“单词字符”为大小写字母、数字和下划线，其它字符均认为是单词分隔符。

输入格式:

输入给出一段非空文本，最后以符号#结尾。输入保证存在至少10个不同的单词。

输出格式:

在第一行中输出文本中所有不同单词的个数。注意“单词”不区分英文大小写，例如“PAT”和“pat”被认为是同一个单词。

随后按照词频递减的顺序，按照词频:单词的格式输出词频最大的前10%的单词。若有并列，则按递增字典序输出。

输入样例：

This is a test.

The word "this" is the word with the highest frequency.

Longlonglonglongword should be cut off, so is considered as the same as longlonglonglonee.  But this_8 is different than this, and this, and this...#
this line should be ignored.

输出样例：（注意：虽然单词`the`也出现了4次，但因为我们只要输出前10%（即23个单词中的前2个）单词，而按照字母序，`the`排第3位，所以不输出。）

23
5:this
4:is


map映射到数组，统计并排序。

#include <iostream>
#include <algorithm>
#include <map>
#include <cstring>
using namespace std;
struct str
{
    char s[16];
    int num;
}ans[10000];
int no = 1;
bool cmp(str a,str b)
{
    if(a.num == b.num)return strcmp(a.s,b.s) < 0;
    return a.num > b.num;
}
int main()
{
    char ch,s[16];
    int c = 0;
    map<string,int> p;
    while((ch = cin.get()) != '#')
    {
        if(ch == '_')
        {
            if(c < 15)s[c ++] = ch;
        }
        else if(isdigit(ch))
        {
            if(c < 15)s[c ++] = ch;
        }
        else if(isalpha(ch))
        {
            if(c < 15)s[c ++] = tolower(ch);
        }
        else
        {
            if(c)
            {
                s[c] = '\0';
                c = 0;
                if(!p[s])
                {
                    strcpy(ans[no].s,s);
                    ans[no].num = 1;
                    p[s] = no ++;
                }
                else
                {
                    ans[p[s]].num ++;
                }
            }
        }
    }
    sort(ans + 1,ans + no,cmp);
    cout<<no - 1<<endl;
    for(int i = 1;i <= no / 10;i ++)
    {
        cout<<ans[i].num<<':'<<ans[i].s<<endl;
    }
}

2019重做：

#include <iostream>
#include <cstdlib>
#include <map>
#include <algorithm>
#include <sstream>
#include <vector>
using namespace std;
map<string,int> mp;
vector<string> vec;
bool cmp(string a,string b) {
    if(mp[a] == mp[b]) return a < b;
    return mp[a] > mp[b];
}
void check(string t) {
    int s = 0;
    for(int i = 0;i <= t.size();i ++) {
        if(!isalnum(t[i]) && t[i] != '_') {
            string temp = "";
            if(i - s > 15) temp = t.substr(s,15);
            else if(i != s) temp = t.substr(s,i - s);
            s = i + 1;
            if(temp == "") continue;
            if(!mp[temp] ++) vec.push_back(temp);
        }
        else t[i] = tolower(t[i]);
    }
}
int main() {
    string s;
    while(getline(cin,s)) {
        istringstream in(s);
        string t;
        while(in>>t) {
            check(t);
        }
        if(s.size() && s[s.size() - 1] == '#') break;
    }
    sort(vec.begin(),vec.end(),cmp);
    cout<<vec.size()<<endl;
    for(int i = 0;(i + 1) * 10 <= vec.size();i ++) {
        cout<<mp[vec[i]]<<':'<<vec[i]<<endl;
    }
}

2020：

import sys
dict = {}
flag = 0
for line in sys.stdin:
    if(flag == 1):
        line = ""
    for i in line.split(' '):
        s = 0
        i = i + " "
        for j in range(len(i)):
            if(not i[j].isalnum() and i[j] != "_"):
                if(i[j] == "#"):
                    flag = 1
                temp = ""
                if(j - s > 15):
                    temp = i[s:s + 15]
                elif(j != s):
                    temp = i[s:j]
                s = j + 1
                if(temp != ""):
                    temp = temp.lower()
                    if(dict.get(temp) == None):
                        dict[temp] = 1
                    else:
                        dict[temp] += 1
dict = sorted(dict.items(),key=lambda d:(d[0]),reverse=False)
dict = sorted(dict,key=lambda d:(d[1]),reverse=True)
len = len(dict)
print(len)
for i in range(int(len / 10)):
    print("%d:%s" % (dict[i][1],dict[i][0]))

发表于 2018-02-12 12:27 杰瑞不辣~ 阅读(321) 评论(0) 编辑收藏举报

输入格式:

输出格式:

输入样例：

输出样例：（注意：虽然单词the也出现了4次，但因为我们只要输出前10%（即23个单词中的前2个）单词，而按照字母序，the排第3位，所以不输出。）

公告

输出样例：（注意：虽然单词`the`也出现了4次，但因为我们只要输出前10%（即23个单词中的前2个）单词，而按照字母序，`the`排第3位，所以不输出。）