如何使用python统计单词出现次数

发布时间：2020-04-28 09:55:42 阅读：1044 作者：小新栏目：编程语言

Python开发者专用服务器限时活动，0元免费领，库存有限，领完即止！点击查看>>

这篇文章主要为大家详细介绍了如何使用python统计单词出现次数，文中示例代码介绍的非常详细，具有一定的参考价值，感兴趣的小伙伴们可以参考一下。

python统计单词出现次数

做单词词频统计，用字典无疑是最合适的数据类型，单词作为字典的key，单词出现的次数作为字典的 value，很方便地就记录好了每个单词的频率，字典很像我们的电话本，每个名字关联一个电话号码。

下面是具体的实现代码，实现了从importthis.txt文件读取单词，并统计出现次数最多的5个单词。

# -*- coding:utf-8 -*-
import io
import re

class Counter:
    def __init__(self, path):
        """
        :param path: 文件路径
        """
        self.mapping = dict()
        with io.open(path, encoding="utf-8") as f:
            data = f.read()
            words = [s.lower() for s in re.findall("\w+", data)]
            for word in words:
                self.mapping[word] = self.mapping.get(word, 0) + 1

    def most_common(self, n):
        assert n > 0, "n should be large than 0"
        return sorted(self.mapping.items(), key=lambda item: item[1], reverse=True)[:n]

if __name__ == '__main__':
    most_common_5 = Counter("importthis.txt").most_common(5)
    for item in most_common_5:
        print(item)

执行效果：

('is', 10)
('better', 8)
('than', 8)
('the', 6)
('to', 5)

关于如何使用python统计单词出现次数就分享到这里了，希望以上内容可以对大家有一定的参考价值，可以学以致用。如果喜欢本篇文章，不妨把它分享出去让更多的人看到。

亿速云「云服务器」，即开即用、新一代英特尔至强铂金CPU、三副本存储NVMe SSD云盘，价格低至29元/月。点击查看>>

向AI问一下细节

如何使用python统计单词出现次数

猜你喜欢

最新资讯

相关推荐

开发者交流群：

相关标签