今天小编给大家分享一下Python异步爬取知乎热榜的方法的相关知识点,内容详细,逻辑清晰,相信大部分人都还太了解这方面的知识,所以分享这篇文章给大家参考一下,希望大家阅读完这篇文章后有所收获,下面我们一起来了解一下吧。
import asyncio
from bs4 import BeautifulSoup
import aiohttp
headers={
'user-agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.198 Safari/537.36',
'referer': 'https://www.baidu.com/s?tn=02003390_43_hao_pg&isource=infinity&iname=baidu&itype=web&ie=utf-8&wd=%E7%9F%A5%E4%B9%8E%E7%83%AD%E6%A6%9C'
}
async def getPages(url):
async with aiohttp.ClientSession(headers=headers) as session:
async with session.get(url) as resp:
print(resp.status) # 打印状态码
html=await resp.text()
soup=BeautifulSoup(html,'lxml')
items=soup.select('.HotList-item')
for item in items:
title=item.select('.HotList-itemTitle')[0].text
try:
abstract=item.select('.HotList-itemExcerpt')[0].text
except:
abstract='No Abstract'
hot=item.select('.HotList-itemMetrics')[0].text
try:
img=item.select('.HotList-itemImgContainer img')['src']
except:
img='No Img'
print("{}\n{}\n{}".format(title,abstract,img))
if __name__ == '__main__':
url='https://www.zhihu.com/billboard'
loop=asyncio.get_event_loop()
loop.run_until_complete(getPages(url))
loop.close()
发现详细链接、图片链接、问题摘要等都在JS里面(CSDN的开发者助手插件确实好用)
正则表达式获取上述信息:
接下来就是详细的代码啦
import asyncio
import json
import re
import aiohttp
headers={
'user-agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/86.0.4240.198 Safari/537.36',
'referer': 'https://www.baidu.com/s?tn=02003390_43_hao_pg&isource=infinity&iname=baidu&itype=web&ie=utf-8&wd=%E7%9F%A5%E4%B9%8E%E7%83%AD%E6%A6%9C'
}
async def getPages(url):
async with aiohttp.ClientSession(headers=headers) as session:
async with session.get(url) as resp:
print(resp.status) # 打印状态码
html=await resp.text()
regex=re.compile('"hotList":(.*?),"guestFeeds":')
text=regex.search(html).group(1)
# print(json.loads(text)) # json换成字典格式
for item in json.loads(text):
title=item['target']['titleArea']['text']
question=item['target']['excerptArea']['text']
hot=item['target']['metricsArea']['text']
link=item['target']['link']['url']
img=item['target']['imageArea']['url']
if not img:
img='No Img'
if not question:
question='No Abstract'
print("Title:{}\nPopular:{}\nQuestion:{}\nLink:{}\nImg:{}".format(title,hot,question,link,img))
if __name__ == '__main__':
url='https://www.zhihu.com/billboard'
loop=asyncio.get_event_loop()
loop.run_until_complete(getPages(url))
loop.close()
以上就是“Python异步爬取知乎热榜的方法”这篇文章的所有内容,感谢各位的阅读!相信大家阅读完这篇文章都有很大的收获,小编每天都会为大家更新不同的知识,如果还想学习更多的知识,请关注亿速云行业资讯频道。
亿速云「云服务器」,即开即用、新一代英特尔至强铂金CPU、三副本存储NVMe SSD云盘,价格低至29元/月。点击查看>>
免责声明:本站发布的内容(图片、视频和文字)以原创、转载和分享为主,文章观点不代表本网站立场,如果涉及侵权请联系站长邮箱:is@yisu.com进行举报,并提供相关证据,一经查实,将立刻删除涉嫌侵权内容。