用requests库和BeautifulSoup4库爬取新闻列表

用requests库和BeautifulSoup4库,爬取校园新闻列表的时间、标题、链接、来源。


import requests
from bs4 import BeautifulSoup

res = requests.get('http://news.gzcc.cn/html/xiaoyuanxinwen/')
res.encoding='utf-8'
soup = BeautifulSoup(res.text,'html.parser')

for news in soup.select('li'):
if len(news.select('.news-list-title'))>0:
title=news.select('.news-list-title')[0].text
url=news.select('a')[0]['href']
time=news.select('.news-list-info')[0].contents[0].text
print(time,title,url)

 

 

posted on 2017-09-27 10:55  31黄智涛  阅读(205)  评论(0编辑  收藏  举报