-
安装必要库:
- 使用
pip安装requests:bash pip install requests - 安装
beautifulsoup4:bash pip install beautifulsoup4 - 安装
scrapy:bash pip install scrapy - 安装
selenium:bash pip install selenium
- 使用
-
发送HTTP请求:
-
使用
requests.get发送 GET 请求,处理常见错误和超时:import requests page = requests.get('https://example.com', timeout=5) if page.status_code == 200: print(page.text) elif page.status_code == 404: print("Page not found") else: print(f"HTTP错误:{page.status_code}")
-
-
处理验证码:
-
短信验证码:模拟输入短信验证码,需要处理动态内容。
-
图像验证码:使用
BeautifulSoup解析页面,定位输入框并模拟输入。 -
示例:
from bs4 import BeautifulSoup import requests url = 'https://www.example.com/login' page = requests.get(url, headers={'User-Agent': 'Mozilla/5. (Windows NT 10.; Win64; x64)'}) soup = BeautifulSoup(page.text, 'html.parser') captcha_input = soup.find('input', {'type': 'text', 'id': 'captcha'}) # 模拟输入验证码 captcha = '123456' captcha_input.send_keys(captcha) login_page = requests.post( url, data={'username': 'user', 'password': 'pass', 'captcha': captcha}, headers={'Referer': 'https://www.example.com/register'})
-
-
保持会话状态:
- 使用
requests.Session来维护 cookies:session = requests.Session() session.get('https://www.example.com/login') # 登录页面 response = session.post('https://www.example.com/member', data={'name': 'John'}) print(response.text)
- 使用
-
处理动态内容:
-
使用
Scrapy或Selenium处理 JavaScript 加载的内容。 -
Scrapy 示例:
from scrapy.spider import Spider, Selector from scrapy.http import Request class MySpider(Spider): def start_requests(self): yield Request('https://www.example.com', callback=self.parse) def parse(self, response): selector = Selector(response.text) links = selector.xpath('//a/@href').extract() for link in links: yield Request(link, callback=self.parse_link) def parse_link(self, response): selector = Selector(response.text) yield { '标题': selector.xpath('//h1/text()').extract_first(), '链接': response.url } # 执行爬虫 from scrapy.cmdline import run_spider run_spider('my_spider')
-
-
使用代理IP:
-
安装
proxychains模块:bash pip install proxychains && proxychains4 python your_script.py -
示例:
import requests # 使用代理IP proxies = { 'http': 'http://10.10.1.10:3128', 'https': 'http://10.10.1.10:108' } requests.get('https://www.example.com', proxies=proxies)
-
-
遵守隐私和政策:
- 遵守 GDPR 和其他隐私法规。
- 确保抓取行为不侵犯网站使用条款,避免滥用技术。
通过以上步骤,您可以在遵守技术和法律规范的前提下,使用 Python 访问国内网站,随着实践,逐步掌握各类请求和处理细节,确保在合法合规的框架内完成数据抓取任务。









